AI Engineer 4 (AI Foundations, LLM Core and Agentic AI)
Job Description
Build foundational AI systems that power real-world experiences for associates and customers. Capital One is hiring an AI Engineer 4 on the Intelligent Foundations and Experiences (IFX) team to design, deploy, and improve production AI capabilities including LLM inference and agentic AI. The role emphasizes end-to-end architecture, reliability, and optimization across scalability, cost, latency, and throughput.
What you’ll do
- Partner with engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how associates work and how customers interact with Capital One.
- Design, develop, test, deploy, and support AI software components, including foundation model training, LLM inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
- Apply and integrate a broad stack of Open Source and SaaS AI technologies, including AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more.
- Introduce state-of-the-art foundation model optimization techniques to improve production performance, including scalability, cost, latency, and throughput.
- Contribute to technical vision and the long-term roadmap for foundational AI systems across Capital One.
- Own end-to-end architecture for complex AI systems with an emphasis on maintainability, observability, and ethical alignment.
- Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift.
- Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines.
- Lead cross-functional technical reviews for new AI deployments, ensuring security, data governance, and compliance standards are met.
- Mentor Principal and Senior Associates on scalable design, performance tuning, and research-to-production translation.
Qualifications
- Bachelor’s degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years developing AI/ML algorithms or technologies, or a Master’s degree plus at least 2 years developing AI/ML algorithms or technologies.
- At least 4 years of programming experience with Python, Go, Scala, CUDA, or Java.
Preferred experience
- Experience leading development of AI systems with tradeoffs across cost, latency, throughput, and accuracy.
- 6 years deploying scalable and responsible AI solutions on cloud platforms such as AWS, Google Cloud, Azure, or equivalent private cloud.
- Experience designing, developing, delivering, and supporting AI services.
- Experience building AI/ML algorithms and services (for example: LLM inference, similarity search and vector databases, guardrails, memory) using Python, C++, C#, Java, CUDA, or Golang.
- Experience optimizing training and inference software for improved hardware utilization, latency, throughput, and cost.
- Experience developing agentic AI systems and agentic workflows.
- Ability to apply new AI research and systems judiciously in production.
- Proficiency designing distributed systems for model training, evaluation, and online inference at petabyte scale.
- Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules.
- Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms.
Tools and technologies you may work with
- AWS Ultraclusters, Huggingface, VectorDBs, PyTorch
- Python, Go, Scala, CUDA, Java
- AWS, Google Cloud, Azure
- C++, C#, Golang
Location and compensation
San Jose, CA (onsite) | USD 215,200 - 245,600 per year. Performance-based incentive compensation may be available, including cash bonus(es) and/or long term incentives (LTI).
Capital One also offers a comprehensive, competitive, and inclusive set of health, financial, and other benefits that support your total well-being.