AI Engineer 4 (AI Foundations)
Job Description
Capital One is building foundation model capabilities to help teams across the company deliver AI-powered experiences in responsible and scalable ways. In the Intelligent Foundations and Experiences (IFX) group, this role will help design, deploy, and evolve foundation model systems that support new ways of working for associates and improved interactions for customers. This onsite position is based in Cambridge, MA.
Role overview
You will build and deploy foundation model systems end to end, including AI software components for training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, evaluation, governance, and observability. The work includes owning architecture and reliability targets through defined AI service-level objectives, while collaborating across engineering, research, and product to bring AI deployments to production.
Responsibilities
- Partner with engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how associates work and how customers interact with Capital One.
- Design, develop, test, deploy, and support AI software components such as foundation model training, LLM inference, agents and multi-agent workflows, similarity search, guardrails, and model evaluation, along with experimentation, governance, and observability.
- Use a broad stack of Open Source and SaaS AI technologies including AWS Ultraclusters, Huggingface, VectorDBs, and PyTorch, with additional tools as needed.
- Introduce foundation model optimization techniques to improve performance, including scalability, cost, latency, and throughput, for production AI systems.
- Help shape the technical vision and long-term roadmap for foundational AI systems at Capital One.
- Own end-to-end architecture for complex AI systems, with a focus on maintainability, observability, and ethical alignment.
- Define and maintain AI reliability service-level objectives covering latency, uptime, and model performance drift.
- Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines.
- Lead cross-functional technical reviews for new AI system deployments, supporting security, data governance, and compliance requirements.
- Mentor Principal and Senior Associates on scalable design, performance tuning, and translating research into production.
Requirements
- Bachelor’s degree in Computer Science/AI/Electrical Engineering/Computer Engineering or related fields, plus at least 4 years of experience developing AI and ML algorithms or technologies; or a Master’s degree in a related field plus at least 2 years of experience.
- At least 4 years of programming experience with Python, Go, Scala, CUDA, or Java.
Technologies
- AWS Ultraclusters
- Huggingface
- VectorDBs
- PyTorch
- AWS, Google Cloud, Azure
- Python, Go, Scala, CUDA, Java, C++, C#, Golang
- GPU, TPU
Team context
- The Intelligent Foundations and Experiences (IFX) team brings Capital One’s AI vision to life.
- The team partners across the company to advance state-of-the-art science and AI engineering and to build and deploy proprietary solutions central to the business and valuable to millions of customers.
- AI models and platforms on IFX empower teams to enhance products with the transformative power of AI in responsible and scalable ways.
Preferred qualifications
- Experience leading development AI systems and making tradeoffs around cost, latency, throughput, and accuracy.
- 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (AWS, Google Cloud, Azure, or equivalent private cloud).
- Experience designing, developing, delivering, and supporting AI services.
- Experience developing AI and ML algorithms or technologies such as LLM inference, similarity search and VectorDBs, guardrails, and memory using Python, C++, C#, Java, CUDA, or Golang.
- Experience applying state-of-the-art techniques to optimize training and inference software for improved hardware utilization, latency, throughput, and cost.
- Experience building agentic AI systems and agentic workflows.
- Experience designing distributed systems for model training, evaluation, and online inference at petabyte scale.
- Experience defining AI model governance processes such as producibility, lineage tracking, and automated retaining schedules.
- Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms.
Compensation and location
- Location: Cambridge, MA (onsite)
- Salary: USD 197,300 - 225,100 per year