AI Engineer 5 (FM Hosting, LLM Inference)
Job Description
Build and scale foundation-model systems for production AI at the center of Capital One’s AI vision. On the Intelligent Foundations and Experiences (IFX) team, you will design and deploy AI software that supports foundation model training and LLM inference, with a focus on strong engineering practices for evaluation, guardrails, observability, and optimization across cost, latency, throughput, and scalability.
You will collaborate with engineers, research scientists, technical program managers, and product managers to deliver AI-powered capabilities that help associates work more effectively and improve how customers interact with Capital One, using responsible and scalable approaches.
What you’ll do
- Design, develop, test, deploy, and support AI software components spanning foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
- Partner across the organization to deliver AI-powered products and proprietary solutions that create high-leverage value for customers.
- Leverage a broad stack of AI technologies, including AWS Ultraclusters, Huggingface, VectorDBs, and PyTorch.
- Develop state-of-the-art foundation model optimization techniques to improve scalability, cost, latency, and throughput in large-scale production systems.
- Contribute to the technical vision and long-term roadmap for foundational AI systems at Capital One.
- Design and optimize multi-model orchestration pipelines that integrate LLMs, vector search, and domain-specific models into unified systems.
- Establish and lead cost-performance governance reviews across AI systems, including tracking GPU utilization, model throughput, and inference cost efficiency.
- Lead team design councils or design review boards to support technical consistency and compliance with AI engineering standards.
- Mentor Principal and Manager-level AI engineers to support cross-domain learning and elevate organizational technical maturity.
What we’re looking for
- Bachelor’s degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 6 years developing AI and ML algorithms or technologies, or a Master’s degree plus at least 4 years developing AI and ML algorithms or technologies.
- At least 6 years programming with Python, Go, Scala, CUDA, or Java.
Tools you’ll use
- AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, AWS
- Google Cloud, Azure
- Python, Go, Scala, CUDA, Java
- C++, C#, Golang
- LLMs, vector search
Benefits
- Incentive compensation, which may include cash bonus(es) and/or long-term incentives (LTI).
- Comprehensive, competitive, and inclusive health, financial, and other benefits designed to support total well-being.
Location and compensation
McLean, VA (onsite)
Salary: USD 229,900 - 262,400 per year
Preferred qualifications
- Experience leading development of AI systems with tradeoff decisions around cost, latency, throughput, and accuracy.
- 7 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g., AWS, Google Cloud, Azure, or equivalent private cloud).
- Experience designing, developing, delivering, and supporting complex AI systems.
- Experience developing AI and ML algorithms or technologies (e.g., LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang.
- Experience building and applying state-of-the-art optimization techniques for training and inference software to improve hardware utilization, latency, throughput, and cost.
- Experience building agentic AI systems and agentic workflows.
- Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers.
- Experience architecting and integrating heterogeneous AI systems, including rule-based, retrieval-augmented, and generative components, into unified production pipelines.
- Experience defining and enforcing standards for ethical AI deployment, including explainability, fairness, and human-in-the-loop review processes.
- Ability to balance model performance and operational cost through dynamic inference strategies and model compression.
- Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs).