This position is no longer accepting applications
Closed on August 28, 2026.
This role is filled — get an email when new Engineering roles open on EngineerJobs.io:
Lead AI Engineer (FM Hosting, LLM Inference)
Get alerted when similar jobs are posted — set up a New Engineering jobs on EngineerJobs.io alert.
See other roles at Capital One.
Job Description
Capital One's IFX team is seeking a Lead AI Engineer to drive foundation model training, large language model inference, and the design, development, and deployment of AI powered products. This onsite role in New York, NY blends hands-on engineering with cross-functional collaboration to redefine how associates work and how customers interact with Capital One. The position offers a salary range of USD 215,200 to 245,600 per year.
Responsibilities
- Collaborate with engineers, research scientists, technical program managers, and product managers across functions to deliver AI driven products that transform how associates work and how customers interact with Capital One.
- Design, build, test, deploy, and maintain AI software components including foundation model training, LLM inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
- Utilize a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more.
- Develop and apply state-of-the-art LLM optimization techniques to improve the performance metrics of large-scale production AI systems, including scalability, cost, latency, and throughput.
- Contribute to the technical vision and long-term roadmap for foundational AI systems at Capital One.
Requirements
- Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields, with at least 4 years of experience developing AI and ML algorithms or technologies.
- Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields, with at least 2 years of experience developing AI and ML algorithms or technologies.
- At least 4 years of experience programming with Python, Go, Scala, or Java.
- 6 years of experience deploying scalable and responsible AI solutions on cloud platforms such as AWS, Google Cloud, Azure, or an equivalent private cloud.
- Experience designing, developing, delivering, and supporting AI services.
- Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang.
- Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost.
- Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production.
Technologies
- Python
- Go
- Scala
- Java
- AWS Ultraclusters
- Huggingface
- VectorDBs
- Nemo Guardrails
- PyTorch
Benefits
- Health benefits
- Financial benefits
- Incentives including performance-based incentive compensation (cash bonuses and/or long-term incentives)