Lead Machine Learning Engineer
Job Description
Join Capital One in Richmond, VA on site to help shape AI driven risk technology. You will collaborate with a multidisciplinary team to deliver AI powered products, with a strong emphasis on responsible and explainable AI. This role offers ongoing opportunities to grow your technical leadership while enjoying comprehensive benefits that support you and your family.
Benefits and culture
We provide a comprehensive suite of benefits, including health and financial security, along with additional programs designed to support your wellbeing and career growth. In this role you will work closely with engineers, data scientists, product managers, and designers to create impactful AI solutions that improve how our associates work and deliver value to customers. You will contribute to a culture that values responsible AI practices, clear governance, and strong collaboration across teams.
Responsibilities
- Collaborate with a cross functional team of engineers, data scientists, product managers and designers to deliver AI driven products that transform workflows and deliver customer value.
- Design, develop, test, deploy and support AI software components using machine learning models, including model evaluation and experimentation, large language model inference, similarity search, guardrails, governance, observability and agentic AI.
- Fine tune, develop and evaluate machine learning and foundation models.
- Work within an Agile, cross functional team to create and enhance software that utilizes state of the art AI and ML capabilities.
- Provide thought leadership and technical vision for the long term roadmap of pioneering AI systems at Capital One.
- Leverage a broad mix of Open Source and SaaS AI technologies.
- Inform ML infrastructure decisions with an understanding of modeling techniques and associated challenges.
- Retrain, maintain and monitor models in production.
- Construct optimized data pipelines to feed machine learning models.
- Ensure code quality and security, maintain model governance from a risk perspective, and apply Responsible and Explainable AI practices.
Requirements
- Bachelor’s Degree
- At least 6 years of experience designing and building data‑intensive solutions using distributed computing (internship experience does not apply)
- At least 4 years of experience programming with Python, Scala or Java
- At least 2 years of experience building, scaling and optimizing ML systems
Technologies
- Python
- Scala
- Java
- scikit-learn
- PyTorch
- Dask
- Spark
- TensorFlow
- AWS Bedrock
- Google Cloud
- Azure
- Retrieval Augmented Generation (RAG)
The ideal candidate
- Enjoys building systems, takes pride in code quality, and is motivated to contribute to banking for good.
- Strong communicator who can explain complex technical concepts to non technical partners across the business, including presentations to larger audiences.
- Keeps up with the latest AI research and can interpret scientific publications to apply novel techniques in production.
- Adapts quickly, brings clarity to large, undefined problems, asks thoughtful questions, and communicates findings concisely. Willing to share new ideas even when they are unproven.
- Highly technical with a solid foundation in engineering and mathematics; ability to leverage hardware, software, and AI to identify optimization opportunities.
- Resilient and capable of forging new paths to achieve business goals when the route is not obvious.
- Continual interest in AI research and systems, with a track record of applying advanced techniques in production environments.