Sr. Lead Machine Learning Engineer
Job Description
Capital One offers a performance-based incentive compensation opportunity and a comprehensive, competitive, and inclusive benefits package focused on health, financial support, and overall well-being. This onsite role is based in McLean, VA and supports work on production ML systems within an Agile environment.
What you’ll do
Work on end-to-end machine learning delivery, from technical design through production monitoring. You will design, build, and deliver machine learning models and components that address real business problems, partnering closely with Product and Data Science teams. Your responsibilities include developing application code alongside ML model development, implementing automated testing, and enabling deployment workflows.
- Collaborate with Product and Data Science to solve real-world business problems using ML models and components.
- Apply ML expertise to drive infrastructure decisions across model choice, data and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation.
- Develop and validate ML models and application code, write and test code, and automate tests and deployment.
- Build and enhance software for state-of-the-art big data and ML applications as part of a cross-functional Agile team.
- Retrain, maintain, and monitor ML models in production.
- Use cloud-based architectures and platforms (leveraging or building as needed) to deliver optimized ML models at scale.
- Construct optimized data pipelines to feed ML models.
- Apply continuous integration and continuous deployment best practices, including test automation and monitoring, to support successful releases.
- Support secure, well-managed code, govern models from a risk perspective, and follow best practices in Responsible and Explainable AI.
Requirements
- Bachelor’s Degree
- 8+ years designing and building data-intensive solutions using distributed computing (internship experience does not apply)
- 4+ years programming with Python, Scala, or Java
- 3+ years building, scaling, and optimizing ML systems
- 2+ years leading teams developing ML solutions
Preferred qualifications
- Master’s or doctoral degree in computer science, electrical engineering, mathematics, or a similar field
- Experience developing and deploying ML solutions in a public cloud such as AWS, Azure, or Google Cloud Platform
- 4+ years on-the-job experience with an industry recognized ML framework such as scikit-learn, PyTorch, Dask, Spark, Kubeflow, or TensorFlow
- 3+ years developing performant, resilient, and maintainable code
- 3+ years of experience with data gathering and preparation for ML models
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
- 3+ years building production-ready data pipelines that feed ML models
- Ability to communicate complex technical concepts clearly to a variety of audiences
Tech stack
Python, Scala, Java, AWS, Azure, Google Cloud Platform, scikit-learn, PyTorch, Dask, Spark, Kubeflow, TensorFlow
Compensation and location
McLean, VA (onsite): USD 229,900 - 262,400 per year
Capital One may consider sponsoring a new qualified applicant for employment authorization for this position. This role is expected to accept applications for a minimum of 5 business days. No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination. Capital One promotes a drug-free workplace. Candidates hired to work in other locations will be subject to the pay range associated with that location.