Capital One is seeking a Machine Learning Engineer to design, build, and deliver machine learning models and reusable components that solve real business needs. In this role, you will help scale multi-tenant ML platforms, deploy and monitor models in production, and strengthen end-to-end pipelines from data preparation to reliable releases.
This is a hands-on engineering position based in McLean, VA (onsite), supporting responsible and explainable AI practices while working within an Agile product and data environment.
What you’ll do
- Design, build, and deliver ML models and components in collaboration with Product and Data Science
- Build and scale massive multi-tenant platforms for large footprint model training and/or serving
- Apply ML modeling knowledge to inform ML infrastructure decisions across model choice, data and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation
- Develop and test application code, build and validate ML models, and automate tests and deployment
- Work in a cross-functional Agile team to create and enhance big data and ML application software
- Retrain, maintain, and monitor models deployed in production
- Leverage or build cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale
- Construct optimized data pipelines that feed ML models
- Use CI/CD best practices, including test automation and monitoring, to support successful deployments
- Ensure code is well-managed to reduce vulnerabilities, that models are well-governed from a risk perspective, and that ML follows Responsible and Explainable AI best practices
- Use programming languages such as Python, Scala, or Java
Required qualifications
- Bachelor’s Degree or higher in Computer Science, Machine Learning, or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 6 years programming with Python, Java, Golang, or C++
- At least 6 years Machine Learning experience using PyTorch or Tensorflow and libraries including Pandas, NumPy, and Scikit-learn
- At least 6 years experience operating large-scale distributed systems such as Spark and Ray for preparing AI/ML data
- At least 4 years deploying and operating Machine Learning solutions in production services in the cloud (AWS, GCP, Azure) and using Kubernetes to manage containerized ML software systems at scale
Technologies
- Python, Scala, Java, Golang, C++
- PyTorch, Tensorflow
- Pandas, NumPy, Scikit-learn
- Spark, Ray
- AWS, GCP, Azure
- Kubernetes
Benefits
- Comprehensive, competitive, and inclusive set of health, financial, and other benefits supporting total well-being
- Performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
Preferred qualifications
- Master’s or doctoral degree in computer science, electrical engineering, mathematics, or a related field
- 5+ years of experience optimizing ML algorithms, configurations, and infrastructure
- 5+ years of experience following software development best practices including source control, testing, code reviews, CI/CD, etc.
- 5+ years of experience building resilient software solutions with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring, alarms, and preparing incident response plans
- 5+ years of experience working with ML techniques including supervised, semi-supervised, unsupervised, and reinforcement learning, and model types such as regression, classification, and clustering
- 5+ years of experience with model architectures including RNNs, CNNs, LSTMs, and Transformers, plus training concepts and evaluating model accuracy and diagnosing issues like underfitting and overfitting
- 5+ years of experience designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
- Ability to communicate complex technical and machine learning concepts clearly to a variety of audiences
Compensation: $229,900 - $262,400 per year for Machine Learning Engineer 5 in McLean, VA.
Additional information: Applications are expected to be accepted for a minimum of 5 business days. No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider qualified applicants with a criminal history in a manner consistent with applicable laws. If you require an accommodation related to searching or applying on the website, contact Capital One Recruiting at 1-800-304-9102 or RecruitingAccommodation@capitalone.com. For technical support or questions about the recruiting process, email Careers@capitalone.com. Capital One does not provide, endorse, or guarantee third-party products or information available through this site.