Capital One is building Agile teams that productionize machine learning applications at scale, with a focus on responsible and explainable AI. As a Sr. Lead Machine Learning Engineer in New York, NY (onsite), you will help design ML architectures end to end, from model development and deployment automation to production monitoring and governance.
This role combines hands-on engineering with leadership in ML systems: you will collaborate with Product and Data Science teams to deliver real-world solutions, while using cloud platforms and CI/CD practices to keep model and application delivery reliable. You’ll also ensure code quality and risk-aware governance across the ML lifecycle.
Responsibilities
- Design, build, and deliver machine learning models and components for business problems in collaboration with Product and Data Science teams
- Guide ML infrastructure decisions based on modeling fundamentals, including model selection, data and feature selection, training, hyperparameter tuning, dimensionality, and validation; address bias/variance tradeoffs
- Solve complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
- Work within a cross-functional Agile team to create and enhance software that supports big data and ML applications
- Retrain, maintain, and monitor models in production
- Leverage or build cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale
- Construct optimized data pipelines to feed ML models
- Use CI/CD best practices, including test automation and monitoring, to support successful deployment of ML models and application code
- Manage code to reduce vulnerabilities, keep models well-governed from a risk perspective, and follow best practices in Responsible and Explainable AI
- Use programming languages such as Python, Scala, or Java
Requirements
- Bachelor’s Degree
- 8+ years of experience designing and building data-intensive solutions using distributed computing (internship experience does not apply)
- 4+ years programming with Python, Scala, or Java
- 3+ years building, scaling, and optimizing ML systems
- 2+ years experience leading teams developing ML solutions
- Master’s or Doctoral Degree in computer science, electrical engineering, mathematics, or a similar field
- Experience developing and deploying ML solutions in a public cloud such as AWS, Azure, or Google Cloud Platform
- 4+ years on-the-job experience with an industry-recognized ML framework such as scikit-learn, PyTorch, Dask, Spark, or XGboost
- 3+ years building performant, resilient, and maintainable code
- 3+ years experience with data gathering and preparation for ML models
- 3+ years of people management experience
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
- 3+ years building production-ready data pipelines that feed ML models
- Ability to clearly communicate complex technical concepts to a variety of audiences
- Experience leveraging interactive AI tooling to accelerate productivity, using capabilities beyond basic code completion
Technologies
Python, Scala, Java, AWS, Azure, Google Cloud Platform, scikit-learn, PyTorch, Dask, Spark, XGboost
Benefits
- Comprehensive, competitive, and inclusive set of health, financial, and other benefits to support your total well-being
- Performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
Salary: USD 250,800 - 286,200 per year.