Sr. Lead Machine Learning Engineer
Job Description
Capital One is hiring a Sr. Lead Machine Learning Engineer to productionize machine learning applications and systems at scale within an Agile delivery environment. The role focuses on building ML architectures, deploying and operating models reliably, and applying governance and responsible AI practices.
Responsibilities
- Design, build, and deliver machine learning models and supporting components that address real-world business needs in collaboration with Product and Data Science teams
- Shape ML infrastructure direction by applying knowledge of modeling techniques and common issues, including model selection, data and feature selection, model training, hyperparameter tuning, dimensionality, bias/variance, and validation
- Develop and test production application code, build and validate ML models, and automate tests and deployment
- Partner within a cross-functional Agile team to create and improve software supporting big data and ML applications
- Retrain, maintain, and monitor ML models in production
- Leverage and/or build cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale
- Construct optimized data pipelines that feed ML models
- Apply CI/CD best practices, including test automation and monitoring, to support dependable releases of ML model and application code
- Manage code to reduce vulnerabilities, ensure risk-governed model practices, and follow Responsible and Explainable AI best practices
- Use programming languages such as Python, Scala, or Java
Required Qualifications
- Bachelor’s Degree
- 8+ years of experience designing and building data-intensive solutions using distributed computing (internship experience does not apply)
- 4+ years of experience programming with Python, Scala, or Java
- 3+ years building, scaling, and optimizing machine learning systems
- 2+ years experience leading teams developing ML solutions
- Experience developing and deploying ML solutions in a public cloud such as AWS, Azure, or Google Cloud Platform
- 4+ years of on-the-job experience with an industry-recognized ML framework: scikit-learn, PyTorch, Dask, Spark, or TensorFlow
- 3+ years experience building performant, resilient, and maintainable code
- 3+ years experience with data gathering and preparation for ML models
- 3+ years of people management experience
- 3+ years experience building production-ready data pipelines that feed ML models
- Ability to communicate complex technical concepts to a variety of audiences
- Experience using interactive AI tooling to accelerate productivity, utilizing capabilities beyond basic code completion
Technologies
- Programming: Python, Scala, Java
- Cloud: AWS, Azure, Google Cloud Platform
- ML frameworks and data/ML tools: scikit-learn, PyTorch, Dask, Spark, TensorFlow
Preferred Qualifications
- Master’s or Doctoral Degree in computer science, electrical engineering, mathematics, or a similar field
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
- Experience using interactive AI tooling to accelerate productivity, utilizing capabilities beyond basic code completion
Compensation and Location
- Location: Richmond, VA (onsite)
- Salary: USD 209,000 - 238,500 per year
Benefits
- Performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
- Comprehensive, competitive, and inclusive set of health, financial and other benefits supporting total well-being
Additional Information
- Expected application acceptance window: minimum of 5 business days
- No agencies please
- Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination
- Capital One promotes a drug-free workplace
- Capital One will consider qualified applicants with a criminal history in accordance with applicable laws
- Accommodation request contact: 1-800-304-9102 or [email protected]
- Technical support and recruiting process questions: [email protected]