Senior Machine Learning Engineer (AI Foundations)
Job Description
Capital One's AI Foundations team is seeking a Senior Machine Learning Engineer to productionize ML applications at scale from its New York presence. The role centers on designing robust ML architectures, ensuring high availability and performance, and driving end-to-end ML engineering work, including data pipelines, cloud infrastructure, CI/CD, and responsible AI practices. This onsite position in New York, NY offers a salary range of USD 176,500 - 201,400 per year.
What you'll do
- Design, build, and deliver ML models and components that address real-world business needs, collaborating with Product and Data Science teams.
- Inform ML infrastructure decisions using modeling techniques and considerations such as model choice, data, feature selection, training, hyperparameters, dimensionality, bias/variance, and validation.
- Address complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment.
- Collaborate within a cross-functional Agile team to create and enhance software powering state-of-the-art big data and ML applications.
- Retrain, maintain, and monitor models in production environments.
- Leverage or build cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale.
- Construct optimized data pipelines to feed ML models.
- Apply CI/CD best practices, including test automation and monitoring, to ensure successful deployment of ML models and application code.
- Ensure code quality and security, maintain model governance, and adhere to Responsible and Explainable AI practices.
- Use programming languages such as Python, Scala, or Java.
Basic Qualifications
- Bachelor’s Degree.
- At least 4 years of experience programming with Python, Scala, or Java (internship experience does not apply).
- At least 3 years of experience designing and building data-intensive solutions using distributed computing.
- At least 2 years of on-the-job experience with an industry-recognized ML framework (scikit-learn, PyTorch, Dask, Spark, or TensorFlow).
- At least 1 year of experience productionizing, monitoring, and maintaining models.
Preferred Qualifications
- 1+ years of experience building, scaling, and optimizing ML systems.
- 1+ years of experience with data gathering and preparation for ML models.
- 2+ years of experience developing performant, resilient, and maintainable code.
- Experience developing and deploying ML solutions in a public cloud such as AWS, Azure, or Google Cloud Platform.
- Master's or doctoral degree in computer science, electrical engineering, mathematics, or a similar field.
- 3+ years of experience with distributed file systems or multi-node database paradigms.
- Contributed to open source ML software.
- Authored or co-authored a paper on an ML technique, model, or proof of concept.
- 3+ years of experience building production-ready data pipelines that feed ML models.
- Experience designing, implementing, and scaling complex data pipelines for ML models and evaluating their performance.
- Experience leveraging interactive AI tooling to accelerate productivity beyond basic code completion.
Technologies
- Python
- Scala
- Java
- scikit-learn
- PyTorch
- Dask
- Spark
- TensorFlow
- AWS
- Azure
- Google Cloud Platform
Benefits
- Health benefits
- Financial benefits
- Performance-based incentive compensation (cash bonuses and/or long-term incentives)