Machine Learning Engineer 4 (Manager, IC)
Backend Developer
Manager
Ai Ml
Artificial Intelligence
Automation
Azure Machine Learning
Big Data
Bigdata
Cloud
Cloud Infrastructure
Cloud Machine Learning
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Cloud Technology
Data & Ai
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Pipeline
Data Platform
Data Processing
Data Science
Deep Learning
DevOps
DevSecOps
Engineering
Engineering Software
Facilities Management
Google Cloud
Information Technology (IT)
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning
Machine Learning Engineer
Machine Learning Engineering
Machine Learning Modeling
Machine Learning Operations
Machine Learning Pipelines
Machine Learning Platform
Platform Engineering
Programming
Programming Language
Programming Languages
PyTorch
Risk Management
scikit-learn
Security Automation
Software Engineering
TensorFlow
Job Description
Capital One is seeking a Machine Learning Engineer 4 (Manager, IC) in Chicago, IL (onsite) to design, build, deploy, and monitor production machine learning models and supporting components at scale. The role partners with Product and Data Science teams and applies cloud and CI/CD practices to deliver reliable ML systems.
What You’ll Do
- Design, build, and deliver ML models and components to solve real-world business problems in collaboration with Product and Data Science teams
- Use expertise in ML modeling techniques and practical challenges to guide infrastructure decisions, including model and data selection, feature selection, training, hyperparameter tuning, dimensionality, bias/variance considerations, and validation
- Address complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
- Work within a cross-functional Agile team to create and improve software supporting big data and ML applications
- Retrain, maintain, and monitor production models
- Leverage and/or build cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale
- Construct optimized data pipelines to supply ML models
- Apply continuous integration and continuous deployment best practices, including test automation and monitoring, to support successful model and application releases
- Maintain well-managed code to reduce vulnerabilities, ensure ML governance from a risk perspective, and follow best practices in Responsible and Explainable AI
- Use programming languages such as Python, Scala, or Java
Required Qualifications
- Bachelor’s Degree or higher in Computer Science, Machine Learning, or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 4 years of experience programming with Python, Java, Golang, or C++
- At least 4 years of Machine Learning experience using industry standard frameworks PyTorch or Tensorflow and libraries such as Pandas, NumPy, Scikit-learn
- At least 4 years of experience using and operating large-scale distributed systems (Spark, Ray) to prepare AI or Machine Learning data
- At least 2 years of experience deploying and operating Machine Learning solutions in production, operating production services in cloud environments (AWS, GCP, Azure), and using Kubernetes to manage large-scale containerized ML systems
Preferred Qualifications
- Master’s or Doctoral Degree in Computer Science, Electrical Engineering, Mathematics, or a related field
- 3+ years of experience optimizing ML algorithms, configurations, and infrastructure
- 3+ years of experience following software development best practices, including source control, testing, code reviews, and CI/CD
- 3+ years of experience building resilient software solutions with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring, alarms, and preparing incident response plans
- 3+ years of experience with Machine Learning techniques and model types, including supervised, semi-supervised, unsupervised, and reinforcement learning; regression, classification, and clustering; as well as architectures such as RNNs, CNNs, LSTMs, and Transformers
- 3+ years of experience with training concepts (loss function, hyperparameters, regularization) and evaluating model accuracy, diagnosing issues (underfitting, overfitting), and addressing them
- 3+ years of experience designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models
- 1+ years of experience as a technical lead developing ML solutions using industry best practices, patterns, and automation
- Authored or co-authored a paper on a ML technique, model, or proof of concept
Technologies
- Python, Scala, Java, Golang, C++
- PyTorch, Tensorflow
- Pandas, NumPy, Scikit-learn
- Spark, Ray
- AWS, GCP, Azure
- Kubernetes
Benefits
- Performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
- Comprehensive, competitive, and inclusive set of health, financial, and other benefits
Compensation
- Chicago, IL: USD 179,400 - 204,700 per year
- New York, NY: USD 215,200 - 245,600 per year