Machine Learning Engineer 5 (Senior Manager, IC)
Manager
Ai Ml
Application Security
Artificial Intelligence
Automation
Big Data
Cloud
Cloud Infrastructure
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Cloud Technology
Data Analysis
Data Platform
Data Science
DevOps
DevSecOps
Engineering
Engineering Software
Facilities Management
Information Technology (IT)
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning Engineer
Machine Learning Engineering
Machine Learning Inference
Machine Learning Operations
Platform Engineering
Programming
Programming Language
Risk Management
Security Automation
Software Security
Job Description
As a Machine Learning Engineer 5 (Senior Manager, IC) in Risk Tech, you will help build and deploy proprietary risk management solutions using advanced AI. The role focuses on designing and scaling ML models and platforms, operating models in production, and partnering with cross-functional teams to deliver big data and ML capabilities.
Location and Work Setting
- McLean, VA (onsite)
- Full-time
Compensation
- McLean, VA: USD 209,000 - 262,400 per year
- Richmond, VA: USD 209,000 - 262,400 per year
Performance-based incentive compensation eligibility may include cash bonus(es) and/or long term incentives (LTI). Incentives may be discretionary or non-discretionary depending on the plan.
Key Responsibilities
- Design, build, and/or deliver ML models and components to solve real-world business problems in collaboration with Product and Data Science teams
- Build and scale massive multi-tenant platforms for ML training and/or serving at scale
- Apply ML modeling expertise to influence ML infrastructure decisions across model choice, data and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation
- Solve complex technical problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
- Collaborate in a cross-functional Agile team to create and enhance software for state-of-the-art big data and ML applications
- Retrain, maintain, and monitor models in production
- Leverage or build cloud-based architectures, technologies, and/or platforms to deliver optimized ML models at scale
- Construct optimized data pipelines to feed ML models
- Use CI/CD best practices, including test automation and monitoring, to support successful deployment of ML models and application code
- Ensure code is well-managed to reduce vulnerabilities, models are well-governed from a risk perspective, and ML follows best practices in Responsible and Explainable AI
- Use programming languages such as Python, Scala, or Java
Required Qualifications
- Bachelor's Degree or higher in Computer Science, Machine Learning, or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 6 years of experience programming with Python, Java, Golang, or C++
- At least 6 years of Machine Learning experience using industry-standard frameworks PyTorch or Tensorflow and libraries (Pandas, NumPy, Scikit-learn)
- At least 6 years of experience using and operating large-scale distributed systems (Spark, Ray) to prepare AI/ML data
- At least 4 years of experience deploying and operating Machine Learning solutions in production and operating production services in the cloud (AWS, GCP, Azure), including using Kubernetes to manage large-scale containerized ML systems
Preferred Qualifications
- Master's or doctoral degree in computer science, electrical engineering, mathematics, or related field
- 5+ years of experience optimizing ML algorithms, configurations, and infrastructure
- 5+ years of experience following software development best practices including source control, testing, code reviews, and CI/CD
- 5+ years of experience building resilient software solutions with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring, alarms, and incident response planning
- 5+ years of experience working with ML techniques (supervised, semi-supervised, unsupervised, reinforcement learning) and model types (regression, classification, clustering)
- 5+ years of experience with model architectures (RNNs, CNNs, LSTMs, Transformers), training concepts (loss function, hyperparameters, regularization), and diagnosing common model issues (underfitting, overfitting)
- 5+ years of experience designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
- Ability to communicate complex technical and machine learning concepts clearly to a variety of audiences
Technologies
- Python, Scala, Java, Golang, C++
- PyTorch, Tensorflow, Pandas, NumPy, Scikit-learn
- Spark, Ray
- AWS, GCP, Azure
- Kubernetes
Application Notes
- This role is expected to accept applications for a minimum of 5 business days.
- No agencies please.
Equal Opportunity and Workplace Standards
- Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws.
- Capital One promotes a drug-free workplace.
- Capital One will consider qualified applicants with a criminal history in a manner consistent with applicable laws regarding criminal background inquiries.
Work Authorization and Sponsorship
Capital One will consider sponsoring a new qualified applicant for employment authorization for this position.
Accommodations and Recruiting Contacts
- For accommodations related to applying, contact Capital One Recruiting at 1-800-304-9102 or [email protected].
- For technical support or questions about the recruiting process, email [email protected].