Machine Learning Engineer 5
Job Description
Capital One is building AI-powered solutions to support responsible, governed risk management, and this role in Risk Tech helps design and deploy machine learning systems that can be used in real business workflows. You will work with the GRC team and partners to deliver proprietary models and scalable platforms, with an emphasis on production operations, explainable AI, and model governance. This position is based in McLean, VA (onsite).
Salary: USD 229,900 - 262,400 per year.
Responsibilities
- Design, build, and deliver ML models and components that address real-world business needs, in collaboration with Product and Data Science teams.
- Build and scale multi-tenant platforms that support large footprint ML model training and/or serving at scale.
- Make ML infrastructure decisions grounded in ML modeling expertise, including model choice, data and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation.
- Solve complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment.
- Contribute within a cross-functional Agile team to create and improve software powering big data and ML applications.
- Retrain, maintain, and monitor models in production.
- Use or build cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale.
- Construct optimized data pipelines to feed ML models.
- Apply CI/CD best practices, including test automation and monitoring, to support reliable deployments of both ML models and application code.
- Maintain well-managed code to reduce vulnerabilities, ensure risk-aware model governance, and follow best practices for Responsible and Explainable AI.
- Use programming languages such as Python, Scala, or Java.
Requirements
- Bachelor's degree or higher in Computer Science, Machine Learning, or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering).
- At least 6 years of experience programming with Python, Java, Golang, or C++.
- At least 6 years of Machine Learning experience using industry standard frameworks PyTorch or Tensorflow and libraries such as Pandas, NumPy, Scikit-learn.
- At least 6 years using and operating large scale distributed systems (Spark, Ray) to prepare AI/ML data.
- At least 4 years deploying and operating Machine Learning solutions in production and operating production services in the cloud (AWS, GCP, Azure), using Kubernetes to manage large scale containerized ML systems.
Technologies
Python, Scala, Java, Golang, C++, PyTorch, Tensorflow, Pandas, NumPy, Scikit-learn, Spark, Ray, AWS, GCP, Azure, Kubernetes, CI/CD, Agile
Benefits
- Performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI).
- Comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being.
Preferred Qualifications
- Master's or doctoral degree in computer science, electrical engineering, mathematics, or related field.
- 5+ years of experience optimizing ML algorithms, configurations, and infrastructure.
- 5+ years of experience following software development best practices including source control, testing, code reviews, CI/CD, etc.
- 5+ years of experience building resilient software solutions with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring, alarms, and preparing incident response plans.
- 5+ years of experience working with Machine Learning techniques (Supervised, semi-supervised, and unsupervised, reinforcement learning, etc.), model types (Regression, Classification, Clustering, etc.), model architectures (RNNs, CNNs, LSTMs, Transformers), training concepts (loss function, hyperparameters, regularization), and evaluating model accuracy and diagnosing common issues (underfitting, overfitting).
- 5+ years of experience designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models.
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents.
- Ability to communicate complex technical and machine learning concepts clearly to a variety of audiences.
Additional Information
- Capital One will consider sponsoring a new qualified applicant for employment authorization for this position.
- This role is expected to accept applications for a minimum of 5 business days.
- No agencies please.
- Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws.
- Capital One promotes a drug-free workplace.
- Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with applicable laws regarding criminal background inquiries.
- If you need an accommodation while applying, contact Capital One Recruiting at 1-800-304-9102 or via email at [email protected].
- For technical support or questions about Capital One's recruiting process, email [email protected].
Location note: Minimum and maximum full-time annual salaries for this role vary by location. For McLean, VA: $229,900 - $262,400. For Richmond, VA: $209,000 - $238,500.