EngineerJobs.io
← Back to all jobs

Job Description

Capital One is building production-grade machine learning capabilities to power real-time decisioning across customers’ credit journeys. In this Agile role, you will help design, develop, and deploy machine learning models and the supporting infrastructure at scale, using modern software practices and Kubernetes-based technologies. You’ll partner closely with Product and Data Science to take ML from model development through reliable operations in production.

Responsibilities

  • Design, build, and deliver machine learning models and components that address real-world business needs in collaboration with Product and Data Science teams
  • Guide machine learning infrastructure choices using expertise across modeling and evaluation topics, including model and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation
  • Solve complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
  • Work within a cross-functional Agile team to create and improve software for big data and state-of-the-art ML applications
  • Retrain, maintain, and monitor models in production
  • Leverage or build cloud-based architectures, technologies, and platforms to deploy optimized ML solutions at scale
  • Construct optimized data pipelines that supply data to ML models
  • Apply CI/CD best practices with test automation and monitoring to support successful releases of ML models and application code
  • Maintain well-managed code to reduce vulnerabilities, ensure risk-governed models, and follow Responsible and Explainable AI best practices
  • Use programming languages such as Python, Scala, or Java

Requirements

  • Bachelor’s Degree
  • 8+ years of experience designing and building data-intensive solutions using distributed computing (internship experience does not apply)
  • 4+ years of experience programming with Python, Scala, or Java
  • 3+ years of experience building, scaling, and optimizing machine learning systems
  • 2+ years of experience leading teams developing ML solutions

Technologies

  • Python, Kubernetes, Scala, Java
  • AWS, Azure, Google Cloud Platform
  • scikit-learn, PyTorch, Dask, Spark, TensorFlow

Preferred Qualifications

  • Master’s or Doctoral Degree in computer science, electrical engineering, mathematics, or a similar field
  • Experience developing and deploying ML solutions in a public cloud such as AWS, Azure, or Google Cloud Platform
  • 4+ years of on-the-job experience with an industry-recognized ML framework such as scikit-learn, PyTorch, Dask, Spark, or TensorFlow
  • 3+ years of experience developing performant, resilient, and maintainable code
  • 3+ years of experience with data gathering and preparation for ML models
  • 3+ years of people management experience
  • ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
  • 3+ years of experience building production-ready data pipelines that feed ML models
  • Ability to communicate complex technical concepts clearly to a variety of audiences
  • Experience leveraging interactive AI tooling to accelerate productivity, using capabilities beyond basic code completion

Location: McLean, VA (onsite)
Salary: USD 229,900 - 262,400 per yearly
Eligible compensation: performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
Benefits: comprehensive, competitive, and inclusive health, financial, and other benefits

Similar Jobs