EngineerJobs.io
← Back to all jobs

Job Description

Benefits and culture

Become part of JPMorganChase in Palo Alto, onsite, as a Lead Machine Learning Engineer - MLOps on the Recommendation Engine team. You will work with a modern MLOps stack to design, train, and deploy ML models, including distributed training on GPU-enabled clusters, real-time and batch inference, and production monitoring and validation. The role offers a competitive salary and a comprehensive benefits package, reflecting a culture that prioritizes collaboration with product, architecture, and engineering peers to deliver scalable, impactful solutions.

  • Salary: USD 164,350 - 260,000 per year
  • Commission-based pay and/or discretionary incentive compensation, paid in cash and/or forfeitable equity
  • Comprehensive health care coverage
  • On-site health and wellness centers
  • Retirement savings plan
  • Backup childcare
  • Tuition reimbursement
  • Mental health support
  • Financial coaching

In this role, you will lead the development, deployment, and ongoing operation of ML models on a robust MLOps platform. Expect collaboration across teams to implement scalable, efficient solutions, with a focus on performance, reliability, and continuous improvement.

Responsibilities

  • Build, deploy, and maintain robust pipelines for distributed training on GPU-enabled clusters to support scalable ML workflows
  • Develop and manage pipelines for high-throughput real-time inference as well as batch inference, prioritizing performance and reliability
  • Implement quantization techniques and deploy large language models (LLMs) to maximize efficiency and resource utilization
  • Oversee the management and optimization of vector databases to support advanced AI and ML applications
  • Establish and maintain comprehensive monitoring and observability pipelines to ensure system health and rapid issue resolution
  • Collaborate with cross-functional teams to integrate new technologies and continuously improve existing infrastructure
  • Partner with product, architecture, and other engineering teams to define scalable and performant technical solutions

Requirements

  • MS in Computer Science or related Engineering field with 4+ years of experience
  • BS in Computer Science or related Engineering field with 6+ years of experience
  • Solid knowledge and extensive experience in Python and in cloud computing, preferably AWS
  • Understanding of quantization techniques such as PTQ, AWQ etc. used to quantize LLMs for accelerating inference on specific GPU architectures
  • Experience in systems engineering fundamentals: caching, CUDA, autoscaling, high throughput, low latency, x-region resilient applications
  • Deep knowledge and passion for data science fundamentals, training and deploying models
  • Experience in monitoring and observability tools to monitor model input/output and features stats
  • Operational experience in big data/ML tools such as Ray, DuckDB, Spark and in training/inference systems such as Ray, vllm/SGLang
  • Solid grounding in engineering fundamentals and analytical mindset

Technologies

  • Python
  • AWS
  • CUDA
  • PTQ
  • AWQ
  • Ray
  • DuckDB
  • Spark
  • vllm
  • SGLang
  • Docker
  • Kubernetes
  • ECS
  • Airflow
  • Kubeflow
  • vector databases

Similar Jobs