Lead Machine Learning Engineer-MLOps
Job Description
Benefits and culture
Become part of JPMorganChase in Palo Alto, onsite, as a Lead Machine Learning Engineer - MLOps on the Recommendation Engine team. You will work with a modern MLOps stack to design, train, and deploy ML models, including distributed training on GPU-enabled clusters, real-time and batch inference, and production monitoring and validation. The role offers a competitive salary and a comprehensive benefits package, reflecting a culture that prioritizes collaboration with product, architecture, and engineering peers to deliver scalable, impactful solutions.
- Salary: USD 164,350 - 260,000 per year
- Commission-based pay and/or discretionary incentive compensation, paid in cash and/or forfeitable equity
- Comprehensive health care coverage
- On-site health and wellness centers
- Retirement savings plan
- Backup childcare
- Tuition reimbursement
- Mental health support
- Financial coaching
In this role, you will lead the development, deployment, and ongoing operation of ML models on a robust MLOps platform. Expect collaboration across teams to implement scalable, efficient solutions, with a focus on performance, reliability, and continuous improvement.
Responsibilities
- Build, deploy, and maintain robust pipelines for distributed training on GPU-enabled clusters to support scalable ML workflows
- Develop and manage pipelines for high-throughput real-time inference as well as batch inference, prioritizing performance and reliability
- Implement quantization techniques and deploy large language models (LLMs) to maximize efficiency and resource utilization
- Oversee the management and optimization of vector databases to support advanced AI and ML applications
- Establish and maintain comprehensive monitoring and observability pipelines to ensure system health and rapid issue resolution
- Collaborate with cross-functional teams to integrate new technologies and continuously improve existing infrastructure
- Partner with product, architecture, and other engineering teams to define scalable and performant technical solutions
Requirements
- MS in Computer Science or related Engineering field with 4+ years of experience
- BS in Computer Science or related Engineering field with 6+ years of experience
- Solid knowledge and extensive experience in Python and in cloud computing, preferably AWS
- Understanding of quantization techniques such as PTQ, AWQ etc. used to quantize LLMs for accelerating inference on specific GPU architectures
- Experience in systems engineering fundamentals: caching, CUDA, autoscaling, high throughput, low latency, x-region resilient applications
- Deep knowledge and passion for data science fundamentals, training and deploying models
- Experience in monitoring and observability tools to monitor model input/output and features stats
- Operational experience in big data/ML tools such as Ray, DuckDB, Spark and in training/inference systems such as Ray, vllm/SGLang
- Solid grounding in engineering fundamentals and analytical mindset
Technologies
- Python
- AWS
- CUDA
- PTQ
- AWQ
- Ray
- DuckDB
- Spark
- vllm
- SGLang
- Docker
- Kubernetes
- ECS
- Airflow
- Kubeflow
- vector databases