AI Engineer, Model Distillation
Job Description
Tesla AI is hiring an AI Engineer, Model Distillation (onsite) in Palo Alto, CA to build and scale distillation systems for autonomy and robotics models.
Responsibilities
- Design and execute distillation pipelines that transfer complex reasoning, vision-language understanding, and action policies from large teacher models to ultra-efficient students for FSD (Full Self-Driving), Optimus, and Digital Optimus
- Develop and iterate on advanced distillation methods, including logit-level, sequence-level, and trajectory-level distillation
- Create novel distillation recipes that combine synthetic data generation, soft-label training, supervised fine-tuning (SFT), and Reinforcement Learning (RL)
- Run experiments to measure distillation scaling behavior across teacher size, student capacity, student architecture choices, data mixtures, and compute allocation, focusing on capability retention, latency, and on-robot performance
- Build and maintain infrastructure for efficient teacher inference, distillation data curation, and distributed student training; identify and resolve compute and memory bottlenecks end-to-end
- Evaluate distilled models for teacher fidelity and product metrics such as task success, robustness, and real-world behaviors; address gaps when students diverge from teachers
- Partner with cross-functional teams to ship distilled models into production while meeting performance, safety, and reliability standards
- Contribute tools and frameworks to keep distillation pipelines reproducible, measurable, and reusable across Tesla AI model families
Requirements
- Deep, proven expertise in deep learning fundamentals, with experience training, compressing, or distilling large-scale language, vision, or multimodal models
- Strong understanding of teacher-student interaction, synthetic data workflows, and tradeoffs between capability, latency, and compute
- In-depth knowledge of loss design such as KL and other soft-target objectives, plus familiarity with modern neural architectures (including mixture of experts and hybrid attention)
- Hands-on experience with post-training methods including distillation, SFT, and policy optimization / RL
- Expertise in distributed computing and large-scale training or inference pipeline engineering
- Proficiency in Python and strong software engineering best practices
- Experience with deep learning frameworks such as PyTorch, TensorFlow, or JAX
- Ability to collaborate effectively in a cross-functional team environment
- Strong problem-solving skills with experience troubleshooting complex issues across data, training, and deployment
Technologies
- Python
- PyTorch
- TensorFlow
- JAX
- Mixture of experts
- Hybrid attention
Compensation and Benefits
- Expected compensation: USD 124,000 - 558,000 per year + cash and stock awards + benefits
- Pay may vary based on market location, job-related knowledge, skills, and experience; total compensation may include additional elements depending on the role
Benefits
- Medical plans with plan options including $0 payroll deduction
- Family-building benefits: fertility, adoption, and surrogacy
- Dental (including orthodontic coverage) and vision plans with $0 paycheck contribution (plan options available)
- Company paid HSA contribution when enrolled in the High-Deductible medical plan with HSA
- Healthcare and Dependent Care Flexible Spending Accounts (FSA)
- 401(k) with employer match
- Employee Stock Purchase Plans and other financial benefits
- Company paid Basic Life and AD&D
- Short-term and long-term disability insurance (90 day waiting period)
- Employee Assistance Program
- Sick and Vacation time (Flex time for salary positions, accrued hours for Hourly positions) and Paid Holidays
- Back-up childcare and parenting support resources
- Voluntary benefits: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
- Weight Loss and Tobacco Cessation Programs
- Tesl a Babies program
- Commuter benefits
- Employee discounts and perks program
What to Expect
- Work on real-world autonomy at large scale by training frontier-scale foundation models on multimodal telemetry, vision, language, and robotic action data
- Study capability transfer from large teacher models into compact student models that must meet strict latency, power, and reliability constraints
- Leverage high GPU resources per engineer to train frontier-scale teachers, generate distillation corpora at volume, and iterate on student architectures
- Run distillation experiments across language, multimodal, and policy models at fidelity and scale; deploy distilled models onto FSD / Optimus in the real world