Staff Machine Learning Engineer
Job Description
Build production-grade AI for game experiences at Unity Technologies SF. This onsite role in Mountain View focuses on taking cutting-edge computer vision and multi-modal research into dependable, scalable systems across cloud, server, and on-device targets, with a team culture centered on measurement, engineering rigor, and cross-functional collaboration.
Responsibilities
- Help set the technical vision and roadmap for computer vision and multi-modal AI models, including transformers, diffusion models, vision-language models, and JEPA-style generative architectures.
- Design and implement models for image and video understanding and generation, including segmentation, detection, and dense prediction, plus multi-modal reasoning over images, text, and 3D inputs.
- Make architecture and delivery decisions across quality, capability, latency, and cost for deployment targets, including training strategy, data pipelines, and evaluation.
- Own the path from research prototype to production: training, fine-tuning, distillation, export, and serving, spanning cloud GPUs to efficient on-device inference where needed.
- Collaborate closely with research scientists to translate novel CV and multi-modal architectures into deployable, well-engineered implementations.
- Build scalable multi-modal inference systems that handle diverse inputs (images, video, text, primitives, and metadata) and generate outputs from semantic predictions to pixel-level generation.
- Rapidly adopt field breakthroughs in areas such as vision-language pretraining and alignment, efficient diffusion (consistency models, flow matching), efficient attention (including FlashAttention and linear-attention variants), and tokenization or representation learning for vision.
- When latency or device constraints require it, apply compression approaches such as quantization, pruning, and knowledge distillation, and work with runtimes including TensorRT, ONNX Runtime, CoreML, and TFLite.
- Lead and mentor ML engineers, define engineering best practices and code review standards, and establish rigorous benchmarking and evaluation methodology.
- Partner with research, platform engineers, product managers, and runtime teams to align ML capabilities with product roadmaps and target-platform constraints.
- Champion a measurement culture by defining KPIs for model quality, accuracy, latency, memory, and cost, and ensuring consistent tracking.
Requirements
- 6+ years of ML engineering experience with substantial depth in computer vision and/or multi-modal modeling.
- Production experience with transformer-based and diffusion-based vision models, including examples such as ViT, CLIP/SigLIP-style encoders, Stable Diffusion, and DETR/SAM-style architectures.
- Strong command of the full model lifecycle, including data curation, training and fine-tuning, evaluation, and serving at scale.
- Familiarity with efficient attention, diffusion samplers, multi-modal fusion, and vision-language alignment techniques.
- Strong Python skills and modern deep-learning tooling (notably PyTorch), plus solid software engineering fundamentals.
- Proven technical leadership, including setting direction, influencing cross-functional partners, and growing engineers.
Technologies
Python, PyTorch, TensorRT, ONNX Runtime, CoreML, TFLite, FlashAttention, ViT, CLIP, SigLIP, Stable Diffusion, DETR, SAM
Benefits
- Comprehensive health, life, and disability insurance
- Commute subsidy
- Employee stock ownership
- Competitive retirement/pension plans
- Generous vacation and personal days
- Support for new parents through leave and family-care programs
- Office food snacks
- Mental Health and Wellbeing programs and support
- Employee Resource Groups
- Global Employee Assistance Program
- Training and development programs
- Volunteering and donation matching program
Additional Information
- Location: Mountain View, CA (onsite)
- Salary: USD 172,200 - 283,900 per year
- Zone A: $218,400 - $283,900; Zone B: $194,100 - $252,300; Zone C: $172,200 - $223,900
- Beyond base salary, the role may be eligible for equity awards and participation in company incentive plans (such as annual discretionary bonuses or sales commissions)
- Final offer amount depends on geographic location, relevant experience, professional background, and skill set
You Might Also Have
- Experience with world-model, video-generation, or neural rendering pipelines (NeRF, 3DGS, or similar)
- Experience deploying models to constrained or on-device targets, including quantization (INT8/INT4/FP16), pruning, distillation, and runtimes such as CoreML, TFLite, ONNX
- Familiarity with mobile SoC accelerators (Apple Neural Engine, Qualcomm Hexagon/Adreno, ARM Mali) or compiler stacks such as MLIR, TVM, or XLA
- Contributions to open-source ML frameworks or peer-reviewed CV/ML research publications
- Background in real-time graphics or game engine pipelines (Metal, Vulkan, OpenGL ES)