EngineerJobs.io
← Back to all jobs

Job Description

Build production-grade AI for game experiences at Unity Technologies SF. This onsite role in Mountain View focuses on taking cutting-edge computer vision and multi-modal research into dependable, scalable systems across cloud, server, and on-device targets, with a team culture centered on measurement, engineering rigor, and cross-functional collaboration.

Responsibilities

  • Help set the technical vision and roadmap for computer vision and multi-modal AI models, including transformers, diffusion models, vision-language models, and JEPA-style generative architectures.
  • Design and implement models for image and video understanding and generation, including segmentation, detection, and dense prediction, plus multi-modal reasoning over images, text, and 3D inputs.
  • Make architecture and delivery decisions across quality, capability, latency, and cost for deployment targets, including training strategy, data pipelines, and evaluation.
  • Own the path from research prototype to production: training, fine-tuning, distillation, export, and serving, spanning cloud GPUs to efficient on-device inference where needed.
  • Collaborate closely with research scientists to translate novel CV and multi-modal architectures into deployable, well-engineered implementations.
  • Build scalable multi-modal inference systems that handle diverse inputs (images, video, text, primitives, and metadata) and generate outputs from semantic predictions to pixel-level generation.
  • Rapidly adopt field breakthroughs in areas such as vision-language pretraining and alignment, efficient diffusion (consistency models, flow matching), efficient attention (including FlashAttention and linear-attention variants), and tokenization or representation learning for vision.
  • When latency or device constraints require it, apply compression approaches such as quantization, pruning, and knowledge distillation, and work with runtimes including TensorRT, ONNX Runtime, CoreML, and TFLite.
  • Lead and mentor ML engineers, define engineering best practices and code review standards, and establish rigorous benchmarking and evaluation methodology.
  • Partner with research, platform engineers, product managers, and runtime teams to align ML capabilities with product roadmaps and target-platform constraints.
  • Champion a measurement culture by defining KPIs for model quality, accuracy, latency, memory, and cost, and ensuring consistent tracking.

Requirements

  • 6+ years of ML engineering experience with substantial depth in computer vision and/or multi-modal modeling.
  • Production experience with transformer-based and diffusion-based vision models, including examples such as ViT, CLIP/SigLIP-style encoders, Stable Diffusion, and DETR/SAM-style architectures.
  • Strong command of the full model lifecycle, including data curation, training and fine-tuning, evaluation, and serving at scale.
  • Familiarity with efficient attention, diffusion samplers, multi-modal fusion, and vision-language alignment techniques.
  • Strong Python skills and modern deep-learning tooling (notably PyTorch), plus solid software engineering fundamentals.
  • Proven technical leadership, including setting direction, influencing cross-functional partners, and growing engineers.

Technologies

Python, PyTorch, TensorRT, ONNX Runtime, CoreML, TFLite, FlashAttention, ViT, CLIP, SigLIP, Stable Diffusion, DETR, SAM

Benefits

  • Comprehensive health, life, and disability insurance
  • Commute subsidy
  • Employee stock ownership
  • Competitive retirement/pension plans
  • Generous vacation and personal days
  • Support for new parents through leave and family-care programs
  • Office food snacks
  • Mental Health and Wellbeing programs and support
  • Employee Resource Groups
  • Global Employee Assistance Program
  • Training and development programs
  • Volunteering and donation matching program

Additional Information

  • Location: Mountain View, CA (onsite)
  • Salary: USD 172,200 - 283,900 per year
  • Zone A: $218,400 - $283,900; Zone B: $194,100 - $252,300; Zone C: $172,200 - $223,900
  • Beyond base salary, the role may be eligible for equity awards and participation in company incentive plans (such as annual discretionary bonuses or sales commissions)
  • Final offer amount depends on geographic location, relevant experience, professional background, and skill set

You Might Also Have

  • Experience with world-model, video-generation, or neural rendering pipelines (NeRF, 3DGS, or similar)
  • Experience deploying models to constrained or on-device targets, including quantization (INT8/INT4/FP16), pruning, distillation, and runtimes such as CoreML, TFLite, ONNX
  • Familiarity with mobile SoC accelerators (Apple Neural Engine, Qualcomm Hexagon/Adreno, ARM Mali) or compiler stacks such as MLIR, TVM, or XLA
  • Contributions to open-source ML frameworks or peer-reviewed CV/ML research publications
  • Background in real-time graphics or game engine pipelines (Metal, Vulkan, OpenGL ES)

Similar Jobs