EngineerJobs.io
← Back to all jobs

Job Description

Apple is hiring a Machine Learning Engineer, Proactive, focused on building and optimizing intelligent search and AI experiences. In this onsite role in Cupertino, you will design and deploy transformer-based and foundation models, along with semantic retrieval and retrieval-augmented generation systems, with a strong emphasis on efficient on-device performance and evaluation.

What you’ll do

  • Build semantic retrieval systems, including embeddings, reranking, and retrieval-augmented generation (RAG), plus models for query understanding, intent prediction, personalization, retrieval, and ranking.
  • Evaluate search relevance and user behavior by designing evaluation methodologies, offline benchmarks, and online metrics for retrieval quality, ranking, personalization, and language model performance.
  • Create scalable experimentation and evaluation pipelines for LLMs and search models, covering model quality, robustness, latency, efficiency, and end-to-end product metrics.
  • Design, train, fine-tune, distill, and optimize transformer-based language models and foundation models for efficient on-device deployment.
  • Develop LLM fine-tuning and post-training approaches, including supervised fine-tuning, instruction tuning, preference optimization, parameter-efficient fine-tuning, and task-specific adaptation.
  • Research and prototype on-device generative AI methods such as knowledge distillation, model compression, quantization, pruning, and low-latency inference.
  • Transfer capabilities from large foundation models into compact on-device models while balancing quality, latency, memory footprint, power consumption, and compute constraints.
  • Partner with engineers, researchers, product managers, and designers to move AI from research into production, shaping technical strategy across projects and exploring applications including foundation models, multimodal AI, agentic retrieval, and personalized intelligence.

Requirements

  • Master’s or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • Experience optimizing machine learning models for resource-constrained environments, including knowledge distillation, model compression, quantization, and pruning.
  • Experience with on-device machine learning or edge AI, including optimizing models for latency, memory, compute, and power constraints.
  • Experience distilling capabilities from large foundation models into smaller language models or task-specific models for efficient inference.
  • Experience building retrieval-augmented generation, vector search, embedding retrieval, neural reranking, or semantic search systems.
  • Experience with query understanding, query rewriting, intent classification, personalized retrieval, learning-to-rank, or recommendation models.
  • Experience with transformer architectures and foundation model families such as BERT, T5, Llama, Gemma, Mistral, or related architectures.
  • Experience evaluating language models, designing AI quality metrics, and building automated and human-in-the-loop evaluation pipelines.
  • Experience building large-scale production search, recommendation, personalization, or generative AI systems.
  • Familiarity with multimodal foundation models, tool use, agentic AI, or agentic retrieval systems.
  • Strong understanding of tradeoffs among model quality, latency, memory, power consumption, privacy, and reliability for production on-device AI systems.
  • Ability to prototype new ideas, conduct rigorous experiments, solve ambiguous technical problems, and translate research into production-quality machine learning solutions.
  • Bachelor degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.
  • Experience training, fine-tuning, or deploying transformer-based models and large language models.
  • Experience with modern deep learning architectures and techniques, including transformers, embeddings, representation learning, and neural ranking.
  • Programming skills in Python and/or C/C++, with production software experience using PyTorch, JAX, or TensorFlow.
  • Ability to work onsite in Cupertino, California in accordance with Apple’s applicable work policies.

Technologies

  • Python
  • C/C++
  • PyTorch
  • JAX
  • TensorFlow
  • Transformer architectures
  • BERT, T5, Llama, Gemma, Mistral
  • Vector search, embeddings, neural reranking
  • Retrieval-augmented generation
  • Mobile inference frameworks
  • On-device machine learning, edge AI

Compensation and benefits

The base pay range for this role is $150,400 to $225,300 per year, depending on skills, qualifications, experience, and location. Apple base pay is one part of a total compensation package.

  • Comprehensive medical and dental coverage
  • Retirement benefits
  • Discounted products and free services
  • Reimbursement for certain educational expenses, including tuition
  • Discretionary bonuses or commission payments and potential relocation eligibility
  • Discretionary employee stock programs, including opportunity to become an Apple shareholder
  • Eligible for discretionary restricted stock unit awards
  • Employee Stock Purchase Plan option to purchase Apple stock at a discount

Note: Benefits, compensation, and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Similar Jobs