Senior Machine Learning Engineer, Proactive
Job Description
Apple is hiring a Senior Machine Learning Engineer (Proactive) in Santa Clara, CA to help build next-generation intelligent search and AI experiences. The role focuses on systems that understand user intent and context while preserving privacy, with strong emphasis on transformer-based modeling and efficient on-device deployment.
Role Overview
You will design, train, fine-tune, optimize, and deploy transformer-based language models, semantic retrieval systems, and ranking models. The work centers on building proactive intelligence that supports intelligent search and AI experiences by connecting query understanding, personalization, retrieval, and model performance with practical evaluation and production metrics.
Responsibilities
- Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems, including models for query understanding, intent prediction, personalization, retrieval, and ranking.
- Evaluate search relevance and user behavior to create evaluation methodologies, offline benchmarks, and online metrics for retrieval quality, ranking, personalization, and language model performance.
- Create scalable experimentation and evaluation pipelines for LLMs and search models, covering model quality, robustness, latency, efficiency, and end-to-end product metrics.
- Design, train, fine-tune, distill, and optimize transformer-based language models and foundation models for efficient on-device deployment.
- Develop LLM fine-tuning and post-training approaches, including supervised fine-tuning, instruction tuning, preference optimization, parameter-efficient fine-tuning, and task-specific adaptation.
- Research and prototype on-device generative AI methods such as knowledge distillation, model compression, quantization, pruning, and low-latency inference.
- Transfer capabilities from large foundation models into compact on-device models while balancing model quality, latency, memory footprint, power consumption, and compute constraints.
- Collaborate with engineers, researchers, product managers, and designers to move AI capabilities from research into production, including work across foundation models, multimodal AI, agentic retrieval, and personalized intelligence.
Required Qualifications
- Master’s or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
- 5+ years of industry or research experience developing machine learning systems.
- Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.
- Experience training, fine-tuning, or deploying transformer-based models and large language models.
- Experience optimizing machine learning models for resource-constrained environments, including knowledge distillation, model compression, quantization, and pruning.
- Experience with on-device machine learning or edge AI, or mobile inference frameworks, including optimization for latency, memory, compute, and power constraints.
- Experience distilling capabilities from large foundation models into smaller language models or task-specific models for efficient inference.
- Experience building retrieval-augmented generation, vector search, embedding retrieval, neural reranking, or semantic search systems.
- Experience with query understanding, query rewriting, intent classification, personalized retrieval, learning-to-rank, or recommendation models.
- Experience working with transformer architectures and foundation model families such as BERT, T5, Llama, Gemma, Mistral, or related architectures.
- Experience evaluating language models, designing AI quality metrics, and building automated and human-in-the-loop evaluation pipelines.
- Experience building large-scale production search, recommendation, personalization, or generative AI systems.
- Familiarity with multimodal foundation models, tool use, agentic AI, or agentic retrieval systems.
- Strong understanding of tradeoffs among model quality, latency, memory, power consumption, privacy, and reliability for production on-device AI systems.
- Ability to prototype new ideas, conduct rigorous experiments, solve ambiguous technical problems, and translate research advances into production-quality machine learning solutions.
- Ability to work onsite in Cupertino, California, in accordance with Apple’s applicable work policies.
Technologies
- Python
- C/C++
- PyTorch
- JAX
- TensorFlow
- Transformers
- BERT
- T5
- Llama
- Gemma
- Mistral
Pay and Benefits
- Base pay: USD 184,700 - 324,800 per year, depending on skills, qualifications, experience, and location.
- Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.
- Eligibility for discretionary restricted stock unit awards, and ability to purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan.
- Medical and dental coverage.
- Retirement benefits.
- Range of discounted products and free services.
- Education reimbursement for certain educational expenses, including tuition, for formal education related to advancing your career at Apple.
- This role might also be eligible for discretionary bonuses or commission payments and relocation.
Note: Apple benefit, compensation, and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program. At Apple, base pay is one component of total compensation and is determined within a range.
Minimum Education
Master degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.