EngineerJobs.io
← Back to all jobs

Job Description

On Apple’s Human-Centered AI team, this role focuses on evaluating and improving Foundation Models and generative AI experiences with an emphasis on how people perceive quality. You will build evaluation frameworks and MLOps pipelines that turn model behavior into measurable insights, helping ensure AI systems are reliable, safe, and aligned with human expectations.

In this onsite position in Seattle, you will partner across engineering, research, product, and Responsible AI teams to refine model performance based on evaluation telemetry and human-centric measures.

Responsibilities

  • Architect and execute evaluation suites for LLMs and multimodal models, identifying edge cases across multi-step reasoning, factuality, adversarial robustness, safety, and alignment.
  • Build deterministic, heuristic, and LLM-assisted evaluation frameworks (including LLM-as-a-judge and reward modeling) to quantify human-perceived quality signals such as helpfulness and hallucination rates.
  • Convert qualitative failure modes into quantifiable loss patterns, programmatic guardrails, and actionable data-mixture adjustments for training and inference.
  • Collaborate with engineering teams to refine model behavior using evaluation telemetry to guide prompt engineering, Retrieval-Augmented Generation (RAG) strategies, and model fine-tuning.
  • Apply advanced ML techniques such as embedding-based clustering, representation learning, and perturbation analysis to map error taxonomies and latent failure manifolds.
  • Develop robust MLOps workflows to define evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI/CD pipelines.
  • Architect scalable, distributed inference and processing pipelines (for example, Ray and vLLM) to support high-throughput evaluation, automated annotation, and output analysis.
  • Define quantitative evaluation frameworks that capture nuanced human factors, including trust calibration, conversational state tracking, and interpretability.
  • Build automated evaluation pipelines that use LLMs to assess outputs at scale and improve correlation with human baseline annotations.
  • Work with ML researchers, software developers, and product managers across Apple to translate product requirements into scalable, reliable, and efficient evaluation infrastructure.

Requirements

  • Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.
  • 5+ years of relevant industry experience in ML Engineering or Applied Research.
  • Advanced proficiency in Python and modern deep learning ecosystems including PyTorch, JAX, and Hugging Face.
  • Experience building scalable ML inference pipelines, model evaluation workflows, and structured rating frameworks for large-scale AI systems.
  • Ability to interpret unstructured model outputs (text, transcripts, embedding spaces) and synthesize qualitative findings into actionable engineering guidance and training objectives.
  • Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
  • Familiarity with AI quality metrics, hallucination detection techniques such as SelfCheckGPT, model alignment methods including RLHF and DPO, and LLM-as-a-judge frameworks such as G-Eval and DeepEval.
  • Experience building internal tools or automated pipelines for ML workflows using MLflow, Weights & Biases, or similar platforms.
  • Strong familiarity with prompt engineering, RAG architectures (vector databases, semantic search), and fine-tuning.
  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field.

Technologies

  • Python, PyTorch, JAX, Hugging Face
  • Ray, vLLM
  • MLflow, Weights & Biases
  • SelfCheckGPT, RLHF, DPO, G-Eval, DeepEval
  • LLM-as-a-judge, Retrieval-Augmented Generation (RAG)
  • Vector databases, semantic search

Benefits

  • Comprehensive medical and dental coverage.
  • Retirement benefits.
  • Range of discounted products and free services.
  • Reimbursement for certain educational expenses, including tuition, for formal education related to advancing your career.
  • Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.
  • Eligible for discretionary restricted stock unit awards.
  • Option to purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan.
  • May be eligible for discretionary bonuses or commission payments and relocation.

Similar Jobs