AI Engineer - Algorithm Evaluation & Agentic Systems
Job Description
On Apple’s DAQ team, you will focus on evaluating and improving advanced visual technologies. In this role, you will work at the intersection of algorithm evaluation and agentic systems, helping translate model behavior into measurable performance and production-ready workflows for advanced computer vision and video understanding.
Working within DAQ, your contribution centers on building evaluation pipelines, diagnosing failures, and architecting autonomous multi-modal solutions that connect experimentation with real-world deployment. You will lead benchmarking and integration of state-of-the-art image and video understanding models, while partnering with core model training teams to deliver data-driven insights for iteration and fine-tuning.
Responsibilities
- Design, build, and scale comprehensive evaluation pipelines for both end-to-end system evaluation and granular component-level testing on complex image and video understanding tasks.
- Analyze model outputs to identify root causes of visual hallucinations, temporal inconsistencies in video, and edge-case failures.
- Build, deploy, and evaluate agentic workflows that use vision models to autonomously solve multi-step user problems such as video summarization and visual search.
- Use component-level evaluation to isolate which parts of an agentic workflow, including tool selection, memory retrieval, and visual reasoning, succeed or fail.
- Lead strategy for curating high-quality schematized datasets and ground-truth benchmarks tailored for evaluating multi-modal capabilities.
- Partner with core model training teams to provide actionable metrics and insights that guide subsequent model training and fine-tuning.
- Lead benchmarking and integration of state-of-the-art models for image and video understanding.
Requirements
- MS and a minimum of 3 years relevant industry experience.
- 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation.
- Deep understanding of core machine learning principles, including probability, statistics, data distributions, and model bias/variance.
- Deep theoretical and practical understanding of computer vision and vision-language models, including how Vision Transformers (ViTs), spatial-temporal modeling, and image/video processing work.
- Proven ability to define robust metrics and KPIs and design rigorous evaluation frameworks for generative AI or foundation models, including custom benchmark creation, automated regression testing, LLM/VLM-as-a-judge methodologies, and human-in-the-loop evaluation.
- Experience building and evaluating LLM/VLM-powered agents, including tool use, multi-step reasoning, planning, and memory management workflows.
- Strong ability to probe ML models for edge cases, hallucinations, and performance bottlenecks in constrained environments, and translate findings into improvement recommendations.
- Strong proficiency in Python and experience with PyTorch, including running inference, extracting embeddings, and building scalable evaluation pipelines.
Technologies
- Python
- PyTorch
- Vision Transformers (ViTs)
- Computer Vision (CV)
- Vision-Language Models (VLMs)
- LLM/VLM-as-a-judge methodologies
- LLM/VLM-powered agents
Benefits
- Comprehensive medical and dental coverage
- Retirement benefits
- Discounted products and free services
- Reimbursement for certain educational expenses, including tuition for formal education related to advancing your career
- Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs
- Discretionary restricted stock unit awards
- Ability to purchase Apple stock at a discount through participation in the Employee Stock Purchase Plan
- Potential discretionary bonuses or commission payments
- Relocation (may be eligible)
What We Value
- Production mindset: correctness, observability, and maintainability
- Ability to reason about system-level tradeoffs, not only model performance
- Ability to balance experimentation speed with engineering rigor
- Comfort working in ambiguous spaces and defining metrics from first principles
- Clear communication of technical findings to both technical and non-technical audiences
Preferred Qualifications
- Demonstrated ability to lead technical evaluation strategies end-to-end, drive architectural decisions for testing infrastructure, and mentor engineers.
- Strong foundation in statistics, including hypothesis testing, confidence intervals, and experimental design.
- Knowledge of reinforcement learning, planning, or decision-making systems.
- Experience evaluating multi-modal or multi-agent systems.
- Prior work on AI reliability, safety, or benchmarking.
Pay & Benefits
- Base pay range: $150,400 to $277,600 per year
- Employees may have opportunities to become shareholders through discretionary employee stock programs
- Eligible for discretionary restricted stock unit awards
- Ability to purchase Apple stock at a discount if participating in the Employee Stock Purchase Plan
- Role might be eligible for discretionary bonuses or commission payments and relocation
- Learn more about Apple Benefits
- Note: Benefit, compensation, and employee stock program eligibility are subject to requirements and terms of the applicable plan or program.
Location: Sunnyvale, CA (onsite)
Minimum experience: 3 years
Education: MS
Compensation: USD 150,400 - 277,600 per yearly