As an AI Evaluation Engineer at Distyl AI, you will build evaluation systems that support Evaluation-Driven Development for AI products deployed in customer environments. The work focuses on creating measurable, iteration-ready signals through evaluation pipelines, golden test suites, and calibrated LLM-based graders.
Role Focus
AI systems deployed in production require evaluation signals that reflect real user needs, domain constraints, and business objectives. In this role, you will design evaluation frameworks and integrate offline and online evaluation pipelines into the system iteration loop so that improvements are guided by measurable outcomes rather than intuition alone.
Responsibilities
- Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments
- Define how system quality is measured within each domain, ensuring evaluation signals align with real user needs, domain constraints, and business objectives
- Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases
- Develop and maintain evaluation pipelines for both offline and online workflows, integrated into system iteration loops to inform prompt design, agent logic, model selection, and release readiness
- Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments and investigating where evaluation signals diverge from real-world outcomes
- Collaborate with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks guide production development and deployment
Requirements
- 2+ years of software engineering experience
- Strong Python engineering skills: write clean, maintainable Python and build evaluation and experimentation pipelines that can run in production; evaluation code is treated with the same rigor as application code
- Experience with Evaluation-Driven or Experiment-Driven Development: familiarity with structured evaluation or experimentation frameworks and the ability to recognize pitfalls such as overfitting to metrics that do not reflect real outcomes
- Ability to translate human judgment into code: work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale
- Systems-oriented mindset: understand evaluation interactions across prompts, agents, data, and deployment, designing evaluation systems that support fast iteration while maintaining trust and safety in production
- AI-native working style: use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration
- Travel: travel between 10-50% of the time, depending on the project, your role, and level of interest in doing so
Technologies
- Python
- LLM-based graders
- Evaluation-Driven Development
- Experiment-Driven Development
- HSA
- FSA
- 401(k)
- Carrot
Work Model
Hybrid collaboration model with 3+ days per week (TuesdayβThursday) in the office.
Travel
Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so.
Compensation
The base salary range for this role is $150,000 - $250,000 per year, depending on experience, location, and level.
Benefits
- Meaningful equity, along with a comprehensive benefits package
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible time off
- Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
- Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
- Complimentary in-office lunches and snacks provided
- Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
- Ownership of high-impact projects across top enterprises
- A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Location
San Francisco, CA (hybrid).