AI Engineer, Evals & Agent Quality
Job Description
Town.com, Inc. is developing an AI assistant and is seeking an AI Engineer to take ownership of evaluation and agent quality. This role focuses on building quality measurement that supports fast iteration, reliable improvements, and regression prevention across the assistant’s experiences.
Key Responsibilities
- Design and build a generalized evaluation system that assesses assistant quality across all user-facing surfaces, including multi-step agent trajectories.
- Create and maintain golden datasets, along with a labeling loop that keeps datasets current and continuously validates improvements while mitigating regressions.
- Develop model routing capabilities and online evaluation tooling to identify, compare, and route to the best-performing models.
- Ensure prompts and system updates are measurable, enabling rapid progress without breaking existing functionality.
- Collaborate with engineers across product teams to instrument quality, connect signals to corrective actions, and maintain an end-to-end feedback loop.
What You Will Bring
- Experience building or owning LLM evaluation systems, or leading offline and online quality measurement at scale.
- A rigorous approach to measurement, including thoughtful test design and quality assurance practices.
- Hands-on familiarity with the evaluation tooling landscape, including off-the-shelf frameworks, plus the judgment to choose appropriate solutions.
- Comfort with model routing and evaluating tradeoffs between different models.
- Ability to ship fixes, not only dashboards and metrics.
- Senior or staff engineering experience working in greenfield environments where the measurement and evaluation systems do not yet exist.
Location & Schedule
The role is based in San Francisco, CA, with five days per week in person at the company’s Financial District office.
Compensation
$250,000 - $300,000 per year.