Join ServiceNow in Mountain View to build the judgment layer for an agent evaluation platform used in stateful, multi-tenant enterprise environments. This role focuses on making LLM judge signals calibrated and actionable, so evaluation outputs can be used to train and optimize agents, not just to report results. You will also shape the runtime, observability, and simulation infrastructure that allow teams to run repeatable evals at production scale.
Onsite location: Mountain View, CA
Compensation: Base pay of $161,300 - $274,200 per year, plus equity (when applicable), variable/incentive compensation, and benefits.
What you will build
- Develop the judgement layer for scoring multi-step agent trajectories, including rubrics, judges, calibration against human labels, and the methodology that makes scores meaningful.
- Own the end-to-end evaluation runtime for multi-turn agent scenarios: stand up the environment and user simulator, run the useragentworld loop, collect transcripts and traces, run validators and scoring, then tear down.
- Engineer production-grade execution with scheduling, retries, high-concurrency execution, and run isolation at production dataset sizes.
- Implement versioned specs, datasets, and reports, with run-to-run comparison as a first-class operation.
- Consolidate current one-off eval workflows into a single orchestration service that acts as one source of truth for scheduling and retry.
- Establish a reliability floor and an SLO for the evaluation harness.
- Make evals self-serve so any team can run an eval without bespoke integration.
Reliability and observability for agent trajectories
- Lead the move to OpenTelemetry-native observability, replacing parallel per-service logging, correlation, and redaction mechanisms used today.
- Define a span data model for agent trajectories including prompts, tool calls, plan updates, and outcomes so trajectory data is queryable without hand reconstruction from logs.
- Provide trace context propagation across async boundaries and long-lived sessions (minutes or hours).
- Ensure full prompts and completions survive the pipeline intact while preventing eval traffic from contaminating its own data.
- Support fault attribution and cross-run diffing to identify which component broke and what changed since the last green run.
- Provide debug surface support and the tracing contract with the team building the agent.
Simulation environment and repeatable runs
- Build the simulation environment with stateful fakes of enterprise systems agents call (ITSM, HR, knowledge bases, inventory), backed by a real datastore that persists changes during a run.
- Implement per-run data injection and programmatic setup/teardown so every run is hermetic and repeatable.
- Create LLM-driven user simulators for open-ended personas and scripted state-machine simulators for deterministic flows.
- Implement contract-testing mocks against real API schemas in CI to prevent simulation fidelity drift as vendor APIs change.
- Develop isolated sandbox environments that reproduce the config, identity, search content, and permissions agents actually read, provisioned from an identical baseline and torn down every run.
Foundation for optimization: Lay the groundwork for using eval signals to optimize the agent, not just measure it.
Required qualifications
- 5+ years building production backend or infrastructure systems.
- Strong in Python or Go (ideally both).
- Experience designing and operating systems that handle real traffic at scale.
- Comfort making a non-deterministic system measurable, with interest in turning fuzzy agent behavior into a signal engineers can gate releases on (no ML background required).
- Comfort with ambiguity; novel problems without textbook solutions.
- Experience in at least 3 of the following:
- Distributed systems
- Orchestration and workflow runtimes
- Observability internals as a builder (OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high-cardinality trace data)
- Concurrent and async programming (Python asyncio, Go concurrency, structured cancellation)
- Data-intensive pipelines (high-volume ingest, schema evolution, sampling and retention trade-offs)
- gRPC/protobuf service and interface design
Technologies
- OpenTelemetry
- Python
- Go
- Python asyncio
- gRPC
- protobuf
- Temporal
- Airflow
- Argo
Benefits
- Base pay of $161,300 - $274,200, plus equity (when applicable), variable/incentive compensation and benefits.
- Health plans, including flexible spending accounts.
- 401(k) Plan with company match.
- ESPP.
- Matching donations.
- Flexible time away plan.
- Family leave programs.
Work personas
ServiceNow approaches work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories assigned depending on the nature of work and assigned work location. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal opportunity employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. Applicants with arrest or conviction records will be considered in accordance with legal requirements.
Accommodations
If you require a reasonable accommodation to complete any part of the application process, or are unable to use the online application and need an alternative method to apply, contact globaltalentss@servicenow.com for assistance.
Export control regulations
For positions requiring access to controlled technology subject to export control regulations (including the U.S. Export Administration Regulations), ServiceNow may need to obtain export control approval from government authorities for certain individuals. Employment is contingent upon ServiceNow obtaining any required export license or other approval from relevant export control authorities.