EngineerJobs.io
← Back to all jobs

Job Description

ServiceNow is seeking a Staff AI Engineer for the Agent Orchestration team in Santa Clara, CA (onsite). The role will build and own key components of an agent execution harness for agentic conversational AI at enterprise scale, with emphasis on reliability, observability, prompt infrastructure, evaluation, and LLM integration.

Role Focus

Own significant parts of the agent orchestration execution harness, including the orchestration layer that routes inputs, manages context, invokes tools, handles retries, and exposes execution state across multi-step agentic workflows. This includes designing runtime behaviors for production reliability and ensuring agent systems can operate safely and predictably in enterprise environments.

Key Responsibilities

  • Design and build the agent execution harness that orchestrates multi-step agentic workflows, including input routing, context management, tool invocation, retry handling, and execution state reporting.
  • Own the runtime’s fault tolerance, latency, and throughput, with designs that prevent silent or unpredictable failures in workflows that cannot fail quietly.
  • Instrument the harness with tracing, cost attribution, and latency visibility to help the team reason about agent behavior in production and detect failures before customers are impacted.
  • Build prompt management systems covering versioning, templating, and systematic evaluation to keep agent behavior stable across model updates and configuration changes.
  • Design and own evaluation frameworks (unit evals, integration evals, and production monitors) to measure agent quality, identify regressions, and support data-informed decisions.
  • Integrate with and abstract over frontier LLMs, including model routing, fallback strategies, and production tradeoffs among cost and latency.
  • Set technical standards via architecture decisions, code reviews, and coaching, with particular focus on agentic design patterns and production AI discipline.
  • Define and enforce where agent logic lives, clarifying distinctions among tool calls, sub-agents, hardcoded paths, and human escalation, and establish standards across the team.

Required Qualifications

  • 7+ years building production software systems with a track record in reliability, performance, and scalability.
  • Hands-on experience shipping generative AI products that production users depend on, beyond prototype work or basic LLM API integration.
  • Strong depth in how large language models work, including failure modes, context constraints, and the impact of prompt design at scale.
  • Practical prompt engineering experience in systematically designing, versioning, and evaluating prompts across model updates and A/B evaluation cycles.
  • Proven eval engineering background, including building and shipping evaluation suites that drive quality decisions in production AI systems.
  • System-level cost and efficiency awareness, including reasoning about model routing, inference cost, and latency tradeoffs in production.
  • Strong software engineering fundamentals, including distributed systems, API design, and testing discipline.
  • Comfort operating in fast-moving, ambiguous, startup-like AI product environments.
  • Experience with multi-agent coordination patterns such as A2A and MCP.
  • Familiarity with agent frameworks including LangChain or LlamaIndex (or similar).
  • Prior experience shipping AI systems in enterprise software.
  • Experience with AI observability tooling, including tracing, cost tracking, and LLM-specific monitoring.
  • Familiarity with cloud-native infrastructure and service observability, including logging, monitoring, reliability engineering, and production troubleshooting.

Technologies

  • LangChain
  • LlamaIndex
  • A2A
  • MCP

Compensation and Benefits

  • Base pay: USD 176,100 - 308,200 per year
  • Equity (when applicable)
  • Variable/incentive compensation and benefits
  • On Target Earnings (OTE) incentive compensation structure for sales positions
  • Health plans, including flexible spending accounts
  • 401(k) Plan with company match
  • ESPP
  • Matching donations
  • Flexible time away plan
  • Family leave programs

Team Overview

The Agent Orchestration team owns the execution core, including the agent harness, orchestration runtime, multi-agent coordination, memory management, and evaluation frameworks that help ensure agents behave correctly in production.

Work Location and Persona

This role is located in Santa Clara, CA (onsite). Work personas may be flexible, remote, or require office presence based on assigned work location. ServiceNow may confirm the distance between a candidate’s primary residence and the closest ServiceNow office using a third-party service to determine eligibility for a work persona.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. Applicants with arrest or conviction records will be considered in accordance with legal requirements.

Accommodations

If you require a reasonable accommodation to complete any part of the application process, or are unable to use the online application and need an alternative method to apply, contact [email protected].

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations (including the U.S. Export Administration Regulations), ServiceNow may need to obtain export control approval from government authorities for certain individuals. Employment is contingent upon ServiceNow obtaining any export license or other approval required by relevant export control authorities.

Similar Jobs