Senior Software Engineer, Agentic Systems
Job Description
Join ServiceNow in Mountain View, CA (onsite) to help build runtime infrastructure for Moveworks AI agents. This Senior Software Engineer role focuses on distributed systems engineering for agent orchestration, session management, event-driven messaging, concurrency, observability, and state and caching layers, supporting agent sessions that can span minutes and multiple LLM calls and tool invocations.
What you’ll build
- Agent orchestration engine using a state machine that coordinates long-running agent sessions, planning, execution, and user interaction across multiple LLM calls and tool invocations
- Distributed session management with lease-based ownership using DynamoDB conditional writes, heartbeat protocols, and crash recovery through checkpointing
- Event-driven message pipeline combining SQS FIFO queues for ordered delivery, Kafka consumers for event processing, and real-time streaming via gRPC and Socket.IO
- Structured concurrency with Python asyncio TaskGroups running multiple concurrent tasks per session (message polling, lease heartbeats, output publishing, orchestrator execution), using fail-fast semantics and graceful cancellation
- Observability infrastructure with OpenTelemetry instrumentation, distributed trace context propagation across async boundaries, and custom span lifecycle management for sessions that span minutes
- Caching and state layers using Redis and DynamoDB KV stores with per-org/per-bot scoping, batch read optimization, and hot-reload configuration
What you bring
- 5+ years building production backend and infrastructure systems
- Strong experience in Python or Go (ideally both)
- Experience designing and operating systems that handle real traffic at scale
- Comfort with ambiguity, since these are novel problems without textbook solutions
- Deep experience in at least 3 of the following areas: distributed systems (consistency models, idempotency, exactly-once delivery, distributed locking or leasing)
- Deep experience in at least 3 of the following areas: concurrent and async programming (Python asyncio, Go goroutines, structured concurrency, cancellation handling)
- Deep experience in at least 3 of the following areas: event-driven architectures (SQS, Kafka, pub/sub, backpressure, delivery guarantees)
- Deep experience in at least 3 of the following areas: database systems for infrastructure (DynamoDB including conditional writes and transactions, Redis including connection pooling and pub/sub)
- Deep experience in at least 3 of the following areas: observability (OpenTelemetry, distributed tracing, span context propagation, Prometheus metrics)
- Deep experience in at least 3 of the following areas: gRPC and protobuf (streaming RPCs, service interface design, error handling patterns)
Technologies
- Python, Go, DynamoDB
- SQS FIFO queues, Kafka
- gRPC, Socket.IO
- Python asyncio, asyncio TaskGroups
- OpenTelemetry, Prometheus
- Redis
- gRPC/protobuf, Protocol Buffers
Work Personas (flexible, remote, or required in office) and eligibility determination may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. ServiceNow is an equal opportunity employer statement (consideration without regard to protected categories; arrest/conviction records per legal requirements). If you require a reasonable accommodation for the application process, contact [email protected] for assistance. Export Control Regulations: for positions requiring access to controlled technology, ServiceNow may need export control approval; all employment is contingent upon obtaining any required export license or other approval.