Benefits and culture
Join a Denver, Colorado on-site role that emphasizes meaningful impact in healthcare through reliable, auditable AI decisioning. You will work alongside a multidisciplinary team spanning AI researchers, modeling and platform engineers, architects, and clinical/domain experts to tackle complex challenges across payers, providers, and life sciences. The compensation reflects a strong base range of $110,700 to $372,900 per year, plus a substantial performance-based incentive designed to grow with the value you help create, backed by a committed, well-capitalized platform. This environment prizes practical engineering, rigorous testing, and clear documentation, with opportunities to influence tool use, grounding, and long-horizon reasoning within real clinical and operational workflows.
Responsibilities
- Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution within regulated operational processes.
- Build stateful workflows using LangGraph and LangChain, including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
- Engineer for long-horizon reliability, enabling multi-step task completion, recovery from compounded errors, planning under uncertainty, and robust tool use when steps fail.
- Develop the reasoning behind regulated decisions with policy- and criteria-grounded outputs, structured proposer/critic/judge-style reviews, and auditable rationales across clinical review, prior authorization, claims integrity, and care management.
- Create end-to-end Retrieval-Augmented Generation pipelines covering ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
- Engine memory and context management, including conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
- Apply context-delivery patterns to ensure agents access the right information at the right time within workflows.
- Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
- Incorporate guardrails and safety controls to minimize hallucinations and unsafe actions.
- Evaluate agents at trajectory and task levels with multi-step task success metrics, failure-mode analysis, sandboxed testing, and complementary retrieval/generation quality checks plus human review.
- Institutionalize healthcare-grade safety with deployment eval gates, escalation models, auditability for regulated decisions, and PHI/HIPAA-aware data handling.
- Integrate with internal and external tools, APIs, enterprise systems, databases, and model providers to operate safely within real business workflows.
- Deliver production-quality code with robust testing, CI/CD, logging, versioning, and documentation; balance quality, safety, latency, cost, and model risk in architecture decisions.
- Collaborate with modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and, when helpful, fine-tuning or reasoning-optimized models.
- Translate ambiguous, high-complexity processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.
Requirements
- Bachelorβs degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
- Proven depth in building and shipping production agentic systems as a primary craft, with a track record of shipped systems, research, model releases, or open-source contributions; strong software and ML fundamentals plus substantial, recent hands-on agentic work.
- Hands-on experience delivering production agent systems with modern orchestration (LangGraph/LangChain or equivalent), including custom orchestration.
- Experience designing and optimizing end-to-end Retrieval-Augmented Generation systems: indexing, retrieval, reranking, grounding, and evaluation.
- Solid understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
- Deep practical knowledge of LLM behavior, including strengths and limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs, plus the evaluation methods used to measure them.
- Experience evaluating and debugging agent behavior through task-success and trajectory analysis, not solely output quality.
- Strong Python engineering skills with modern software practices: testing, CI/CD, version control, API integration; experience implementing observability, tracing, and debugging for LLM-based systems in production.
- Hands-on experience with at least one frontier model platform (eg, Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (eg, Llama via vLLM), including production tool use and agent capabilities.
- Ability to travel 0-50 percent, depending on client work and engagement requirements.
- Limited immigration sponsorship may be available.
Technologies
- LangGraph, LangChain, vLLM, Llama via vLLM
- Pinecone, Weaviate, Milvus
- Anthropic, Google, OpenAI
- Python, FHIR
The Team
Deloitte brings together AI researchers, modeling and platform engineers, architects, clinical and domain specialists, and product leaders to build, deploy, and operate vertical AI systems across software, data, models, and cloud infrastructure. The work spans the healthcare sectorβpayers, providers, and life sciencesβand involves genuinely hard reasoning problems, nuanced operational workflows, and a high bar for quality and safety.