Agentic AI Engineer β Healthcare AI
Job Description
What Deloitte Offers
Join a platform-backed environment in New York City that empowers you to build end-to-end agentic AI systems for healthcare decisioning. Youβll work at the intersection of payers, providers, and life sciences, translating complex operational processes into robust, scalable AI solutions. The role is onsite in New York, with a base compensation range of USD 110,700 to 372,900 per year, plus a substantial performance-based incentive designed to grow with the value you help create. Deloitte provides a startup-style upside within a proven, well-capitalized platform, backed by collaboration across AI researchers, modeling and platform engineers, architects, clinical and domain experts, and product leaders. Expect a culture that emphasizes rigor, safety, and auditable decision-making in one of the most complex operating environments in healthcare. The role supports healthcare-grade safety and PHI/HIPAA-aware data handling, with clear deployment gates, oversight models, and traceability for regulated decisions. Location is onsite in New York, with travel up to 50% based on project needs.
Responsibilities
- Architect and deliver agentic AI systems that perform multi-step reasoning, planning, tool use, and workflow execution within highly regulated healthcare processes.
- Develop stateful, reusable workflows using LangGraph and LangChain, incorporating branching, retries, self-correction, human-in-the-loop checkpoints, and modular orchestration patterns.
- Design for long-horizon reliability, enabling multi-step task completion, recovery from cascading errors, planning under uncertainty, and robust tool usage when steps fail.
- Build the reasoning layer for regulated decisions with policy-grounded outputs, structured proposer/critic/judge reviews, and auditable rationales across clinical review, prior authorization, claims integrity, and care management.
- Construct end-to-end Retrieval-Augmented Generation pipelines, covering ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
- Engineer memory and context management, including conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
- Apply contemporary context-delivery patterns so agents access the right information at the right time within real business workflows.
- Establish observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
- Implement guardrails, safety controls, and failure-handling to minimize hallucinations and unsafe actions.
- Evaluate agents at trajectory and task levels, including multi-step task success, failure-mode analysis, sandboxed testing, and integrated metrics for retrieval and generation quality, with automated checks and human reviews.
- Engineer healthcare-grade safety through deployment eval gates, human oversight and escalation paths, and auditable, traceable decisions with PHI/HIPAA-aware data handling.
- Develop integrations with internal and external tools, APIs, enterprise systems, databases, and model providers to ensure safe operation within authentic business workflows.
- Deliver production-quality code with strong testing, CI/CD, logging, versioning, and documentation; make architecture decisions balancing quality, safety, latency, cost, and model risk.
- Collaborate with modeling and post-training engineers to improve tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and, where helpful, fine-tuned or reasoning-optimized models.
- Translate ambiguous operational processes into robust system logic and reusable AI patterns, keeping current with agentic-system advances and translating research into practical engineering decisions.
Requirements
- Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
- Proven deep experience building and shipping production-grade agentic systems; we value substantial, recent hands-on agentic work and a track record of delivered systems, releases, and open source contributions.
- Hands-on experience building production agent systems with modern orchestration, including LangGraph or LangChain or equivalent with custom orchestration.
- Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
- Strong memory and context management knowledge, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
- Deep practical understanding of LLM behavior, including strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs; familiarity with evaluation methods.
- Experience evaluating and debugging agent behavior with task-success and trajectory analysis, beyond surface-level output quality.
- Strong Python engineering skills and modern software practices: testing, CI/CD, version control, API integration; experience implementing observability, tracing, and debugging for LLM-based systems in production.
- Hands-on experience with at least one frontier model platform (Anthropic, Google, OpenAI) or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
- Ability to travel up to 50% depending on client work and industry needs.
- Limited immigration sponsorship may be available.
Technologies
LangGraph, LangChain, Python, Llama via vLLM, Anthropic, Google, OpenAI, Pinecone, Weaviate, Milvus, FHIR
The Team
Deloitte brings together AI researchers, modeling and platform engineers, architects, clinical and domain specialists, and product leaders to build, deploy, and operate verticalized AI systems across software, data, models, and cloud infrastructure - engineered for one of the most complex operating environments in the world. The work spans the healthcare industry - payers, providers, and life sciences - and involves genuinely hard reasoning problems, nuanced operational workflows, and a high bar
Compensation
The base salary is benchmarked to leading technology companies rather than traditional consulting scales, with a substantial performance-based incentive designed to grow with the value you help create. The estimated base range is USD 110,700 to 372,900 per year, not adjusted for geographic differential; actual base pay depends on skills, experience, and level. Eligible for additional incentives tied to performance and impact.