EngineerJobs.io
← Back to all jobs

Job Description

Deloitte's Healthcare AI initiative is expanding its capabilities to deliver agentic decisioning tools that operate across clinical and operational workflows. This onsite role in Raleigh, NC focuses on building end-to-end systems powered by large language models and retrieval-augmented pipelines, deployed within an AI-first program to accelerate healthcare decision making.

The Agentic AI Engineer will translate complex, regulated processes into robust, reusable architectures that reason across steps, use tools, and execute multi-stage workflows while maintaining safety, compliance, and auditability in live clinical and payer environments.

Responsibilities

  • Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution within regulated operational processes.
  • Develop stateful workflows with LangGraph and LangChain, including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
  • Engineerto support long-horizon reliability, ensuring multi-step task completion, recovery from compounded errors, planning under uncertainty, and robust tool use when steps fail.
  • Build the reasoning behind regulated decisions with policy- and criteria-grounded outputs, structured proposer/critic/judge-style review, and auditable rationales for high-stakes decisions across clinical review, prior authorization, claims integrity, and care management.
  • Develop end-to-end Retrieval-Augmented Generation (RAG) pipelines, including ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
  • Engineer memory and context management for conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
  • Apply contemporary context-delivery patterns to ensure agents access the right information at the right time.
  • Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
  • Apply guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
  • Evaluate agents at the trajectory and task level with multi-step task success metrics, failure-mode analysis, and sandboxed testing alongside retrieval and generation quality metrics, automated checks, and human review.
  • Engineer healthcare-grade safety through deployment eval gates, human oversight and escalation models, and auditable traceability for regulated decisions with PHI/HIPAA-conscious data handling.
  • Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers to operate safely within real business workflows.
  • Deliver production-quality code with rigorous testing, CI/CD, logging, versioning, and documentation; balance quality, safety, latency, cost, and model risk in architecture decisions.
  • Collaborate with modeling and post-training engineers to improve tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and fine-tuning where appropriate.
  • Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
  • Proven depth building and shipping production agentic systems as a primary craft, with a track record of shipped systems, research, model releases, or open-source work over multiple years.
  • Strong hands-on experience building production agent systems with modern orchestration frameworks such as LangGraph or LangChain (or equivalents), including custom orchestration.
  • Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
  • Solid understanding of memory and context management, including context windows, retrieval-driven assembly, persistent memory, and high-signal context selection.
  • Deep practical knowledge of LLM behavior, including strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs, plus methods to evaluate them.
  • Experience evaluating and debugging agent behavior beyond surface output quality, focusing on task success and trajectory analysis.
  • Strong Python engineering skills with modern software practices: testing, CI/CD, version control, API integration; experience implementing observability, tracing, and debugging for production LLM-based systems.
  • Hands-on experience with at least one frontier model platform (Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
  • Ability to travel 0-50 percent, depending on client work and engagement needs.
  • Limited immigration sponsorship may be available.

Technologies

  • LangGraph, LangChain, Python, vLLM, Llama, Pinecone, Weaviate, Milvus, LoRA, QLoRA, RAG, Anthropic, Google, OpenAI, FHIR

Preferred Qualifications

  • Experience with multi-agent systems and agent collaboration patterns.
  • Familiarity with vector databases and retrieval infrastructure such as Pinecone, Weaviate, or Milvus.
  • Exposure to model adaptation and fine-tuning techniques such as LoRA or QLoRA.
  • Understanding of traditional NLP concepts including tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
  • Experience operating in highly regulated, high-stakes, or operationally complex environments; healthcare exposure or standards such as FHIR is a plus, not a requirement.
  • Demonstrated habit of staying current with AI research, benchmarks, and emerging engineering patterns.

Details

  • Location: Raleigh, NC (onsite)
  • Salary: USD 134,500 - 265,100 per year
  • Education: Bachelor's degree

Similar Jobs