EngineerJobs.io
← Back to all jobs

Job Description

Deloitte is seeking an Agentic AI Engineer focused on Healthcare AI to design, build, and operationalize end-to-end agentic systems for healthcare decisioning across payers, providers, and life sciences. This onsite role in Jersey City, NJ requires ownership of architecture through production and delivering into live clinical and operational settings within the initial months.

Responsibilities

  • Develop and deploy agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution within complex, regulated operations.
  • Construct persistent workflows with LangGraph and LangChain, incorporating branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
  • Prioritize long-horizon reliability, ensuring multi-step task completion, recovery from compounded errors, planning under uncertainty, and robust tool use when steps fail.
  • Engineer the reasoning for regulated decisions, delivering policy-grounded outputs with auditable rationales and a proposer/critic/judge style review across clinical review, prior authorization, claims integrity, and care management.
  • Build end-to-end RAG pipelines, including ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
  • Design memory and context management, covering conversational state, persistent memory, retrieval-driven context assembly, and token-efficient context selection.
  • Adopt modern context-delivery patterns to ensure agents access the right information at the right time.
  • Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
  • Apply guardrails and safety controls to reduce hallucinations and unsafe actions in agent behavior.
  • Evaluate agents at both trajectory and task levels, including multi-step task success, failure modes, regression analysis, sandboxed testing, and quality metrics for retrieval and generation.
  • Incorporate healthcare-grade safety measures with deployment eval gates, human oversight and escalation models, and auditability for regulated decisions and PHI/HIPAA-compliant data handling.
  • Develop integrations with internal and external tools, APIs, enterprise systems, databases, and model providers to operate safely within real business workflows.
  • Deliver production-quality code with rigorous testing, CI/CD, logging, versioning, and documentation; make architecture decisions balancing quality, safety, latency, cost, and model risk.
  • Collaborate with modeling and post-training engineers to enhance tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and selective fine-tuning.
  • Translate ambiguous and complex operational processes into robust system logic and reusable AI patterns, staying current with advances in agentic systems and translating research into practical engineering decisions.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
  • Proven experience shipping production agentic systems, with a focus on real deployments; strong software and ML fundamentals and substantial recent hands-on agentic work.
  • Hands-on experience building production agent systems with modern orchestration such as LangGraph and LangChain or equivalents, including custom orchestration.
  • Experience designing and optimizing end-to-end RAG systems, including indexing, retrieval, reranking, grounding, and evaluation.
  • Deep understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
  • Strong grasp of LLM behavior, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs; familiarity with evaluation methods.
  • Experience evaluating and debugging agent behavior, focusing on task success and trajectory analysis.
  • Strong Python engineering skills with modern software practices, including testing, CI/CD, version control, API integration, and production observability/tracing for LLM-based systems.
  • Hands-on experience with frontier model platforms (Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
  • Ability to travel up to 50 percent, depending on client engagements; limited immigration sponsorship may be available.

Technologies

  • LangGraph
  • LangChain
  • Python
  • Llama via vLLM (open-weight/self-hosted models)
  • Pinecone
  • Weaviate
  • Milvus
  • Anthropic
  • Google
  • OpenAI
  • FHIR

Preferred Qualifications

  • Experience with multi-agent systems and agent collaboration patterns.
  • Familiarity with vector databases and retrieval infrastructure such as Pinecone, Weaviate, or Milvus.
  • Exposure to model adaptation and fine-tuning techniques such as LoRA or QLoRA.
  • Foundational NLP knowledge including tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
  • Experience operating in regulated, high-stakes environments; healthcare exposure with workflows or standards such as FHIR is a plus, not a requirement.
  • Proven habit of staying current with AI research, benchmarks, and emerging engineering patterns.

Compensation

The base salary range is $134,500 to $265,100 per year, not adjusted for geographic differential. The role includes a substantial performance-based incentive opportunity designed to grow with the value you help create, offering startup-style upside within a supported, well-capitalized platform. Actual pay depends on your skills, experience, and level.

Education

Bachelor's degree in a relevant field as listed in Requirements.

Similar Jobs