Join Deloitte in Miami as an Agentic AI Engineer focused on Healthcare AI. You will design, build, and operationalize LLM- and SLM-powered decisioning systems that support payers, providers, and life sciences, owning agent systems end to end and deploying them into live clinical and operational settings within your first months. This onsite role blends a collaborative, cross-disciplinary culture with the opportunity to work on mission-critical workflows and scalable AI solutions.
Responsibilities
- Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, regulated processes.
- Build stateful workflows with LangGraph and LangChain, including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
- Engineer for long-horizon reliability: multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when steps fail.
- Develop the reasoning behind regulated decisions with policy- and criteria-grounded outputs, proposer/critic/judge style review, and auditable rationales for high-stakes healthcare decisions.
- Create end-to-end Retrieval-Augmented Generation pipelines: ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
- Engineer memory and context management including conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
- Apply context-delivery patterns to ensure agents access the right information at the right time within business workflows.
- Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
- Incorporate guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
- Evaluate agents at trajectory and task levels with sandboxed testing, alongside retrieval- and generation-quality metrics and automated checks with human review.
- Ensure healthcare-grade safety through deployment eval gates, human oversight and escalation models, and auditable traceability for regulated decisions, including PHI/HIPAA-aware data handling.
- Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers to operate within real business workflows safely.
- Deliver production-quality code with solid testing, CI/CD, logging, versioning, and documentation; balance quality, safety, latency, cost, and model risk in architectural decisions.
- Collaborate with modeling and post-training engineers to improve tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and, where helpful, fine-tuning or reasoning-optimized models.
- Translate ambiguous, high-complexity processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.
Requirements
- Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
- Proven track record shipping production agentic systems as a primary craft, demonstrating deep software and ML fundamentals along with substantial, recent hands-on agentic work.
- Strong hands-on experience building production agent systems with modern orchestration frameworks (LangGraph/LangChain or equivalents), including custom orchestration.
- Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
- Solid understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
- Practical knowledge of LLM behavior, including strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs; with proven evaluation methods.
- Experience evaluating and debugging agent behavior with task-success and trajectory analysis, not just output quality.
- Strong Python engineering skills and modern software practices: testing, CI/CD, version control, API integration; experience implementing observability, tracing, and debugging for production LLM-based systems.
- Hands-on experience with at least one frontier model platform (Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
- Ability to travel 0-50% on average, depending on client engagements and project scope.
- Limited immigration sponsorship may be available.
Technologies
- LangGraph, LangChain, Pinecone, Weaviate, Milvus, Llama via vLLM
- Anthropic, Google, OpenAI
- Python
The Team
Deloitte brings together AI researchers, modeling and platform engineers, architects, clinical and domain specialists, and product leaders to build, deploy, and operate vertical AI systems across software, data, models, and cloud infrastructure. The healthcare focus spans payers, providers, and life sciences, tackling genuinely hard reasoning problems, nuanced workflows, and a high bar for quality and safety.
Compensation
The base salary is benchmarked to leading technology companies rather than traditional consulting scales and includes a substantial performance-based incentive opportunity designed to grow with the value you help create. This startup-style upside is supported by a committed, well-capitalized platform. The estimated base salary range is $110,700-$372,900 per year, not adjusted for geographic differential; actual base pay depends on skills, experience, and level. You may also be eligible for additional compensation aligned with performance and impact.