EngineerJobs.io
← Back to all jobs

Job Description

The Senior AI Engineer will own LLM-powered systems and agentic workflows from early scoping to production operation. The role focuses on accuracy-evaluable automation, deterministic orchestration, privacy and security baselines, tracing and auditability, and day-to-day system delivery in a multi-tenant environment.

Role Overview

You will deliver end-to-end LLM-powered capabilities in Python and TypeScript, partnering with product managers and data scientists to scope ambiguous customer problems, define evaluation and guardrail strategies, and operate systems that meet release gates. You will also mentor engineers with shared responsibility in an on-call rotation.

Key Responsibilities

  • Own delivery of LLM-powered systems, progressing from an ambiguous customer problem through scoping with product managers and data scientists to production operation.
  • Define the evals and guardrails strategy for your product area, including release gates, production signals watched by the team, and simulation of changes against historical claims before shipping.
  • Improve team effectiveness with agentic coding tools by refining specifications, instruction files, hooks, and permission policies; ensure generated code meets the same review standard before merge.
  • Design NLP pipelines to transform unstructured claim data into structured outputs customers can act on, selecting between semantic search, classical NLP, or LLM approaches.
  • Build and operate durable, event-driven, multi-tenant systems where agents run as steps within deterministic workflows, including failure modes, replay semantics, and cost profiling prior to load.
  • Own tracing and audit for your area so every model call and agent step can be reconstructed and cited back to source material.
  • Own how claims data is modeled and handled across relational, document, and graph stores.
  • Set the privacy and security baseline for systems in your area, covering tenant isolation, least-privilege access for agents and their tools, redaction and retention, and what personal and medical data may reach a model provider.
  • Author design documents for complex systems to obtain feedback and buy-in, review engineering designs and code with consistent rigor, and mentor engineers in your scope.
  • Participate in a shared on-call rotation for systems you own.

Required Qualifications

  • 6+ years of professional software engineering experience, including an LLM-powered system you architected, shipped to external customers, and carried through at least one model or prompt migration.
  • Strong Python and TypeScript practices at the level of setting team conventions, including async model usage, typing strictness, and test strategy for shared codebases, plus profiling and performance fixes.
  • Proven experience establishing an evals and guardrails approach in a product area built by others, including how human approvals work and how production corrections re-enter the eval loop.
  • Ability to clearly explain how you constrain an agentic coding tool’s context and verify output quality in production work, including examples across tools such as Claude Code, Cursor, or GitHub Copilot.
  • NLP pipeline experience blending LLMs, semantic search, and classical NLP, including a defendable retrieval design and an example where classical techniques beat an LLM on cost, latency, or reliability.
  • Experience designing durable, event-driven systems with deterministic orchestration around non-deterministic model steps, including compensating actions, idempotency, backpressure, and event-sourced replay.
  • Production experience making data-store decisions among PostgreSQL, MongoDB, Neo4j, SQL Server, and Cosmos DB, with the ability to defend the data model, partition key, consistency approach, and tenant isolation decisions.
  • Experience architecting and operating on Azure or AWS, with CI/CD and infrastructure as code, plus hands-on time with a durable workflow engine such as Temporal or the Durable Task Scheduler.
  • Authorship of technical design documents for systems spanning multiple teams, leading to shared feedback and adoption of resulting designs.
  • Clear communication skills for both technical and non-technical audiences, including translating customer problems and trade-offs.

Technologies

Python, TypeScript, Claude Code, Cursor, GitHub Copilot, PostgreSQL, MongoDB, Neo4j, SQL Server, Cosmos DB, Azure, AWS, CI/CD, infrastructure as code, Temporal, Durable Task Scheduler, LangGraph, Microsoft Agent Framework, Model Context Protocol (MCP), Agent2Agent (A2A) protocol, LangSmith, Arize, Braintrust, OpenTelemetry, React, A/B tests

Benefits

  • 401K Match
  • Paid time off
  • Annual Incentive Plan Performance Bonus
  • Comprehensive health insurance
  • Adoption Assistance
  • Tuition Reimbursement
  • Wellness Programs
  • Stock Purchase Plan options
  • Employee Resource Groups

Even Better If You Have

  • React and TypeScript experience to deliver full-stack features end to end, including chat, streaming, and citation-heavy AI experiences.
  • Experience with agent frameworks including LangGraph and the Microsoft Agent Framework, and familiarity with how agents reach tools through MCP, A2A, command-line tools, and microservices.
  • Experience with LLM observability and evaluation tooling such as LangSmith, Arize, or Braintrust, plus OpenTelemetry-based tracing.
  • Experience designing A/B tests for AI features.
  • Exposure to the insurance, claims, or automotive repair domains.

What Success Looks Like

  • Within the first six months, model changes in your area are gated by evals you designed, regressions are caught before customers see issues, and accuracy results can be shown as held or improved.
  • You handle ambiguity and convert it into a roadmap, shipped system, and measurable outcomes for customers.
  • Your team ships more than its headcount suggests by applying rigorous specifications and reviews that protect output quality in production.
  • Your systems are designed with known failure modes, and when failures occur, traces identify the exact source of the problem.
  • Privacy and security reviews for your area find no new issues beyond what was planned in the original design.
  • You establish standards that continue to work after any single project ends.
  • Engineers you mentor ship LLM-powered features with evals and guardrails in place without needing direct involvement.
  • You bring product, data science, and customer teams into a single aligned scope, and you can debate trade-offs without needing to win.
  • You prefer problems where the answer can be measured rather than argued, aligned with ambitious, collaborative, and empathetic values.

Interview Policy & Privacy Notice

  • A video interview is required for this position, and video interviews are transcribed.
  • Transcriptions are retained and may be reviewed by CCC and our recruiters.
  • Candidates may not use generative AI or automated assistance during interviews unless explicitly allowed by the interview team for a specific exercise.
  • Our Job Applicant Privacy Notice is available HERE.

Similar Jobs