Agentic AI Engineer
Job Description
Catapult Sports is building athlete performance intelligence powered by production specialist AI agents and multi-agent orchestration. This role is focused on delivering recommendations that practitioners can understand and trust, with confidence calibration, human-in-the-loop escalation, and measurable evaluation and observability in production. This is an onsite position in New York, NY.
Compensation
The target Total Compensation range for this position is $107,250 - $214,500 per year, inclusive of base salary and a target incentive plan (which may include equity, commission, or other bonus structures). Your specific compensation within this range will be determined by factors such as your geographic location, relevant experience, and job-related skills.
Responsibilities
- Design and ship specialist AI agents that use memory, tools, data, and multi-step reasoning.
- Build multi-agent orchestration that routes work between specialist agents, manages dependencies, and synthesizes conflicting outputs.
- Develop systems that evaluate confidence, uncertainty, and consequence before recommendations reach a practitioner.
- Implement human-in-the-loop escalation to determine when the system should answer, ask for more information, or defer to a human.
- Create workflows that convert sport scientist expertise into validated, versioned, testable agent capabilities.
- Build evaluation, observability, and regression testing so agent performance can be measured and improved in production.
- Partner with domain experts to ensure outputs are grounded, traceable, and actionable.
Requirements
- Personally shipped a production agentic AI system used by real users.
- Hands-on experience with memory or persistent state.
- Hands-on experience with tool use or tool calling.
- Hands-on experience with multi-step reasoning or workflows.
- Hands-on experience with production deployment and operation.
- Built or substantially contributed to a production multi-agent system.
- Understands agent routing and orchestration, specialist agent composition, and dependency-aware workflows.
- Understands parallel and sequential execution, conflicting agent outputs, and response synthesis.
- Experience with LangGraph, AutoGen, CrewAI, or equivalent frameworks is valuable.
- Hands-on experience calibrating probabilistic ML or AI systems, including comfort with Platt scaling, isotonic regression, and Expected Calibration Error (ECE), plus reliability and calibration curves and confidence/uncertainty estimation.
- 5+ years of professional experience in applied ML, AI, or software engineering.
- Strong Python and software engineering fundamentals.
- Experience building and operating production systems.
- Experience with production RAG and reranking.
- Experience with foundation-model fine-tuning or domain adaptation, including LoRA and PEFT.
- Experience with LLM observability and drift detection.
- Experience with evaluation harnesses and automated regression testing.
- Experience with human-in-the-loop architectures, confidence thresholds, and escalation models.
- Experience with causal or counterfactual reasoning.
- Experience with Go/Golang.
- Experience with AWS, including ECS, EC2, Lambda, SNS, or SQS.
- Experience with GraphQL, REST, or gRPC.
- Experience with PostgreSQL or MongoDB.
- Experience working with sport scientists, clinicians, or other domain experts is a plus. Familiarity with workload, readiness, recovery, biomechanics, or athlete performance data will help.
Technologies
Python, LangGraph, AutoGen, CrewAI, Platt scaling, isotonic regression, Expected Calibration Error (ECE), LoRA, PEFT, Go/Golang, AWS, ECS, EC2, Lambda, SNS, SQS, GraphQL, REST, gRPC, PostgreSQL, MongoDB
Benefits
- Generous paid leave and recognized company holidays
- Health, Dental, and Vision insurance
- 401(k) retirement plan with company match
- Opportunity to participate in Catapult’s comprehensive benefits package
What Success Looks Like
- Build a platform where specialist agents investigate complex performance questions, use the right evidence, assess uncertainty, and produce recommendations a practitioner can understand and trust.
- The system knows when not to answer.
- Every recommendation is grounded, calibrated, traceable, and escalation-aware.
- The practitioner remains responsible for the decision, with AI that makes it better informed, faster, and more defensible.
Before You Apply
- Production agentic AI
- Production multi-agent orchestration
- Hands-on confidence calibration
Catapult does not expect every candidate to have every preferred skill. If you have core experience and are excited by the problem, please apply.
All offers of employment are subject to Catapult’s positive prehire check. To find out more, please contact the Talent Partner for this role.