Rearc is building production-grade, enterprise-ready AI/ML systems, and this role is a 100% hands-on engineering position focused on shipping real outcomes. You’ll design, build, and deploy AI-driven capabilities that go beyond prototypes, with an emphasis on end-to-end delivery: evaluation, monitoring, and integration into business workflows.
Location: Remote
Compensation: USD 95,000 - 209,000 per year
Experience: 4+ years
What you’ll do
- Design and implement AI agents, including RAG pipelines, orchestration workflows, and tool invocation
- Build evaluation frameworks to measure accuracy, latency, cost, and reliability
- Implement observability and monitoring across the AI system lifecycle
- Integrate with multiple AI providers and develop abstraction layers for multi-model architectures
- Optimize AI systems for performance, cost, and scalability
- Build and deploy AI-powered applications tightly coupled with real business workflows
- Integrate AI systems into existing enterprise platforms and APIs
- Debug and optimize live production systems
- Collaborate closely with client and internal engineering teams
- Participate in technical design discussions with a focus on implementation
What you bring
- 4+ years building and deploying AI/ML systems in production (beyond demos or experimentation)
- A track record of architecting, building, and successfully shipping AI/ML or software solutions using modern AI-assisted workflows
- Strong understanding of AI system evaluation and measurement: offline metrics, online monitoring, LLM-as-judge processes, regression testing, and cost/latency tracking
- Practical judgment in retrieval and agent design trade-offs, with the ability to explain and choose between RAG, agent loops, and workflows as needed
- Hands-on experience with LLM platforms such as OpenAI, Anthropic, Google Vertex, or similar, plus orchestration/harness patterns
- Python proficiency, forming the foundation of engineering work
- Backend engineering skills including building and deploying APIs, working with Docker, and navigating cloud-native environments (containers and basic infrastructure)
- Strong software engineering fundamentals: production-grade, maintainable code
- Experience with CI/CD pipelines, infrastructure as code, and production observability
- Ability to debug and optimize systems already in production
- Strong communication skills, including explaining technical trade-offs to non-technical stakeholders
Technologies you may work with
Python, OpenAI, Anthropic, Google Vertex, Docker, CI/CD pipelines, infrastructure as code, LLM-as-judge processes, RAG, FastAPI, Pydantic, PostgreSQL, MySQL, DuckDB, DSPy, MLflow, promptfoo, RAGAS, Claude SDK, OpenAI SDK, TypeScript, Go, Databricks, AWS, Azure, GCP
Preferred experience
- Familiarity with prompt optimization or evaluation tools (DSPy, MLflow, promptfoo, RAGAS, etc.)
- LLMOps/MLOps experience building robust, monitored, self-healing AI systems
- Experience with harness engineering (for example, developing on/with Goose, Pi, Claude Code, Codex)
- Databricks experience (preferred)
- Experience with cloud platforms such as AWS, Azure, or GCP
- Experience with FastAPI, Pydantic, PostgreSQL, MySQL, or DuckDB
- Experience using the Claude SDK or OpenAI SDK
- Additional programming languages beyond Python (TypeScript or Go are strong positives)
- Experience mentoring or upskilling fellow engineers