Lead Principal Machine Learning Engineer
Job Description
Based in Seattle, WA onsite, Oracle seeks a Lead Principal ML Engineer to define, build, and operate production-grade agentic AI platforms on Oracle Cloud Infrastructure, providing technical leadership and architectural direction across distributed AI systems.
Responsibilities
- Provide senior technical ownership for OCI AI platform capabilities, including agent execution, inference pipelines, model serving, AI workflow orchestration, evaluation, and observability.
- Design and deliver scalable agentic AI systems with reasoning, planning, tool usage, workflow execution, multi-step task orchestration, and secure human-in-the-loop escalation.
- Develop production-ready services for tool invocation, agent memory, context management, Model Context Protocol integration, vector retrieval, multi-agent coordination, policy enforcement, and evaluation.
- Lead architecture across distributed services optimized for low latency, high throughput, GPU efficiency, reliability, cost, operability, and secure multi-tenant operation.
- Define service boundaries, APIs, data models, state management strategies, consistency tradeoffs, failure modes, SLIs/SLOs, rollout plans, and operational readiness criteria for AI platform services.
- Drive cross-functional technical strategy spanning infrastructure, platform, security, data, and application engineering, translating goals into multi-quarter roadmaps with milestones.
- Securely and reliably integrate AI agents with enterprise APIs, cloud services, databases, identity systems, secrets management, and external systems.
- Establish AgentOps and LLMOps practices for tracing, monitoring, evaluation suites, regression testing, experimentation, safety guardrails, prompt and tool versioning, and production reliability.
- Evaluate and operationalize emerging technologies in generative AI, agentic workflows, inference optimization, long-context systems, reasoning models, AI developer tooling, and agent-first development.
- Advance engineering excellence through code and design reviews, test strategy, deployment automation, incident analysis, documentation, and AI-assisted development using tools like Codex, Claude Code, Cursor, Copilot, or similar systems.
- Mentor staff and senior engineers, raise architectural standards, and influence OCI engineering practices without direct managerial authority.
- Own critical production outcomes including reliability, performance, security posture, cost efficiency, and supportability for delivered systems.
Requirements
- Bachelor's, master's, or doctoral degree in computer science, AI/ML, engineering, or a related field, or equivalent practical experience.
- 12+ years of professional software engineering experience with substantial ownership of production systems, or equivalent impact at senior/principal levels.
- Proven track record as Staff, Senior Staff, Principal, or equivalent technical leader influencing architecture and execution across multiple teams.
- Extensive experience designing, building, and operating large-scale distributed systems, cloud services, infrastructure platforms, or AI/ML platform services.
- Hands-on experience with production AI systems, agentic applications, autonomous workflows, tool-using agents, multi-step orchestration, or multi-agent setups.
- Practical experience with orchestration frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar ecosystems.
- Strong understanding of LLM application patterns, including prompt design, structured outputs, function/tool calling, context management, RAG, memory, tool safety, and evaluation.
- Proficiency in Python with ability to contribute production-grade code, code reviews, tests, and debugging in complex distributed environments.
- Deep expertise with Kubernetes, Docker, cloud-native infrastructure, service-to-service communication, scalability, fault tolerance, observability, and performance analysis.
- Experience defining SLIs/SLOs, production readiness criteria, incident response practices, monitoring, tracing, experiments, and reliability programs for AI or distributed systems.
- Strong understanding of AI safety, governance, security, and operational risks for autonomous or semi-autonomous systems, including data handling, access control, auditability, and human accountability.
- Excellent written and verbal communication, with demonstrated ability to lead technical direction, resolve ambiguity, and influence senior stakeholders.