This position is no longer accepting applications
Closed on August 17, 2026.
This role is filled — get an email when new Engineering roles open on EngineerJobs.io:
Lead Principal Machine Learning Engineer
Ai Ml
Ai Platform
Artificial Intelligence
Cloud
Cloud Native
DevOps
Engineering
Machine Learning Engineer
Ml Ops
Oci
Oracle Cloud
Platform Engineering
Programming Language
Programming Languages
Technical Lead
View similar jobs
Get alerted when similar jobs are posted — set up a New Engineering jobs on EngineerJobs.io alert.
See other roles at Oracle.
Job Description
Based in Seattle, WA onsite, Oracle seeks a Lead Principal ML Engineer to define, build, and operate production-grade agentic AI platforms on Oracle Cloud Infrastructure, providing technical leadership and architectural direction across distributed AI systems.
Responsibilities
- Provide senior technical ownership for OCI AI platform capabilities, including agent execution, inference pipelines, model serving, AI workflow orchestration, evaluation, and observability.
- Design and deliver scalable agentic AI systems with reasoning, planning, tool usage, workflow execution, multi-step task orchestration, and secure human-in-the-loop escalation.
- Develop production-ready services for tool invocation, agent memory, context management, Model Context Protocol integration, vector retrieval, multi-agent coordination, policy enforcement, and evaluation.
- Lead architecture across distributed services optimized for low latency, high throughput, GPU efficiency, reliability, cost, operability, and secure multi-tenant operation.
- Define service boundaries, APIs, data models, state management strategies, consistency tradeoffs, failure modes, SLIs/SLOs, rollout plans, and operational readiness criteria for AI platform services.
- Drive cross-functional technical strategy spanning infrastructure, platform, security, data, and application engineering, translating goals into multi-quarter roadmaps with milestones.
- Securely and reliably integrate AI agents with enterprise APIs, cloud services, databases, identity systems, secrets management, and external systems.
- Establish AgentOps and LLMOps practices for tracing, monitoring, evaluation suites, regression testing, experimentation, safety guardrails, prompt and tool versioning, and production reliability.
- Evaluate and operationalize emerging technologies in generative AI, agentic workflows, inference optimization, long-context systems, reasoning models, AI developer tooling, and agent-first development.
- Advance engineering excellence through code and design reviews, test strategy, deployment automation, incident analysis, documentation, and AI-assisted development using tools like Codex, Claude Code, Cursor, Copilot, or similar systems.
- Mentor staff and senior engineers, raise architectural standards, and influence OCI engineering practices without direct managerial authority.
- Own critical production outcomes including reliability, performance, security posture, cost efficiency, and supportability for delivered systems.
Requirements
- Bachelor's, master's, or doctoral degree in computer science, AI/ML, engineering, or a related field, or equivalent practical experience.
- 12+ years of professional software engineering experience with substantial ownership of production systems, or equivalent impact at senior/principal levels.
- Proven track record as Staff, Senior Staff, Principal, or equivalent technical leader influencing architecture and execution across multiple teams.
- Extensive experience designing, building, and operating large-scale distributed systems, cloud services, infrastructure platforms, or AI/ML platform services.
- Hands-on experience with production AI systems, agentic applications, autonomous workflows, tool-using agents, multi-step orchestration, or multi-agent setups.
- Practical experience with orchestration frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar ecosystems.
- Strong understanding of LLM application patterns, including prompt design, structured outputs, function/tool calling, context management, RAG, memory, tool safety, and evaluation.
- Proficiency in Python with ability to contribute production-grade code, code reviews, tests, and debugging in complex distributed environments.
- Deep expertise with Kubernetes, Docker, cloud-native infrastructure, service-to-service communication, scalability, fault tolerance, observability, and performance analysis.
- Experience defining SLIs/SLOs, production readiness criteria, incident response practices, monitoring, tracing, experiments, and reliability programs for AI or distributed systems.
- Strong understanding of AI safety, governance, security, and operational risks for autonomous or semi-autonomous systems, including data handling, access control, auditability, and human accountability.
- Excellent written and verbal communication, with demonstrated ability to lead technical direction, resolve ambiguity, and influence senior stakeholders.