Principal Machine Learning Engineer
Job Description
Oracle, Seattle onsite invites you to join the Generative AI Services team as a Principal Machine Learning Engineer. In this role, you will lead the architecture, design, and development of distributed, scalable AI infrastructure for training, fine-tuning, and inference. You will collaborate with partner teams to deliver reliable production deployments and contribute to open source frameworks such as vLLM and SGLang, strengthening Oracle Cloud Infrastructure's footprint in the AI ecosystem.
What you will do
- Lead the architecture, design, and development of distributed, scalable, and high-performance systems supporting AI model training, fine-tuning, and inference. Align these systems with production requirements and reliability goals.
- Build and optimize next-generation AI infrastructure that powers large-scale generative AI workloads, focusing on throughput, latency, and operational efficiency.
- Lead the analysis of model architectures to drive improvements in performance, efficiency, and scalability, evaluating trade-offs and selecting effective strategies.
- Leverage cutting-edge technologies to develop state-of-the-art AI systems and onboard frontier models, coordinating with research and product teams to translate breakthroughs into production capabilities.
- Benchmark, diagnose, troubleshoot, and resolve issues across the AI model lifecycle, including training, fine-tuning, and serving, to ensure reliable and scalable production deployments.
- Contribute to open-source frameworks such as vLLM and SGLang and strengthen Oracle Cloud Infrastructure through collaboration with internal teams and the broader ecosystem.
- Lead a team of senior and junior engineers, guiding them to deliver the roadmap on time with high quality while fostering mentorship and engineering excellence.
Technologies
- vLLM
- SGLang
Similar Jobs
J