Machine Learning Engineer
Agent Orchestration
Agentic Ai
Agentic Ai Orchestration
Agentic Systems
Ai Agent
Ai Agent Platform
Ai Engineer
Ai Ml
Artificial Intelligence
Data Science Ml
DevOps
Generative AI
Llm Agents
Machine Learning Engineer
Machine Learning Evaluation
Machine Learning Infrastructure
Machine Learning Models
Machine Learning Pipelines
Ml Ops
Multi Agent Orchestration
NLP
Programming
Programming Language
Programming Languages
Systems Engineer
Job Description
UST is seeking a Machine Learning Engineer (ML Engineer II) to build and optimize multi-agent systems for real-world performance. This onsite role in Oregon focuses on taking agent workloads from application design through inference runtime behavior and down to accelerator hardware and middleware layers.
What you’ll do
- Develop a multi-agent workload at the application level, then profile and optimize its execution all the way through the inference runtime, OMIX/middleware, compute runtime, Linux GPU driver, and accelerator hardware.
- Improve key performance metrics including latency, throughput, tokens/sec, memory footprint, and accelerator utilization.
- Select and tune models and inference engines based on agent workload characteristics, hardware capabilities, and performance requirements.
What you bring
- Hands-on development experience with multi-agent workloads.
- Experience with agent frameworks such as LangGraph/LangChain, AutoGen, and CrewAI.
- Strong understanding of LLMs, SLMs, and multimodal models, including Transformer architecture, attention, tokenization, context windows, and KV cache concepts.
- Hands-on experience selecting, evaluating, and deploying models for different agentic workload requirements.
- Knowledge of model formats and optimization workflows, including ONNX, OpenVINO IR, safe tensors, and quantization (FP16, BF16, INT8, INT4).
- Experience with inference engines and runtimes such as OpenVINO, ONNX Runtime, vLLM, llama.cpp, or TGI.
- Understanding of prefill vs. decode, batching, continuous batching, speculative decoding, KV-cache management, and memory optimization.
- Solid grasp of CPU/GPU execution, device placement, and heterogeneous inference.
- Ability to benchmark and compare models, inference engines, and runtime configurations, then identify performance bottlenecks.
- Knowledge of GPU memory, kernel execution, synchronization, device selection, and host/device data movement.
- Familiarity with middleware and accelerator compute runtimes such as OMIX/OneAPI/SYCL, Level Zero, or OpenCL.
- Python strength, plus working knowledge of C/C++.
- Experience with Git/GitHub, Linux shell, and debugging tools.
- Practical use of GitHub Copilot for development, debugging, and code generation.
- Docker and container fundamentals.
- Ability to benchmark and profile AI workloads.
- Understanding of latency, throughput, tokens/sec, GPU utilization, memory bandwidth, and CPU/GPU bottlenecks, including identifying whether a performance issue comes from the agent, model, inference engine, runtime, or driver.
Core technologies
- LangGraph, LangChain, AutoGen, CrewAI
- Python, C/C++, Git, GitHub, GitHub Copilot
- Docker
- ONNX, Open VINO IR, safe tensors
- FP16, BF16, INT8, INT4
- Open VINO, ONNX Runtime, vLLM, llama.cpp, TGI
- OMIX, OneAPI, SYCL, Level Zero, OpenCL
Role location
Oregon (onsite)
Compensation
$82,000 - $123,000 USD per year
Benefits
- Full-time employees accrue a minimum of 10 days of paid vacation per year.
- 6 days of paid sick leave each year (pro-rated for new hires throughout the year).
- 10 paid holidays.
- Eligible for paid bereavement leave and jury duty.
- 401(k) Retirement Plan with employer matching.
- Medical, dental, and vision insurance for employees and dependents residing in the US.
- Company-paid Employee Only benefits: basic life insurance, accidental death and disability insurance, and short- and long-term disability benefits.
- Regular employees may purchase additional voluntary short-term disability benefits.
- Eligible to participate in a Health Savings Account (HSA).
- Eligible to participate in a Flexible Spending Account (FSA) for healthcare, dependent child care, and/or commuting expenses.
Additional skills
- Agentic AI
- Python
- Linux
- MCP3
- RAG
- Kubernetes