EngineerJobs.io
← Back to all jobs

Job Description

UST is seeking a Machine Learning Engineer (ML Engineer II) to build and optimize multi-agent systems for real-world performance. This onsite role in Oregon focuses on taking agent workloads from application design through inference runtime behavior and down to accelerator hardware and middleware layers.

What you’ll do

  • Develop a multi-agent workload at the application level, then profile and optimize its execution all the way through the inference runtime, OMIX/middleware, compute runtime, Linux GPU driver, and accelerator hardware.
  • Improve key performance metrics including latency, throughput, tokens/sec, memory footprint, and accelerator utilization.
  • Select and tune models and inference engines based on agent workload characteristics, hardware capabilities, and performance requirements.

What you bring

  • Hands-on development experience with multi-agent workloads.
  • Experience with agent frameworks such as LangGraph/LangChain, AutoGen, and CrewAI.
  • Strong understanding of LLMs, SLMs, and multimodal models, including Transformer architecture, attention, tokenization, context windows, and KV cache concepts.
  • Hands-on experience selecting, evaluating, and deploying models for different agentic workload requirements.
  • Knowledge of model formats and optimization workflows, including ONNX, OpenVINO IR, safe tensors, and quantization (FP16, BF16, INT8, INT4).
  • Experience with inference engines and runtimes such as OpenVINO, ONNX Runtime, vLLM, llama.cpp, or TGI.
  • Understanding of prefill vs. decode, batching, continuous batching, speculative decoding, KV-cache management, and memory optimization.
  • Solid grasp of CPU/GPU execution, device placement, and heterogeneous inference.
  • Ability to benchmark and compare models, inference engines, and runtime configurations, then identify performance bottlenecks.
  • Knowledge of GPU memory, kernel execution, synchronization, device selection, and host/device data movement.
  • Familiarity with middleware and accelerator compute runtimes such as OMIX/OneAPI/SYCL, Level Zero, or OpenCL.
  • Python strength, plus working knowledge of C/C++.
  • Experience with Git/GitHub, Linux shell, and debugging tools.
  • Practical use of GitHub Copilot for development, debugging, and code generation.
  • Docker and container fundamentals.
  • Ability to benchmark and profile AI workloads.
  • Understanding of latency, throughput, tokens/sec, GPU utilization, memory bandwidth, and CPU/GPU bottlenecks, including identifying whether a performance issue comes from the agent, model, inference engine, runtime, or driver.

Core technologies

  • LangGraph, LangChain, AutoGen, CrewAI
  • Python, C/C++, Git, GitHub, GitHub Copilot
  • Docker
  • ONNX, Open VINO IR, safe tensors
  • FP16, BF16, INT8, INT4
  • Open VINO, ONNX Runtime, vLLM, llama.cpp, TGI
  • OMIX, OneAPI, SYCL, Level Zero, OpenCL

Role location

Oregon (onsite)

Compensation

$82,000 - $123,000 USD per year

Benefits

  • Full-time employees accrue a minimum of 10 days of paid vacation per year.
  • 6 days of paid sick leave each year (pro-rated for new hires throughout the year).
  • 10 paid holidays.
  • Eligible for paid bereavement leave and jury duty.
  • 401(k) Retirement Plan with employer matching.
  • Medical, dental, and vision insurance for employees and dependents residing in the US.
  • Company-paid Employee Only benefits: basic life insurance, accidental death and disability insurance, and short- and long-term disability benefits.
  • Regular employees may purchase additional voluntary short-term disability benefits.
  • Eligible to participate in a Health Savings Account (HSA).
  • Eligible to participate in a Flexible Spending Account (FSA) for healthcare, dependent child care, and/or commuting expenses.

Additional skills

  • Agentic AI
  • Python
  • Linux
  • MCP3
  • RAG
  • Kubernetes

Similar Jobs