Staff Machine Learning Engineer
Job Description
The Staff Machine Learning Engineer role is a senior individual contributor position focused on end-to-end delivery of complex machine learning systems and partner integrations for Xometry’s DFM AI + IQE initiative. This hybrid role in Denver, CO will own AI/ML architecture spanning real-time, low-latency pipelines and the MLOps and observability layer connecting Xometry’s platform with Solid Edge, NX, Designcenter, and Teamcenter.
Key Responsibilities
- Lead with technical depth and own the full lifecycle from requirements through release, ensuring high-quality, on-time delivery across complex, cross-functional efforts
- Own the partner integration AI/ML plane by architecting and building the high-performance AI/ML layer of the embedded DFM AI + IQE integration with Teamcenter and Designcenter
- Define the real-time machine learning serving architecture and the low-latency signal path delivering DFM and pricing feedback directly into the designer’s environment
- Set model data contracts for inputs and outputs, and implement mission-critical MLOps, governance, and observability for a public-marketplace partner integration
- Develop cloud-based production systems for real-time endpoints and MLOps, integrated with Xometry’s broader systems and infrastructure
- Address complex, cross-domain technical challenges by evaluating variable factors and meeting business and technical objectives
- Identify opportunity areas, take ownership of new processes and solutions, and build multi-quarter roadmaps aligned to key technical goals
- Apply best practices in automated testing, parallel and distributed computing, and secure software development for ML systems
- Partner with engineers, product managers, data scientists, and business stakeholders to translate requirements into robust technical solutions
- Support team execution through design reviews, code reviews, and technical mentorship
- Stay current with advances in ML/AI and bring relevant tools, approaches, and frameworks into production
Required Qualifications
- Bachelor’s degree in a STEM field (or equivalent experience) plus 6-8 years of machine learning engineering experience, including owning and delivering complex ML systems in production
- Deep expertise in ML and AI, including Gradient Boosting, Deep Learning, and/or Generative AI frameworks, with a focus on backend scalability and reusability
- Hands-on experience deploying real-time ML products at scale in cloud environments (AWS strongly preferred), including auto-scaling, monitoring, and alerting
- Strong proficiency in Python and advanced ML/AI frameworks such as TensorFlow or PyTorch
- Solid software engineering foundations, including data structures and algorithms
- Demonstrated MLOps experience, including model monitoring, data and concept drift detection, and automated retraining and redeployment pipelines
- Experience with CI/CD pipelines (for example, GitHub Actions), test-driven development, and infrastructure as code (for example, Terraform)
- Ability to profile and optimize existing ML deployments for latency and throughput
- Capability to work independently on new and ambiguous assignments, determine methods and procedures, and communicate effectively across engineering, product, and business audiences
- Experience with state-of-the-art modeling approaches, including transformers, self-supervised pre-training, large language models (LLMs), or generative AI
- Knowledge of containers, container orchestration (Kubernetes), and cloud-native distributed systems
- Background in manufacturing, supply chain, or marketplace environments is a plus
Tools and Technologies
- Python, TensorFlow, PyTorch, Gradient Boosting, Deep Learning, Generative AI frameworks
- AWS, CI/CD pipelines, GitHub Actions, test-driven development, Terraform
- Transformers, self-supervised pre-training, large language models (LLMs)
- Containers, Kubernetes, cloud-native distributed systems
- Solid Edge, NX, Designcenter, Teamcenter
- MLOps, model monitoring, data drift detection, concept drift detection, automated retraining and redeployment pipelines
- Infrastructure as code
Location and Salary
- Location: Denver, CO (hybrid)
- Compensation: USD 200,000 - 220,000 per year
Benefits
- 401(k) match
- Medical, dental and vision insurance
- Life and disability insurance
- Generous paid time off including vacation, sick leave, floating and fixed holidays, maternity and bonding leave
- EAP and other wellbeing resources