Staff Machine Learning Engineer
Job Description
Lead the design and delivery of mission-critical ML systems for Xometry’s AI/ML integration, with emphasis on real-time serving and low-latency data flows into partner tooling.
Responsibilities
- Own end-to-end delivery from requirements gathering through release, ensuring high-quality, on-time outcomes across complex, cross-functional initiatives
- Architect and build partner integration ML, delivering a high-performance AI/ML layer for the embedded DFM AI + IQE integration with Teamcenter and Designcenter
- Design real-time ML serving architecture and a low-latency signal path that returns DFM and pricing feedback directly into the designer environment
- Define data contracts for model inputs and outputs; implement MLOps, governance, and observability for a mission-critical public-marketplace partner integration
- Develop cloud-based production systems including real-time endpoints and MLOps integrated with Xometry’s broader systems and infrastructure
- Handle cross-domain technical challenges by evaluating variable factors and aligning solutions with both business and technical objectives
- Surface opportunity areas, drive new processes and solutions, and build multi-quarter roadmaps for key technical goals
- Apply best practices across automated testing, parallel and distributed computing, and secure software development for ML systems
- Collaborate with engineers, product managers, data scientists, and business stakeholders to translate requirements into robust technical solutions
- Conduct design reviews, code reviews, and provide technical mentorship to raise team capability
- Stay current with ML/AI advances and introduce relevant new approaches, tools, and frameworks
Requirements
- Bachelor’s degree in a STEM field (or equivalent experience) plus 6-8 years of experience in machine learning engineering, with proven ownership of complex production ML systems
- Strong expertise in ML and AI technologies, including Gradient Boosting, Deep Learning, and/or Generative AI frameworks, with emphasis on backend scalability and reusability
- Hands-on experience deploying real-time ML products at scale in cloud environments (with AWS strongly preferred), including auto-scaling, monitoring, and alerting
- Advanced proficiency in Python and ML/AI frameworks such as TensorFlow or PyTorch
- Solid software engineering fundamentals, including data structures and algorithms
- Experience with MLOps: model monitoring, data and concept drift detection, and automated retraining plus redeployment pipelines
- Proficiency with CI/CD pipelines (e.g., GitHub Actions), test-driven development, and infrastructure as code (e.g., Terraform)
- Demonstrated ability to profile and optimize existing ML deployments for latency and throughput
- Capability to operate independently on new and ambiguous assignments, select methods and procedures, and communicate effectively across engineering, product, and business audiences
- Experience with modern modeling approaches including transformers, self-supervised pre-training, large language models (LLMs), or generative AI
- Knowledge of containers, Kubernetes, and cloud-native distributed systems
- Manufacturing, supply chain, or marketplace background is a plus, with emphasis that curiosity and drive matter
Technologies
- Python
- TensorFlow
- PyTorch
- Gradient Boosting
- Deep Learning
- Generative AI frameworks
- AWS
- CI/CD pipelines
- GitHub Actions
- Test-driven development
- Terraform
- MLOps
- Model monitoring
- Data and concept drift detection
- Auto-scaling
- Containers
- Kubernetes
- Transformers
- Self-supervised pre-training
- Large language models (LLMs)
- Solid Edge
- NX
- Designcenter
- Teamcenter
- Parallel and distributed computing
Benefits
- 401(k) match
- Medical, dental and vision insurance
- Life and disability insurance
- Generous paid time off including vacation, sick leave, floating and fixed holidays
- Maternity and bonding leave
- EAP and other wellbeing resources