Staff Machine Learning Engineer, ML Platform
Ai Ml
Artificial Intelligence
Automation
Big Data
Cloud
Cloud Infrastructure
Cloud Native
Cloud Platform
Cloud Platforms
Cloud Technology
Data Analysis
Data Integration
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Database
Databricks Mlflow
DevOps
DevSecOps
Engineering
Engineering Software
Facilities Management
Kubernetes
Machine Learning
Machine Learning Engineering
Message Queue
Ml Ops
MongoDB
Platform Engineering
Programming
Programming Language
Programming Languages
RabbitMQ
Risk Management
Security Automation
Technical Lead
Job Description
Braze is hiring a Staff Machine Learning Engineer for the Predictive and Generative AI (PGAI) team to own the ML platform underneath production ML and drive reliability at global scale.
Responsibilities
- Identify and lead major initiatives that change production ML operations, including replatforming queueing and orchestration, overhauling deployment and cloud identity, or retiring infrastructure components
- Ship high-impact infrastructure and ML platform changes at high velocity, with direct hands-on ownership of complex initiatives from design through production
- Own the platform technical vision and production quality bar, including direction for training, deployment, serving, and observability
- Lead incident response for ML systems and drive reliability and cost improvements to keep the platform efficient at scale
- Coordinate and deliver cross-team initiatives leveraging shared infrastructure, deployment tooling, and data systems owned with partner teams
- Improve engineering quality through design review, code review, and production readiness practices for ML systems
- Mentor other senior engineers and data scientists
- Translate technical decisions into customer and business outcomes and represent the team’s technical perspective to product and engineering leadership
Requirements
- 8+ years building and operating distributed systems in production, with depth in deployment and operations
- Hands-on production experience with ML workloads
- Proven technical leadership: owned team direction, led multi-quarter initiatives across team boundaries, and grew senior engineers while maintaining high personal output
- Deep working knowledge of Kubernetes and cloud infrastructure, including identity and access management, networking, and the cost profile of workloads
- Strong verbal and written communication skills, able to build consensus and drive forward decision-making
- Bonus: Experience with queueing and orchestration systems such as Celery, RabbitMQ, Kafka, or Ray
- Bonus: Experience with ML platform tooling such as MLflow or other model registry tools, feature stores, or ML observability
- Bonus: Experience with Braze stack: Python, Ruby on Rails, MongoDB, Redis, Kubernetes
- Bonus: Experience operating under compliance regimes such as SOX or HIPAA
- Bonus: Experience in customer engagement, personalization, or marketing technology domains
Technology Focus
- Celery, RabbitMQ, Kafka, Ray
- MLflow
- Python, Ruby on Rails
- MongoDB, Redis
- Kubernetes
- SOX, HIPAA
Location and Compensation
- Location: Chicago, IL (hybrid)
- Salary: USD 184,000 - 314,000 per year
- Minimum experience: 8 years
Benefits
- Competitive compensation that may include equity
- Retirement and Employee Stock Purchase Plans
- Flexible paid time off
- Comprehensive benefit plans covering medical, dental, vision, life, and disability
- Family services including fertility benefits and equal paid parental leave
- Professional development supported by formal career pathing, learning platforms, and a yearly learning stipend
- Curated in-office employee experience designed to foster community, team connections, and innovation
- Opportunities to give back, including annual Volunteer Week and donation matching
- Employee Resource Groups
- Collaborative, transparent, and fun culture recognized as a Great Place to Work®