Machine Learning Engineer, Reliability
Application Security
Artificial Intelligence
Automation
Cloud
Cloud Infrastructure
Cloud Native
Cloud Platform
Cloud Platforms
Cloud Technology
Data Science
DevOps
DevSecOps
Engineering
Engineering Software
Facilities Management
Generative Ai Security
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning
Machine Learning Engineer
Machine Learning Engineering
Machine Learning Pipelines
Ml Ops
Monitoring
Performance Engineering
Platform Engineering
Programming
Reliability Engineering
Risk Management
Security
Security Automation
Security Engineer
Site Reliability Engineering
Software Engineering
Software Security
Job Description
Own reliability, security, and safety for fal’s fleet of generative media model APIs, combining ML engineering with site reliability practices.
Responsibilities
- Own availability, latency, and throughput SLOs for a large fleet of generative media model APIs handling production traffic at scale
- Build monitoring, alerting, and observability to detect ML-specific failures, output quality degradation, pipeline breakage, and model regressions before customers are impacted
- Harden model deployment workflows using canary releases, shadow testing, automated rollbacks, and validation gates for safer model version shipping
- Drive the model fleet’s security posture through secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns
- Operationalize safety systems for generative media, including content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without harming performance
- Lead incident response for model API outages and degradations, conduct postmortems, and drive engineering actions to prevent recurrence
- Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic
- Partner with model and infrastructure teams to ensure reliability, security, and safety requirements are integrated into new model onboarding to the platform
Requirements
- 5+ years of professional experience, including 2 years operating production ML or high-scale API systems, ideally with on-call ownership
- Production experience supporting diffusion models
- Strong systems fundamentals: distributed systems, networking, observability, and incident management
- Working knowledge of modern generative models (including diffusion and transformers) and how they fail in production
- Familiarity with security and safety practices for ML systems, including abuse prevention and content safety or trust & safety experience (strong plus)
- Bias toward automation, measurement, and blameless postmortems
Technologies
- Python
- torch
- diffusers
- Kubernetes
- fal Python SDK
Location & Employment
- Remote (APAC)
- Must be based in India, Australia, or New Zealand
- Full time
- Department: EngineeringML
- Team focus: ML Engineering / Site Reliability Engineering
Additional context: Access to a massive GPU cluster for inference and evaluation; work alongside a team focused on quickly iterating on and deploying new AI breakthroughs while maintaining reliability.