EngineerJobs.io
← Back to all jobs

Job Description

Own reliability, security, and safety for fal’s fleet of generative media model APIs, combining ML engineering with site reliability practices.

Responsibilities

  • Own availability, latency, and throughput SLOs for a large fleet of generative media model APIs handling production traffic at scale
  • Build monitoring, alerting, and observability to detect ML-specific failures, output quality degradation, pipeline breakage, and model regressions before customers are impacted
  • Harden model deployment workflows using canary releases, shadow testing, automated rollbacks, and validation gates for safer model version shipping
  • Drive the model fleet’s security posture through secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns
  • Operationalize safety systems for generative media, including content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without harming performance
  • Lead incident response for model API outages and degradations, conduct postmortems, and drive engineering actions to prevent recurrence
  • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic
  • Partner with model and infrastructure teams to ensure reliability, security, and safety requirements are integrated into new model onboarding to the platform

Requirements

  • 5+ years of professional experience, including 2 years operating production ML or high-scale API systems, ideally with on-call ownership
  • Production experience supporting diffusion models
  • Strong systems fundamentals: distributed systems, networking, observability, and incident management
  • Working knowledge of modern generative models (including diffusion and transformers) and how they fail in production
  • Familiarity with security and safety practices for ML systems, including abuse prevention and content safety or trust & safety experience (strong plus)
  • Bias toward automation, measurement, and blameless postmortems

Technologies

  • Python
  • torch
  • diffusers
  • Kubernetes
  • fal Python SDK

Location & Employment

  • Remote (APAC)
  • Must be based in India, Australia, or New Zealand
  • Full time
  • Department: EngineeringML
  • Team focus: ML Engineering / Site Reliability Engineering

Additional context: Access to a massive GPU cluster for inference and evaluation; work alongside a team focused on quickly iterating on and deploying new AI breakthroughs while maintaining reliability.

Similar Jobs