Machine Learning Engineer, Reliability
Remote • Remote • FullTime
Posted 1mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Machine Learning Engineer
About the job
fal is building the generative media ecosystem for the next generation of AI products, providing the infrastructure, tools, and model access needed to scale from idea to production. As generative media reshapes industries, fal is becoming the foundation for ambitious teams. This hybrid ML Engineering / Site Reliability Engineering role will own the reliability, security, and safety of fal's generative media model APIs, ensuring they remain available, performant, secure, and safe for thousands of developers and enterprises. You will address model-specific failure modes, such as degraded output quality, drift, unsafe generations, and abuse patterns, as critical reliability concerns alongside uptime and latency.
Responsibilities
- Own availability, latency, and throughput SLOs for a large fleet of generative media model APIs.
- Build monitoring, alerting, and observability systems to detect ML-specific failures and output quality degradation.
- Harden model deployment workflows using canary releases, shadow testing, and automated rollbacks.
- Drive the security posture of the model fleet, including secure model serving and abuse detection.
- Operationalize safety systems for generative media, including content moderation and safety classifiers.
- Lead incident response for model API outages and degradations, conduct postmortems, and prevent recurrence.
- Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads.
- Partner with model and infrastructure teams to integrate reliability, security, and safety requirements into new model onboarding.
Requirements
- 3+ years of professional experience, with 1 year operating production ML or high-scale API systems, ideally with on-call ownership.
- Strong systems fundamentals in distributed systems, networking, observability, and incident management.
- Working knowledge of modern generative models (diffusion, transformers) and their production failure modes.
- Familiarity with security and safety practices for ML systems, abuse prevention, content safety, or trust & safety engineering.
- A bias toward automation, measurement, and blameless postmortems.