Member of Technical Staff - Model Serving / API Backend Engineer
$180k - $300k • San Francisco (United States) • FullTime
Posted 2y ago
Job Location
San Francisco (United States)
Tech Stack
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Black Forest Labs is at the forefront of generative AI, known for foundational technologies like Latent Diffusion and Stable Diffusion. We are building the next generation of creative tools used by millions worldwide. This role is crucial for bridging the gap between cutting-edge research and production-ready systems, ensuring that our advanced models can be efficiently deployed and experienced by users. You will be instrumental in accelerating the pace at which research breakthroughs become usable APIs and demos, directly impacting inference speed, API performance under load, and the overall user experience of our models.
Responsibilities
- Transform research checkpoints into production-ready inference services.
- Design and maintain high-performance APIs serving millions of requests.
- Optimize inference latency and throughput across GPU infrastructure.
- Build scalable serving architectures capable of handling unpredictable traffic.
- Enhance reliability, monitoring, and observability for model-serving systems.
- Prototype and ship demos showcasing new capabilities rapidly.
- Collaborate with researchers to expedite the transition from idea to live endpoints.
- Manage distributed systems and task queues under variable load.
- Implement monitoring and observability for production ML systems.
- Debug performance bottlenecks across model, infrastructure, and network layers.
Requirements
- Proven experience building and operating systems at meaningful scale.
- Understanding the distinction between research prototypes and production systems.
- Comfort navigating ambiguity, making tradeoffs, and improving systems under real-world constraints.
- Strong judgment regarding performance, reliability, and cost tradeoffs.
- Experience scaling APIs or ML systems under load.
- Comfort working in fast-moving, research-adjacent environments.
- Demonstrated ownership from system design through debugging and deployment.
- Experience building and operating ML inference services in production.
- Experience designing scalable API architectures with async processing.
- Experience optimizing GPU workloads (batching, quantization, compilation, CUDA).
- Experience with real-time or low-latency inference systems.
- Experience with TensorRT, reduced precision, layer fusion, or model compilation techniques.
- Experience with CI/CD and automated testing for ML systems.
- Experience with security best practices for API and model serving.
Benefits
- Monthly in-person week for remote employees.
- Coverage of reasonable travel costs for in-person meetings.