Member of Technical Staff - Model Serving / API Backend Engineer

$180k - $300k San Francisco (United States) FullTime

Posted 2y ago

Job Location

San Francisco (United States)

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Black Forest Labs is at the forefront of generative AI, known for foundational technologies like Latent Diffusion and Stable Diffusion. We are building the next generation of creative tools used by millions worldwide. This role is crucial for bridging the gap between cutting-edge research and production-ready systems, ensuring that our advanced models can be efficiently deployed and experienced by users. You will be instrumental in accelerating the pace at which research breakthroughs become usable APIs and demos, directly impacting inference speed, API performance under load, and the overall user experience of our models.

Responsibilities

  • Transform research checkpoints into production-ready inference services.
  • Design and maintain high-performance APIs serving millions of requests.
  • Optimize inference latency and throughput across GPU infrastructure.
  • Build scalable serving architectures capable of handling unpredictable traffic.
  • Enhance reliability, monitoring, and observability for model-serving systems.
  • Prototype and ship demos showcasing new capabilities rapidly.
  • Collaborate with researchers to expedite the transition from idea to live endpoints.
  • Manage distributed systems and task queues under variable load.
  • Implement monitoring and observability for production ML systems.
  • Debug performance bottlenecks across model, infrastructure, and network layers.

Requirements

  • Proven experience building and operating systems at meaningful scale.
  • Understanding the distinction between research prototypes and production systems.
  • Comfort navigating ambiguity, making tradeoffs, and improving systems under real-world constraints.
  • Strong judgment regarding performance, reliability, and cost tradeoffs.
  • Experience scaling APIs or ML systems under load.
  • Comfort working in fast-moving, research-adjacent environments.
  • Demonstrated ownership from system design through debugging and deployment.
  • Experience building and operating ML inference services in production.
  • Experience designing scalable API architectures with async processing.
  • Experience optimizing GPU workloads (batching, quantization, compilation, CUDA).
  • Experience with real-time or low-latency inference systems.
  • Experience with TensorRT, reduced precision, layer fusion, or model compilation techniques.
  • Experience with CI/CD and automated testing for ML systems.
  • Experience with security best practices for API and model serving.

Benefits

  • Monthly in-person week for remote employees.
  • Coverage of reasonable travel costs for in-person meetings.

About Black Forest Labs

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.