Generative AI Inference Engineer
Remote • United States
Posted 4mo ago
Job Location
United States
Tech Stack
Remote Work Policy
Fully remote
Categories
Machine Learning Engineer
About the job
We are seeking passionate Machine Learning Engineers to join our Inference team, focusing on the creative applications of generative AI models. The ideal candidate will have substantial experience developing and running inference for multi-modal models. A deep understanding of diffusion model architectures and familiarity with workflow tools like ComfyUI are a big plus. You will be expected to leverage and push the boundaries of state-of-the-art inference optimization techniques for multi-modal generative models. This role offers the opportunity to work alongside top researchers and engineers, utilizing cutting-edge high-performance computing resources to make a significant impact in the rapidly evolving field of generative AI.
Responsibilities
- Lead the design and development of customer-facing multi-modal ML inference systems.
- Build inference systems for next-generation models, focusing on optimization, model tuning, and deployment.
- Partner with cloud providers to deliver hosted inference solutions.
- Act as a strategic thought partner for leaders on driving business impact through machine learning.
- Contribute to bringing new models and pipelines into existence.
- Prototype and productionize inference platform improvements and new features.
Requirements
- 7+ years of experience productionizing machine learning systems, including inference pipeline development.
- Expert-level knowledge of writing and running Python services at scale.
- 5+ years of experience with the Python scientific stack and PyTorch.
- Experience with at least one high-performance inference framework (e.g., Triton, TensorRT).
- Deep understanding of Diffusion Architectures.
- Experience profiling and optimizing deep neural networks on Nvidia GPUs using tools like NVIDIA Nsight.
- Experience with Python-based image manipulation/encoding/decoding frameworks (e.g., OpenCV).
- Experience deploying to cloud orchestration systems (e.g., Kubernetes) and cloud providers (AWS, GCP, Azure).
- Experience with Docker.
- Ability to rapidly prototype solutions and iterate under tight product deadlines.
- Strong communication, collaboration, and documentation skills.
- Experience with the open-source ML ecosystem (e.g., HuggingFace, W&B).