Machine Learning Engineer - Inference
$160k - $230k • Remote • San Francisco
Posted 1mo ago
About the job
Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models and ensuring they run efficiently and effectively at scale. You will collaborate closely with AI researchers and engineers to create cutting-edge AI solutions and shape the future of AI inference.
Responsibilities
- Design and build production systems for the inference engine, ensuring reliability and performance at scale.
- Develop and optimize runtime inference services for large-scale AI applications.
- Collaborate with researchers, engineers, product managers, and designers to deliver new features and research capabilities.
- Conduct design and code reviews to maintain high quality standards.
- Create services, tools, and developer documentation to support the inference engine.
- Implement robust and fault-tolerant systems for data ingestion and processing.
Requirements
- 3+ years of experience writing high-performance, well-tested, production-quality code.
- Proficiency with Python and PyTorch.
- Demonstrated experience in building high-performance libraries and tooling.
- Excellent understanding of low-level operating system concepts including multi-threading, memory management, networking, storage, performance, and scale.
- Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum (preferred).
- Knowledge of AI inference techniques such as speculative decoding (preferred).
- Knowledge of CUDA/Triton programming (preferred).
- Knowledge of Rust, Cython, and compilers (nice to have).
Benefits
- Competitive compensation
- Startup equity
- Health insurance
- Other competitive benefits