Machine Learning Engineer
$160k - $220k • Remote • San Francisco
Posted 1mo ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Categories
Machine Learning Engineer
About the job
Together AI is seeking an ML Engineer to develop systems and APIs for customer inference and fine-tuning of LLMs. The ideal candidate will have experience implementing runtime systems for large-scale AI/ML model inference, including the largest LLMs. This role involves designing and building production systems for reliability and performance at scale, partnering with cross-functional teams, and improving system efficiency and stability.
Responsibilities
- Design and build production systems for Together Cloud inference and fine-tuning APIs.
- Enable reliability and performance at scale for inference and fine-tuning APIs.
- Partner with researchers, engineers, product managers, and designers to launch new features.
- Analyze and improve efficiency, scalability, and stability of system resources.
- Conduct design and code reviews.
- Create services, tools, and developer documentation.
- Create testing frameworks for robustness and fault-tolerance.
- Participate in an on-call rotation for critical incidents.
Requirements
- 5+ years of experience writing high-performance, well-tested, production-quality code.
- Bachelor’s degree in computer science or equivalent industry experience.
- Familiarity with the LLM inference ecosystem, including frameworks and engines (e.g., vLLM, SGLang, TRT).
- Demonstrated experience building large-scale, fault-tolerant, distributed systems (storage, search, computation).
- Expert-level programming in Python, Go, Rust, or C/C++.
- Experience implementing runtime inference services at scale or similar.
Benefits
- Competitive compensation
- Startup equity
- Health insurance
- Other competitive benefits