Machine Learning Engineer

$160k - $220k Remote San Francisco

Posted 1mo ago

Remote Work Policy

Fully remote

Categories

Machine Learning Engineer

About the job

Together AI is seeking an ML Engineer to develop systems and APIs for customer inference and fine-tuning of LLMs. The ideal candidate will have experience implementing runtime systems for large-scale AI/ML model inference, including the largest LLMs. This role involves designing and building production systems for reliability and performance at scale, partnering with cross-functional teams, and improving system efficiency and stability.

Responsibilities

  • Design and build production systems for Together Cloud inference and fine-tuning APIs.
  • Enable reliability and performance at scale for inference and fine-tuning APIs.
  • Partner with researchers, engineers, product managers, and designers to launch new features.
  • Analyze and improve efficiency, scalability, and stability of system resources.
  • Conduct design and code reviews.
  • Create services, tools, and developer documentation.
  • Create testing frameworks for robustness and fault-tolerance.
  • Participate in an on-call rotation for critical incidents.

Requirements

  • 5+ years of experience writing high-performance, well-tested, production-quality code.
  • Bachelor’s degree in computer science or equivalent industry experience.
  • Familiarity with the LLM inference ecosystem, including frameworks and engines (e.g., vLLM, SGLang, TRT).
  • Demonstrated experience building large-scale, fault-tolerant, distributed systems (storage, search, computation).
  • Expert-level programming in Python, Go, Rust, or C/C++.
  • Experience implementing runtime inference services at scale or similar.

Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Other competitive benefits

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.