Machine Learning, Platform Engineer

$160k - $250k Remote San Francisco

Posted 1mo ago

Remote Work Policy

Fully remote

Categories

Machine Learning Engineer

About the job

Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems. This role is part of a team dedicated to enabling custom models and dedicated inference on Together's platform. The team is responsible for building a container platform, optimizing autoscaling, minimizing cold starts, achieving the best end-to-end model performance, and providing a best-in-class developer experience with great tooling. The work often involves video or audio generation across the stack, including CUDA kernels, PyTorch optimization, inference engines, container orchestration, and queueing theory.

Responsibilities

  • Work on multi-cluster orchestration, portfolio optimization, predictive autoscaling, control panes, model bring-up, model optimization, APIs for managing deployments, inference worker SDKs, and CLI tools.
  • Analyze and improve the robustness and scalability of existing distributed systems, APIs, databases, and infrastructure.
  • Partner with product teams to understand functional requirements and deliver solutions that meet business needs.
  • Write clear, well-tested, and maintainable software and Infrastructure as Code (IaC) for both new and existing systems.
  • Conduct design and code reviews, create developer documentation, and develop testing strategies for robustness and fault tolerance.

Requirements

  • 5+ years of demonstrated experience in building large scale, fault tolerant, distributed systems.
  • Experience running serverless inference platforms, doing model bring-up on short notice, being on call, or running a cloud provider is a very big plus.
  • Good taste and ability to thoughtfully discuss how what you’ve built has failed over time.
  • Experience designing, analyzing and improving efficiency, scalability, and stability of various system resources.
  • Excellent understanding of low-level operating systems concepts including concurrency, networking and storage, performance and scale.
  • Expert-level programmer in one or more of Python, Golang, Rust, C++, or Haskell.
  • Proficiency in writing and maintaining Infrastructure as Code (IaC) using tools like Terraform.
  • Experience with Kubernetes internals or other container orchestration systems.
  • Sound judgment for when to use and when to not use LLMs for code.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.

Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Other competitive benefits

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.