Member of Technical Staff, ML Platform
Remote • Remote • FullTime
Posted 9h ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Runway is seeking an ML infrastructure engineer to take ownership of model evaluation processes end-to-end. This role is crucial as every decision regarding model training, user deployment, and promising architectures relies on robust evaluation. You will be responsible for designing and building systems that generate samples at scale, score them using automated metrics and human annotations, track results, and provide insights to researchers quickly. You will define operational metrics, confidence levels, and decision-making criteria for model deployment. Embedded within research teams as part of the ML Platform, you will set the technical direction for evaluation across the company, significantly impacting the quality of Runway's models.
Responsibilities
- Own the evaluation platform end-to-end, including tooling and systems for generation, annotation, review, and adaptation at scale.
- Define CLIs, APIs, GUIs, and storage layers to ensure a seamless, fast, sophisticated, and collaborative evaluation process.
- Collaborate with research teams in video, image, audio, agents, and robotics to understand measurement needs and build generalized platform solutions.
- Establish company-wide standards for model evaluation, covering reproducibility, metric definitions, reporting, and decision-making criteria.
- Support the broad adoption and integration of the evaluation platform across Runway's research efforts, including training and production model serving.
- Contribute to the ML Platform team's tools and systems for training and serving frontier models.
Requirements
- 5+ years of experience building ML infrastructure or data platforms in production environments, with a focus on evaluation, experimentation, or benchmarking systems.
- Strong proficiency in Python and PyTorch, with hands-on experience running large batch GPU workloads on Kubernetes.
- Experience designing data pipelines and storage for large volumes of media or model outputs, emphasizing versioning and reproducibility.
- Familiarity with experimental statistics, including paired comparisons, confidence intervals, multiple-comparison pitfalls, and inter-rater agreement.
- Comfort building internal tools end-to-end, from command line to browser interfaces.
- Ability to lead a broad technical area by gathering requirements, setting direction, making tradeoffs, and driving a roadmap.
- Familiarity with the full model development lifecycle: data, training, evaluation, and serving.
- Self-starter capable of working embedded with research teams and moving quickly.
- Strong systems thinking and a pragmatic approach to production reliability.
- Humility and open-mindedness, with a willingness to learn from others.