Software Engineer, Productivity - Inference Runtime
San Francisco • FullTime
Posted 4mo ago
About the job
We are seeking a highly autonomous and ownership-driven engineer to join our Developer Productivity team, focusing on OpenAI's Inference Runtime. This role is crucial for scaling the engineering systems, safeguards, and developer workflows that enable our teams to innovate rapidly without compromising reliability or performance. You will operate at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. Your work will involve developing tooling and operational foundations for model launches, inference optimizations, cloud integrations, and large-scale deployments within our rapidly evolving inference stack. A primary focus will be enhancing the tooling and infrastructure for deploy gates of inference engine images, ensuring correctness, numerical soundness, regression-free performance, and optimized metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You will also contribute to hardening pre-production issue detection systems, reducing noise from flaky tests, and improving automation for triage, debugging, and escalation, alongside enhancing observability, rollout safety, release automation, and developer self-service tooling.
Responsibilities
- Improve systems ensuring inference engine releases are correct, performant, and regression-free by evolving tooling and infrastructure for deploy gate validation.
- Bring rigor to release, validation, branching, and deployment processes across the inference stack.
- Improve canary, async, and large-scale validation workflows for inference systems.
- Harden CI, testing, and validation infrastructure to make failures actionable and trustworthy.
- Reduce noisy or flaky failures caused by infrastructure instability, GPU scheduling, or test environment issues.
- Build automation for failure triage, ownership detection, debugging, and escalation.
- Partner closely with inference teams, research developer productivity, engine acceleration, and infrastructure teams to improve release quality and rollout safety.
- Reduce developer friction in testing, debugging, and release workflows to enable engineers to move faster with confidence.
Requirements
- Strong experience with CI/CD systems, testing infrastructure, release tooling, developer productivity, or large-scale build and validation systems.
- Excited by high-impact infrastructure where small regressions in correctness, latency, or reliability meaningfully affect production systems.
- Care about building systems engineers can trust.
- Strong developer empathy and enjoyment in improving workflows, reducing friction, and making engineers more effective.
- Demonstrate high ownership, proactively identify problems, drive improvements, and follow issues through resolution.
- Comfortable working in Python-heavy environments and debugging complex distributed systems.
- Enjoy building automation that reduces manual triage, improves signal quality, and scales operational effectiveness.
- Comfortable operating in ambiguous areas without a fully predefined roadmap.
- Enjoy partnering closely with engineers to understand workflows, pain points, and operational challenges.
- Pragmatic, collaborative, and motivated by helping teams move faster with more confidence.
- Excited to learn about large-scale inference systems.
- Python experience is highly relevant.
- C++ experience is helpful, but not required.
- Prior inference experience is not required.
- Strong instincts around developer productivity, testing, release engineering, and automation.
- Technically curious, comfortable navigating ambiguous, cross-functional operational problems.
- Motivated to improve the reliability, safety, and developer experience of large-scale production infrastructure.