Senior Support Engineer - San Francisco

San Francisco FullTime

Posted 6mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

OpenAI is seeking a Senior Support Engineer to join our Technical Support team in San Francisco. This role involves collaborating directly with strategic enterprise accounts and product teams to solve complex customer challenges. You will be a key technical troubleshooting expert, providing guidance and resolving difficult issues for our customers and internal engineering teams. The position focuses on designing and running operational processes to monitor top strategic customers and a 24x7 response team, working closely with Infrastructure and Engineering teams to ensure a seamless customer experience at scale. This is a low-volume, high-difficulty role focused on the success of innovative, disruptive, and high-scale AI solutions built on the OpenAI API platform.

Responsibilities

  • Serve as a primary technical and troubleshooting expert for OpenAI's API platform, acting as the final escalation point before the core Engineering team.
  • Identify and implement opportunities to scale support operations using automation and AI technologies.
  • Configure and utilize advanced monitoring and alerting systems to detect customer-impacting issues in real-time.
  • Partner with engineering to conduct reliability reviews and ensure operational readiness for new features, launches, and customer requirements.
  • Design and refine incident response processes and documentation for strategic customers, engineering, and support teams.
  • Analyze operational metrics and incident root cause analyses to identify and implement improvements in monitoring, alerting, and support workflows.
  • Provide support coverage during holidays and weekends as needed.

Requirements

  • Bachelor's degree in Computer Science or a related field, with a strong software engineering foundation.
  • 8+ years of experience in technical operations roles (e.g., SRE/NOC), including designing monitoring systems and resolving production issues in mission-critical environments.
  • Proven track record of troubleshooting complex technical problems at the systems level.
  • Deep familiarity with modern monitoring, alerting, and observability practices, including hands-on experience with metrics, logging, and tracing for distributed systems (SLIs/SLOs, alert tuning, dashboard creation).
  • Experience leading incident response for high-severity outages, including real-time coordination, root cause analysis, and driving follow-ups.
  • Proficiency in scripting or software engineering (e.g., Python) for task automation and tool integration.
  • Solid understanding of cloud infrastructure and distributed systems fundamentals, including cloud services, load balancers, databases, and containerized applications.
  • Effective cross-functional collaboration skills in a high-trust environment.
  • Strong communication skills to explain technical issues to both technical and non-technical stakeholders.
  • Ability to coordinate efforts across teams and provide updates during ongoing incidents.

Benefits

  • Relocation assistance for new employees.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.