Senior Support Engineer - Toronto

Remote Ontario - Remote FullTime

Posted 1mo ago

Job Location

Ontario - Remote

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

OpenAI is seeking a Senior Support Engineer to join its Technical Support team in Toronto. This role involves collaborating directly with strategic enterprise accounts and product teams to solve complex customer issues and provide technical guidance. You will be a key technical troubleshooting expert for OpenAI's API platform, acting as the final point of escalation before the core Engineering team. The position requires designing and running operational processes for monitoring top customers and a 24x7 response team, working closely with Infrastructure and Engineering teams to ensure customer success at scale. This is a low-volume, high-difficulty role focused on supporting innovative AI solutions built on the OpenAI API platform.

Responsibilities

  • Serve as a primary technical and troubleshooting expert for OpenAI's API platform.
  • Proactively identify and implement opportunities to scale support operations using automation and AI.
  • Configure and utilize advanced monitoring and alerting workflows to detect customer-impacting issues in real-time.
  • Collaborate with engineering on reliability reviews and preparedness for new features and launches.
  • Ensure operational readiness, including monitoring, alerting, and fallback plans, for all changes.
  • Design and refine incident response processes and documentation across teams.
  • Analyze operational metrics and incident root cause analyses to identify and implement improvements.
  • Provide support coverage during holidays and weekends as needed.

Requirements

  • Bachelor's degree in Computer Science or a related field, with a strong software engineering foundation.
  • 8+ years of experience in technical operations roles (e.g., SRE/NOC), including designing monitoring systems and resolving production issues in fast-paced, mission-critical environments.
  • Proven track record of troubleshooting complex technical problems at the systems level.
  • Deep familiarity with modern monitoring, alerting, and observability practices.
  • Hands-on experience setting up or managing metrics, logging, and tracing for distributed systems.
  • Proven experience leading incident response for high-severity outages or service disruptions.
  • Strong skills in scripting or software engineering (e.g., Python) for automation and tool integration.
  • Solid understanding of cloud infrastructure and distributed systems fundamentals.
  • Comfort working with cloud services, load balancers, databases, and containerized applications.
  • Effective cross-functional collaboration skills in a high-trust environment.
  • Strong communication skills to explain technical issues to both technical and non-technical stakeholders.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.