Software Engineer, Host Assurance
Remote • San Francisco • FullTime
Posted 9d ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability.
Responsibilities
- Design, build, and operate components of the Host Assurance platform that establish trust in bare-metal & VM hosts before they are eligible for production use.
- Ensure hosts are verifiably trustworthy from delivery and installation through secure bootstrap and readiness to join orchestration systems.
- Build and improve systems such as machine identity, certificate issuance and enrollment, HSM-backed or key-management-backed trust services, host attestation, measurement, and baseline verification tooling.
- Validate delivered hardware and firmware against vendor claims and continuously detect and manage drift over time.
- Eliminate insecure bootstrap patterns while preserving deployment throughput and operational reliability.
- Own software components through implementation, operational readiness, launch, and ongoing operation, including defining readiness criteria, validating behavior under load and failure, establishing monitoring and actionable alerts, planning staged rollouts and tested recovery procedures, and maintaining runbooks.
- Participate in production support via on-call rotation.
- Help define observable, testable security properties for host platforms and improve the telemetry and validation needed to enforce them in practice.
- Work across different deployment models and provider boundaries while maintaining a consistent bar for host trust outcomes.
Requirements
- Strong software engineering experience building and operating reliable production systems at scale.
- Depth in distributed systems, backend or platform engineering, operating systems, or infrastructure security.
- Experience with PKI, HSMs, machine identity, attestation, secure boot, or hardware security is valuable but not required.
- Ability to reason across services, APIs, and host software, and learn boot, firmware, and hardware trust mechanisms.
- Apply sound judgment to trust boundaries, correctness, and safe failure behavior.
- Experience writing production quality code and ability to reason clearly about failure modes, operational safety, and long-term maintainability.
- Experience supporting critical production systems at scale, such as being on-call.
- Balance rigor with pragmatism, and care about making strong security controls deployable in real-world environments.
- Self-directed, low ego, and willing to work across disciplines to solve the most important problems.
- Enjoy building in ambiguous spaces where the architecture is still emerging, stakes are at all time high, and the future is being built.