Infrastructure engineer
New York City, NY • FullTime
Posted 13h ago
About the job
At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.
Responsibilities
- Automate operational tasks and infrastructure management with Python or Go.
- Design scalable, fault-tolerant infrastructure across AWS, GCP, and Azure, working fluently across Kubernetes, Helm, and Terraform.
- Run agents in your daily loop to investigate incidents, draft infrastructure changes, write runbooks, scaffold tooling, and review PRs.
- Build agentic setups where humans and digital teammates work as one team with shared skills and context.
- Encode recurring infra tasks as internal skills for teammates (human or agent).
- Lead incident response, post-mortems, and root-cause analyses.
- Own the reliability, performance, and efficiency of WRITER's core services end-to-end.
- Define and uphold SLOs and error budgets.
- Balance critical weekly work with 6-12 month platform direction.
- Provide expert guidance on system design for reliability, performance, and scalability.
- Connect the infra agenda to product and revenue context.
Requirements
- 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company.
- Experience running containerization in production with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred).
- Good proficiency in Python or Go for automation and tooling.
- AI is part of how you ship; agentic tooling is in your daily loop.
- Have built or adopted AI-assisted workflows others now use.
- Have strong opinions on where AI is unreliable.
- Demonstrated ability to challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems.
- Reason from constraints and failure modes.