Infrastructure engineer (UK)
Remote • London, UK • FullTime
Posted 2mo ago
About the job
At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.
Responsibilities
- Bring deep focus to one problem at a time, with the breadth to move between SRE, DevOps, Infrastructure, and Platform work.
- Automate operational tasks and infrastructure management with Python or Go.
- Design scalable, fault-tolerant infrastructure across AWS, GCP, and Azure, working fluently across Kubernetes, Helm, Terraform, and supporting cloud and AI tooling.
- Use AI agents in your daily workflow to investigate incidents, draft infrastructure changes, write runbooks, scaffold tooling, and review PRs.
- Build agentic setups where humans and digital teammates work as one team with shared skills, context, and on-call workflows.
- Encode recurring infra tasks as internal skills for teammates (human or agent).
- Lead incident response, post-mortems, and root-cause analyses, tracing failures to the underlying problem and applying learnings back into the architecture.
- Own the reliability, performance, and efficiency of WRITER's core services end-to-end, defining and upholding SLOs and error budgets.
- Carry the on-call pager and stand behind the outcome metric.
- Balance critical weekly work with the 6–12-month platform direction, shaping observability, cost, and reliability investments.
- Collaborate with product, security, and engineering peers, providing expert guidance on system design for reliability, performance, and scalability.
- Connect the infra agenda to product and revenue context.
Requirements
- 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company.
- Experience running containerization in production with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred).
- Good proficiency in Python or Go for automation and tooling.
- AI is part of how you ship; agentic tooling is in your daily loop, and you have strong opinions on where it's unreliable.
- Demonstrated ability to challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems.
- Reason from constraints and failure modes, naming tradeoffs in business terms.