Engineering Manager, Cloud Infrastructure
Remote • Foster City, CA • FullTime
Posted 16h ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Replit is seeking a hands-on Engineering Manager to lead their Cloud Infrastructure team. This role focuses on the shared infrastructure as code, networking, storage, compute, and service mesh platforms that are critical for Replit's product and platform teams. You will lead and grow an existing engineering team responsible for building and operating these foundational elements, including Kubernetes, shared edge networking, service mesh, and workload identity. This is a platform-building role with production accountability, requiring comfort with deep technical dives and incident management, while fostering a team that can operate autonomously.
Responsibilities
- Own the cloud platform roadmap, translating product, platform, reliability, and security needs into actionable outcomes.
- Lead the team in building and maintaining IaC interfaces for services, cells, connectivity, identities, and shared resources to make infrastructure repeatable and self-service.
- Ensure the availability, upgrades, isolation, recovery, and incident remediation of the platforms built by the team.
- Maintain clear SLOs, sustainable on-call coverage, and ownership for critical systems.
- Stay technically engaged by reviewing designs, debugging failures, and using AI tools for prototyping and automation.
- Apply rigorous review and verification to AI-generated infrastructure changes.
- Build and grow a high-ownership engineering team through coaching, performance management, and strategic hiring.
Requirements
- Demonstrated engineering management experience, including leading and developing engineers, making prioritization and performance decisions, hiring, and delivering through a team.
- Software-oriented infrastructure depth, with experience building and operating cloud platforms or distributed systems.
- Proficiency in reasoning across infrastructure code, Kubernetes, networking, service identity, and stateful dependencies.
- Experience with safe-change practices and production judgment, including owning migrations and incidents.
- Ability to explain failure modes and rollback limits, and to simplify systems when necessary.
- Understanding of platform-product dynamics and engineering judgment, including creating adopted interfaces and balancing tradeoffs.
Benefits
- Competitive Salary & Equity
- 401(k) Program with a 4% match (US Only)
- Health, Dental, Vision and Life Insurance
- Short Term and Long Term Disability
- Paid Parental, Medical, Caregiver Leave
- Flexible Time Off (FTO) + Holidays
- Commuter Benefits (In-Office & US Only)
- Monthly Wellness Stipend
- Autonomous Work Environment
- In Office Set-Up Reimbursement (In-Office Only)
- Quarterly Team Gatherings
- In Office Amenities (In-Office Only)