Staff Infrastructure Engineer
Remote • Foster City, CA • FullTime
Posted 2mo ago
About the job
Replit is seeking a Staff Infrastructure Engineer to join their Infrastructure Engineering team. This role focuses on ensuring the reliability, scalability, and performance of Replit's platform, which serves millions of developers globally. The engineer will bridge development and operations, implementing automation and best practices to enable efficient scaling and high availability. The position involves proactively identifying and resolving reliability issues, designing robust monitoring solutions, automating operational tasks, and mentoring the engineering team on reliability principles.
Responsibilities
- Architect, build, and improve automation to eliminate toil and operational work.
- Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi.
- Create self-healing systems that can automatically respond to common failure scenarios.
- Collaborate with core infrastructure and product teams to performance tune and optimize cloud deployments (Kubernetes, Docker, GCP).
- Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency.
- Design and implement improvements to build, test, and deployment systems to enhance software delivery.
- Partner with service owners to understand pain points and implement build/test/deploy enhancements.
- Create and maintain centralized tooling and automation for the engineering lifecycle.
- Debug difficult technical problems to make systems and products more robust and operable.
- Review feature and system designs, ensuring security, scale, and operational integrity.
- Educate and mentor the engineering team to improve system reliability.
- Write high-quality, well-tested code, including building pipelines for third-party integrations.
Requirements
- 8-10 years of experience in Infrastructure Engineering, DevOps, Systems Engineering, or Site Reliability Engineering.
- Strong programming skills in languages like Python or Go.
- Ability to write high-quality, well-tested code.
- Deep understanding of distributed systems, including designing, building, scaling, and maintaining production services and service-oriented architectures.
- Experience with container orchestration platforms (Kubernetes) and cloud-native technologies.
- Proven track record of implementing and maintaining monitoring/observability solutions.
- Strong debugging and performance tuning skills.
- Strong incident management skills with experience leading incident response and critical thinking under pressure.
- Experience with infrastructure as code (e.g., Terraform) and configuration management tools.
- Excellent written and verbal communication skills, with the ability to explain technical concepts clearly.
- Strong interpersonal skills, with experience working with engineers of all levels.
- Willingness to debug and improve any layer of the stack.
- Passion for making software creation accessible.
Benefits
- Competitive Salary & Equity
- 401(k) Program with a 4% match (US Only)
- Health, Dental, Vision and Life Insurance
- Short Term and Long Term Disability
- Paid Parental, Medical, Caregiver Leave
- Flexible Time Off (FTO) + Holidays
- Commuter Benefits (In-Office Only)
- Monthly Wellness Stipend
- Autonomous Work Environment
- In Office Set-Up Reimbursement (In-Office Only)
- Quarterly Team Gatherings
- In Office Amenities (In-Office Only)