Senior Software Engineer - Core Cloud Platform
Remote • San Francisco Office (Fremont St) • FullTime
Posted 13h ago
About the job
Lambda, The Superintelligence Cloud, is seeking a Senior or Staff Software Engineer to build and operate the distributed systems powering their GPU cloud. This role involves developing APIs, workflows, stateful controllers, schedulers, and operational tooling to transform large-scale GPU infrastructure into reliable, secure, customer-facing cloud services. You will be a full-cycle engineer, owning systems from design through deployment, on-call, and continuous improvement. This position is ideal for engineers passionate about cloud infrastructure, distributed systems, operational excellence, and solving complex problems across software and infrastructure domains. Senior engineers will lead complex work within a team, while Staff engineers will also shape cross-team architecture and enhance the effectiveness of other teams.
Responsibilities
- Design, build, and operate services, APIs, control planes, and platform capabilities for Lambda's AI cloud.
- Solve distributed systems problems related to state, consistency, concurrency, scheduling, failure recovery, and lifecycle management.
- Manage the full engineering lifecycle: problem framing, architecture, implementation, testing, rollout, observability, on-call, and continuous improvement.
- Enhance system availability, latency, throughput, efficiency, security, and operability.
- Implement durable engineering improvements based on incidents and near misses, including automation, testing, guardrails, and backstops.
- Collaborate with product, infrastructure, networking, storage, security, and SRE teams to resolve dependencies and achieve customer-centric outcomes.
- Utilize AI-assisted development tools judiciously, verifying correctness, security, and maintainability independently.
- Contribute to technical standards, design and code reviews, and mentorship; at Staff level, lead cross-team architecture and elevate the organization's technical capabilities.
Requirements
- 7+ years of professional software engineering experience or equivalent impact in building production systems.
- Proficiency in at least one general-purpose language (primarily Go and Python), with a strong understanding of concurrency, error handling, and testing.
- Experience designing, building, and operating backend services, distributed systems, infrastructure, or platform capabilities at scale.
- Practical understanding of system design, data models, APIs, failure modes, performance, and trade-offs for reliable production software.
- Proven track record of owning complex work through delivery and operation, including testing, staged rollout, monitoring, incident response, and root-cause improvement.
- Demonstrated ability to align cross-functional partners and build consensus on decisions and trade-offs.
Benefits
- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan