Sr. Staff Software Engineer, Managed Platform Services
$250k - $300k • San Francisco, CA - US • FullTime
Posted 2mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Crusoe is seeking Sr. Staff Software Engineers to act as senior floating technical leaders within the Managed Platform Services (MAPS) organization. This role involves deploying expertise across various critical areas, including infrastructure scale-out, performance and reliability enhancements, new product exploration, customer-facing platform development, and cross-functional initiatives. The objective is to transform ambiguity into clarity, prototypes into products, and individual team solutions into organization-wide leverage, driving the company's mission to accelerate the abundance of energy and intelligence.
Responsibilities
- Design and drive automation for new on-prem server and site bring-up during rapid data center expansion.
- Conduct deep dives into system components to identify and resolve performance and reliability issues.
- Establish platform-wide benchmarks and architect for 10x scale with high availability and fault isolation.
- Build Proofs of Concept (POCs) to evaluate the viability of new platform capabilities, including AI platforms and managed databases.
- Analyze AI-native customer profiles to identify features and insights for the platform.
- Partner with customer success and solution engineering to define missing customer-facing functionality and tooling.
- Investigate and scope high-leverage architectural shifts, such as migrating managed services between orchestration platforms.
- Drive cross-organizational initiatives, like service isolation and air-gapped solutions.
- Utilize AI to augment development and build tools that multiply organizational output.
- Coach senior engineers and introduce systems to uplevel the team.
- Create an environment that encourages learning from failure and fosters collaboration, responsiveness, and excellence.
Requirements
- Hands-on expertise in designing and operating distributed systems at scale, including sharding, replication, consensus, load balancing, and concurrency.
- Comfort operating across the entire technology stack, from infrastructure bring-up to customer-facing features, and adapting to shifting priorities.
- Proven ability to rapidly build high-fidelity prototypes, explore multiple technical approaches, and determine when a solution is 'good enough'.
- Experience using AI tools to accelerate development, research, and decision-making, and a drive to build AI-powered tooling for others.
- Proficiency in benchmarking, profiling, and fixing performance issues.
- Customer-facing on-call experience, with a focus on preventing incidents.
- Ability to engage in ambiguous product spaces, ask pertinent questions, and translate requirements into actionable increments.
- Demonstrated ability to drive technical outcomes across organizational boundaries without formal authority.
- Exemplary communication skills, providing the right level of context and ensuring understanding.