Software Engineer, AI Infra
Palo Alto HQ • FullTime
Posted 1mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
We are seeking a Staff/Lead Software Engineer, AI Infrastructure, to play a critical role in building and scaling the core infrastructure that powers Pika’s AI capabilities. In this position, you will lead the design and implementation of GPU infrastructure, AI model serving APIs, and general AI infrastructure execution, enabling cutting-edge machine learning features that drive our products. You will be responsible for architecting robust, distributed systems optimized for high-performance AI workloads, large-scale GPU orchestration, and low-latency, reliable API serving. Your work will directly impact the way users experience and interact with generative AI at scale. As a senior technical leader, you’ll also mentor engineers, drive best practices, and set the technical vision for AI infrastructure at Pika.
Responsibilities
- Design, develop, and maintain scalable GPU infrastructure for training and serving AI models
- Architect and optimize high-throughput, low-latency APIs for AI model serving and inference
- Lead the orchestration, scheduling, and efficient utilization of GPU resources across clusters
- Build and support robust systems for model deployment, monitoring, scaling, and reliability
- Collaborate with ML, backend, and platform engineering teams to deliver AI-powered product features
- Drive technical direction, code reviews, and mentorship across the AI Infrastructure team
Requirements
- 5+ years of experience as a software engineer working on systems infrastructure
- Hands-on experience with ML serving and GPU orchestration
- Deep knowledge of distributed systems, Kubernetes, and cloud-native infrastructure (AWS/GCP/Azure)
- Proven expertise in building and optimizing APIs for large-scale AI model serving (TensorFlow Serving, Triton, TorchServe, or similar)
- Familiarity with challenges of high-throughput, scalable GPU fleet management, scheduling, and efficient model execution
- Proficiency in backend languages such as Python, Go, or C++
- Experience optimizing for performance and reliability
- Ownership mentality and the drive to solve complex problems independently
- Excellent communication, collaborative, and mentorship skills
Benefits
- Competitive salary in the AI industry
- Equity in a rapidly growing team
- Comprehensive health benefits
- Monthly stipends
- Company retreats