Software Engineer - Training/Inference (C++)
$180k - $440k • Palo Alto, CA
Posted 5d ago
Remote Work Policy
On-site
Categories
Applied AI Engineer
About the job
SpaceXAI is building a high-performance inference platform that serves Grok to millions of users daily with exceptional speed and reliability. As a Member of Technical Staff - Inference, you will be responsible for designing and optimizing large-scale model serving systems from end-to-end. This includes distributed infrastructure components like global KV cache, continuous batching, and load balancing, as well as deep low-level optimizations such as GPU kernels, quantization, and speculative decoding. This is a high-impact role where your contributions will directly influence the speed and reliability of user interactions with Grok at a massive scale.
Responsibilities
- Architect and implement scalable distributed infrastructure for model serving, including load balancing, auto-scaling, batch scheduling, and global KV cache.
- Optimize latency and throughput of model inference under real production workloads.
- Build reliable, high-concurrency serving systems with 100% uptime, 0% error rate, and excellent tail latency.
- Benchmark, fine-tune, and accelerate inference engines, including low-level GPU kernel work and code generation.
- Develop custom tools for tracing, replaying, and fixing issues across the full stack, from orchestration to GPU kernels.
- Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates.
- Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.
Requirements
- Deep low-level systems programming experience in C/C++ or Rust.
- Experience with large-scale, high-concurrency production serving.
- Experience with GPU inference engines (e.g., vLLM, SGLang, Triton, TensorRT-LLM).
- Strong background in system optimizations such as batching, caching, load balancing, and parallelism.
- Proficiency in low-level inference optimizations, including GPU kernels and code generation.
- Experience with algorithmic inference optimizations like quantization, speculative decoding, distillation, and low-precision numerics.
- Experience with testing, benchmarking, and ensuring the reliability of inference services.
- Experience designing and implementing CI/CD infrastructure for inference.
Benefits
- Equity
- Comprehensive medical, vision, and dental coverage
- Access to a 401(k) retirement plan
- Short & long-term disability insurance
- Life insurance
- Various other discounts and perks