Software Engineer - Training/Inference (C++)

$180k - $440k Palo Alto, CA

Posted 5d ago

Job Location

Palo Alto, CA

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

SpaceXAI is building a high-performance inference platform that serves Grok to millions of users daily with exceptional speed and reliability. As a Member of Technical Staff - Inference, you will be responsible for designing and optimizing large-scale model serving systems from end-to-end. This includes distributed infrastructure components like global KV cache, continuous batching, and load balancing, as well as deep low-level optimizations such as GPU kernels, quantization, and speculative decoding. This is a high-impact role where your contributions will directly influence the speed and reliability of user interactions with Grok at a massive scale.

Responsibilities

  • Architect and implement scalable distributed infrastructure for model serving, including load balancing, auto-scaling, batch scheduling, and global KV cache.
  • Optimize latency and throughput of model inference under real production workloads.
  • Build reliable, high-concurrency serving systems with 100% uptime, 0% error rate, and excellent tail latency.
  • Benchmark, fine-tune, and accelerate inference engines, including low-level GPU kernel work and code generation.
  • Develop custom tools for tracing, replaying, and fixing issues across the full stack, from orchestration to GPU kernels.
  • Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates.
  • Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.

Requirements

  • Deep low-level systems programming experience in C/C++ or Rust.
  • Experience with large-scale, high-concurrency production serving.
  • Experience with GPU inference engines (e.g., vLLM, SGLang, Triton, TensorRT-LLM).
  • Strong background in system optimizations such as batching, caching, load balancing, and parallelism.
  • Proficiency in low-level inference optimizations, including GPU kernels and code generation.
  • Experience with algorithmic inference optimizations like quantization, speculative decoding, distillation, and low-precision numerics.
  • Experience with testing, benchmarking, and ensuring the reliability of inference services.
  • Experience designing and implementing CI/CD infrastructure for inference.

Benefits

  • Equity
  • Comprehensive medical, vision, and dental coverage
  • Access to a 401(k) retirement plan
  • Short & long-term disability insurance
  • Life insurance
  • Various other discounts and perks

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.