Member of Technical Staff (AI Inference Engineer)
San Francisco • FullTime
Posted 5mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
We are seeking an engineer to join our team responsible for building and running the inference engine behind Perplexity's queries. This role involves deploying dozens of model architectures at scale, managing tight latency and cost budgets, and working with a stack including Rust, Python, CUDA, and CuTe DSL. You will contribute to supporting new models, migrating GPU kernels, developing a Rust-native serving runtime, optimizing performance, and enhancing reliability and observability.
Responsibilities
- Support transformer-based retrieval, text-generation, and multimodal models in inference infrastructure.
- Migrate in-house CUDA kernels to NVIDIA's CuTe DSL for future hardware compatibility.
- Develop an internal Rust-based inference server to handle growing traffic and address Python limitations.
- Profile and resolve performance bottlenecks across the inference pipeline.
- Build dashboards, alerts, and automated remediation for production systems.
- Respond to and learn from production incidents to improve system reliability.
Requirements
- Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar).
- Understanding of modern LLM architectures and ability to deploy them reliably in production.
- Experience building and operating production distributed systems under load, especially performance-critical ones.
- Comfortable working across Rust, Python, and CUDA/CuteDSL.
- Ability to own problems end-to-end, from research to production debugging.
- Self-directed and capable of thriving in fast-moving environments.
- 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
- Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
- Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
- Understanding of common LLM architectures and inference optimization techniques.