AI Performance Engineer
$60k - $300k • Sunnyvale • FullTime
Posted 1mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Applied Intuition is seeking a performance engineer to specialize in making large-scale machine learning workloads fast and cost-efficient within the datacenter. This role focuses on optimizing distributed training runs across multiple nodes and high-throughput batch inference for processing vast amounts of real-world autonomy logs. The primary goal is to improve throughput, cluster efficiency, and reduce cost per unit of data processed, directly impacting the company's iteration speed. You will be responsible for identifying and resolving performance bottlenecks across the entire stack, from accelerators to ML frameworks and data infrastructure, working collaboratively with various engineering teams to achieve significant improvements in training time and processing costs.
Responsibilities
- Profile and optimize distributed training end-to-end, including data loading, preprocessing, augmentation, kernel execution, gradient communication, and checkpointing.
- Optimize large-scale offline and batch inference over petabyte-scale sensor logs, focusing on batching, scheduling, quantization, graph optimization, and accelerator saturation.
- Establish performance models for workloads, quantify performance gaps, and prioritize optimization opportunities.
- Improve multi-node scaling efficiency by addressing sharding, parallelism, collective communication, interconnect utilization, and memory/kernel bottlenecks.
- Drive cluster throughput by reducing GPU idle time caused by input pipeline stalls, I/O issues, scheduling gaps, stragglers, and failure recovery.
- Develop benchmarking, observability, and regression-detection tooling to maintain performance integrity.
- Collaborate with cross-functional engineers to solve complex data and compute problems at scale.
- Contribute to a team culture that values collaboration, technical excellence, and innovation.
Requirements
- Hands-on ML performance engineering experience, including profiling, roofline analysis, throughput optimization, and root-cause investigation.
- Experience with distributed multi-node training at scale (e.g., FSDP, DeepSpeed, Megatron, NCCL) and diagnosing scaling inefficiencies.
- Deep familiarity with GPU or accelerator performance concepts such as memory bandwidth, kernel launch overhead, occupancy, quantization, and collective communication.
- Experience with high-throughput or batch inference systems (e.g., NVIDIA Triton, TensorRT, ONNX Runtime, Ray).
- Fluency in Python and proficiency in C++ or another systems language.
- Excellent debugging, analytical, and problem-solving skills.
- Deep understanding of machine learning foundations and the ability to develop novel technical solutions.