AI Infrastructure Engineer, pAGI
Remote • San Francisco • FullTime
Posted 11h ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
We are seeking an AI Systems Engineer to join our pAGI Infra team, responsible for building and operating the systems that power large-scale model training and evaluation. This role involves owning projects from identifying bottlenecks and designing solutions to deployment and operation, combining distributed systems engineering, performance optimization, and close collaboration with researchers. You will contribute to critical infrastructure projects, such as building shared grading services, improving resource allocation, or bringing new training stacks into production, directly impacting the speed and reliability of research advancements.
Responsibilities
- Build and operate infrastructure for large-scale training and evaluation, enhancing reliability, throughput, and resource efficiency.
- Develop shared inference and grading platforms with automated capacity management, health monitoring, and performance visibility.
- Improve compute scheduling and resource allocation to minimize idle GPU time and ensure quick workload recovery from failures.
- Diagnose bottlenecks across training, inference, and orchestration, collaborating with teams to boost end-to-end performance.
- Create self-service tools, automated validation, and observability features to assist researchers in launching experiments, diagnosing issues, and comparing results with reduced manual effort.
Requirements
- Strong software engineering fundamentals and experience building or operating large-scale distributed systems.
- Experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling.
- Highly self-motivated with the ability to take ownership of open-ended problems.
- Proficiency in debugging across system boundaries and using measurements to drive performance and reliability improvements.