Principal AI Ops Architect, GPS
Doha, Qatar; London, UK
Posted 2mo ago
Job Location
Doha, Qatar; London, UK
Tech Stack
Remote Work Policy
On-site
Categories
AI Infrastructure Engineer
About the job
Scale's Global Public Sector team is dedicated to leveraging AI to tackle significant challenges within the public sector worldwide. This involves creating custom AI applications impacting millions, generating high-quality training data for national LLMs, and providing AI upskilling and advisory services. As a Principal AI Ops Architect, you will be instrumental in designing and developing the production lifecycle for full-stack AI applications. Your role will encompass ensuring end-to-end system reliability, real-time inference observability, sovereign data orchestration, secure software integration, and the resilient cloud infrastructure necessary for international government partners. At Scale, we empower the public sector to transform operations and enhance citizen services through advanced technology, and we are looking for individuals ready to shape the future of AI in this domain.
Responsibilities
- Own the production outcome and long-term performance/reliability of deployed AI use cases.
- Oversee the end-to-end health of the platform, ensuring seamless integration of AI core with all full-stack components.
- Build automated systems to monitor model performance and data drift across geographically dispersed environments.
- Manage the technical lifecycle within diverse international regulatory frameworks.
- Lead incident response for production issues in mission-critical environments and implement preventative measures.
- Translate deep technical performance metrics into clear insights for senior government officials.
- Collaborate with Engineering and ML teams to ensure field learnings influence future technical architecture and decisions.
Requirements
- 6+ years in a high-impact technical role (SRE, FDE, or MLOps) with public sector experience.
- Familiarity with international government security standards and sovereign AI deployment complexities.
- Proven experience maintaining production-grade applications with a deep understanding of the full request lifecycle.
- Proficiency in coding and modern AI infrastructure, including Kubernetes, vector databases, agentic development, and LLM observability tools.
- Demonstrated ownership of production deployments and proactive problem-solving.
- Understanding of the critical importance of reliability in public sector AI applications.
- Ability to explain system performance degradation and resolution strategies to high-ranking officials.