Prometheus Jobs
3 open roles mentioning Prometheus
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AI
Together AI is seeking a Staff Engineer to design and deliver multi-petabyte storage systems optimized for large-scale AI training and inference workloads. You will architect high-performance parallel filesystems and object stores, integrate cutting-edge technologies, and drive significant cost optimization. The role involves building Kubernetes-native storage operators and self-service platforms for automated provisioning and multi-tenancy. You will focus on optimizing data paths, designing multi-tier caching architectures, and tuning parallel filesystems for AI applications. This is a research-driven role within a company focused on lowering the cost of modern AI systems through co-design of software, hardware, algorithms, and models.
$250k - $300k
Applied AI Engineer, Site Reliability Engineer - EMEA
Mistral AI
Mistral AI is seeking a founding engineer for its Applied AI Site Reliability Engineering (SRE) sub-team. This role is crucial for building and operating a framework that ensures the reliability and sustainability of Mistral's AI solutions across all customer accounts, whether hosted by Mistral or the customer. You will operate in four key modes: BUILD (designing for a fleet of platforms, proactive reliability, authoring runbooks, implementing observability), RUN (operating Tier-1 customer environments, ensuring SLO compliance, managing incidents), ENABLE (productizing deployment, security, and scaling of Applied AI solutions), and SECURE (owning security operations, leading CVE response, and implementing supply-chain integrity controls). This is a framework-first, fleet management role focused on structurally solving problems for all customers, not just individual ones. The team values people and outputs, direct feedback, low ego, and high standards in a fast-paced, unstructured environment.
Staff Infrastructure Software Engineer, Enterprise AI
Scale AI
Scale GP is seeking a Senior or Staff Infrastructure Engineer to lead the engineering of the 'paved road' for knowledge retrieval and inference engines, defining deployment standards for Agentic workflows at scale. This role bridges complex AI orchestration with world-class infrastructure, ensuring platform reliability for enterprise agents. The ideal candidate is passionate about deep technical work, mentoring, and setting long-term technical strategy while maintaining a hands-on delivery focus. You will architect and implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in regulated industries.
$252k - $315k