Kubernetes Jobs
410 open roles mentioning Kubernetes
Senior Software Engineer, Developer Experience
Harvey
Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This role is on the Developer Experience team, which is crucial for maintaining Harvey's rapid growth by building CI/CD systems, internal frameworks, and platform-level guardrails. The team enables all Harvey engineers to ship ambitious AI features rapidly and safely. As a Software Engineer on this team, you will create frameworks and systems to maximize the velocity and efficiency of every engineer at Harvey, working across the entire tech stack. You will also have a unique opportunity to apply AI directly to developer productivity, shaping the future of software development within an AI-native company.
$175k - $250k
Software Engineer, Backend (New-York)
Mistral AI
Mistral AI is seeking passionate and skilled backend software engineers to join our new team in NYC. As a Backend Engineer, you will contribute to the development of core systems powering our AI platform, including AI Studio, Le Chat, and Mistral Code. Your work will shape how millions of users and developers interact with our AI platform at scale, focusing on building reliable, high-performance backend systems and APIs. This role directly impacts the developer and user experience, making it more engaging, efficient, and intuitive. We welcome engineers from mid-level to senior and staff levels who are eager to learn, collaborate, and ship impactful products.
Site Reliability Engineer, Inference Infrastructure
Cohere
Cohere is seeking a Site Reliability Engineer to join the Model Serving team. This role is crucial for developing, deploying, and operating the AI platform that delivers Cohere's large language models via API endpoints. You will work closely with various teams to deploy optimized NLP models into production environments, ensuring low latency, high throughput, and high availability. The position also offers the opportunity to interact with customers and create customized deployments to meet their specific needs, contributing to the widespread adoption of AI.
Staff Software Engineer, Inference Infrastructure
Cohere
Cohere is seeking Members of Technical Staff to join the Model Serving team. This role focuses on developing, deploying, and operating the AI platform that delivers Cohere's large language models via API endpoints. You will work closely with various teams to deploy optimized NLP models into production environments, ensuring low latency, high throughput, and high availability. The position also offers the opportunity to interact with customers and build customized deployments to meet their specific needs.
Forward Deployed Engineer - Systems
Modal
Modal is seeking an experienced Forward Deployed Engineer (FDE) to partner with our sales team and drive technical sales success. As an FDE, you will be the technical voice in our sales process, working directly with Account Executives to help enterprise customers understand how Modal can transform their AI/ML infrastructure. You will partner with Account Executives to identify, qualify, and close strategic enterprise opportunities, lead technical discovery sessions with prospective customers to understand their current infrastructure, pain points, and requirements, and design and present compelling technical solutions that demonstrate how Modal addresses customer needs. You will also architect migration paths from existing cloud infrastructure (AWS, GCP, Azure) to Modal's serverless platform, conduct technical demos, experiments, and proof-of-concepts that showcase Modal's capabilities, and navigate complex technical evaluations and address security, compliance, and integration concerns. Additionally, you will build trusted advisor relationships with technical decision-makers, collaborate with product and engineering teams to communicate customer feedback and influence product roadmap, and support contract negotiations by providing technical expertise.
Security Engineer, Cloud
Rogo.ai
Rogo is seeking a Staff Security Engineer to lead the design and implementation of cloud security architecture across AWS and GCP. This is a hands-on role for an engineer experienced in building and operating secure cloud platforms at scale, using code, systems design, and automation to solve security challenges. You will own the technical direction of cloud security, focusing on secure primitives, large-scale Terraform authoring, identity and network architecture, and embedding security into the core platform. The role requires a senior technical leader who is also tactical, writing production code, reviewing infrastructure changes, and providing pragmatic security solutions.
Forward Deployed Engineer - AI Engineer
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We are seeking a core member for our Applied AI team to lead Forward Deployed Engineering efforts with enterprise customers. This role involves translating advanced AI research into high-impact, real-world applications, owning the technical strategy and delivery of agentic systems from discovery to production launch.
Product Manager, Forge
Mistral AI
Mistral AI is seeking a talented and experienced Product Manager to define and execute the strategy for Forge, a product that empowers customers to build, fine-tune, and deploy custom AI models at scale. Forge transforms cutting-edge research into enterprise-ready capabilities by supporting model fine-tuning, reinforcement learning, and post-training workflows. This role operates at the intersection of research and product, enabling customers to train specialized models for real-world business value. You will collaborate closely with applied AI scientists and research engineers to translate frontier techniques into scalable and reliable solutions, shaping a 0-1 product with significant business impact and defining the future of how organizations train and deploy AI models.
Forward Deployed Engineer, Infrastructure Specialist (Europe)
Cohere
Cohere is a leading security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products. We are seeking engineers to join our team and contribute to the widespread adoption of AI. This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications, acting as a bridge between our core North product and client engineering teams. You will be at the forefront of solving complex problems and securely integrating AI into critical sectors like finance, healthcare, and telecommunications, working with esteemed clients and focusing on Agentic AI.
$20k - $40k
Software Engineer, New Grad
Mistral AI
Mistral AI is seeking early-career software engineers, including new graduates or those with up to 1-2 years of experience, to join their software engineering teams. In this role, you will contribute to building and enhancing the core systems that power Mistral's products, such as AI Studio and Applications, as well as supporting operations like SRE, data, and security. You will play a key part in shaping how users and developers interact with the AI platform at scale, working closely with experienced engineers who will mentor and support your growth. This is an opportunity for enthusiastic graduates to learn, collaborate, and deliver impactful products in a fast-paced environment.
Senior Software Engineer, Site Reliability Engineer
Harvey
Harvey is transforming legal and professional services by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. We are a fast-scaling company with strong product-market fit, seeking ambitious individuals to help build a generational company. Our team operates with intensity, ownership, and a commitment to our mission, valuing decisiveness, simplicity, and continuous improvement. As a Software Engineer on the Site Reliability team, you will ensure the reliability, scalability, and performance of our legal AI platform, owning the systems that keep our platform fast, secure, and always on. Your work will be crucial in maintaining platform resilience as we grow across 50+ regions.
$200k - $260k
Staff Software Engineer, Site Reliability Engineer
Harvey
Why Harvey At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come. This is a rare chance to help build a generational company at a true inflection point. With 1500+ customers in 60+ countries, strong product-market fit, and world-class investor support, we’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth — personal, professional, and financial — is unmatched. Our team moves fast, takes ownership, and is deeply committed to the mission — operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job's Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we're never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we'd love to build with you. At Harvey, the future of professional services is being written today — and we’re just getting started. Role Overview As a Staff Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability, and performance of our legal AI platform. You’ll join a high-leverage team that sits at the intersection of infrastructure and product, owning the systems that keep our platform fast, secure, and always on. From scaling across 50+ regions to automating mission-critical operations, your work will ensure that Harvey remains resilient as we grow. If you’re passionate about building robust systems and reducing complexity through automation, we’d love to work with you. This role is based in San Francisco, CA. We use an in-person work model and offer relocation assistance to new employees. What You’ll Do - Design, implement, and manage monitoring, alerting, and infrastructure resources (compute, storage, networking) across 50+ global regions - Lead incident management processes, including postmortems, root cause analyses, and driving actionable improvements - Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention - Establish best practices for security, compliance, and reliability and collaborate across teams to drive these principles throughout the software lifecycle - Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality - Provide technical mentorship and leadership, promoting best practices and fostering team growth What You Have - 10+ years of experience in Site Reliability Engineering or similar roles supporting production environments, with proven ability to mentor and guide technical teams - Expertise in infrastructure as code(IaC) tools (Pulumi, Terraform, CloudFormation, etc.) - Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident response practices (PagerDuty, IncidentIO, etc.) - Proficiency with cloud infrastructure platforms (Azure, GCP, AWS, etc.) - Strong programming skills (Python, Bash, Go, or similar languages) - Proven track record of diagnosing complex system problems and implementing durable solutions - Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles - Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence Compensation Range $238,000 - $290,000 USD Depending on your location, an Applicant Privacy Notice may apply to you. You can find all of our Applicant Privacy Notices [here]. #LI-AN2 Harvey is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made by emailing [email protected]
$238k - $290k
Senior ML Systems Engineer, Frameworks & Tooling
Cohere
Cohere is seeking a Senior ML Systems Engineer to join their team and build, maintain, and evolve the training framework that powers their frontier-scale language models. This role is ideal for someone passionate about large-scale training, distributed systems, and HPC infrastructure, offering the opportunity to design and maintain core components for fast, reliable, and scalable model training. You will also build tooling to connect research ideas to thousands of GPUs, working across the full stack of ML systems with significant autonomy and impact.
Software Engineer, Fleet Management
OpenAI
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience.
Member of Technical Staff, LLM Infrastructure
fireworks ai
As a Software Engineer on the AI Infrastructure team, you will help design the core systems that power Fireworks AI’s generative AI platform. You will build infrastructure and tools that ensure the reliability, performance, quality, and availability of our AI system. Your mission is to make Fireworks AI the most reliable and user-friendly generative AI platform in the world. You will partner closely with our cloud infrastructure, product, and performance teams to deliver infrastructure that bridges the gap between our customers and the ultra-performant proprietary Fireworks inference engine.
Senior Backend Engineer
Legora
Legora is redefining how legal work gets done with an AI-native workspace that helps legal professionals move faster, think more clearly, and operate with sharper precision. We analyze thousands of documents in minutes and power end-to-end workflows, cutting through complexity so teams can focus on judgment, strategy, and outcomes. As a Senior Backend Engineer, you will build, review, and ship amazing products that assist legal professionals globally. You will have the opportunity to work across the entire tech stack with clear ownership and autonomy.
Full Stack Software Engineer, Gov
OpenAI
We are seeking Software Engineers to join a nimble team driving the deployment of OpenAI’s technology into new environments and infrastructure that power critical missions in the public sector. You will work cross-functionally with product, security, and compliance teams to build the functionality needed to deliver a scalable, reliable platform. You’ll also partner directly with customers to design and build new products and features that create real-world impact. This role offers both breadth and technical depth, giving you the opportunity to shape the future of OpenAI’s technology where it matters most.
Software Engineer - Model Products
Baseten
Baseten is seeking a Software Engineer to join their Model Performance team, focusing on the infrastructure that powers hosted API endpoints for cutting-edge open-source models. This role involves working on distributed systems, model serving, and developer experience to ensure models running on the Baseten platform are fast, reliable, and cost-efficient. You will contribute to defining how developers interact with AI models at scale, joining a high-impact team at the intersection of product, model performance, and infrastructure.
Applied AI Engineer, Prototyping
Mistral AI
Mistral AI is seeking an Applied AI Engineer to join their Proto Team, which acts as the technical pre-sales arm of the GTM organization. This role operates at the intersection of product and customer, translating Mistral's AI technology into real-world solutions that deliver value quickly. The engineer will build high-impact, full-stack AI solutions for global customers on short timelines, owning the end-to-end execution from scoping to deployment. This position also involves collaborating with GTM, product, and engineering teams, contributing to internal tools and product improvements, and testing new capabilities to inform product direction. The ideal candidate is a fast-moving, highly curious builder who thrives in complexity and ambiguity, possesses strong problem-solving skills, and has a hacker mindset with deep engineering instincts.
AI Scientist - Warsaw
Mistral AI
Mistral AI is a pioneering company dedicated to democratizing AI through high-performance, optimized, open-source models, products, and solutions. We aim to simplify tasks, save time, and enhance learning and creativity by integrating cutting-edge AI into daily working life. Our comprehensive platform serves both enterprise and personal needs, featuring offerings like Le Chat, La Plateforme, Mistral Code, and Mistral Compute. We are a dynamic, collaborative, and diverse team passionate about AI's potential to transform society, driving innovation from our distributed teams across France, USA, UK, Germany, and Singapore. Join us to shape the future of AI and make a meaningful impact.