Kubernetes Jobs
408 open roles mentioning Kubernetes
Senior Solution Engineer
Lambda
Lambda, a leader in AI cloud infrastructure, is seeking a Senior Solution Engineer to join their growing team. This role is crucial in enabling customers to achieve their business goals with AI infrastructure by partnering with leading AI researchers and enterprise engineering teams. The ideal candidate will design, scale, and optimize high-performance GPU cloud solutions, turning complex compute challenges into seamless, production-ready AI infrastructure. This position requires a customer-first mindset and technical mastery to drive growth and customer success.
Deployed Engineer (Houston)
Langchain
LangChain is seeking a Deployed Engineer to join their team, focusing on making intelligent agents ubiquitous. This role involves working on challenging applied AI problems, building production systems that real teams depend on, and directly shaping the future of AI agents in the real world. You will collaborate closely with customer engineering teams, guiding them through the entire lifecycle of AI agent development and deployment, from pre-sales evaluations to post-deployment advisory work. The feedback loop is rapid, the impact is tangible, and your contributions will be instrumental in how AI agents are adopted and operated globally.
$150k - $250k
Deployed Engineer (Dallas)
Langchain
LangChain is seeking a Deployed Engineer to join their team, focusing on building and running AI agents in production. This role involves partnering closely with customer engineering teams to transform prototypes into reliable, production-ready AI systems. You will work on challenging applied AI problems, contributing to the development of systems that real teams depend on, with a fast feedback loop and visible impact. The work directly shapes how AI agents are built and adopted in the real world, bridging the gap between engineering, product, and go-to-market strategies.
$200k - $250k
Deployed Engineer (Austin)
Langchain
LangChain is seeking a Deployed Engineer to join their team, focusing on making intelligent agents ubiquitous. This role involves working on challenging applied AI problems, building production systems that real teams depend on, and directly shaping the future of AI agents. You will collaborate closely with customer engineering teams, moving beyond prototypes to create reliable, production-ready AI agents. The feedback loop is rapid, the impact is tangible, and your contributions will significantly influence how AI agents are developed and deployed globally.
$150k - $250k
Developer Support Engineer (Singapore)
Braintrust
Braintrust is seeking Developer Support Engineers, both mid-level and senior, who are passionate about helping developers overcome technical challenges and achieve their goals. In this remote role based in Singapore, you will troubleshoot issues, identify workarounds, ship fixes, and document findings to accelerate developer progress. This position combines technical problem-solving, developer empathy, and close collaboration with Engineering, Solutions, and Product teams. If you excel at solving complex problems, articulating technical concepts clearly, and improving the developer experience, we encourage you to apply.
Member of Technical Staff, AI Training Infrastructure
fireworks ai
Fireworks is seeking a Training Infrastructure Engineer to design, build, and optimize the infrastructure that powers large-scale model training operations. This role is crucial for developing high-performance AI training infrastructure, requiring collaboration with AI researchers and engineers to create robust training pipelines, optimize distributed training workloads, and ensure reliable model development. The position offers the opportunity to solve hard problems at the forefront of AI infrastructure, build what's next with bleeding-edge technology, and have a direct impact on the future of AI within a fast-growing, passionate team.
Founding Developer Relations
Litellm
LiteLLM is seeking its first Developer Relations hire to foster the growth of its open-source AI gateway. This role is crucial for helping developers discover LiteLLM, understand its capabilities, and successfully deploy AI applications in production. You will be instrumental in building the developer journey, establishing how the company educates developers, engages with the community, supports product launches, and translates developer feedback into improved documentation, examples, and product decisions. This is a hands-on technical individual-contributor position where you will spend significant time building with LiteLLM and creating resources to empower other developers. You should be comfortable with coding, explaining technical concepts, and discussing developer challenges.
Senior Software Engineer, Production Engineering
Harvey
Harvey is seeking a Production Engineer to help build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth. The ideal candidate will have a systems-thinking mindset and a passion for building simple, reliable, and scalable systems.
$161k - $242k
Staff Software Engineer, Production Engineering
Harvey
Harvey is seeking a Production Engineer to build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.
$231k - $340k
Staff Software Engineer, Production Engineering
Harvey
Harvey is seeking a Production Engineer to build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.
$231k - $340k
Senior Software Engineer, Production Engineering
Harvey
Harvey is seeking a Production Engineer to help build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth. The ideal candidate will have a systems-thinking mindset and a passion for building simple, reliable, and scalable systems.
$161k - $242k
Senior Production Engineer, Managed Cloud
crusoe
Crusoe is seeking a Senior Production Engineer to join their Managed Cloud team. This role is crucial for ensuring the reliability and scalability of Crusoe's AI-optimized cloud platform, delivering a seamless cloud experience at scale. The engineer will be responsible for building and delivering highly available, performant, and cost-efficient AI infrastructure to support compute-intensive, latency-sensitive workloads for customers.
$170k - $205k
Lead Site Reliability Engineer
Glean
Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. Your role is pivotal in ensuring our services meet stringent Service Level Objectives (SLOs) and in building resilient, automated production environments in the cloud. You'll lead a team and be responsible for products globally, providing technical leadership to key projects and empowering your team to do the same. Much of our software development focuses on building infrastructure to scale our operations in a hybrid cloud environment and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale and fast growth which are unique to Glean, while using your expertise in coding, algorithms, problem-solving, and SRE practices. We keep Glean applications up and running, ensuring our customers have the best and most reliable experience possible.
$200k - $260k
ML Platform Engineer
synthesia.io
Synthesia is seeking an Engineer to join its ML Platform team. This team is responsible for building and operating the systems that enable researchers and product teams to train, serve, and deploy generative models efficiently and reliably. The role involves working on research infrastructure, production serving systems, internal tooling, and platform interfaces, with a growing focus on making these systems automation-friendly and agent-oriented. This is a hands-on individual contributor role with significant ownership, where you will help shape the evolution of the ML platform as it scales.
Member of Engineering (Data & Analytics)
poolside
Poolside is building a world where AI powers economically valuable work and scientific progress, aiming to accelerate AGI development by enhancing the developer experience with agentic systems, coding assistants, and frontier models. This role is part of the Model Factory team, responsible for building and maintaining the data and analytics systems that underpin model development, such as agent trajectory stores and experiment configuration registries. The work involves close collaboration with researchers to understand their needs and improve data and analytics, directly contributing to faster model development cycles. This position offers the opportunity to own foundational systems, build at scale, and work on greenfield projects as well as extending existing systems to meet demanding performance and reliability requirements.
Senior Cloud Support Engineer
crusoe
Crusoe Cloud is revolutionizing high-performance computing by offering sustainable, low-cost GPU compute power. As a Cloud Support Engineer, you will be the primary point of contact for technical support, ensuring our customers can seamlessly utilize Crusoe Cloud for advancements in fields like AI/ML, physics simulations, and computational biology. This role directly impacts Crusoe's mission by enabling customers to accelerate their research and development, contributing to a more sustainable future. You will be involved in exciting projects, working with cutting-edge technologies and collaborating with a talented team to solve complex challenges. The ideal candidate is a highly motivated and experienced technical professional with a passion for customer success, a deep understanding of cloud technologies, and a commitment to Crusoe's values.
Tech Lead Manager, Inference
lumalabs
Luma is seeking a Tech Lead Manager for its Inference team to own the entire inference serving stack, encompassing routing, scheduling, and fleet-wide orchestration across thousands of GPUs, multiple clouds, and hardware vendors. This is a hands-on role where at least half of your time will be dedicated to architecting and building core platform components, making critical design decisions, and debugging complex incidents. You will also be responsible for leading, growing, and developing the inference engineering team, including hiring, coaching, and managing on-call rotations. The role involves setting the technical roadmap for serving infrastructure, owning platform SLOs and economics, and partnering with research to deploy new architectures and integrate serving into online RL and evaluation loops. The ideal candidate has extensive experience operating large-scale inference fleets and a genuine desire to remain hands-on in building and improving the serving stack.
$30k - $60k
Staff AI Infrastructure Engineer
lumalabs
Luma is seeking a Staff AI Infrastructure Engineer to own the reliability of its extensive GPU fleet, focusing on scheduling, efficiency, and resilience. This role requires deep systems knowledge to drive company-wide reliability and technical leadership. The position involves close-to-the-metal work with kernels, containers, schedulers, networking, storage, and GPU behavior, addressing challenges posed by high demand. You will also be responsible for setting technical standards and growing the engineering team.
$30k - $60k
Software Engineer, Inference
lumalabs
Luma is seeking a Software Engineer to own the serving of their models. This role involves integrating new architectures into the inference engine, scaling deployments across thousands of machines, and optimizing GPU fleet utilization while meeting internal service level objectives (SLOs). The work focuses on large-scale inference systems, including scheduling, fleet management, deployment pipelines, and reliability across various clusters and hardware providers. This position is ideal for a strong systems engineer experienced with model serving and Kubernetes at scale, rather than pure modeling.
$30k - $60k
Research Scientist / Engineer – Reinforcement Learning Infrastructure
lumalabs
Luma is seeking a Research Scientist / Engineer to build the systems that enable reinforcement learning (RL) at frontier scale. This role involves coupling policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems to transform model behavior into learning signals. RL is crucial for Luma's models to evolve from capable to useful. Operating RL at scale is a complex systems challenge, encompassing training, rollout generation, environment execution, and reward computation across thousands of GPUs, demanding speed, stability, and correctness. This position is ideal for someone with hands-on experience in post-training LLMs with RL, building environments and verifiers, and debugging large-scale asynchronous rollout pipelines.
$30k - $60k