PyTorch Jobs
223 open roles mentioning PyTorch
Product Manager, Platform
fireworks ai
Fireworks is seeking a Product Manager to focus on their core inference product and general platform. This role is central to enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. You will be responsible for setting strategy, writing product specifications, engaging directly with customers running production traffic, and collaborating across inference, infrastructure, and go-to-market teams to deliver significant customer impact. Example challenges include optimizing rate limits and preventing fraudulent usage while maintaining a smooth experience for legitimate users.
Product Manager, Training
fireworks ai
Fireworks is seeking a Product Manager to focus on their AI training platform. In this role, you will be instrumental in shaping the strategy, defining product specifications, and directly engaging with customers to understand their real-world tuning challenges. You will collaborate closely with the training, research, and inference teams to deliver impactful solutions that enhance customer AI development. Example initiatives include driving growth in self-serve training usage and improving the speed and ease of training for enterprise clients.
Forward Deployed Engineer (Training)
Baseten
Baseten is seeking a Forward Deployed Engineer (Training) to work directly with leading AI companies, taking ownership of their technical outcomes on the Baseten platform. This role involves tackling complex challenges in serving and improving AI models at scale, spanning the entire model lifecycle from inference to post-training and the systems that connect them. You will act as a technical advisor, guiding customers from initial problem framing through to production deployment, and ensuring the quality and performance of their AI workloads through rigorous evaluation and optimization.
Member of Technical Staff - Reliability Engineering
fireworks ai
Fireworks AI is seeking a Member of Technical Staff focused on Reliability Engineering to ensure the dependable operation of their AI platform. This role involves working across cloud infrastructure, AI systems, and product teams to guarantee seamless integration, graceful failure handling, and robust performance under load. You will be instrumental in defining reliability standards, owning the reliability toolchain, and ensuring a positive customer experience by addressing failures that span across systems. The position requires a proactive approach to incident management, a commitment to reducing operational toil through automation, and strong collaboration with various engineering teams.
Partnerships Lead
fireworks ai
Fireworks AI is seeking a high-ownership sales operator to drive sourced pipeline and revenue through the Microsoft Azure channel. This quota-carrying, field-facing role involves mapping Microsoft's ISV and enterprise field organization, activating co-sell motions with key personnel, and building repeatable sales plays to guide customers to Fireworks AI via Azure Foundry. You will own the forecast and the number, with significant influence over how Fireworks engages Microsoft's field teams and builds a scalable co-sell motion. This is an opportunity for someone who thrives on creating from scratch and is accountable to achieving targets in a dynamic environment.
Staff AI Scientist
fiddler-ai
Fiddler is building trust into AI, especially with the rise of Generative AI and Agents. Our platform helps organizations deploy trustworthy and transparent AI solutions by monitoring, evaluating, securing, analyzing, and improving AI applications. We partner with AI-first organizations to establish responsible AI practices, enabling engineering teams and business stakeholders to understand AI outcomes. Joining Fiddler means making an impact by ensuring AI applications at production scale have operational transparency and security. This is an opportunity to be a trailblazer in the rapidly innovating AI and ML industry, contributing to AI Observability.
$220k - $260k
Member of Technical Staff, Agentic Environments
Cohere
Cohere is seeking a senior engineer to join our team, focusing on the practical challenges of deploying AI systems at scale in production environments. This hands-on, engineering-driven role involves working with frontier AI models, building scalable solutions, and bridging research concepts with real-world implementations. You will contribute to both engineering and research efforts, designing and writing high-performing software for model training, developing new tools to support LLM research, and collaborating with various engineering and scientific teams. We provide access to world-class compute resources, data, and talent to enable you to do your best work.
Machine Learning Engineer, API Multicloud
OpenAI
OpenAI is seeking Machine Learning Engineers to join its API Multicloud team, focusing on extending OpenAI's API platform into strategic cloud environments, starting with AWS. This role involves building and improving AI systems that help strategic partners adapt OpenAI models for cloud-native use cases. You will operate at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure, spanning post-training workflows, evaluation, data pipelines, and API/infrastructure integration. The ideal candidate will enjoy working with external technical partners, diagnosing issues, and translating learnings into platform improvements, collaborating closely with Research, Applied, Safety Systems, and infrastructure teams.
Product Designer
fireworks ai
Fireworks is seeking a Product Designer passionate about transforming complex AI infrastructure into intuitive developer experiences. This role involves owning the design process end-to-end, from model deployment to AI agent creation on the Fireworks platform. The focus is on simplifying powerful technologies, requiring close collaboration with engineers and PMs who have a deep understanding of systems. The ideal candidate will be excited about console UIs, API ergonomics, and enhancing the developer's journey to first inference.
Senior Staff Software Engineer, DC Infrastructure
crusoe
Crusoe is seeking a highly skilled Software Engineer to join their Data Center Infrastructure Engineering team. This role focuses on developing software for managing a fleet of GPU servers and the data centers that house them. The position involves creating and implementing advanced diagnostic, observability, automation, and repair tools for high-performance GPU compute clusters. The ideal candidate will be a hands-on problem solver, comfortable working independently, and will play a crucial role in maintaining the health and scalability of Crusoe's rapidly expanding GPU fleet.
$250k - $300k
Staff Software Engineer, DC Infrastructure
crusoe
Crusoe is seeking a highly skilled and motivated Software Engineer to join its Data Center Infrastructure Engineering team. This role focuses on developing software for managing a fleet of GPU servers and the data centers that house them. The position involves creating and implementing advanced diagnostic, observability, automation, and repair tools for high-performance GPU compute clusters. The ideal candidate will be a hands-on problem solver, comfortable working independently, and will play a crucial role in maintaining the health and scalability of Crusoe's rapidly expanding GPU fleet.
$215k - $260k
AI Performance Engineer
Applied Intuition
Applied Intuition is seeking a performance engineer to specialize in making large-scale machine learning workloads fast and cost-efficient within the datacenter. This role focuses on optimizing distributed training runs across multiple nodes and high-throughput batch inference for processing vast amounts of real-world autonomy logs. The primary goal is to improve throughput, cluster efficiency, and reduce cost per unit of data processed, directly impacting the company's iteration speed. You will be responsible for identifying and resolving performance bottlenecks across the entire stack, from accelerators to ML frameworks and data infrastructure, working collaboratively with various engineering teams to achieve significant improvements in training time and processing costs.
$60k - $300k
Machine Learning Data Scientist, Forecasting
OpenAI
OpenAI is seeking a senior Machine Learning Data Scientist to lead forecasting initiatives within the Strategic Finance team. This role involves building and scaling robust, interpretable, and production-ready forecasting systems to predict key business metrics like user growth, revenue, and compute consumption. You will be a founding member of the Forecasting pillar, collaborating closely with product managers, researchers, engineers, and finance leaders to operationalize insights, influence company strategy, and build foundational forecasting capabilities. This is a highly cross-functional position requiring technical excellence, product intuition, and business acumen.
Member of Technical Staff, Enterprise Foundations
fireworks ai
Enterprise Foundations is responsible for building the core capabilities that large enterprises require to operate their business on the Fireworks platform. This involves addressing complex needs related to organizational structure, user access control, usage metering, billing, data security, and deployment within customer-owned cloud environments. The role offers end-to-end ownership of features, starting from customer conversations, evolving into changes in core data models and APIs, and extending through the control plane, training and inference stacks, SDK, CLI, and console. The work is foundational, often involving new primitives that impact other teams' code and must be rolled out to production without disruption. This position is closely tied to revenue, with requirements stemming directly from enterprise deals and customer interactions, offering a unique opportunity to influence product direction and build solutions for scaled companies.
GTM Engineer
fireworks ai
Fireworks is seeking a GTM Engineer to design, operate, and enhance the systems that drive our revenue engine. This role bridges GTM systems architecture and field execution, ensuring our tools and automations boost productivity for sales and marketing teams. You will collaborate closely with sellers, marketers, and GTM leadership to ensure our systems provide crucial signals, reduce manual tasks, and accelerate team velocity. The ideal candidate is passionate about inventing the future of sales platforms and applying world-class AI to go-to-market execution.
Research Engineer, Mid-Training
Cognition
We are an applied AI lab building end-to-end software agents, known for creating Devin, the first AI software engineer. Our team is composed of highly talented individuals with backgrounds in competitive programming and leadership roles at cutting-edge AI companies. We are tackling significant global challenges and developing AI capable of real-world reasoning. This role focuses on the critical 'mid-training' phase, bridging pre-training and post-training to refine raw model capabilities. You will be instrumental in shaping our models' fundamental abilities by owning late-stage training decisions, including data mix and quality, annealing schedules, context length extension, capability injection, and synthetic data strategies.
Senior Software Engineer (DCIE)
crusoe
Crusoe is seeking a highly skilled Software Engineer to join their Data Center Infrastructure Engineering team. This role focuses on developing software for managing a fleet of GPU servers and the data centers that house them. The position involves creating and implementing advanced diagnostic, observability, automation, and repair tools for high-performance GPU compute clusters. The ideal candidate is a hands-on problem solver who can work independently and play a critical role in maintaining the health and scalability of Crusoe's rapidly growing GPU fleet.
$170k - $205k
Member of Technical Staff, Research
fireworks ai
Fireworks is seeking a Member of Technical Staff for its Research team to push the boundaries of generative AI. This role involves advancing LLMs and multimodal systems through foundational research, focusing on enhancing model efficiency, accuracy, and scalability to shape high-performance AI infrastructure. You will collaborate with experts in deep learning, distributed systems, and optimization to translate cutting-edge research into practical applications and influence how leading companies build and deploy AI.
Senior Solution Engineer
Lambda
Lambda, a leader in AI cloud infrastructure, is seeking a Senior Solution Engineer to join their growing team. This role is crucial in enabling customers to achieve their business goals with AI infrastructure by partnering with leading AI researchers and enterprise engineering teams. The ideal candidate will design, scale, and optimize high-performance GPU cloud solutions, turning complex compute challenges into seamless, production-ready AI infrastructure. This position requires a customer-first mindset and technical mastery to drive growth and customer success.
Researcher, Frontier Cybersecurity Risks
OpenAI
As a Researcher for cybersecurity risks, you will help design and implement an end-to-end mitigation stack to reduce severe cyber misuse across OpenAI’s products. This role requires strong technical depth and close cross-functional collaboration to ensure safeguards are enforceable, scalable, and effective. You’ll contribute directly to building protections that remain robust as products, model capabilities, and attacker behaviors evolve. Models are becoming increasingly capable—moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. As we push toward AGI, cybersecurity becomes one of the most important and urgent frontiers: the same systems that can accelerate productivity can also accelerate exploitation.