PyTorch Jobs
109 open roles mentioning PyTorch
Endpoint Engineer, IT
thinkingmachines
Thinking Machines Lab is seeking an Endpoint Engineer to join their IT team. This role will focus on building secure infrastructure and efficient processes for an all-Mac environment, managing the endpoint fleet as a distributed platform with production-engineering practices. The team utilizes version-controlled workflows for endpoint configurations, security policies, scripts, and software deployments, incorporating testing, review, staged rollouts, and rollback capabilities. The Endpoint Engineer will collaborate closely with IT, Security, and Identity teams to ensure a secure and reliable employee computing experience.
$200k - $325k
Safety Operations Lead
thinkingmachines
Thinking Machines Lab is seeking an Operations Analyst to focus on ensuring product safety while supporting rapid product iteration. This role involves close collaboration with product engineers, security, researchers, and designers to integrate safety and integrity into the entire product lifecycle. Daily tasks will include managing the moderation queue and developing tools and policies to enhance the speed and accuracy of future moderation efforts.
$190k - $300k
Engineering Manager, GPU Infrastructure
Cohere
Cohere is a leading enterprise AI company building cutting-edge foundation AI models and end-to-end products. The GPU Clusters team is central to Cohere's infrastructure, responsible for building and operating the superclusters that power our frontier AI models. This role involves enabling research and development at the intersection of cutting-edge hardware, distributed systems, and AI research. As an Engineering Manager, you will lead a team of highly motivated engineers passionate about GPU infrastructure and AI, fostering a culture of technical excellence and innovation in a remote-first environment. This is a unique opportunity to shape the infrastructure powering the next generation of AI.
Software Engineer, Full Stack
thinkingmachines
Thinking Machines Lab is seeking a full stack engineer to build and ship products from prototype to scale, and to maintain tools that accelerate research and product teams. This role involves working across frontend and backend components, and contributing to the reliability, observability, and security of production systems. The company's mission is to empower humanity through advancing collaborative general intelligence, building a future where everyone has access to AI tools for their unique needs.
$350k - $475k
Software Engineer, Full Stack, Tinker
thinkingmachines
Thinking Machines Lab is seeking a full-stack engineer to develop and deploy the products and services that Tinker users engage with daily. This role involves working across frontend, backend, and infrastructure to build the Tinker console, developer tools, and other essential components for the platform. Tinker is a fine-tuning API that enables researchers and developers to customize frontier AI models using their own data and algorithms, managing the underlying infrastructure to provide flexibility and access to advanced capabilities.
$350k - $475k
Software Engineer, Developer Productivity, AI Tools
thinkingmachines
Thinking Machines Lab is seeking a developer productivity engineer to enhance internal software development processes, focusing on safety, speed, and user experience. This role will concentrate on AI tools and coding agents, collaborating with platform, security, and product engineers to build cutting-edge tooling for AI-assisted software development and significantly accelerate the inner development loop. The position involves both establishing company-wide platforms and assisting individual developers in optimizing their workflows.
$350k - $475k
Software Engineer, Data Infrastructure
thinkingmachines
Thinking Machines Lab is seeking an engineer to join a high-impact team focused on data infrastructure. This role is crucial for architecting and scaling the core systems that power distributed training pipelines, multimodal data catalogs, and intelligent processing of petabytes of data. You will work directly with researchers to accelerate experiments, develop new datasets, enhance infrastructure efficiency, and derive key insights from our data assets. If you are passionate about distributed systems, large-scale data mining, and building foundational tools from the ground up, we encourage you to apply.
$350k - $475k
Site Reliability Engineer (SRE)
thinkingmachines
Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to ensure the end-to-end reliability of their Tinker platform. This role involves working closely with engineers and research teams to enhance the robustness and resilience of every system layer. The SRE will be instrumental in maintaining and improving the infrastructure that supports custom AI model fine-tuning, ensuring a seamless experience for researchers and developers.
$350k - $475k
Research Engineer, Infrastructure, Inference
thinkingmachines
Thinking Machines Lab is seeking an infrastructure research engineer to design, optimize, and scale the systems that power large AI models. The goal is to make inference faster, more cost-effective, more reliable, and more reproducible, enabling research teams to focus on advancing model capabilities. This role is crucial for ensuring that every experiment, evaluation, and deployment runs smoothly at scale, with a focus on performant and efficient model inference for both real-world applications and research acceleration.
$350k - $475k
Research Engineer, Developer Experience, Tinker
thinkingmachines
Thinking Machines Lab is seeking a Research Engineer focused on developer experience to build and enhance their Tinker platform. This role involves working hands-on with users to understand their challenges and translate them into product improvements. You will be responsible for creating and updating documentation, adding library features, prototyping integrations, and ensuring users can smoothly customize frontier AI models. This position acts as a crucial link between Tinker users and the internal research and infrastructure teams, surfacing user patterns to inform product and infrastructure priorities and sharing learnings through various channels.
$350k - $475k
Research Engineer, Code RL (Reinforcement Learning)
Anthropic
We are seeking a Research Engineer for our Code RL team, focused on advancing AI models' capabilities in writing, editing, testing, debugging, and shipping real software. This role involves designing RL environments, coding tasks, and reward signals, as well as running training experiments on frontier models. You will diagnose model performance, improve pipeline speed and reliability, and contribute to areas like agentic coding behaviors, code correctness, and autonomous engineering. The position blends cutting-edge research with practical engineering to build high-quality, scalable AI systems.
Research Engineer, Discovery
Anthropic
As a Research Engineer on our team, you will work end-to-end across the entire model stack, identifying and addressing key infrastructure blockers on the path to scientific AGI. You should have familiarity with elements of language model training, evaluation, and inference, and be eager to quickly dive into and get up to speed in areas where you are not yet an expert. This may include performance optimization, distributed systems, VM/sandboxing/container deployment, and large-scale data pipelines. Join us in our mission to develop advanced AI systems that push the frontiers of science and benefit humanity.
Research Engineer, Interpretability
Anthropic
The Interpretability team at Anthropic is dedicated to understanding how large language models work, believing that a mechanistic understanding is key to making advanced AI systems safe and reliable. This role involves building and maintaining the specialized infrastructure for interpretability research, akin to performing 'neuroscience' on neural networks. The work spans the entire lifecycle of a production language model, from pretraining and inference to performance optimization, pushing the boundaries of hardware and software to address critical bottlenecks. As interpretability research matures and is applied to safety audits on frontier models, engineering and infrastructure have become crucial, making this role directly impactful on one of AI's most significant open problems.
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic
As a Research Engineer within Reinforcement Learning, you will collaborate with a diverse group of researchers and engineers to advance the capabilities and safety of large language models. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to the research direction. You'll work on fundamental research in reinforcement learning, creating 'agentic' models via tool use for open-ended tasks such as computer use and autonomous software generation, improving reasoning abilities in areas such as mathematics, and developing prototypes for internal use, productivity, and evaluation.
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic
As a Research Engineer within Reinforcement Learning, you will collaborate with a diverse group of researchers and engineers to advance the capabilities and safety of large language models. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to the research direction. You'll work on fundamental research in reinforcement learning, creating 'agentic' models via tool use for open-ended tasks such as computer use and autonomous software generation, improving reasoning abilities in areas such as mathematics, and developing prototypes for internal use, productivity, and evaluation.
Research Engineer, Performance RL (Reinforcement Learning)
Anthropic
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to safely write correct, fast code for accelerators. You'll need to know accelerator performance well to turn it into tasks and signals models can learn from. Specifically, you will: - Invent, design and implement RL environments and evaluations. - Conduct experiments and shape our research roadmap. - Deliver your work into training runs. - Collaborate with other researchers, engineers, and performance engineering specialists across and outside Anthropic. You may be a good fit if you: - Have expertise with accelerators (CUDA, ROCm, Triton, Pallas), ML framework programming (JAX or PyTorch). - Have worked across the stack – kernels, model code, distributed systems. - Know how to balance research exploration with engineering implementation. - Are passionate about AI's potential and committed to developing safe and beneficial systems. Strong candidates may also have: - Experience with reinforcement learning. - Experience porting ML workloads between different types of accelerators. - Familiarity with LLM training methodologies. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $350,000—$850,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings. How we're different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.
Research Engineer, Pretraining
Anthropic
Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.
Research Engineer, Pretraining Scaling
Anthropic
Anthropic's ML Performance and Scaling team is responsible for training the company's production pretrained models, a critical function that directly shapes the future of Anthropic and its mission to build safe, beneficial AI systems. As a Research Engineer on this team, you will ensure that frontier models train reliably, efficiently, and at scale. This role bridges the gap between research and engineering, involving work across the entire production training stack, including performance optimization, hardware debugging, experimental design, and launch coordination. During model launches, the team operates in close collaboration, addressing production issues that require immediate attention.
Research Engineer, Pretraining Scaling - London
Anthropic
Anthropic's ML Performance and Scaling team is responsible for training our production pretrained models, a critical function that directly shapes the company's future and its mission to build safe, beneficial AI systems. As a Research Engineer on this team, you will ensure our frontier models train reliably, efficiently, and at scale. This demanding, high-impact role requires deep technical expertise and a passion for large-scale ML systems, operating at the boundary between research and engineering. You will work across the entire production training stack, including performance optimization, hardware debugging, experimental design, and launch coordination, responding to critical production issues during launches.
Research Engineer/Research Scientist, Pre-training
Anthropic
Anthropic is seeking a Research Engineer to join its Pre-training team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, contributing to the creation of safe, steerable, and trustworthy AI systems. The team is dedicated to ensuring that transformative AI systems are aligned with human interests and societal benefit.