Transformers Jobs
19 open roles mentioning Transformers
Machine Learning Infrastructure Engineer, Safeguards Research
Anthropic
Anthropic's Safeguards team is responsible for developing systems that detect and mitigate misuse of AI models. This role focuses on building and owning the infrastructure that supports the research efforts of this team. You will create the tooling, pipelines, and abstractions that enable researchers to run experiments, train detection methods, and select detections for launch efficiently and reliably. This position bridges the gap between research and production, ensuring fast iteration for researchers and dependable results for detection systems as models evolve. The ideal candidate will have a proven ability to solve large-scale systems and data problems and a strong desire to deepen their machine learning expertise.
TPU Kernel Engineer
Anthropic
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. As a TPU Kernel Engineer, you will be responsible for identifying and addressing performance issues across various ML systems, including research, training, and inference. A significant part of this role involves designing and optimizing kernels specifically for the TPU, and providing feedback to researchers on how model changes affect performance. We are looking for individuals with a strong background in solving large-scale systems problems and performing low-level optimizations.
ML/Research Engineer, Safeguards
Anthropic
Anthropic is seeking ML Engineers and Research Engineers to join the Safeguards ML team. The primary focus of this role is to develop systems that detect and mitigate misuse of AI systems, ranging from individual policy violations to sophisticated coordinated attacks. You will build defenses to ensure product safety as AI capabilities advance, protect user well-being, and guarantee appropriate model behavior across various contexts. This work is crucial for Anthropic's Responsible Scaling Policy commitments.
Performance Engineer
Anthropic
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. As a Performance Engineer, you will identify and solve novel systems problems that arise from running machine learning (ML) algorithms at scale. You will develop systems to optimize the throughput and robustness of our largest distributed systems. This role requires a strong track record of solving large-scale systems problems and a desire to become an expert in ML.
Engineering Manager, Inference
Anthropic
Anthropic's performance and scaling teams are dedicated to optimizing compute resources for both inference and training. As an Engineering Manager on these teams, you will lead efforts to identify and eliminate bottlenecks, develop resilient and scalable solutions, and enhance system efficiency. You will also provide clear direction, focus, and context to your team within a fast-paced, evolving environment. This role requires a leader who can guide engineering initiatives to improve model performance and scale inference and training systems, while also contributing technically to the team's stack.
Engineering Manager, GPU (ML Accelerator)
Anthropic
Anthropic's performance and scaling teams focus on making the most efficient and impactful use of our compute resources for inference and training. As an Engineering Manager on these teams, you will be responsible for identifying and removing bottlenecks, building robust and durable solutions, and maximizing the efficiency of our systems. You will also help bring clarity, focus, and context to your teams in a fast-paced, dynamic environment.
Machine Learning Engineer, Reliability
Fal
fal is building the generative media ecosystem for the next generation of AI products, providing the infrastructure, tools, and model access needed to scale from idea to production. As generative media reshapes industries, fal is becoming the foundation for ambitious teams. This hybrid ML Engineering / Site Reliability Engineering role will own the reliability, security, and safety of fal's generative media model APIs, ensuring they remain available, performant, secure, and safe for thousands of developers and enterprises. You will address model-specific failure modes, such as degraded output quality, drift, unsafe generations, and abuse patterns, as critical reliability concerns alongside uptime and latency.
Research Engineer, Core ML
Together AI
This research engineering role focuses on translating new Reinforcement Learning (RL) algorithms, scheduling methods, and inference optimizations into production-grade systems that power Together's API. The Core ML team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems, building and maintaining high-performance inference and RL engines at production scale. The goal is to significantly improve model speed, cost-efficiency, and capabilities through RL-based post-training. This position requires a blend of algorithmic understanding and systems engineering, with opportunities to work across the entire stack from RL algorithms and training engines to kernels and serving systems, ultimately driving measurable improvements in latency, throughput, cost, and model quality at scale.
$200k - $280k
AI Researcher, Core ML (Turbo)
Together AI
The Turbo team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems. We are responsible for building and managing the systems that power Together's API, focusing on high-performance inference and RL/post-training engines capable of operating at production scale. Our core mission is to advance the frontiers of efficient inference and RL-driven training, aiming to make models significantly faster and more cost-effective to run, while simultaneously enhancing their capabilities through RL-based post-training methods. This role involves working across the entire stack, from RL algorithms and training engines to kernels and serving systems, to develop and refine state-of-the-art models using RL pipelines. We value individuals with deep expertise in one area and a strong willingness to collaborate and grow across others.
$200k - $280k
Frontier Agents Intern (Fall 2026)
Together AI
The Agents team investigates how to build, align, and scale frontier AI systems capable of complex, multi-step tasks and workflows across text and speech, with a focus on agentic and scientific domains. This role sits at the intersection of agent capabilities, human-computer interaction, and infrastructure, exploring areas like post-training methods for agentic behavior and developing evaluation frameworks for open-ended tasks. As a research intern, you will tackle challenges in alignment, reliability, and scalability, potentially working on new training recipes for self-learning and long-horizon reasoning, curating datasets, studying failure modes, or building scalable agent infrastructure.
Research Engineer, Materials Science
Google DeepMind
Google DeepMind is seeking a Research Engineer to join their materials science team. This role involves accelerating the discovery of new functional materials by integrating artificial intelligence, computational simulation, and automated experimentation. You will collaborate with a diverse interdisciplinary team of domain experts, ML researchers, and engineers. The work focuses on pioneering research in various scientific domains, enabling the validation of early ideas and building infrastructure for promising research lines. You will contribute your scientific domain knowledge to the team's collective expertise.
$141k - $202k
Tech Lead Manager- MLRE, ML Systems
Scale AI
Scale's LLM post-training platform team builds our internal distributed framework for large language model training, powering MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs. This platform also serves as the underlying training framework for the data quality evaluation pipeline. You will work closely with Scale’s ML teams and researchers to build the foundation platform which supports all our ML research and development works, optimizing it to enable next generation LLM training, inference, and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you!
$265k - $331k
Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI
Scale AI
Scale is seeking a Machine Learning Systems Research Engineer to join their Enterprise ML Research Lab. This role will focus on building algorithms for a next-generation Agent RL training platform, supporting large-scale training, and integrating state-of-the-art technologies to optimize ML systems. You will collaborate with other ML researchers and engineers who apply these algorithms to client use cases, including AI cybersecurity firewalls and healthtech search models. If you are passionate about shaping the future of AI, this is an exciting opportunity to contribute to cutting-edge advancements in enterprise GenAI.
$265k - $331k
ML Research Engineer, ML Systems
Scale AI
Scale's ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. This platform powers MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs, as well as data quality evaluation. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development, optimizing it to enable the next generation of LLM training, inference, and data curation. If you are excited about shaping the future of AI via fundamental innovations, we would love to hear from you!
$190k - $237k
Applied AI, Evaluation Engineer
Mistral AI
Mistral AI is seeking an Evaluation Engineer to join our customer-facing Applied AI team. This role is crucial for ensuring our AI solutions are production-ready by designing methodologies, building infrastructure, and defining evaluation standards across various industries and use cases. You will bridge the gap between research and customer needs, creating evaluation frameworks that measure LLM performance in real-world scenarios, moving beyond standard benchmarks to address domain-specific risks and requirements. This position offers a unique opportunity to impact the deployment of cutting-edge AI by directly contributing to its measurable success for enterprise clients.
Audio Inference Engineer, Model Efficiency
Cohere
Cohere is seeking an Audio Inference Engineer focused on Model Efficiency to join a fast-growing team of researchers and engineers. The mission of this team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will advance core audio model serving metrics, including latency, throughput, and quality by diving deep into systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You will collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference.
Member of Technical Staff, Model Efficiency
Cohere
Cohere is seeking a Member of Technical Staff focused on Model Efficiency to join a fast-growing team of researchers and engineers. This role is instrumental in pushing the boundaries of LLM inference efficiency, developing techniques to improve model execution in production for lower latency and higher throughput. You will work across the inference stack, identifying bottlenecks and developing optimizations, collaborating with modeling and systems teams to ship meaningful improvements. Opportunities exist to build expertise in advanced performance techniques like GPU/CUDA optimizations and model execution strategies for large-scale architectures.
Inference Engineer
cartesia
Cartesia is seeking an Inference Engineer to advance its mission of building real-time multimodal intelligence. This role involves designing and building a low-latency, scalable, and reliable model inference and serving stack for cutting-edge foundation models. You will collaborate closely with research and product engineering teams to serve products efficiently and reliably, and design robust inference infrastructure with monitoring. This position offers significant autonomy to shape products and influence how AI is applied across various devices and applications.
Open-Source Software, Machine Learning Engineer
Mistral AI
Mistral AI is democratizing AI through high-performance, optimized, open-source models, products, and solutions. We are a dynamic, collaborative team passionate about AI's potential to transform society, with teams distributed globally. We are seeking an Open-Source Software, Machine Learning Engineer to join our OSS team, which is embedded within our Science team. This role is critical in helping turn research breakthroughs into tangible solutions and improving Mistral's open-source ecosystem by open-sourcing state-of-the-art models and maintaining our publicly available libraries.