Transformers Jobs
29 open roles mentioning Transformers
Member of Engineering (Post-training)
poolside
Poolside is building a company to create Artificial General Intelligence, aiming to accelerate software development through agentic systems, coding assistants, and frontier models. This role is part of the Applied Research team, focused on transforming pre-trained Large Language Models (LLMs) into well-aligned and highly capable AI systems specifically for coding and software development. You will be involved in building data pipelines and environments for agentic use cases, researching and implementing post-training algorithms, and designing experiments to test hypotheses, with access to significant GPU resources.
Embedded AI Engineer – Android Automotive (On-Device Intelligence)
Applied Intuition
Applied Intuition is seeking an Embedded AI Engineer to build on-device intelligence for a next-generation Android Automotive platform. This role is responsible for the entire lifecycle of embedded ML systems, ensuring models perform predictably and safely within production environments, adhering to real-world constraints like latency, thermal limits, and functional safety. The company is a leader in powering the future of physical AI, serving industries such as automotive, defense, and construction with solutions for tools, infrastructure, operating systems, and autonomy.
$60k - $300k
Machine Learning Engineer, Integrity
OpenAI
As a Machine Learning Engineer in OpenAI's Integrity team, you will work on state-of-the-art models and classifiers, experiment with new architectures and approaches, and advance our capabilities in content and user understanding. This role offers the chance to transform research breakthroughs into tangible solutions that enhance the trust and safety of our platform, with a focus on training LLMs and building ML models. You will be instrumental in designing and deploying advanced machine learning models to solve real-world problems, bringing OpenAI's research from concept to implementation and creating AI-driven applications with direct impact.
Applied AI, Evaluation Engineer
Mistral AI
Mistral AI is seeking an Evaluation Engineer to join our customer-facing Applied AI team. This role is crucial for ensuring our AI solutions are production-ready by designing methodologies, building infrastructure, and defining evaluation standards across various industries and use cases. You will bridge the gap between research and customer needs, creating evaluation frameworks that measure LLM performance in real-world scenarios, moving beyond standard benchmarks to address domain-specific risks and requirements. This position offers a unique opportunity to impact the deployment of cutting-edge AI by directly contributing to its measurable success for enterprise clients.
Software Engineer - Model Performance Systems
Baseten
Baseten is seeking Software Engineers to join their team in a specialized, high-impact role at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. This position involves not only building automated performance monitoring and diagnostic tools for next-generation AI infrastructure but also defining the roadmap, driving key technical decisions, and taking full ownership of the future direction of this work. The role offers the opportunity to shape the platform that engineers use to deploy cutting-edge AI models into production.
Audio Inference Engineer, Model Efficiency
Cohere
Cohere is seeking an Audio Inference Engineer focused on Model Efficiency to join a fast-growing team of researchers and engineers. The mission of this team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will advance core audio model serving metrics, including latency, throughput, and quality by diving deep into systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You will collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference.
Member of Technical Staff, Model Efficiency
Cohere
Cohere is seeking a Member of Technical Staff focused on Model Efficiency to join a fast-growing team of researchers and engineers. This role is instrumental in pushing the boundaries of LLM inference efficiency, developing techniques to improve model execution in production for lower latency and higher throughput. You will work across the inference stack, identifying bottlenecks and developing optimizations, collaborating with modeling and systems teams to ship meaningful improvements. Opportunities exist to build expertise in advanced performance techniques like GPU/CUDA optimizations and model execution strategies for large-scale architectures.
Inference Engineer
cartesia
Cartesia is seeking an Inference Engineer to advance its mission of building real-time multimodal intelligence. This role involves designing and building a low-latency, scalable, and reliable model inference and serving stack for cutting-edge foundation models. You will collaborate closely with research and product engineering teams to serve products efficiently and reliably, and design robust inference infrastructure with monitoring. This position offers significant autonomy to shape products and influence how AI is applied across various devices and applications.
Open-Source Software, Machine Learning Engineer
Mistral AI
Mistral AI is democratizing AI through high-performance, optimized, open-source models, products, and solutions. We are a dynamic, collaborative team passionate about AI's potential to transform society, with teams distributed globally. We are seeking an Open-Source Software, Machine Learning Engineer to join our OSS team, which is embedded within our Science team. This role is critical in helping turn research breakthroughs into tangible solutions and improving Mistral's open-source ecosystem by open-sourcing state-of-the-art models and maintaining our publicly available libraries.