Spark Jobs

17 open roles mentioning Spark

Staff+ Software Engineer, Account Compromise

7d ago
Anthropic

Anthropic

Anthropic's Safeguards organization is responsible for building systems that ensure Claude is safe to use at scale. The Account Compromise team focuses on protecting users from losing account control and preventing compromised accounts from being used to abuse the platform. This role involves owning the architecture for detecting, containing, and remediating account compromise across Claude and its Developer Platform, making key design decisions for other engineers. The work requires deep technical expertise, adversarial problem-solving, threat modeling, and the ability to scope and lead complex, multi-month projects from ambiguous beginnings. You will operate with significant autonomy, driving alignment with various teams and shaping Anthropic's global approach to account security.

London, UK onsite
AnthropicPythonClaude +3 more

Engineering Manager, Model Infrastructure

9d ago
Harvey

Harvey

Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This role offers a rare chance to help build a generational company at a true inflection point, with strong product-market fit and world-class investor support. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As the Engineering Manager for Model Infrastructure, you will lead the team responsible for the platform powering every model request across Harvey, partnering closely with AI Research, Product Engineering, and Infrastructure to ensure reliability, scalability, and cost-efficiency. This is a strategic engineering organization, critical for every product capability, and will evolve to build the infrastructure for Harvey to train, evaluate, deploy, and operate its own frontier AI models.

$260k - $340k

San Francisco hybrid FullTime
OpenAIAnthropicAzure +5 more

Member of Technical Staff - Data Platform

10d ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We are building open models that empower individuals to control their intelligence and shape the future of AI. This role focuses on building and operating the company-wide foundations platform, providing reliable, scalable developer infrastructure, SRE capabilities, and high-throughput data ingestion tooling. The team designs and implements core data systems and pipelines that power research, training, and production environments, enabling high-velocity experimentation, reliable model development, and scalable production workflows by unifying ingestion, processing, and orchestration across the data lifecycle.

New York, NY onsite FullTime
AirflowSparkKafka

Member of Technical Staff - Engineering Lead, Data Ingestion

10d ago
Reflection ai

Reflection ai

Reflection's Data team is responsible for building the training corpora that our frontier models learn from. The ingestion layer is the critical machinery that transforms raw data from the open web, licensed sources, and other large-scale origins into well-structured, versioned, and auditable datasets for pre-training. As the Data Ingestion Lead, you will provide front-line leadership for the team responsible for this layer, overseeing web crawl, data ingestion pipelines, and data lakes. You will build, mentor, and grow a team of data ingestion engineers, guide technical and architectural decisions across crawling, extraction, and corpus storage/delivery, and collaborate closely with research, data quality, and data partnerships teams. You will also remain hands-on to contribute technically and maintain a deep understanding of the team's work, as subtle decisions at the ingestion layer significantly impact model performance, safety, and failure points.

San Francisco, CA onsite FullTime
AirflowSpark

Software Engineer, Data Infrastructure

13d ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking an engineer to join a high-impact team focused on data infrastructure. This role is crucial for architecting and scaling the core systems that power distributed training pipelines, multimodal data catalogs, and intelligent processing of petabytes of data. You will work directly with researchers to accelerate experiments, develop new datasets, enhance infrastructure efficiency, and derive key insights from our data assets. If you are passionate about distributed systems, large-scale data mining, and building foundational tools from the ground up, we encourage you to apply.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +5 more

Research Engineer, Discovery

16d ago
Anthropic

Anthropic

As a Research Engineer on our team, you will work end-to-end across the entire model stack, identifying and addressing key infrastructure blockers on the path to scientific AGI. You should have familiarity with elements of language model training, evaluation, and inference, and be eager to quickly dive into and get up to speed in areas where you are not yet an expert. This may include performance optimization, distributed systems, VM/sandboxing/container deployment, and large-scale data pipelines. Join us in our mission to develop advanced AI systems that push the frontiers of science and benefit humanity.

San Francisco, CA onsite
AnthropicAWSDocker +5 more

Research Scientist, Data

29d ago
pika

pika

Pika is seeking a staff or lead-level Research Engineer, Data to architect and scale data engineering systems for their advanced multimodal foundation models. This role is crucial for strengthening research teams by building, optimizing, and owning large-scale data pipelines and ML data curation. The goal is to ensure foundation models have access to high-quality, diverse datasets, enabling millions of creators. If you are passionate about data infrastructure and innovative research-engineering, this is an opportunity to make a significant impact.

Palo Alto HQ onsite FullTime
AWSAzurePython +4 more

Software Engineer, Research Acceleration

1mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking engineers to build the libraries and tools that accelerate research. You will own internal infrastructure, including evaluation libraries, RL training libraries, and experiment tracking platforms, and build systems that compound research velocity over time. This is a collaborative role where you will work directly with researchers to identify bottlenecks and pain points. Success means researchers trust your systems to just work and find them a delight to use.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +4 more

Research Infrastructure Engineer, Research Acceleration

1mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking engineers to build the libraries and tools that accelerate research. You will own internal infrastructure, including evaluation libraries, RL training libraries, and experiment tracking platforms, to build systems that compound research velocity over time. This is a collaborative role where you will work directly with researchers to identify bottlenecks and pain points. Success means researchers trust your systems to just work and find them a delight to use.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +4 more

Software Engineer, Data Infrastructure

1mo ago
Scale AI

Scale AI

Scale AI is seeking a highly skilled and motivated Software Engineer to join our dynamic Public Sector Engineering team. You will play a critical role in supporting Scale’s government customers by scoping and developing onsite solutions. Your expertise will be instrumental in designing and implementing systems that can handle interactions with existing customer systems to help our products integrate into existing customer workflows. We are looking for an exceptional Senior Software Engineer to architect and build the foundational data infrastructure that will serve as the brain of a project ecosystem. You will be responsible for designing highly novel data models and processing pipelines capable of handling massive quantities of output data from complex simulations. At the core of this role is the challenge of building a foundational data ensemble—a unified architecture that seamlessly aggregates, structures, and stages diverse sources of simulation outputs and user inputs. Your systems will manage enormous batch throughput jobs with strict, minimal latency requirements, ensuring that downstream AI systems and language models have the exact context they need to actionably reason over complex, multi-dimensional scenarios.

$216k - $300k

New York, NY; Washington, DC onsite
PythonGoRust +8 more

Member of Technical Staff - Data Ingestion Engineer

1mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. As a Member of Technical Staff on the Data Team, you will be instrumental in building and operating the ingestion systems that transform large-scale data sources, such as the open web, into reliable corpora for training frontier AI models. This role involves owning the machinery for data acquisition, extraction, normalization, versioning, and delivery to pre-training pipelines. You will collaborate closely with world-class researchers, contributing to the critical link between data collection and its impact on model performance. This position is ideal for engineers who excel at building robust distributed systems while also enjoying experimentation, reasoning about data acquisition tradeoffs, and iterating quickly based on measurable outcomes.

San Francisco, CA onsite FullTime
Spark

Member of Technical Staff - Web Crawl Engineer

1mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. As a Member of Technical Staff on the Data Team, you will be instrumental in building and operating large-scale web crawling systems. These systems are crucial for discovering, acquiring, and processing internet content, directly influencing the capabilities of frontier AI systems. You will own the infrastructure for web-scale data collection, from initial URL discovery to distributed crawling and dataset delivery, working closely with researchers to identify and acquire high-value content efficiently. This role is perfect for engineers passionate about distributed systems, large-scale crawler optimization, and tackling the unique challenges of modern web data collection.

San Francisco, CA onsite FullTime
Spark

Backend Software Engineer — Data Platform & AI Data Products

1mo ago
Together AI

Together AI

You will join the Data Platform team, responsible for building the backend services and data products that power how data moves through the company. This involves creating core platform primitives like high-quality event streams, reliable access layers, and developer-friendly APIs and tools. The goal is to enable teams across the organization to self-serve their data needs and ship faster. You will contribute to backend services that derive value from company data and enhance the self-serve capabilities of the data platform, allowing product and engineering teams to easily create and operate event-driven architectures, publish/consume streams, define access models, and manage data products end-to-end. Additionally, you will work on LLM-adjacent services, including prompt categorization, enrichment, and metadata systems, transforming raw telemetry into trusted, usable products with guidance from experienced engineers.

$120k - $170k

San Francisco remote
PythonGoRust +9 more

Staff Software Engineer, Data Platform

2mo ago
Scale AI

Scale AI

Scale is at the forefront of the AI revolution, developing data engines and technologies that power the world's leading LLMs and generative models. This role is on the Platform Engineering team, responsible for the foundational data infrastructure that supports these cutting-edge AI products. You will lead the design and development of core data storage, streaming, caching, and indexing platforms, gaining exposure to the rapidly evolving AI landscape across various industries. The work involves driving architecture, implementation, and reliability of these critical systems, collaborating with stakeholders, and mentoring junior engineers.

$252k - $315k

San Francisco, CA; New York, NY remote
KubernetesPythonFine-Tuning +10 more

Software Engineer, Systems Generalist

3mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking generalist infrastructure and systems engineers to build the core systems powering their foundation models and support internal research and product development teams. This high-impact role involves architecting and scaling critical infrastructure across the full technical stack, solving complex distributed systems problems, and building robust, scalable platforms. You will work directly with researchers to accelerate experiments, improve infrastructure efficiency, and enable key insights across models, products, and data assets.

$350k - $475k

San Francisco onsite
OpenAIMistralKubernetes +5 more

Senior Member of Technical Staff, Synthetic Data

8mo ago
Cohere

Cohere

Cohere is seeking a Senior Machine Learning Engineer specializing in synthetic data to develop and manage the synthetic data pipeline crucial for advanced language models. This role involves end-to-end management of synthetic data, including pipeline optimization, data analysis and generation, and conducting data ablations and model evaluations. You will transform diverse web and code data using generative models to enhance token efficiency and model quality, bridging research and engineering to improve throughput and accelerator utilization. This position is key to Cohere's mission of delivering efficient and reliable language capabilities and driving innovation in natural language processing.

Toronto remote FullTime
CoherePythonNLP +1 more

Software Engineer, Data

1y ago
Mistral AI

Mistral AI

Mistral AI is a pioneering company focused on democratizing AI through high-performance, open-source models and solutions. We are seeking passionate and skilled software engineers to join our dynamic, collaborative team. In this role, you will design, build, and maintain our data infrastructure, ensuring data accuracy, accessibility, and security. Your contributions will be crucial in enabling our science teams to enhance AI model quality and empowering business users to make informed decisions.

Paris remote Full-time
MistralPythonSQL +8 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.