Spark Jobs

61 open roles mentioning Spark

AI Engineer

25d ago
B

Baseten

Baseten is seeking an AI Engineer to join their Training Product team. This role involves building AI-driven product features for customers training frontier models and enhancing Baseten's internal AI capabilities by transforming manual workflows into agentic ones. You will collaborate directly with research engineers, taking concepts from internal research to customer-ready products. This is a hands-on position offering significant autonomy, where you will identify key problems, develop reliable AI systems with appropriate harnesses and guardrails, and take ownership of the outcomes. The ideal candidate has a proven track record of shipping agents and is eager to make that their primary focus.

San Francisco hybrid FullTime
LangChainPythonGo +4 more

Forward Deployed Engineer (Training)

26d ago
B

Baseten

Baseten is seeking a Forward Deployed Engineer (Training) to work directly with leading AI companies, taking ownership of their technical outcomes on the Baseten platform. This role involves tackling complex challenges in serving and improving AI models at scale, spanning the entire model lifecycle from inference to post-training and the systems that connect them. You will act as a technical advisor, guiding customers from initial problem framing through to production deployment, and ensuring the quality and performance of their AI workloads through rigorous evaluation and optimization.

San Francisco hybrid FullTime
KubernetesPyTorchDeep Learning +2 more

Procurement Lead

28d ago
B

Baseten

Baseten is seeking a Procurement Lead to establish and build the company's procurement function from the ground up. This is a foundational role with high ownership, focused on defining and implementing scalable purchasing workflows, vendor relationships, and procurement policies. The goal is to make procurement a source of leverage for the business, freeing up business owners' time and capacity while building necessary guardrails for growth. This role requires close partnership with Finance, Legal, Security, Compliance, IT, and business owners across the company, with success measured by team support and unblocked progress.

San Francisco hybrid FullTime
Spark

Senior Data Engineer

1mo ago
R

Runpod

Runpod is seeking a Senior Data Engineer to join their remote-first team. This role is crucial for building and maintaining the AI Developer Cloud platform that serves over a million developers. You will be responsible for owning data pipelines end-to-end, from ingestion to data products, utilizing technologies like Snowflake, dbt, Dagster, Terraform, and AWS. The position also involves working in an AI-forward environment, delegating tasks to and extending the capabilities of autonomous AI agents, and critically reviewing their output. This is an opportunity to shape the future of AI infrastructure for the next generation of developers.

$175k - $220k

Remote - USA remote FullTime
AWSPythonGo +5 more

Software Engineer - AI Developer Productivity

1mo ago
B

Baseten

Baseten is seeking an AI Developer Productivity Engineer to build the internal platform that empowers engineers to work in an AI-first manner. This role involves creating the agent configurations, tooling, and evaluation frameworks that make AI agents competent and easy to use within our codebase. You will be responsible for making the "good path" for AI adoption the easiest path, focusing on shipping infrastructure and measuring its impact. The goal is to enable teams to adopt these AI tools because they are superior to self-assembled solutions, not due to mandates. You will write the playbook for an AI-first Software Development Lifecycle (SDLC) at Baseten, driving adoption through excellent developer experience, clear documentation, and low friction.

San Francisco hybrid FullTime
DockerKubernetesPython +3 more

Software Engineer - Observability

1mo ago
B

Baseten

Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of this team, you will be pivotal in building and shaping the observability experience for our internal and external customers, directly impacting the reliability and operational excellence of Baseten's product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing significantly. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that issues can be detected, diagnosed, and resolved rapidly, even as the systems become more complex.

San Francisco hybrid FullTime
PythonGoRust +1 more

Member of Technical Staff (Product Data Scientist, Search Quality)

1mo ago
P

Perplexity AI

Perplexity is seeking an experienced Product Data Scientist to drive the advancement of search technologies. This role involves identifying strong and sensitive signals from user behavior to enhance the efficiency of A/B experiment data analysis. You will contribute to shaping the product roadmap and accelerating user adoption through data-driven insights, hypothesis validation via A/B testing, and the design of new pipelines for improved ranking quality. This includes discovering new signals, producing metrics, and constructing data labeling pipelines utilizing both human and LLM feedback.

Belgrade onsite FullTime
PythonSQLSpark

AI Inference Engineer

1mo ago
B

Baseten

As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. This is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in.

San Francisco hybrid FullTime
DockerPythonSpark

Engineering Manager, Model Infrastructure

1mo ago
Harvey

Harvey

Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This role offers a rare chance to help build a generational company at a true inflection point, with strong product-market fit and world-class investor support. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As the Engineering Manager for Model Infrastructure, you will lead the team responsible for the platform powering every model request across Harvey, partnering closely with AI Research, Product Engineering, and Infrastructure to ensure reliability, scalability, and cost-efficiency. This is a strategic engineering organization, critical for every product capability, and will evolve to build the infrastructure for Harvey to train, evaluate, deploy, and operate its own frontier AI models.

$260k - $340k

San Francisco hybrid FullTime
OpenAIAnthropicAzure +5 more

Member of Technical Staff - Data Platform

1mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We are building open models that empower individuals to control their intelligence and shape the future of AI. This role focuses on building and operating the company-wide foundations platform, providing reliable, scalable developer infrastructure, SRE capabilities, and high-throughput data ingestion tooling. The team designs and implements core data systems and pipelines that power research, training, and production environments, enabling high-velocity experimentation, reliable model development, and scalable production workflows by unifying ingestion, processing, and orchestration across the data lifecycle.

New York, NY onsite FullTime
AirflowSparkKafka

Member of Technical Staff - Engineering Lead, Data Ingestion

1mo ago
Reflection ai

Reflection ai

Reflection's Data team is responsible for building the training corpora that our frontier models learn from. The ingestion layer is the critical machinery that transforms raw data from the open web, licensed sources, and other large-scale origins into well-structured, versioned, and auditable datasets for pre-training. As the Data Ingestion Lead, you will provide front-line leadership for the team responsible for this layer, overseeing web crawl, data ingestion pipelines, and data lakes. You will build, mentor, and grow a team of data ingestion engineers, guide technical and architectural decisions across crawling, extraction, and corpus storage/delivery, and collaborate closely with research, data quality, and data partnerships teams. You will also remain hands-on to contribute technically and maintain a deep understanding of the team's work, as subtle decisions at the ingestion layer significantly impact model performance, safety, and failure points.

San Francisco, CA onsite FullTime
AirflowSpark

Software Engineer, Data Infrastructure

1mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking an engineer to join a high-impact team focused on data infrastructure. This role is crucial for architecting and scaling the core systems that power distributed training pipelines, multimodal data catalogs, and intelligent processing of petabytes of data. You will work directly with researchers to accelerate experiments, develop new datasets, enhance infrastructure efficiency, and derive key insights from our data assets. If you are passionate about distributed systems, large-scale data mining, and building foundational tools from the ground up, we encourage you to apply.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +5 more

Base Labs Fellowship

2mo ago
B

Baseten

Base Labs, a research initiative by Baseten, is seeking researchers for its Fellowship program. This program offers full-time, 3-month research opportunities in San Francisco, providing funding, mentorship, and support to produce rigorous, published research. The goal is to expose researchers to frontier AI research within an industry setting, contributing to both the open-source ecosystem and Baseten's technical roadmap. Fellows will work closely with senior researchers, with the potential for a full-time offer for exceptional performers.

San Francisco hybrid Temporary
GoSpark

Research Scientist, Data

2mo ago
pika

pika

Pika is seeking a staff or lead-level Research Engineer, Data to architect and scale data engineering systems for their advanced multimodal foundation models. This role is crucial for strengthening research teams by building, optimizing, and owning large-scale data pipelines and ML data curation. The goal is to ensure foundation models have access to high-quality, diverse datasets, enabling millions of creators. If you are passionate about data infrastructure and innovative research-engineering, this is an opportunity to make a significant impact.

Palo Alto HQ onsite FullTime
AWSAzurePython +4 more

Software Engineer, Research Acceleration

2mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking engineers to build the libraries and tools that accelerate research. You will own internal infrastructure, including evaluation libraries, RL training libraries, and experiment tracking platforms, and build systems that compound research velocity over time. This is a collaborative role where you will work directly with researchers to identify bottlenecks and pain points. Success means researchers trust your systems to just work and find them a delight to use.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +4 more

Research Infrastructure Engineer, Research Acceleration

2mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking engineers to build the libraries and tools that accelerate research. You will own internal infrastructure, including evaluation libraries, RL training libraries, and experiment tracking platforms, to build systems that compound research velocity over time. This is a collaborative role where you will work directly with researchers to identify bottlenecks and pain points. Success means researchers trust your systems to just work and find them a delight to use.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +4 more

Software Engineer, Data Infrastructure

2mo ago
Scale AI

Scale AI

Scale AI is seeking a highly skilled and motivated Software Engineer to join our dynamic Public Sector Engineering team. You will play a critical role in supporting Scale’s government customers by scoping and developing onsite solutions. Your expertise will be instrumental in designing and implementing systems that can handle interactions with existing customer systems to help our products integrate into existing customer workflows. We are looking for an exceptional Senior Software Engineer to architect and build the foundational data infrastructure that will serve as the brain of a project ecosystem. You will be responsible for designing highly novel data models and processing pipelines capable of handling massive quantities of output data from complex simulations. At the core of this role is the challenge of building a foundational data ensemble—a unified architecture that seamlessly aggregates, structures, and stages diverse sources of simulation outputs and user inputs. Your systems will manage enormous batch throughput jobs with strict, minimal latency requirements, ensuring that downstream AI systems and language models have the exact context they need to actionably reason over complex, multi-dimensional scenarios.

$216k - $300k

New York, NY; Washington, DC onsite
PythonGoRust +8 more

Member of Technical Staff - Data Ingestion Engineer

2mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. As a Member of Technical Staff on the Data Team, you will be instrumental in building and operating the ingestion systems that transform large-scale data sources, such as the open web, into reliable corpora for training frontier AI models. This role involves owning the machinery for data acquisition, extraction, normalization, versioning, and delivery to pre-training pipelines. You will collaborate closely with world-class researchers, contributing to the critical link between data collection and its impact on model performance. This position is ideal for engineers who excel at building robust distributed systems while also enjoying experimentation, reasoning about data acquisition tradeoffs, and iterating quickly based on measurable outcomes.

San Francisco, CA onsite FullTime
Spark

Member of Technical Staff - Web Crawl Engineer

2mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. As a Member of Technical Staff on the Data Team, you will be instrumental in building and operating large-scale web crawling systems. These systems are crucial for discovering, acquiring, and processing internet content, directly influencing the capabilities of frontier AI systems. You will own the infrastructure for web-scale data collection, from initial URL discovery to distributed crawling and dataset delivery, working closely with researchers to identify and acquire high-value content efficiently. This role is perfect for engineers passionate about distributed systems, large-scale crawler optimization, and tackling the unique challenges of modern web data collection.

San Francisco, CA onsite FullTime
Spark

Backend Software Engineer — Data Platform & AI Data Products

3mo ago
Together AI

Together AI

You will join the Data Platform team, responsible for building the backend services and data products that power how data moves through the company. This involves creating core platform primitives like high-quality event streams, reliable access layers, and developer-friendly APIs and tools. The goal is to enable teams across the organization to self-serve their data needs and ship faster. You will contribute to backend services that derive value from company data and enhance the self-serve capabilities of the data platform, allowing product and engineering teams to easily create and operate event-driven architectures, publish/consume streams, define access models, and manage data products end-to-end. Additionally, you will work on LLM-adjacent services, including prompt categorization, enrichment, and metadata systems, transforming raw telemetry into trusted, usable products with guidance from experienced engineers.

$120k - $170k

San Francisco remote
PythonGoRust +9 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.