Spark Jobs
61 open roles mentioning Spark
AI Engineer
Baseten
Baseten is seeking an AI Engineer to join their Training Product team. This role involves building AI-driven product features for customers training frontier models and enhancing Baseten's internal AI capabilities by transforming manual workflows into agentic ones. You will collaborate directly with research engineers, taking concepts from internal research to customer-ready products. This is a hands-on position offering significant autonomy, where you will identify key problems, develop reliable AI systems with appropriate harnesses and guardrails, and take ownership of the outcomes. The ideal candidate has a proven track record of shipping agents and is eager to make that their primary focus.
Forward Deployed Engineer (Training)
Baseten
Baseten is seeking a Forward Deployed Engineer (Training) to work directly with leading AI companies, taking ownership of their technical outcomes on the Baseten platform. This role involves tackling complex challenges in serving and improving AI models at scale, spanning the entire model lifecycle from inference to post-training and the systems that connect them. You will act as a technical advisor, guiding customers from initial problem framing through to production deployment, and ensuring the quality and performance of their AI workloads through rigorous evaluation and optimization.
Procurement Lead
Baseten
Baseten is seeking a Procurement Lead to establish and build the company's procurement function from the ground up. This is a foundational role with high ownership, focused on defining and implementing scalable purchasing workflows, vendor relationships, and procurement policies. The goal is to make procurement a source of leverage for the business, freeing up business owners' time and capacity while building necessary guardrails for growth. This role requires close partnership with Finance, Legal, Security, Compliance, IT, and business owners across the company, with success measured by team support and unblocked progress.
Senior Data Engineer
Runpod
Runpod is seeking a Senior Data Engineer to join their remote-first team. This role is crucial for building and maintaining the AI Developer Cloud platform that serves over a million developers. You will be responsible for owning data pipelines end-to-end, from ingestion to data products, utilizing technologies like Snowflake, dbt, Dagster, Terraform, and AWS. The position also involves working in an AI-forward environment, delegating tasks to and extending the capabilities of autonomous AI agents, and critically reviewing their output. This is an opportunity to shape the future of AI infrastructure for the next generation of developers.
$175k - $220k
Software Engineer - AI Developer Productivity
Baseten
Baseten is seeking an AI Developer Productivity Engineer to build the internal platform that empowers engineers to work in an AI-first manner. This role involves creating the agent configurations, tooling, and evaluation frameworks that make AI agents competent and easy to use within our codebase. You will be responsible for making the "good path" for AI adoption the easiest path, focusing on shipping infrastructure and measuring its impact. The goal is to enable teams to adopt these AI tools because they are superior to self-assembled solutions, not due to mandates. You will write the playbook for an AI-first Software Development Lifecycle (SDLC) at Baseten, driving adoption through excellent developer experience, clear documentation, and low friction.
Software Engineer - Observability
Baseten
Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of this team, you will be pivotal in building and shaping the observability experience for our internal and external customers, directly impacting the reliability and operational excellence of Baseten's product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing significantly. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that issues can be detected, diagnosed, and resolved rapidly, even as the systems become more complex.
Member of Technical Staff (Product Data Scientist, Search Quality)
Perplexity AI
Perplexity is seeking an experienced Product Data Scientist to drive the advancement of search technologies. This role involves identifying strong and sensitive signals from user behavior to enhance the efficiency of A/B experiment data analysis. You will contribute to shaping the product roadmap and accelerating user adoption through data-driven insights, hypothesis validation via A/B testing, and the design of new pipelines for improved ranking quality. This includes discovering new signals, producing metrics, and constructing data labeling pipelines utilizing both human and LLM feedback.
AI Inference Engineer
Baseten
As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. This is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in.
Engineering Manager, Model Infrastructure
Harvey
Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This role offers a rare chance to help build a generational company at a true inflection point, with strong product-market fit and world-class investor support. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As the Engineering Manager for Model Infrastructure, you will lead the team responsible for the platform powering every model request across Harvey, partnering closely with AI Research, Product Engineering, and Infrastructure to ensure reliability, scalability, and cost-efficiency. This is a strategic engineering organization, critical for every product capability, and will evolve to build the infrastructure for Harvey to train, evaluate, deploy, and operate its own frontier AI models.
$260k - $340k
Member of Technical Staff - Data Platform
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We are building open models that empower individuals to control their intelligence and shape the future of AI. This role focuses on building and operating the company-wide foundations platform, providing reliable, scalable developer infrastructure, SRE capabilities, and high-throughput data ingestion tooling. The team designs and implements core data systems and pipelines that power research, training, and production environments, enabling high-velocity experimentation, reliable model development, and scalable production workflows by unifying ingestion, processing, and orchestration across the data lifecycle.
Member of Technical Staff - Engineering Lead, Data Ingestion
Reflection ai
Reflection's Data team is responsible for building the training corpora that our frontier models learn from. The ingestion layer is the critical machinery that transforms raw data from the open web, licensed sources, and other large-scale origins into well-structured, versioned, and auditable datasets for pre-training. As the Data Ingestion Lead, you will provide front-line leadership for the team responsible for this layer, overseeing web crawl, data ingestion pipelines, and data lakes. You will build, mentor, and grow a team of data ingestion engineers, guide technical and architectural decisions across crawling, extraction, and corpus storage/delivery, and collaborate closely with research, data quality, and data partnerships teams. You will also remain hands-on to contribute technically and maintain a deep understanding of the team's work, as subtle decisions at the ingestion layer significantly impact model performance, safety, and failure points.
Software Engineer, Data Infrastructure
thinkingmachines
Thinking Machines Lab is seeking an engineer to join a high-impact team focused on data infrastructure. This role is crucial for architecting and scaling the core systems that power distributed training pipelines, multimodal data catalogs, and intelligent processing of petabytes of data. You will work directly with researchers to accelerate experiments, develop new datasets, enhance infrastructure efficiency, and derive key insights from our data assets. If you are passionate about distributed systems, large-scale data mining, and building foundational tools from the ground up, we encourage you to apply.
$350k - $475k
Base Labs Fellowship
Baseten
Base Labs, a research initiative by Baseten, is seeking researchers for its Fellowship program. This program offers full-time, 3-month research opportunities in San Francisco, providing funding, mentorship, and support to produce rigorous, published research. The goal is to expose researchers to frontier AI research within an industry setting, contributing to both the open-source ecosystem and Baseten's technical roadmap. Fellows will work closely with senior researchers, with the potential for a full-time offer for exceptional performers.
Research Scientist, Data
pika
Pika is seeking a staff or lead-level Research Engineer, Data to architect and scale data engineering systems for their advanced multimodal foundation models. This role is crucial for strengthening research teams by building, optimizing, and owning large-scale data pipelines and ML data curation. The goal is to ensure foundation models have access to high-quality, diverse datasets, enabling millions of creators. If you are passionate about data infrastructure and innovative research-engineering, this is an opportunity to make a significant impact.
Software Engineer, Research Acceleration
thinkingmachines
Thinking Machines Lab is seeking engineers to build the libraries and tools that accelerate research. You will own internal infrastructure, including evaluation libraries, RL training libraries, and experiment tracking platforms, and build systems that compound research velocity over time. This is a collaborative role where you will work directly with researchers to identify bottlenecks and pain points. Success means researchers trust your systems to just work and find them a delight to use.
$350k - $475k
Research Infrastructure Engineer, Research Acceleration
thinkingmachines
Thinking Machines Lab is seeking engineers to build the libraries and tools that accelerate research. You will own internal infrastructure, including evaluation libraries, RL training libraries, and experiment tracking platforms, to build systems that compound research velocity over time. This is a collaborative role where you will work directly with researchers to identify bottlenecks and pain points. Success means researchers trust your systems to just work and find them a delight to use.
$350k - $475k
Software Engineer, Data Infrastructure
Scale AI
Scale AI is seeking a highly skilled and motivated Software Engineer to join our dynamic Public Sector Engineering team. You will play a critical role in supporting Scale’s government customers by scoping and developing onsite solutions. Your expertise will be instrumental in designing and implementing systems that can handle interactions with existing customer systems to help our products integrate into existing customer workflows. We are looking for an exceptional Senior Software Engineer to architect and build the foundational data infrastructure that will serve as the brain of a project ecosystem. You will be responsible for designing highly novel data models and processing pipelines capable of handling massive quantities of output data from complex simulations. At the core of this role is the challenge of building a foundational data ensemble—a unified architecture that seamlessly aggregates, structures, and stages diverse sources of simulation outputs and user inputs. Your systems will manage enormous batch throughput jobs with strict, minimal latency requirements, ensuring that downstream AI systems and language models have the exact context they need to actionably reason over complex, multi-dimensional scenarios.
$216k - $300k
Member of Technical Staff - Data Ingestion Engineer
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. As a Member of Technical Staff on the Data Team, you will be instrumental in building and operating the ingestion systems that transform large-scale data sources, such as the open web, into reliable corpora for training frontier AI models. This role involves owning the machinery for data acquisition, extraction, normalization, versioning, and delivery to pre-training pipelines. You will collaborate closely with world-class researchers, contributing to the critical link between data collection and its impact on model performance. This position is ideal for engineers who excel at building robust distributed systems while also enjoying experimentation, reasoning about data acquisition tradeoffs, and iterating quickly based on measurable outcomes.
Member of Technical Staff - Web Crawl Engineer
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. As a Member of Technical Staff on the Data Team, you will be instrumental in building and operating large-scale web crawling systems. These systems are crucial for discovering, acquiring, and processing internet content, directly influencing the capabilities of frontier AI systems. You will own the infrastructure for web-scale data collection, from initial URL discovery to distributed crawling and dataset delivery, working closely with researchers to identify and acquire high-value content efficiently. This role is perfect for engineers passionate about distributed systems, large-scale crawler optimization, and tackling the unique challenges of modern web data collection.
Backend Software Engineer — Data Platform & AI Data Products
Together AI
You will join the Data Platform team, responsible for building the backend services and data products that power how data moves through the company. This involves creating core platform primitives like high-quality event streams, reliable access layers, and developer-friendly APIs and tools. The goal is to enable teams across the organization to self-serve their data needs and ship faster. You will contribute to backend services that derive value from company data and enhance the self-serve capabilities of the data platform, allowing product and engineering teams to easily create and operate event-driven architectures, publish/consume streams, define access models, and manage data products end-to-end. Additionally, you will work on LLM-adjacent services, including prompt categorization, enrichment, and metadata systems, transforming raw telemetry into trusted, usable products with guidance from experienced engineers.
$120k - $170k