Member of Technical Staff - Multimodal Understanding

$180k - $440k Palo Alto, CA

Posted 5d ago

Job Location

Palo Alto, CA

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

SpaceXAI is seeking a Member of Technical Staff to join their multimodal team and advance the understanding and generation of AI across image, video, audio, and text. This role involves working across the full stack, from data curation and pre-training to alignment, infrastructure, and end-to-end product experiences. You will collaborate with various teams to deliver cutting-edge multimodal reasoning, world modeling, tool use, agentic behaviors, and human-AI collaboration capabilities. The goal is to build models that can perceive, reason about, and interact with the world in real-time at an unprecedented level.

Responsibilities

  • Design, build, and optimize large-scale distributed systems for multimodal pre-training, post-training, inference, data processing, and tokenization.
  • Develop high-throughput pipelines for multimodal data acquisition, preprocessing, filtering, generation, and management.
  • Advance multimodal capabilities including spatial-temporal compression, cross-modal alignment, world modeling, reasoning, and real-time video processing.
  • Drive data quality through curation, filtering techniques, analysis, and scalable pipelines.
  • Create evaluation frameworks, internal benchmarks, reward models, and metrics for real-world usage and human-AI synergy.
  • Innovate on algorithms, modeling approaches, hardware/software/algorithm co-design, and scaling paradigms.
  • Build research tooling, user-friendly interfaces, prototypes, demos, and full-stack applications.
  • Enable reasoning, tool calling, agentic behaviors, and seamless real-time interactions across the AI stack.

Requirements

  • Hands-on experience with multimodal pre-training, post-training, or fine-tuning (vision, audio, video, or cross-modal).
  • Expert-level proficiency in Python, with strong experience in JAX, PyTorch, or XLA.
  • Proven track record building or optimizing large-scale distributed ML systems.
  • Deep experience designing and running data pipelines at scale, especially for noisy multimodal data.
  • Strong fundamentals in evaluation design, benchmarks, reward modeling, or RL techniques.
  • Proactive self-starter with a passion for pushing multimodal AI frontiers.
  • Willingness to own end-to-end initiatives and deliver breakthrough user experiences.
  • Familiarity with state-of-the-art in multimodal LLMs, scaling laws, tokenizers, compression techniques, reasoning, or agentic systems.
  • Proficiency in Rust and/or C++ for performance-critical components.
  • Hands-on work with large-scale orchestration tools such as Spark, Ray, or Kubernetes.
  • Background building full-stack tooling, performant interfaces, or real-time research demos/apps.
  • Passion for end-to-end user experience in interactive, real-time multimodal AI systems.

Benefits

  • Equity
  • Comprehensive medical, vision, and dental coverage
  • Access to a 401(k) retirement plan
  • Short & long-term disability insurance
  • Life insurance
  • Various other discounts and perks

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.