Senior Software Engineer, Machine Learning Infrastructure & Automation

$170k - $230k • Remote • Remote - USA • FullTime

Posted 6h ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Machine Learning Engineer

About the job

fal is building the generative media ecosystem for the next generation of AI products, providing the infrastructure, tools, and model access needed to move from idea to production at scale. This role focuses on empowering fal's ML team by developing automation, infrastructure, and developer tooling to streamline the development, testing, and deployment of generative AI models. You will own and enhance CI/CD systems for a growing collection of ML models and inference pipelines, aiming to eliminate manual work, accelerate development cycles, and build reliable systems for confident model shipping. This high-impact role involves close collaboration with Applied ML and ML Performance teams, building solutions from automated model validation and performance benchmarking to AI-powered development workflows.

Responsibilities

  • Own and maintain ML CI/CD infrastructure, including automated testing, validation, and deployment pipelines for ML models and inference services.
  • Accelerate development cycles by reducing CI execution times through parallelization, caching, test selection, and efficient compute resource utilization.
  • Build automated model validation systems to test model outputs, detect quality regressions, and validate changes across various models, GPU architectures, and configurations.
  • Automate performance benchmarking to detect regressions in inference latency, throughput, GPU utilization, and cost.
  • Develop automated pricing and deployment checks to validate model pricing, billing configurations, API schemas, and deployments before production.
  • Extend agentic engineering workflows with AI-powered automation to diagnose CI failures, identify regressions, propose fixes, and streamline engineering tasks.
  • Improve deployment reliability through automated safeguards, verification, rollback mechanisms, and monitoring.
  • Identify and automate repetitive tasks across the ML team to allow engineers to focus on model development and performance improvements.

Requirements

  • 5+ years of software engineering experience with Python proficiency.
  • Experience building production infrastructure and developer tooling.
  • Experience designing and operating CI/CD systems using GitHub Actions or comparable technologies.
  • Deep understanding of automated testing, build systems, dependency management, caching, and parallel execution.
  • Experience with containerized workloads, Docker, and cloud infrastructure.
  • Ability to design reliable distributed systems and debug complex infrastructure failures.
  • Strong understanding of observability (logs, metrics, tracing, alerting).
  • Passion for developer productivity and demonstrated ability to eliminate manual processes through automation.
  • Comfort working independently, identifying high-impact problems, and building end-to-end solutions.
  • Experience with ML infrastructure, PyTorch, GPU workloads, or model-serving systems.
  • Familiarity with NVIDIA GPU architectures and multi-GPU environments.
  • Experience building automated inference benchmarks or ML quality evaluation frameworks.
  • Experience with agentic coding tools (e.g., Codex, Claude Code) or building custom AI engineering agents.
  • Experience optimizing CI/CD pipelines at scale, including distributed test execution and ephemeral compute environments.
  • Experience developing internal developer platforms or infrastructure-as-code tooling.

Benefits

  • Health, dental, and vision insurance (US)
  • Regular team events and offsites
  • Access to fal's massive GPU cluster for testing and development

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.