Applied Machine Learning Engineer, EMEA
London • FullTime
Posted 11d ago
Job Location
London
Tech Stack
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Machine Learning Engineer
About the job
Fireworks is seeking an Applied Machine Learning Engineer for the EMEA region. In this role, you will be the technical owner of customer engagements, embedding within client teams to understand their specific needs and challenges. You will be responsible for the entire lifecycle of a customer's deployment on the Fireworks platform, from initial scoping and model selection to ensuring production readiness, performance, and cost-efficiency. This position emphasizes first principles thinking and requires a deep understanding of software engineering, machine learning techniques, and infrastructure optimization.
Responsibilities
- Own customer engagements from scoping through production as the sole technical owner.
- Embed within customer teams to understand their domain, data, and constraints.
- Perform fine-tuning and post-training tasks, including data preparation, evaluation, and production model deployment.
- Optimize model serving performance and cost on the Fireworks platform.
- Accurately assess and scope project feasibility before commitment.
- Provide actionable product feedback to engineering teams based on customer needs.
Requirements
- Demonstrated expertise in at least two of the following areas: software engineering, machine learning, or infrastructure and performance.
- For ML expertise: formal training (PhD, PhD in progress, or strong MSc) and experience post-training LLMs.
- For infrastructure expertise: experience with GPU infrastructure, distributed serving, or inference engines.
- Founder or founding engineer experience.
- Experience in a startup or fast-paced environment.
- Strong Python skills and at least one systems language.
- Fluency with AI-assisted and agentic engineering.
- Experience with SFT, LoRA, RLHF, RLVR, and distillation.
- Ability to design evaluations that reflect real tasks and drive training decisions.
- Understanding of GPU capabilities, serving and inference optimization techniques (quantization, speculative decoding, batching, KV cache, latency vs throughput, multi-GPU/node serving, capacity planning).
Benefits
- Opportunity to solve hard problems at the forefront of AI infrastructure.
- Work with bleeding-edge technology impacting AI adoption globally.
- High ownership and impact in a fast-growing team.
- Learn from world-class engineers and AI researchers.