Multimodal Generative AI Researcher
Remote • Remote
Posted 6mo ago
Job Location
Remote
Tech Stack
Remote Work Policy
Fully remote
Categories
AI Research Engineer
About the job
We are seeking a Research Scientist with deep expertise in training and fine-tuning large Vision-Language and Language Models (VLMs / LLMs) for downstream multimodal tasks. You will be instrumental in advancing models that reason across vision, language, and 3D, translating research breakthroughs into scalable engineering solutions. This role involves designing and fine-tuning large-scale VLMs/LLMs and hybrid architectures for complex tasks like visual reasoning, retrieval, 3D understanding, and embodied interaction.
Responsibilities
- Design and fine-tune large-scale VLMs/LLMs and hybrid architectures for tasks such as visual reasoning, retrieval, 3D understanding, and embodied interaction.
- Build robust, efficient training and evaluation pipelines, including data curation, distributed training, mixed precision, and scalable fine-tuning.
- Conduct in-depth analysis of model performance, including ablations, bias/robustness checks, and generalization studies.
- Collaborate with research, engineering, and 3D/graphics teams to transition models from prototype to production.
- Publish impactful research and contribute to establishing best practices for multimodal model adaptation.
Requirements
- PhD (or equivalent experience) in Machine Learning, Computer Vision, NLP, Robotics, or Computer Graphics.
- Proven track record in fine-tuning or training large-scale VLMs/LLMs for real-world downstream tasks.
- Strong engineering mindset with the ability to design, debug, and scale training systems end-to-end.
- Deep understanding of multimodal alignment and representation learning (e.g., vision-language fusion, CLIP-style pre-training, retrieval-augmented generation).
- Familiarity with recent trends including video-language and long-context VLMs, spatio-temporal grounding, agentic multimodal reasoning, and Mixture-of-Experts (MoE) fine-tuning.
- Awareness of 3D-aware multimodal models using NeRFs, Gaussian splatting, or differentiable renderers for grounded reasoning and 3D scene understanding.
- Hands-on experience with PyTorch/DeepSpeed/Ray and distributed or mixed-precision training.
- Excellent communication skills and a collaborative mindset.