Source description
About the role
PhD in Machine Learning, Computer Science, Computer Vision, Robotics, or a related field, with first-author publications at top-tier venues (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, RSS, CoRL).
Research experience with state-of-the-art video generative models and world models (e.g., Cosmos-3, LTX 2.3, Self-Forcing, Lingbot-World, or comparable systems).
Deep expertise in at least one of the following areas:
Full-stack data pipelines — large-scale video data pipelines and/or simulation data collection; annotation and filtering workflows for video / world model training.
Model training & infrastructure — training large-scale diffusion transformers on large GPU clusters.
Rendering engines & simulation — Unreal Engine and Blueprint-based gym environments, game-engine integration, building interactive simulated environments.
World action models & robotics — world action models / video action models, action-conditioned video generation, world-model applications in robotics.
Strong systems and engineering expertise in deep learning frameworks such as PyTorch.
Highly proficient with modern AI coding agents and web-based coding tools (e.g., Claude Code, Codex, Cursor), and skilled at leveraging them to dramatically accelerate research workflows.
Exceptional problem-solving skills and the ability to navigate ambiguity in rapidly evolving research areas.
More at Institute of Foundation Models
Related open roles
AI Research Internship - WM
San Francisco Bay Area · Onsite
Research Scientist, Agentic Data & Benchmarking
San Francisco Bay Area · Onsite
Research Scientist - Vision Language Model
United States · Onsite
Research Engineer - The Diffusion LLM Team
San Francisco Bay Area · Onsite
Research Scientist - Agents
United States · Onsite
Research Scientist - Speech/Audio Machine Learning
Paris · Onsite
