Source-linked AI summary
Robostral Navigate
Abdelaziz Bounhar, Abhijeet Somani, Aditi Kabra, Adrian Valente, Adrien Petralia, Adrien Sade, Alan Jeffares, Albert Jiang, Aleksandr Timashov, Alexandre Cahill, Alexandre Gavaudan, Alexandre Laval, Alexandre Sablayrolles, Amelie Heliou, Amos You, Andre Jonasson, Andrew Bai, Andrew Ehrenberg, Andrew Zhao, Angele Lenglemetz, Anmol Agarwal, Antonia Calvi, Arata Suzuki, Arjun Majumdar, Arthur Fournier, Artjom Joosen, Avinash Sooriyarachchi, Aylin Guliz Akkus, Aysenur Karaduman, Baptiste Bout, Baptiste Roziere, Baudouin De Monicault, Benjamin Holzschuh, Benjamin Lefaudeux, Benjamin Tibi, Bernhard Stadlbauer, Blazej Osinski, Camille Le Scao, Chaoran Yu, Charlotte Cronjager, Chen-Yo Sun, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clemence Lanfranchi, Corentin Barreau, Corentin Sautier, Cristiana-Diana Diaconu, Cyprien Courtot, Daniel Marczak, Darius Dabert, Diego de Las Casas, Dominik Nuss, Dylan Rubini, Dzmitry Soupel, Elizaveta Demyanenko, Elliot Chane-Sane, Emilien Fugier, Emmanuel Gottlob, Erik Aas, Etienne Goffinet, Etienne Millon, Eujeong Choi, Fabian Paischer, Fabian Schlager, Faruk Ahmed, Federico Baldassarre, Filip Szatkowski, Florian Wiesner, Gabrielle Berrada, Gaetan Ecrepont, Gaetan Lepage, Gaspard Blanchet, Gaspard Donada-Vidal, Gauthier Delerce, Gauthier Guinet, Genevieve Hayes, Georgii Novikov, Giada Pistilli, Gianluca Galletti, Guillaume Breton, Guillaume Kunsch, Guillaume Lample, Guillaume Martin, Guillaume Raille, Gunjan Dhanuka, Gunshi Gupta, Han Zhou, Harshil Shah, Hasan Furkan Vural, Hedi Hadiji, Hope McGovern, Hugo Cisneros, Hugo Thimonier, Indraneel Mukherjee, Ivan Cuevas Salazar, Jacques Sun, Jan Ludziejewski, Jason Rute, Jean Quentin, Jean-Hadrien Chabran, Jean-Malo Delignon, Jie Zhang, Joachim Studnia, Joep Barmentlo, Johannes Brandstetter, John Harvill, Jonas Amar, Jonas Schweizer, Josephine Delas, Josselin Somerville, Julien Denize, Julien Tauran, Kartik Khandelwal, Khyathi Raghavi Chandu, Kilian Tep, Kush Jain, Larissa Laich, Laura Calem, Laurence Aitchison, Laurent Callot, Laurent Fainsin, Leo Cotteleer, Leonard Blier, Lingxiao Zhao, Louis Martin, Louis Serrano, Lucile Saulnier, Ludovic Ho Fuh, Luis Montero, Maarten Buyl, Manon Chossegros, Marcin Mozejko, Margaret Jennings, Markus Hennerbichler, Martin Alexandre, Mathieu Poiree, Mathieu Schmitt, Mathilde Guillaumin, Matthieu Andre, Matthieu Dinot, Matthieu Futeral, Maurits Bleeker, Mauro Comi, Max Mynter, Maxim Berman, Maxime Darrin, Maxime Louis, Maximilian Augustin, Maximilian Muller, Melina Jingting Laimon, Mert Unsal, Mia Chiquier, Michael Pilcer, Michal Pietruszka, Michal Zajac, Mikhail Biriuchinskii, Minh-Quang Pham, Minwoo Kang, Morgane Riviere, Namit Katariya, Nathan Grinsztajn, Nathan Simpson, Neeraj Aggarwal, Neha Gupta, Ola Mysiak, Oliver Leicht, Olivier Bousquet, Olivier Duchenne, Parag Jain, Patricia Wang, Patrick Blies, Patrick von Platen, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Pavan Kumar Reddy, Pavel Kuksa, Philippe Pinel, Philomene Chagniot, Pierre Stock, Pierre-Andre Savalle, Piotr Milos, Prateek Gupta, Pravesh Agrawal, Quentin Desreumaux, Quentin Torroba, Quercus Hernandez, Ram Ramrakhya, Randall Isenhour, Ranjit Parva, Raul Perez Pelaez, Reinhard Sonnleitner, Remi Delacourt, Richard Kurle, Rishi Shah, Rob Romijnders, Rohin Arora, Romain Sauvestre, Roman Soletskyi, Rosalie Millner, Rupert Menneer, Sagar Vaze, Samuel Barry, Samuel Belkadi, Samuel Humeau, Sanchit Gandhi, Sandeep Subramanian, Sarthak Mittal, Saskia Adaime, Sean Cha, Sebastian Kaltenbach, Shashwat Dalal, Shashwat Verma, Sherif Waly, Shrimai Prabhumoye, Siddhant Waghjale, Siddharth Gandhi, Simon Lepage, Simon Sorg, Soham Ghosh, Sophie Marbach, Srijan Mishra, Stanislas Lange, Steve Hong, Sumukh Aithal, Szymon Antoniak, Tarun Kumar Vangani, Teven Le Scao, Theo Cachet, Thibaut Lavril, Thomas Chabal, Thomas Coste, Thomas Defard, Thomas Foubert, Thomas Robert, Thomas Wang, Tianyu Zhang, Tim Lawson, Timothee Lacroix, Tobias Kronlachner, Tom Bewley, Tom Edwards, Tomas Hodan, Tuhin Das, Tyler Wang, Ulrick BLE, Umar Jamil, Umberto Tomasini, Valentin Mace, Van Phung, Vedant Nanda, Victor Jouault, Victor Letzelter, Victor Paltz, Victor Poucheret, Vincent Maladiere, Vincent Pfister, Virgile Richard, Vladislav Bataev, Wassim Bouaziz, Wen Ding Li, William Havard, William Marshall, Xinghui Li, Xingran Guo, Xinyu Yang, Yann Dreze, Yassine El Ouahidi, Yassir Bendou, Yihan Wang, Yimu Pan, Yves Martin des Taillades, Zaccharie Ramzi, Zhenlin Xu, Zsofia Csakany
TL;DR
Existing embodied-navigation systems often require depth, multiple cameras, LiDAR, or pre-built maps, limiting hardware compatibility. Robostral Navigate uses monocular RGB input and image-space waypoint prediction, achieving state-of-the-art performance on R2R-CE and RxR-CE, including 77.4% success on R2R-CE.
Problem
High-performing embodied-navigation systems often require depth, LiDAR, multiple cameras, or pre-built maps, limiting compatible robots and requiring added cost and calibration.
Method
Robostral Navigate is an 8B vision-language model that predicts image-space navigation waypoints from monocular RGB frames, using efficient prefix-cached training and reinforcement learning.
Results
Robostral Navigate achieves state-of-the-art performance on both R2R-CE and RxR-CE, including a 77.4% success rate on R2R-CE validation unseen.
Takeaways & Limitations
The results demonstrate that combining large-scale simulation, efficient training, and reinforcement learning can produce strong embodied navigation with minimal sensing requirements.
Abstract
from arXiv · showhide
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.
1 Introduction
Robostral Navigate is an 8B vision-language model designed for scalable embodied navigation using only monocular RGB images. It combines simulation-scale training and efficient episode packing with strong benchmark performance across continuous-environment navigation tasks.
- Scalable sensing: Robostral Navigate uses only a single RGB camera while achieving state-of-the-art performance on R2R-CE and RxR-CE benchmarks.Its monocular input supports deployment across diverse robotic platforms, from warehouse wheels to delivery drones.
- Scalable training: 2.4 million trajectories across 350k scenes are generated entirely in simulation, eliminating costly real-world data collection.The training recipe is designed to scale without relying on real-world trajectory collection.
- Scalable training: 22× fewer training tokens are required through prefix caching that packs entire episodes into single sequences while preserving learning signals.The approach transforms month-long training runs into days using attention masking.
- Training refinement: 4% additional success-rate improvement comes from online reinforcement learning on curated hard trajectories, targeting exploration and failure recovery.The reinforcement-learning stage uses CISPO after supervised training.
- Benchmark results: 77.4% success on R2R-CE validation unseen exceeds Qwen-RobotNav-4B by 10.5 points and Qwen-RobotNav-8B by 5.3 points despite using neither depth nor multiple cameras.The cited baselines achieve 66.9% and 72.1%, respectively.
- Benchmark results: 75.1% success and 68.7% SPL on english-only RxR-CE validation unseen establish a new state of the art while outperforming all single-camera baselines.The result demonstrates superior performance on the cited continuous-environment benchmark.
2 Modeling
Robostral Navigate separates navigation into coarse waypoint prediction and low-level motor control. Its monocular RGB VLM predicts visually grounded or metric waypoints, which are translated into trajectories by a diffusion policy and embodiment-specific controller.
- System decomposition: The system decomposes navigation into instruction understanding, scene reasoning, and coarse planning, followed by motor-command generation for efficient obstacle-avoiding movement.Separate models address the high-level planning and low-level control challenges.
- Waypoint prediction: Robostral Navigate consumes language instructions and histories of monocular RGB frames, producing waypoint pixel coordinates or displacement commands with an 8B spatial-grounding VLM.Visual history is encoded as tokens appended to the instruction, enabling progress tracking and reasoning about visited locations.
- Low-level control: A approximately 121M-parameter diffusion transformer converts camera-frame waypoints into 10 Hz action chunks, and an embodiment-specific controller converts them into 100 Hz motor commands.The diffusion policy uses the VLM prediction, robot dimensions, and current and VLM-context RGB frames.
- Waypoint modes: The preferred pointing mode predicts image coordinates (u, v) and arrival heading change ∆θ, while a metric fallback predicts local displacements (∆x, ∆y, ∆θ) when the destination is outside view.Jointly predicting all five quantities supports co-training; approximately 10% of training data has invisible destinations.
- Cross-robot generalization: Randomizing robot height, radius, camera placement, and pitch during training reduces dependence on a particular camera setup and supports cross-platform transfer.The randomized ranges are 0.4–1.8 m height, 0.15–0.45 m radius, 70–100% camera placement height, and 0–25° camera pitch.
3 Training
Robostral Navigate combines large-scale simulation data, prefix-cached sequence training with tree-based attention masking, and online reinforcement learning to improve efficient, deployment-relevant navigation training. This recipe produces diverse trajectories, reduces training cost, addresses exposure bias, and focuses learning on difficult visual navigation cases.
- Simulation-based data generation: 2.4 million trajectories across 350k scenes provide simulated training data spanning diverse indoor and outdoor environments, layouts, objects, lighting, and architectural styles.Farthest point sampling diversifies start and goal positions, including trajectories of varying length and some spanning multiple floors.
- Attention masking: Tree-based attention masking prevents training-time access to previous ground-truth actions, reducing reliance on information unavailable during deployment.Vanilla causal attention would expose previous actions that are often highly informative of future actions, especially with farthest-visible waypoints.
- Prefix-caching training: 22× fewer training tokens and training time reduced from months to days result from prefix-sharing entire episodes while retaining all action-prediction targets.The formulation reduces token complexity from O(T 2) to O(T) by encoding each history element once and computing all timestep losses in one forward pass.
- Online reinforcement learning: Online reinforcement learning with CISPO lets the agent learn from consequences of its own actions and states visited during simulator interaction, addressing exposure bias and covariate shift.The reward is −max(2, dist_to_goal), with geodesic distance measured in meters; clipping the penalty near the goal encourages STOP.
- Training tasks and visual curriculum: 35k failed-policy tasks concentrate computation on complex, ambiguous, and long-horizon episodes, while scene-contiguous ordering creates an effective visual curriculum that outperforms random shuffling.Tasks are selected by rolling out the SFT policy and retaining trajectories that fail to reliably reach the goal.
4 Results
Robostral Navigate achieves state-of-the-art navigation performance on the R2R-CE and RxR-CE validation-unseen benchmarks using only a single RGB camera. Online reinforcement learning further improves success over supervised fine-tuning, including on seen, unseen, and long-horizon instruction-following environments.
- Evaluation setup: The evaluation uses R2R-CE and RxR-CE, where agents follow unseen natural-language instructions through unseen photorealistic environments between predicted waypoints.A Habitat pathfinder handles navigation between Robostral Navigate’s predicted waypoints; reported metrics include NE, SR, OS, and SPL.
- R2R-CE: 77.4% SR and 74.2% SPL on R2R-CE validation unseen establish state-of-the-art performance across all metrics.Robostral Navigate exceeds Qwen-RobotNav-4B by 10.5 SR points and Qwen-RobotNav-8B by 5.3 SR points while using no auxiliary sensors.
- R2R-CE: 79.4% SR on R2R-CE validation seen demonstrates strong performance in seen environments after online reinforcement learning.The post-RL checkpoint improves over the SFT baseline by 3.47% on seen environments and 4.03% on unseen environments.
- RxR-CE: 75.1% SR, 68.7% SPL, and 3.47 m NE on RxR-CE show strong success, efficiency, and final-position accuracy.Compared with Qwen-RobotNav-8B, Robostral Navigate gains 1.7 SR points and 5.2 SPL points using a single RGB camera.
- RxR-CE: 75.1% SR after online reinforcement learning improves on the SFT baseline’s 71.2% by 3.9% on RxR-CE.The improvement demonstrates reinforcement learning gains on long-horizon instruction-following tasks.
5 Conclusion
Robostral Navigate is an 8B vision-language model for embodied navigation that achieves state-of-the-art performance on R2R-CE and RxR-CE while requiring only a single RGB camera. The work combines large-scale simulation, efficient training, and reinforcement learning to minimize sensing requirements and support scalable embodied navigation.
- Conclusion: Robostral Navigate achieves state-of-the-art performance on the R2R-CE and RxR-CE embodied-navigation benchmarks.It is an 8B vision-language model designed for embodied navigation.
- Conclusion: The policy predicts navigation actions via pointing and requires only a single RGB camera for deployment.This design minimizes sensor assumptions for embodied navigation.
- Conclusion: Large-scale simulation, prefix-caching training, and reinforcement learning together produce state-of-the-art navigation with minimal sensing requirements.The authors present Robostral Navigate as a first step toward a general-purpose embodied agent.