# language and reasoning annotation: 323 catalogue datasets

Rows 1 to 100, sorted by hours, datasets that do not state it last.

Next page: https://datasets.gurasees.com/kinds/language-and-reasoning-annotation.md?offset=100

| dataset | kind | robot class | hours | episodes | year | licence | access | id |
|---|---|---|---|---|---|---|---|---|
| [Robo-Dopamine GRM Dataset and Bench](https://datasets.gurasees.com/datasets/robo-dopamine-grm-dataset-and-bench.md) | language and reasoning annotation | cross embodiment | 3,400 |  | 2026 | open | gated | robo-dopamine-grm-dataset-and-bench |
| [UrbanNav](https://datasets.gurasees.com/datasets/urbannav.md) | language and reasoning annotation | human egocentric | 1,500 | 3,000,000 | 2025 | not stated | not stated | urbannav |
| [MolmoAct2 re-annotated DROID, Bridge, RT-1 and BC-Z datasets](https://datasets.gurasees.com/datasets/molmoact2-re-annotated-droid-bridge-rt-1-and-bc-z-datasets.md) | language and reasoning annotation | cross embodiment | 329 | 74,604 | 2026 | open | open | molmoact2-re-annotated-droid-bridge-rt-1-and-bc-z-datasets |
| [thinking_bridge_orig (Qwen3-VL reasoning on Bridge)](https://datasets.gurasees.com/datasets/thinking-bridge-orig-qwen3-vl-reasoning-on-bridge.md) | language and reasoning annotation | single arm | 105 | 53,192 | 2026 | not stated | open | thinking-bridge-orig-qwen3-vl-reasoning-on-bridge |
| [nuReasoning](https://datasets.gurasees.com/datasets/nureasoning.md) | language and reasoning annotation | vehicle | 105 |  | 2026 | unclear | gated | nureasoning |
| [STRIDE-QA](https://datasets.gurasees.com/datasets/stride-qa.md) | language and reasoning annotation | vehicle | 100 |  | 2025 | not stated | open | stride-qa |
| [CoVLA](https://datasets.gurasees.com/datasets/covla.md) | language and reasoning annotation | vehicle | 80 |  | 2024 | custom terms | gated | covla |
| [PEEK VLM-labeled BRIDGE_v2](https://datasets.gurasees.com/datasets/peek-vlm-labeled-bridge-v2.md) | language and reasoning annotation | single arm | 76.8 | 38,660 | 2025 | open | open | peek-vlm-labeled-bridge-v2 |
| [doPlan](https://datasets.gurasees.com/datasets/doplan.md) | language and reasoning annotation | vehicle | 50.9 |  | 2026 | not stated | not stated | doplan |
| [SnapMoGen](https://datasets.gurasees.com/datasets/snapmogen.md) | language and reasoning annotation | other | 43.7 |  | 2025 | custom terms | open | snapmogen |
| [BABEL](https://datasets.gurasees.com/datasets/babel.md) | language and reasoning annotation | other | 43 |  | 2021 | not stated | not stated | babel |
| [Honda HAD (Advice Dataset)](https://datasets.gurasees.com/datasets/honda-had-advice-dataset.md) | language and reasoning annotation | vehicle | 32 |  | 2019 | non commercial | on request | honda-had-advice-dataset |
| [HumanML3D](https://datasets.gurasees.com/datasets/humanml3d.md) | language and reasoning annotation | other | 28.6 |  | 2022 | unclear | not stated | humanml3d |
| [DeepThinkVLA libero_cot](https://datasets.gurasees.com/datasets/deepthinkvla-libero-cot.md) | language and reasoning annotation | simulation | 7.6 | 1,693 | 2025 | open | open | deepthinkvla-libero-cot |
| [RBM-1M (Robometer)](https://datasets.gurasees.com/datasets/rbm-1m-robometer.md) | language and reasoning annotation | cross embodiment | 0.5 | 571 | 2026 | open | open | rbm-1m-robometer |
| [OmniAction](https://datasets.gurasees.com/datasets/omniaction.md) | tactile force audio | other |  | 141,162 | 2025 | non commercial | open | omniaction |
| [MolmoAct Midtraining Mixture](https://datasets.gurasees.com/datasets/molmoact-midtraining-mixture.md) | language and reasoning annotation | single arm |  |  | 2025 | open | open | molmoact-midtraining-mixture |
| [EO-Data1.5M](https://datasets.gurasees.com/datasets/eo-data1-5m.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | open | open | eo-data1-5m |
| [MolmoAct Pretraining Mixture](https://datasets.gurasees.com/datasets/molmoact-pretraining-mixture.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | open | open | molmoact-pretraining-mixture |
| [ViFailback](https://datasets.gurasees.com/datasets/vifailback.md) | language and reasoning annotation | bimanual |  | 5,202 | 2025 | open | open | vifailback |
| [RoboReward](https://datasets.gurasees.com/datasets/roboreward.md) | language and reasoning annotation | cross embodiment |  |  | 2026 | open | open | roboreward |
| [Cap3D](https://datasets.gurasees.com/datasets/cap3d.md) | language and reasoning annotation | other |  |  | 2023 | open | open | cap3d |
| [RoboFAC](https://datasets.gurasees.com/datasets/robofac.md) | language and reasoning annotation | cross embodiment |  | 9,440 | 2025 | open | open | robofac |
| [ShareRobot](https://datasets.gurasees.com/datasets/sharerobot.md) | language and reasoning annotation | cross embodiment |  | 51,403 | 2025 | not stated | open | sharerobot |
| [PointWorld-DROID](https://datasets.gurasees.com/datasets/pointworld-droid.md) | video for world models | single arm |  |  | 2026 | not stated | open | pointworld-droid |
| [VSI-Bench](https://datasets.gurasees.com/datasets/vsi-bench.md) | language and reasoning annotation | sensor only |  |  | 2024 | open | open | vsi-bench |
| [3D-LLM data](https://datasets.gurasees.com/datasets/3d-llm-data.md) | language and reasoning annotation | sensor only |  |  | 2023 | not stated | open | 3d-llm-data |
| [RefSpatial](https://datasets.gurasees.com/datasets/refspatial.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | refspatial |
| [RoboBench (embodied brain benchmark)](https://datasets.gurasees.com/datasets/robobench-embodied-brain-benchmark.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | open | open | robobench-embodied-brain-benchmark |
| [EMMOE-100](https://datasets.gurasees.com/datasets/emmoe-100.md) | language and reasoning annotation | mobile manipulator |  |  | 2025 | open | open | emmoe-100 |
| [MMScan](https://datasets.gurasees.com/datasets/mmscan.md) | language and reasoning annotation | sensor only |  |  | 2024 | not stated | open | mmscan |
| [TraceSpatial (TraceSpatial-Trace)](https://datasets.gurasees.com/datasets/tracespatial-tracespatial-trace.md) | language and reasoning annotation | cross embodiment |  |  | 2026 | open | open | tracespatial-tracespatial-trace |
| [MolmoER / Molmo2-ER training corpus](https://datasets.gurasees.com/datasets/molmoer-molmo2-er-training-corpus.md) | language and reasoning annotation | cross embodiment |  |  | 2026 | non commercial | open | molmoer-molmo2-er-training-corpus |
| [VST (Visual Spatial Tuning) data](https://datasets.gurasees.com/datasets/vst-visual-spatial-tuning-data.md) | language and reasoning annotation | sensor only |  |  | 2025 | not stated | open | vst-visual-spatial-tuning-data |
| [VSI-590K](https://datasets.gurasees.com/datasets/vsi-590k.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | vsi-590k |
| [EQA_DATASET (DoYangTan)](https://datasets.gurasees.com/datasets/eqa-dataset-doyangtan.md) | language and reasoning annotation | other |  |  | 2026 | not stated | open | eqa-dataset-doyangtan |
| [RoboPoint data](https://datasets.gurasees.com/datasets/robopoint-data.md) | language and reasoning annotation | sensor only |  |  | 2024 | open | open | robopoint-data |
| [InternSpatial](https://datasets.gurasees.com/datasets/internspatial.md) | language and reasoning annotation | sensor only |  |  | 2025 | no derivatives | open | internspatial |
| [SenseNova-SI-800K](https://datasets.gurasees.com/datasets/sensenova-si-800k.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | sensenova-si-800k |
| [VLA-IT (InstructVLA)](https://datasets.gurasees.com/datasets/vla-it-instructvla.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | not stated | open | vla-it-instructvla |
| [Robo2VLM-1](https://datasets.gurasees.com/datasets/robo2vlm-1.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | open | open | robo2vlm-1 |
| [RoomTour3D](https://datasets.gurasees.com/datasets/roomtour3d.md) | language and reasoning annotation | human egocentric |  |  | 2024 | open | open | roomtour3d |
| [3DSRBench](https://datasets.gurasees.com/datasets/3dsrbench.md) | language and reasoning annotation | sensor only |  |  | 2024 | open | open | 3dsrbench |
| [RefSpatial-Bench and RefSpatial-Expand-Bench](https://datasets.gurasees.com/datasets/refspatial-bench-and-refspatial-expand-bench.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | refspatial-bench-and-refspatial-expand-bench |
| [Embodied-R1.5 SFT Dataset](https://datasets.gurasees.com/datasets/embodied-r1-5-sft-dataset.md) | language and reasoning annotation | other |  |  | 2026 | open | open | embodied-r1-5-sft-dataset |
| [PointArena (Point-Bench)](https://datasets.gurasees.com/datasets/pointarena-point-bench.md) | language and reasoning annotation | sensor only |  |  | 2025 | not stated | open | pointarena-point-bench |
| [Embodied CoT features and demos for LIBERO](https://datasets.gurasees.com/datasets/embodied-cot-features-and-demos-for-libero.md) | language and reasoning annotation | simulation |  | 3,917 | 2025 | open | open | embodied-cot-features-and-demos-for-libero |
| [EmbodiedMemory-Bench](https://datasets.gurasees.com/datasets/embodiedmemory-bench.md) | language and reasoning annotation | simulation |  | 2,554 | 2026 | non commercial | open | embodiedmemory-bench |
| [Reason-RFT CoT Dataset](https://datasets.gurasees.com/datasets/reason-rft-cot-dataset.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | reason-rft-cot-dataset |
| [FineVLA-Data and RoboFine-bench](https://datasets.gurasees.com/datasets/finevla-data-and-robofine-bench.md) | language and reasoning annotation | cross embodiment |  | 47,159 | 2026 | open | open | finevla-data-and-robofine-bench |
| [InstructPart](https://datasets.gurasees.com/datasets/instructpart.md) | language and reasoning annotation | sensor only |  |  | 2025 | not stated | open | instructpart |
| [WGO-Bench](https://datasets.gurasees.com/datasets/wgo-bench.md) | language and reasoning annotation | other |  |  | 2026 | non commercial | open | wgo-bench |
| [DrivingVQA](https://datasets.gurasees.com/datasets/drivingvqa.md) | language and reasoning annotation | other |  |  | 2025 | open | open | drivingvqa |
| [Behavior-Skill](https://datasets.gurasees.com/datasets/behavior-skill.md) | language and reasoning annotation | mobile manipulator |  | 10,000 | 2026 | open | open | behavior-skill |
| [SpatialRGPT-Bench](https://datasets.gurasees.com/datasets/spatialrgpt-bench.md) | language and reasoning annotation | sensor only |  |  | 2024 | not stated | open | spatialrgpt-bench |
| [Cosmos-Reason1 SFT dataset and benchmark](https://datasets.gurasees.com/datasets/cosmos-reason1-sft-dataset-and-benchmark.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | open | open | cosmos-reason1-sft-dataset-and-benchmark |
| [Embodied-Reasoner](https://datasets.gurasees.com/datasets/embodied-reasoner.md) | language and reasoning annotation | simulation |  | 9,390 | 2025 | not stated | open | embodied-reasoner |
| [MMSI-Bench](https://datasets.gurasees.com/datasets/mmsi-bench.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | mmsi-bench |
| [RoboInter-VQA](https://datasets.gurasees.com/datasets/robointer-vqa.md) | language and reasoning annotation | cross embodiment |  |  | 2026 | not stated | open | robointer-vqa |
| [RoboSpatial-Home](https://datasets.gurasees.com/datasets/robospatial-home.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | robospatial-home |
| [PointMotionBench](https://datasets.gurasees.com/datasets/pointmotionbench.md) | language and reasoning annotation | other |  |  | 2026 | not stated | not stated | pointmotionbench |
| [X-Planner benchmark](https://datasets.gurasees.com/datasets/x-planner-benchmark.md) | language and reasoning annotation | cross embodiment |  | 1,500 | 2026 | not stated | open | x-planner-benchmark |
| [SPAR-7M](https://datasets.gurasees.com/datasets/spar-7m.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | spar-7m |
| [SAGE-3D (InteriorGS scenes and VLN data)](https://datasets.gurasees.com/datasets/sage-3d-interiorgs-scenes-and-vln-data.md) | navigation and mobile | wheeled or navigation |  | 2,000,000 | 2025 | open | open | sage-3d-interiorgs-scenes-and-vln-data |
| [SR-3D-Bench](https://datasets.gurasees.com/datasets/sr-3d-bench.md) | language and reasoning annotation | human egocentric |  |  | 2026 | open | open | sr-3d-bench |
| [WM-ABench](https://datasets.gurasees.com/datasets/wm-abench.md) | language and reasoning annotation | simulation |  |  | 2025 | open | open | wm-abench |
| [NVIDIA PhysicalAI-Traffic-Anomaly-Reasoning](https://datasets.gurasees.com/datasets/nvidia-physicalai-traffic-anomaly-reasoning.md) | language and reasoning annotation | sensor only |  |  | 2026 | open | open | nvidia-physicalai-traffic-anomaly-reasoning |
| [allenai/MolmoAct2-SO100_101-Dataset](https://datasets.gurasees.com/datasets/allenai-molmoact2-so100-101-dataset.md) | teleop | low cost arm |  |  | 2026 | open | open | allenai-molmoact2-so100-101-dataset |
| [NaviTrace](https://datasets.gurasees.com/datasets/navitrace.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | open | open | navitrace |
| [EmbodiedEval](https://datasets.gurasees.com/datasets/embodiedeval.md) | language and reasoning annotation | simulation |  |  | 2025 | open | open | embodiedeval |
| [RoadSocial](https://datasets.gurasees.com/datasets/roadsocial.md) | language and reasoning annotation | sensor only |  |  | 2025 | non commercial | gated | roadsocial |
| [Vlaser data (Vlaser-6M)](https://datasets.gurasees.com/datasets/vlaser-data-vlaser-6m.md) | language and reasoning annotation | cross embodiment |  |  | 2026 | open | open | vlaser-data-vlaser-6m |
| [RACER augmented RLBench](https://datasets.gurasees.com/datasets/racer-augmented-rlbench.md) | language and reasoning annotation | single arm |  | 10,159 | 2024 | open | open | racer-augmented-rlbench |
| [Kimodo Human Motion Generation Benchmark](https://datasets.gurasees.com/datasets/kimodo-human-motion-generation-benchmark.md) | language and reasoning annotation | other |  |  | 2026 | open | open | kimodo-human-motion-generation-benchmark |
| [KITScenes LongTail](https://datasets.gurasees.com/datasets/kitscenes-longtail.md) | language and reasoning annotation | vehicle |  |  | 2026 | non commercial | gated | kitscenes-longtail |
| [libero-r-datasets](https://datasets.gurasees.com/datasets/libero-r-datasets.md) | language and reasoning annotation | single arm |  |  | 2026 | not stated | open | libero-r-datasets |
| [Where2Place](https://datasets.gurasees.com/datasets/where2place.md) | language and reasoning annotation | sensor only |  |  | 2024 | open | open | where2place |
| [Language_Tactile](https://datasets.gurasees.com/datasets/language-tactile.md) | language and reasoning annotation | sensor only |  |  | 2026 | not stated | open | language-tactile |
| [PixMo-Points](https://datasets.gurasees.com/datasets/pixmo-points.md) | language and reasoning annotation | sensor only |  |  | 2024 | open | open | pixmo-points |
| [ERIQ](https://datasets.gurasees.com/datasets/eriq.md) | language and reasoning annotation | other |  |  | 2025 | open | open | eriq |
| [PhysBench](https://datasets.gurasees.com/datasets/physbench.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | physbench |
| [Robo2VLM Reasoning (ManipulationVQA-60k)](https://datasets.gurasees.com/datasets/robo2vlm-reasoning-manipulationvqa-60k.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | open | open | robo2vlm-reasoning-manipulationvqa-60k |
| [FSD-Dataset and VABench](https://datasets.gurasees.com/datasets/fsd-dataset-and-vabench.md) | language and reasoning annotation | single arm |  |  | 2025 | not stated | open | fsd-dataset-and-vabench |
| [Embodied-R1 Dataset (Embodied-Points-200K)](https://datasets.gurasees.com/datasets/embodied-r1-dataset-embodied-points-200k.md) | language and reasoning annotation | cross embodiment |  |  | 2025 | custom terms | open | embodied-r1-dataset-embodied-points-200k |
| [OmniSpatial](https://datasets.gurasees.com/datasets/omnispatial.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | omnispatial |
| [VQASynth SpaceLLaVA](https://datasets.gurasees.com/datasets/vqasynth-spacellava.md) | language and reasoning annotation | sensor only |  |  | 2024 | open | open | vqasynth-spacellava |
| [NILS relabels of Fractal and Bridge](https://datasets.gurasees.com/datasets/nils-relabels-of-fractal-and-bridge.md) | language and reasoning annotation | cross embodiment |  |  | 2024 | not stated | open | nils-relabels-of-fractal-and-bridge |
| [EmbodiedBench](https://datasets.gurasees.com/datasets/embodiedbench.md) | language and reasoning annotation | simulation |  |  | 2025 | not stated | open | embodiedbench |
| [TraceSpatial-Bench](https://datasets.gurasees.com/datasets/tracespatial-bench.md) | language and reasoning annotation | sensor only |  |  | 2025 | open | open | tracespatial-bench |
| [Grounded 3D-LLM dataset](https://datasets.gurasees.com/datasets/grounded-3d-llm-dataset.md) | language and reasoning annotation | sensor only |  |  | 2024 | not stated | open | grounded-3d-llm-dataset |
| [BridgeEQA](https://datasets.gurasees.com/datasets/bridgeeqa.md) | language and reasoning annotation | sensor only |  |  | 2026 | open | open | bridgeeqa |
| [PARTNR episodes](https://datasets.gurasees.com/datasets/partnr-episodes.md) | language and reasoning annotation | simulation |  | 111,652 | 2024 | non commercial | open | partnr-episodes |
| [RONAR (RoboNar)](https://datasets.gurasees.com/datasets/ronar-robonar.md) | language and reasoning annotation | mobile manipulator |  |  | 2024 | open | open | ronar-robonar |
| [DriveLM](https://datasets.gurasees.com/datasets/drivelm.md) | language and reasoning annotation | vehicle |  |  | 2023 | non commercial | gated | drivelm |
| [RynnBrain-Bench](https://datasets.gurasees.com/datasets/rynnbrain-bench.md) | language and reasoning annotation | human egocentric |  |  | 2026 | open | open | rynnbrain-bench |
| [VitaSet](https://datasets.gurasees.com/datasets/vitaset.md) | language and reasoning annotation | single arm |  |  | 2025 | open | open | vitaset |
| [ERQA+ (ERQA Plus)](https://datasets.gurasees.com/datasets/erqa-erqa-plus.md) | language and reasoning annotation | sensor only |  |  | 2026 | open | open | erqa-erqa-plus |
| [MobileVLA-CoT](https://datasets.gurasees.com/datasets/mobilevla-cot.md) | language and reasoning annotation | legged |  |  | 2026 | not stated | open | mobilevla-cot |
| [PRISM-100K (DreamVu)](https://datasets.gurasees.com/datasets/prism-100k-dreamvu.md) | language and reasoning annotation | human egocentric |  |  | 2026 | non commercial | gated | prism-100k-dreamvu |
| [Open Spatial Dataset (SpatialRGPT)](https://datasets.gurasees.com/datasets/open-spatial-dataset-spatialrgpt.md) | language and reasoning annotation | sensor only |  |  | 2024 | not stated | open | open-spatial-dataset-spatialrgpt |

Everything that exists: https://datasets.gurasees.com/browse.md
