# Robot dataset index

> Search and list robotics datasets. Two lists behind one address. The catalogue: 3,494 datasets researched one at a time, each with size, licence, access and a link. The Hub list: all 86,424 robotics uploads on the Hugging Face Hub, typed from their own metadata (robot, episodes, hours, licence). Between them they list 2,211,667 hours of data. Read only, free, no key, no login. Compiled October 2026.

This page is for agents. People: https://datasets.gurasees.com/

## Connect

Use the MCP endpoint if you can call tools. It is the shortest route.

- MCP endpoint: `https://datasets.gurasees.com/mcp` (streamable HTTP, no authentication)
- Claude Code: `claude mcp add --transport http robot-datasets https://datasets.gurasees.com/mcp`
- Any MCP client, as JSON: `{"type": "http", "url": "https://datasets.gurasees.com/mcp"}`
- ChatGPT: Plugins, then the plus button, then "Create custom MCP server", server URL `https://datasets.gurasees.com/mcp`, authentication "No authentication"
- OpenAI Responses API tool: `{"type": "mcp", "server_label": "robot_datasets", "server_url": "https://datasets.gurasees.com/mcp", "require_approval": "never"}`
- JSON API: `curl "https://datasets.gurasees.com/api/search?q=humanoid+teleoperation+over+100+hours&limit=5"`, described at https://datasets.gurasees.com/openapi.json
- No tools, only page fetches: every page linked below is plain markdown

## MCP tools

- `search_datasets`: the whole question as a sentence. Returns the catalogue's answer and the matching Hub uploads
- `list_datasets`: catalogue datasets by exact filters, lean rows, paged. For "all of one kind"
- `list_hub_uploads`: Hub uploads by robot, kind, episodes, hours and licence, any sort, paged to the end
- `get_dataset`: one dataset in full by id or Hub repository id, with its copies on the Hub
- `describe_index`: what kinds, robot classes and robots exist, with counts. Call it before listing
- `search`: the same search under the name ChatGPT deep research requires
- `fetch`: one dataset as text, under the name ChatGPT deep research requires

## Pages

- Everything that exists, with counts and a link to each list: https://datasets.gurasees.com/browse.md
- Search both lists with a sentence: https://datasets.gurasees.com/search.md?q=humanoid+teleoperation+over+100+hours
- Hub uploads for one robot, largest first: https://datasets.gurasees.com/hub.md?robot=so101&kind=recording&sort=episodes
- The same for an exact robot family, as counted on the browse page: https://datasets.gurasees.com/hub.md?robot_family=SO-101&sort=episodes
- One dataset: https://datasets.gurasees.com/datasets/droid.md
- A Hub upload: https://datasets.gurasees.com/datasets/lerobot/pusht.md

## How a sentence is read

Conditions in the sentence become hard filters: "over 500 hours", "under 100 episodes", "since 2025", "between 2020 and 2023", "open licence", "not gated", "on Hugging Face", "not on Hugging Face". Topics (humanoid, tactile, egocentric) and robot names (Franka, SO-101, Unitree G1) steer the ranking. A negated topic or name ("no simulation", "not Unitree") removes those rows. "largest", "newest", "most downloaded", "smallest" and "oldest" set the sort.

Every answer says how it was read. Check it. If a condition was missed, pass it as an explicit filter and not in the sentence.

## To list all of one thing

A search returns the best matches. A complete list needs filters and paging.

- Catalogue, one kind: https://datasets.gurasees.com/kinds/teleop.md (or `list_datasets`, or `/api/search?category=teleop&sort=hours&limit=100&offset=0`)
- Hub, one robot: https://datasets.gurasees.com/hub.md?robot=so101&kind=recording&sort=episodes (or `list_hub_uploads`, or `/api/hub?robot=so101&kind=recording&limit=500&offset=0`)
- Every list states its total and prints the address of the next page. Follow it until the page says it is the end.

## What the fields mean, and how far to trust them

Catalogue rows were each listed by one researcher from one source page and not cross checked. Confirm a licence at the dataset's own page before relying on it.

- `hours`, `episodes`: what the source says exists. Empty when it does not say.
- `hours_claimed`: a figure that is announced, planned, unreleased or stated only by a third party. Never used by a filter or a sort.
- `licence_class` is what you may do with it. `access_class` is whether you can get it. They are separate, and an open licence on a gated dataset is common.
- `checked` false: the source page could not be opened.
- `link_kind`: whether the link is the dataset, a paper about it, or a listing page.

Hub rows were read by no person. `robot`, `episodes` and `hours` come from the upload's own `meta/info.json` when it has one. `kind` is inferred from the repository name: `evaluation rollout` for names starting with eval, `test or scratch`, `empty` for zero episodes, `recording` otherwise. `category` follows from that. A fifth of uploads state no robot.

## Limits

300 requests a minute for one caller. A refusal is status 429 with `Retry-After` in seconds. Catalogue lists hold up to 100 rows a page, Hub lists up to 500.
