# GigaBrain-0.5 pre-training corpus

From the catalogue. Listed by one researcher from one source page.

- id: `gigabrain-0-5-pre-training-corpus`
- kind: synthetic or generated
- robot or device: not stated (class: cross embodiment)
- how collected: world model generated video plus real robot data
- size as stated: 10,000 hours claimed, not released
- hours: not stated | hours only claimed: 10,000 | episodes: not stated
- year: 2026
- organisation: GigaAI
- licence: not stated (class: not stated)
- access: not released (class: not stated)
- link (paper): https://arxiv.org/abs/2602.12099
- note: Over 6,000 of the hours are generated video; the source of the 4,000 real hours is not given and may overlap GigaBrain-0 data and public sets.

Confirm the licence at the link before relying on it.
