Model Support

The right data for your model

VLA models, world-action models and world models each demand different data. We capture first-person and multi-camera stereo data in real scenes, annotated with action primitives, task instructions and state changes — aligned directly with model training needs.

10,000h+
hours of first-person stock
100+
standardized task types
4-5
camera stereo network
6+
device form factors
Paradigms

Three model paradigms, three data recipes

From VLA to world models, each paradigm needs its own data shape — so we organize capture and annotation accordingly.

01

VLA: Vision-Language-Action

Taking vision and language instructions as input and directly outputting action sequences, VLA training requires strictly aligned vision-language-action trajectories: frame sequences, action primitives, natural-language instructions and temporal segmentation.

  • Frame indices strictly aligned with the video timeline, traceable frame by frame
  • Action-primitive labels paired with natural-language task instructions
  • Complete scene and station IDs, re-mixable by task family
02

World-Action Models

A new paradigm jointly modeling future video and future actions. Dyna-2, pre-trained on over one million hours of first-person human manipulation video, demonstrated human-to-robot transfer scaling laws — real human operation video is itself a pre-training asset.

  • Head-mounted mono (GoPro 13+) matches the first-person capture shape
  • Daily manipulation task families — cooking, tidying, assembly — captured continuously
  • Long real-scene sequences, scalable by task family
03

World / Physical World Models

Learning how scenes evolve with interaction: how objects respond to contact and how states change. Training needs multi-view spatial information and operation sequences with real causal structure.

  • 4–5 camera stereo networks restore full spatial information
  • "Production-meaning" rule: every action produces a visible state change
  • Environmental diversity: varied tablecloths, materials and staging
Mapping

From model needs to data supply

# Paradigm Key Data Need Our Support Recommended Datasets
01 VLA Models Aligned vision-language-action trajectories Full annotation: frames + action primitives + instructions Industrial Line Ops · Home Service Interactions
02 World-Action Models Large-scale first-person human video 10,000h+ stock, scalable by task family Multimodal Stock Library
03 World Models Scene evolution & physical interaction Stereo network + state-change annotation Retail Shelf Ops · Robot Arm Teleop
04 Post-training / Fine-tuning Small amounts of high-quality robot data Pilot → batch delivery with QC reports Custom Collection & R&D
Why Fit

Why our data fits model training

01

First-person Shape

Head-mounted mono rigs closely match human operation recording, ready as first-person pre-training corpora.

02

Stereo Spatial Info

Multi-camera networks remove blind spots, supporting spatial understanding and reconstruction tasks.

03

Complete Supervision

Action primitives, temporal segmentation, task instructions and state changes — annotation specs open for review.

04

Scalable Supply

6 own bases and a multi-form device fleet, scaling by task family with controllable delivery.

A data plan tailored to your model

Tell us your model paradigm and training stage — we will recommend stock datasets or design a targeted collection plan.

Book a solution call
Note: Dyna-2 is a world-action model published by Dyna Robotics, pre-trained on over one million hours of first-person human video; cited here only as a paradigm example, with no affiliation.