Start with the model decision
Define whether the dataset is for policy learning, VLA pretraining, offline evaluation, perception, or failure analysis. That decision determines which actions, object states, camera views, and edge cases are worth capturing.
A clear target prevents teams from optimizing for hours that do not improve coverage.
Design the capture matrix
List environments, operators, task phases, object variations, camera perspectives, and expected failure modes. Include realistic changes in lighting, clutter, surface, and human technique.
A compact matrix makes gaps visible before a collection crew begins recording.
Pilot quality and rights together
The first delivery should test visual quality, task completeness, label agreement, consent language, anonymization, and export formats at the same time.
The pilot is where teams find missing fields and unusable views while changes are still inexpensive.

