A large dataset is not automatically a useful dataset.
For physical AI, the quality of captured motion can determine whether a model learns meaningful manipulation behavior or simply learns a visually convincing approximation.
This creates what we call the fidelity floor: the minimum level of information a dataset must preserve to remain useful for manipulation learning.
Finger Articulation Matters
Human manipulation is highly dependent on the fingers.
The difference between pushing, grasping, pinching, rotating, and repositioning an object can involve only small changes in finger configuration.
A capture system that records only coarse hand positions may lose these distinctions.
For robotics, that information can be critical.
Finger articulation helps describe:
- Grip formation
- Finger-object contact
- Grasp transitions
- Tool interaction
- Object rotation
- Fine manipulation
- Release behavior
6-DoF Motion
Position alone does not fully describe physical movement.
A useful representation needs to capture six degrees of freedom:
3D position + 3D orientation.
This is particularly important when understanding how a hand approaches an object, rotates a tool, aligns with a surface, or moves an object through space.
Small orientation changes can completely change the outcome of a manipulation task.
Contact Is a Signal
Some of the most important moments in a physical task happen when objects touch.
A hand approaches a surface.
A finger presses a button.
A tool contacts a material.
A container is grasped.
An object is released.
These contact events establish relationships between movement and the physical environment.
If the dataset does not preserve enough information around these events, a model may see what happened without understanding when and why the interaction occurred.
Convenience vs. Fidelity
There is always a temptation to simplify data collection.
Fewer sensors can make deployment easier. Lower-resolution tracking can reduce complexity. Video-only capture can dramatically increase scale.
But every simplification introduces the possibility of losing information.
The right question is therefore not:
“How little data can we capture?”
It is:
“What is the minimum data required to preserve the behavior?”
Designing for the Learning Objective
A high-quality capture system should be designed around downstream model requirements.
For manipulation, that means preserving the signals that describe:
perception → approach → contact → manipulation → release.
When these relationships remain intact, datasets become more than collections of videos. They become structured representations of physical behavior.
That is the fidelity floor.