
The structure
beneath the motion.
Every movement leaves evidence of intention, constraint, and consequence.
A hand reaches. An object moves. A plan encounters resistance. Within that sequence lies something deeper than visual recognition: the relationship between what an agent perceives, what it does, and what changes because of it.
We are building the datasets and World Action Model to study that relationship. Our ambition is physical intelligence grounded in the world as it is experienced, not only as it is described.
The physical world has a hidden curriculum.
The real task unfolds between the instructions.
An instruction names an objective. It does not contain every perceptual judgment, spatial adjustment, or recovery decision required to achieve it.
It exists in the correction before an object slips, the decision to change a grasp, the recognition that a process has failed, and the sequence of actions that brings it back on course.
Our research begins with a hypothesis: continuous human work contains learnable structure that can help machines represent physical state, anticipate action effects, and develop more transferable behavior.
We are building the infrastructure to make that structure accessible, and the models to investigate what it can teach.
Datasets built from the world at work.
Human experience is our starting point.
Our dataset program centers on authorized industrial experience, with an initial focus on warehouse and logistics workflows: receiving, picking, packing, replenishment, inventory handling, palletization, and loading.
Through Dataset Capture and Dataset OS, our approach connects egocentric observation with episode structure, hand-object interactions, task states, and outcome evidence. The objective is not simply to accumulate footage. It is to preserve the information that makes an experience useful for learning.
- SET / 01
Continuous egocentric experience
We prioritize the continuity of work: how a task begins, how attention shifts, how objects move through a process, and how an operator responds when conditions change. Our collection strategy seeks variation in object presentation, workspace configuration, occlusion, task order, and operational context. The research question is what remains transferable across those variations.
- SET / 02
Structured human-object interaction
Our annotation program is designed around temporally linked entities and events: hand keypoints, object tracks, interaction phases, spatial relationships, and task boundaries. We want to represent more than the presence of a hand and an object. We want to preserve their evolving relationship throughout an operation, including where the evidence becomes ambiguous.
- SET / 03
Operationally grounded outcomes
Where authorized and available, operational events provide context beyond the camera: scans, inventory movements, workflow transitions, and completion records. A visible movement and a completed task are different labels. Our approach preserves that distinction, linking physical observations to process evidence rather than treating appearance alone as proof of success.
- SET / 04
Failure, interruption, and recovery
We seek the moments when the expected sequence breaks: an incorrect item, an obstructed path, an incomplete transfer, an interruption, or a correction. These sequences are central to our research into adaptive behavior. We want to understand not only how an action succeeds, but how a system can recognize that its current strategy is no longer working.
The exception is part of the experience.
From recordings to World-Action Episodes.
The unit of learning is not just a frame. It is a transition.
Our World-Action Episode specification connects observations across time with the entities, actions, state changes, and evidence needed to interpret them.
- Observation
- Time-indexed video, available sensor streams, viewpoint information, and synchronization metadata.
- Interaction
- Hand and object tracks, estimated keypoints, interaction phases, visibility, and occlusion.
- Task state
- Instructions, substeps, object relationships, process context, and available operational events.
- Outcome
- Completion evidence, partial outcomes, failures, interventions, corrections, and unresolved states.
- Provenance
- Source lineage, annotation methods, confidence and review status, permitted use, and release version.
This is a collection-dependent specification, not a claim that every episode contains every modality.
Estimated geometry remains distinguishable from measured signals. Inferred events retain their uncertainty and review status. Missing evidence remains missing rather than becoming an artificial label. These distinctions are part of the representation itself.
We want the structure to become richer without the evidence becoming less honest.

Introducing our World Action Model research.
Beyond recognizing the world. Toward acting within it.
Dataset's World Action Model, or WAM, is our research program for connecting perception, prediction, and action.
The program combines industrial video, robot-grounded data, and simulation to investigate world prediction, action learning, memory, and recovery. These are active development directions, not a claim of a completed general-purpose autonomous system.
Can a model learn a representation of the world that is useful not only for describing what happened, but for deciding what to do next?
We are pursuing a modular approach that connects vision-language-action policies with predictive world representations, embodiment-specific interfaces, and governed memory.
- WAM / 01
Represent the state that matters.
We are investigating representations that connect visual observations with task context, object relationships, and interaction history. The objective is to encode the information relevant to action: what has changed, what remains unresolved, which constraints are active, and which parts of the environment are uncertain.
- WAM / 02
Predict the consequences of action.
Our world-model research explores action-conditioned prediction in latent space. Rather than making photorealistic video generation the default objective, we are interested in representations that preserve task-relevant future structure: object configurations, progress, likely failure, and the consequences of candidate actions. A compelling future image is not our endpoint. A prediction that helps improve a decision is.
- WAM / 03
Generate actions grounded in an embodiment.
Our action-learning direction centers on vision-language-action policies, temporally extended action sequences, and embodiment-specific adaptation. Human experience provides a source of interaction structure. Robot-grounded trajectories and controlled simulation provide a means to investigate how that structure connects to executable behavior. Transfer across bodies is a research problem to be measured, not an assumption built into the name.
- WAM / 04
Use memory without confusing it with understanding.
We are exploring memory as a mechanism for retrieving relevant demonstrations, contextualizing unfamiliar situations, and informing recovery. The research challenge is not simply to remember more. It is to retrieve experience that is applicable to the current state, distinguish useful precedent from misleading similarity, and recognize when prior experience is insufficient.
- WAM / 05
Recover when prediction meets reality.
Our goal is a closed loop in which observed outcomes are compared with expected outcomes, discrepancies are detected, and behavior can be reconsidered within defined constraints. Recovery, intervention, and abstention belong inside the research agenda. They are not separate from competent action. The model we are working toward should not only propose a next move. It should help determine whether that move still makes sense.
Two coupled questions.
Action selection and action-conditioned prediction, stated separately.

Here, zt represents the current state inferred from available observations and context; g is the goal; e describes the embodiment; and mt is relevant memory. The policy proposes an action sequence over horizon H, while the predictive model estimates its consequences. This formulation expresses our research direction, not a claim that every component has been jointly trained or validated.
The deeper question lies in their interaction: whether prediction can improve action selection, whether action experience can improve representation, and whether both can remain useful when the environment changes.
What should happen next? What would change if it did? What evidence would tell us we were wrong?
Generalization must survive contact with evidence.
Our evaluation agenda extends beyond reproducing familiar demonstrations.
We are interested in held-out tasks, unfamiliar object arrangements, new operating conditions, and supported embodiment changes under explicitly limited adaptation budgets.
We intend to evaluate action quality alongside task completion, recovery success, intervention frequency, constraint violations, uncertainty calibration, and inference cost. Comparisons against pretrained and adaptation baselines are necessary to isolate what Dataset's data and methods actually contribute.
Human observation, simulated interaction, and physical robot execution remain distinct evidence categories. Training progress or simulation results do not, by themselves, establish real-world robotic capability.
A convincing demonstration is a beginning. A reproducible result is a stronger foundation.
Better models should reveal what to observe next.
Our long-term research strategy connects data acquisition with model evaluation.
Authorized experience informs dataset construction. Controlled experiments expose weaknesses in representation, prediction, and behavior. Those weaknesses can guide targeted collection: the missing variation, the underrepresented interaction, the recovery sequence the model has not learned.
The goal is a more deliberate learning cycle, in which new data addresses specific uncertainty rather than simply increasing volume.
Dataset Capture, Dataset OS, DataFace, and Dataset API form the surrounding product architecture for collection, governed data production, research access, and integration. Any use of customer data or model-failure traces remains subject to the permissions governing that material.
We want each experiment to make the next question sharper.
There is more intelligence in ordinary work than we have learned to represent.
It is present in the adjustment that prevents a mistake, the sequence that keeps a process moving, and the judgment that tells a person to stop and reconsider.
We are building a research program around making that experience legible to machines without stripping away its context, uncertainty, or consequence.
Our ambition is to help move physical intelligence beyond recognition and repetition, toward systems that can anticipate, adapt, and act with a deeper understanding of the situations they encounter.
The structure is there. We are learning how to read it.
Build with Dataset.
We welcome collaboration with AI laboratories, robotics teams, academic researchers, and industrial partners working on multimodal representation learning, world models, action policies, and embodied evaluation.
Dataset availability, annotation coverage, and permitted uses vary by collection and release. WAM capabilities described on this page are under research and development.
← Back to Dataset