Use Case

Robotics & Manipulation

Robots learn manipulation best from human demonstrations. Payana turns real-world, first-person footage of people doing tasks into robot-ready training episodes - the imitation-learning data behind robotics foundation models.

Egocentric (head- or body-mounted) video is the closest human analogue to what a robot's own camera sees: the workspace, the objects, and the hands manipulating them, all from the actor's point of view. Payana captures this footage with wearable cameras while workers simply do their jobs - no scripting, no staging, no lab setup.

Each recording flows through an automated episode pipeline: a vision-language model narrates the action, detection and open-vocabulary models find every object and tool, segmentation produces pixel-accurate masks, speech becomes transcripts, and hand-pose estimation extracts a 7-dimensional end-effector state per frame - position, orientation, and gripper aperture, the same action space robot policies are trained in.

The result is demonstration data that pairs observation with action at every timestep - what the person saw and what their hands did - which is precisely the supervision imitation learning and behavior-cloning policies need. Long recordings of repeated work are automatically segmented into individual task cycles, so one shift of footage becomes hundreds of training episodes.

Every episode is human-reviewed before delivery, and off-task segments (breaks, interruptions, camera fumbles) are flagged by automatic curation so they never contaminate a policy's training set.

How it works

  1. 1

    Capture

    Workers wear head-mounted or body cameras and do their normal tasks; voice notes record intent as it happens.

  2. 2

    Episode pipeline

    A 7-model GPU stack annotates every second: actions, objects, pixel-accurate masks, transcripts, and 7-D hand/end-effector state.

  3. 3

    Segment & review

    Long recordings auto-split into task cycles; humans review every episode and curation drops off-task segments.

  4. 4

    Robot-format export

    Episodes export as LeRobot v3 and RLDS / Open X-Embodiment, with train/valid/test splits.

What you get

  • Egocentric capture with per-second annotation
  • 7-D hand / end-effector state per frame
  • Automatic segmentation of long footage into task episodes
  • Human review of every episode + off-task curation
  • LeRobot v3 and RLDS / Open X-Embodiment exports

Episodes export as LeRobot v3 and RLDS / Open X-Embodiment style datasets - the formats robotics foundation models and imitation-learning frameworks load natively - plus Parquet and JSONL when you need custom pipelines.

Frequently asked questions

Imitation learning (behavior cloning) trains a robot policy from demonstrations: sequences that pair what was observed at each timestep with the action that was taken. Payana produces this from human work - egocentric video paired with per-frame hand/end-effector state - so robots can learn manipulation from people rather than from expensive robot teleoperation.

Have footage to annotate?

Tell us your task and target format - we handle capture, annotation, review and delivery end to end.

Book an intro call