Robots will work for hours. Train them on hours.

Physical Data records people doing real jobs, start to finish, for one to four hours at a time. They wear a head camera and two wrist cameras and talk through what they're doing: the plan, the mistakes, the fixes, the changes of mind. We do it in the same homes again and again. You get the video, the transcript and the labels, in your format.

Request a sample session
Rec Session 0412 Kitchen 07 · Run 23 of 24 Job: reset a kitchen after a dinner for eight
01:47:12/ 03:14:08
PlanPhase 05 of 09
    NarrationNow · wipe
      MemoryRecalled from earlier
        CamerasSynced · 30 fps
        Head1920×1080
        Wrist L1280×720
        Wrist R1280×720
        Illustrated session, built from our task design for a kitchen reset. Drag the timeline to scrub. Arcs on the timeline connect a thing seen to the moment it was recalled
        What's in a session

        Motor skills learn from minutes. Judgment learns from hours.

        Short, clean demonstrations teach a robot to grasp and place. They do not teach it to run a kitchen for three hours, remember what it left soaking, or change the plan when the trash is full. That is what we record.

        Hours, not minutes.

        Whole jobs, start to finish, one to four hours, no cuts. The session above is a kitchen after a dinner party: 3 h 14 min, nine phases, subtasks.

        1 h 2 h 3 h 30 s typical demonstration 15 min longest published robot memory 3 h 14 min one Physical Data session
        Drawn to scale on a 4 hour line

        Mistakes, kept.

        Perfect demonstrations show success, never the way back to it. Nothing is re-shot. Every slip stays in, and the recovery is labeled, so a model can learn to notice and fix, not only to repeat.

        Reasoning, out loud.

        Collectors talk through the job in their own words: what's next, what they remember, when they change the plan and why. narrated decisions in this session, each aligned to the frame it happened in.

        Same place, many times.

        The same job in the same home on different days, with different clutter, light and people. Then other homes. A model learns what can happen in a place, not one lucky run.

        Run lengthMistake keptPlan changed
        How a session is run

        Five steps. The third one is the whole point.

        01

        Brief

        You tell us the job, the environment and what you want narrated. We turn it into a one-page brief a collector reads once.

        02

        Rig

        A head camera and one on each wrist, time-synced. Light enough that nobody works differently because they're wearing it.

        03

        Record

        The collector does the whole job, one to four hours, no cuts, talking through it the way you'd train a new hire. When something goes wrong, they keep going. Nothing is re-shot.

        04

        Repeat

        The same home again on other days, then other homes. Same job, different mess. Twenty or more runs per place is normal.

        05

        Label and deliver

        Transcripts aligned to video, subtask spans, mistakes, recoveries, plan changes, memory recalls, hand keypoints. Checked by a second person before it ships.

        What you get

        One folder per session. Nothing you have to guess at.

        session_0412/
        ├─ head.mp4          1920×1080 · 30 fps · 3 h 14 min
        ├─ wrist_l.mp4       1280×720 · synced to head
        ├─ wrist_r.mp4
        ├─ narration.jsonl   {t, text, type}
        │                    plan decision memory replan mistake recovery
        ├─ subtasks.jsonl    {t0, t1, phase, label}
        ├─ events.jsonl      {t, type, note, links}
        │                    mistake → recovery · memory → seen_at
        ├─ hands.npz         21 keypoints × 2 hands, per frame
        ├─ sync.json         per-camera offsets, dropped frames
        └─ meta.json         home_id, run_id, collector_id, consent_id, checks
        Formats
        RLDS, LeRobot, or the schema you already use. We match your existing data before we scale.
        Narration
        Transcribed and timestamped to the frame. Tagged by type. Spoken in the collector's own words, never written after the fact.
        Labels
        Subtask spans, phase boundaries, mistakes linked to their recoveries, memory recalls linked to what was seen earlier, plan changes with before and after order.
        Quality
        Every session reviewed by a second person. Sync drift, dropped frames and untranscribed stretches are listed in meta.json, not hidden.
        Consent and privacy
        Written consent from the collector and the household. Bystander faces and documents blurred. Stricter rules on request.
        Sample
        One full session, free, on a job you choose, so you can try it in your pipeline before committing to anything.
        Straight answers

        Things people ask on the first call.

        Do you cut out the mistakes?
        No. That is most of the reason to buy this. A dropped glass, a wrong cupboard, a plan that had to change: it stays in, and the recovery is labeled. If you want a clean-only cut as well, we can deliver both.
        Is the narration scripted?
        No. Collectors get a one-page brief and a rule: talk like you're training a new hire. Say what you're about to do, what you're checking, and why you're changing your mind. The words are theirs.
        Who are the collectors?
        People we recruit, train and pay ourselves, working in their own or partner homes. Our team is based in India, which is what makes hundreds of hours affordable. We can run in other countries when a project needs it.
        Can we choose the jobs and the environments?
        Yes, that is the point. Kitchens, laundry, whole-apartment cleaning, packing, restocking, a workshop, a small shop. You choose the job and how many homes and repeats you want. If a job is unusual we will tell you honestly whether we can do it well.
        How fast can we see data?
        A sample session within two weeks of a brief. After that, delivery runs weekly.
        How is it priced?
        Per delivered hour, with volume tiers. The sample is free. We will put a number in writing before you commit to anything.

        Which job should your robot learn next?

        I'm Bhuwan. I've run data collection and annotation teams for AI labs before, and I started Physical Data because the data robots need next doesn't exist yet: long, messy and narrated. We want robots in people's homes as much as you do. Tell me the job, and we'll record it.

        Bhuwan Sai Kosuru
        Write to me directly

        Tell us the task, the environment and the format you want. We reply within five hours with a plan for a free sample session.