Physical AI
First-person task video for robot learning.
People doing real tasks in homes, kitchens, workshops and warehouses, filmed in India from a camera worn on the head. We label the video to your specification, check every clip against it, and deliver it in the format your team loads. You are billed per accepted hour of video.
Nothing is recorded yet. This page sets out the specification we will record to and the check every clip will pass before delivery.
What we record.
First-person video: the task seen from where the person doing it is looking, with their hands and the things they handle in frame.
The camera
Worn on the head and pointed where the wearer looks. If your specification needs a different rig, such as a wrist camera, we tell you in discovery whether we can supply it.
The settings
Homes, home and commercial kitchens, workshops and warehouses, each recorded only with permission: the owner's for a workplace, and that of every adult who lives in a home. Your specification sets the mix.
The tasks
Also set by your specification. For example: preparing food, washing up, folding laundry, sorting and packing items, restocking shelves, assembling parts, using hand tools.
The people
Adults who know the task, each receiving a flat fee. Your specification sets the number of wearers in a batch.
Whole tasks
Each clip is one attempt at one task, recorded from start to finish. If your specification asks for attempts that go wrong, we record those too and label them as failed attempts. A failed attempt that is labelled correctly is data you asked for, not a defect.
The specification.
What comes with every clip. Your specification adds to this list or narrows it, and what is agreed goes into the contract as the standard your sample is scored against.
Labels
- Hands:
- the joints of each hand, and each moment a hand touches, grasps or releases an object.
- Actions:
- every step of the task as a segment, with a start, an end and a name from your list.
- Narration:
- a plain English description of each step, aligned to the video.
- Objects:
- boxes or masks on the objects your specification names, followed through the clip.
- Outcome:
- whether the task succeeded, by the test in your specification.
Formats
Delivered in the format your team loads, with the video, labels and metadata in one package. Send us a sample file and we will tell you in writing whether we can match it.
With every clip
Task and setting IDs, a pseudonymous wearer ID, city, date, the outcome label, the camera and its settings, the calibration file, every blurred area, and the IDs of the releases behind the clip.
Calibration and timing
Every camera is calibrated, and its calibration file ships with every clip; a clip without it fails our check. Every frame carries a timestamp, and motion sensor readings, where the camera records them, are aligned to the video.
Checked by us before it reaches you
Every clip is checked by us against your specification before it is delivered: the task complete from start to finish; the video in focus, with no stretch too blurred to label; no dropped frames; the video, timestamps and sensor readings in sync; the calibration present; and the labels checked against the task and your list. A clip that fails our check is re-recorded or relabelled before you see it. The result of every check ships with the batch, so your team can verify each one.
A datasheet with every batch
Where and when the batch was recorded (without identifying anyone), how it was labelled, the checks it passed, and what it does not cover.
Not in the standard set
Depth and full 3D capture need a different rig. Ask in discovery.
Billed per accepted hour of video.
You are billed only for what you accept. The unit is the clip: one attempt at one task, accepted or rejected whole. The price is per accepted hour of video, so an accepted clip is billed by its length. It measures video delivered, not time we spend.
What counts towards the invoice
A clip is what your loader calls an episode. Its length runs between the task's start and end markers. Setup, calibration and idle footage are not counted, and a task filmed by more than one camera counts once.
Your specification is the standard
It goes into the contract before we record: the tasks, the settings, the labels, the format and the checks a clip must pass.
Our check comes first
Every clip in a batch has passed our check against your specification before the batch reaches you, so your sample tests work we have already checked.
You draw the sample and score it
From every batch, at random, spread across tasks and settings if you want. Your scorer marks each sampled clip pass or fail against the standard. No clip is invoiced until you have accepted it.
A failed clip is reworked at our cost
It is re-recorded or relabelled, goes through the same test again, and reaches the invoice only once you accept it. It is billed once, however many times it was reworked. If a batch fails the sample, the whole batch is held until every clip is rechecked and you have scored a fresh sample.
Exclusive or shared, your choice
You choose before recording. Exclusive: the batch is yours alone, and it is never sold or licensed to anyone else. Shared: the same data may also be licensed to other buyers. The licence states which you chose and for what term.
One line on the invoice
The accepted hours of video in the batch, at the price per accepted hour of video agreed before recording. The licence sets what you may do with the video. It is part of that price, never a separate charge. A disputed clip stays off the invoice while it is settled.
Every batch comes back with its measures, counted by the clip: acceptance rate, rework rate, turnaround, volume delivered, how much of the batch you scored, and disputed clips.
How we measureConsent before we record.
We have not recorded anyone yet. Before we do, all of this will be in place, and you can read the release and the licence before you commit to anything.
A release from everyone recorded
Every person recorded signs a release before recording, in a language they read. It covers the recording, its use to train AI models, and its licence to you and, if the data is shared, to other buyers.
Anyone else in view
Anyone else visible has signed a release or is blurred.
No children
No one under 18 is recorded. If a child appears in any footage, that stretch is deleted and never delivered.
Permission for every place
Every place is recorded with permission: the owner's for a workplace, and that of every adult who lives in a home.
Faces, voices and screens
Faces are blurred before delivery unless the person agreed and your specification needs them. Hands are never blurred. Audio is off, or kept to narration. Screens, documents and labels that show personal details are blurred.
Every clip traceable
Each clip carries the IDs of the releases behind it, so any clip can be traced to its consent record.
Withdrawal
Anyone recorded can withdraw. We then delete their footage from everything we hold and tell you which delivered clips it affects. The licence sets what happens to those clips.
Who sees the footage
Before we record for you, we set up access so that only the people named for your work can see the footage. Every access is logged, and delivery goes into storage you control.
Nothing bought in
Every clip is recorded new, to the specification of the buyer who commissioned it. No footage is scraped or bought from others.
Recorded in India.
Variety
Homes, kitchens and workplaces in India differ widely in layout, tools, light and how the same task is done. That range shows a model more of the ways a task can go.
For US settings
The rooms, fittings and appliances will be Indian ones. If your robot is headed for US settings, send us the objects it must handle: each scene has the objects on your list, or we tell you in discovery which ones we cannot source.
The same check in every setting
However much the settings differ, every clip is checked by us against your specification before it is delivered.
Start with one batch.
A first batch is how you judge us. We agree the specification with you in discovery, record one fixed batch to it, and you score that batch against the standard. You can stop there.
Discovery
We go through what you need and tell you in writing whether we can record it.
One fixed batch
Recorded and labelled to the agreed specification, in one or two settings. Its size and terms are agreed in discovery.
You score it
Against the standard in the contract: the same test every later batch has to pass.
You decide
Carry on at the same standard, change the specification, or stop. The specification is yours to keep either way.
Questions about the video.
Anything else, write to hello@xpusystems.com.
Which format will the video come in?
The one your team loads. Send us a sample file and we will tell you in writing whether we can match it.
Can we have the video to ourselves?
Yes, if you choose exclusive data. Before recording, you choose one of two options. Exclusive: the batch is yours alone, and it is never sold or licensed to anyone else. Shared: the same data may also be licensed to other buyers. The licence states which you chose and for what term.
What does a person on camera agree to?
A written release, signed before they are recorded, in a language they read. It covers the recording, its use to train AI models, and its licence to you and, if the data is shared, to other buyers. Each clip carries the IDs of its releases, and you can read the release before you commit.
Will blurring get in the way of our model?
Hands are never blurred. Faces are, by default, and every blurred area is marked in the clip's metadata so your pipeline can leave it out. If your model needs faces, say so in discovery; each person then has to agree to it in their release.
What if someone withdraws after we have the video?
We delete their footage from everything we hold and tell you which delivered clips it affects. The licence sets what happens to those clips on your side, and you can read it before you commit.
What do you keep after delivery?
For exclusive data, only what rework needs, deleted once you accept the batch, plus the consent records behind each clip. For shared data, we also keep the batch itself, because it may be licensed to other buyers. Either way, the licence states what we keep.
How big is the first batch?
One fixed batch in one or two settings. Its size is agreed in discovery, so that your scorer has enough to judge the work.
How is the price set?
Per accepted hour of video, set in discovery and agreed in writing before recording starts. What moves it: the labels you need, how many settings and wearers, how hard the tasks are, and whether the data is exclusive or shared. Once agreed, it does not change for that batch.
Where do you record?
In India, in homes, kitchens, workshops and warehouses. A place is recorded only with permission: the owner's for a workplace, and that of every adult who lives in a home. The cities and settings for your batch are agreed in discovery, and each clip's metadata records its city and setting.
Do you have a SOC 2 report?
Not yet. Before we record for you, we set up access so that only the people named for your work can see the footage. Every access is logged, and delivery goes into storage you control.
How soon can you start?
Recording starts once your specification is agreed and the releases and licence are in place. You get a start date for the first batch in discovery, and the turnaround for later batches goes into the contract beside the standard.
Tell us what you are training.
A rough answer is enough. We work out the rest with you in discovery.