Buying and Selling Teleoperation Data
Teleoperation data is the workhorse of the robot-data market. How it is captured and priced, how to judge its quality, and how buying and selling it works.
A buyer asks three vendors for five hundred hours of a humanoid stacking dishes. The quotes come back at very different numbers, and the cheapest batch turns out to be nearly worthless. Same task, same hour count, wildly different value. That gap is teleoperation data seen as a market good, where the label on the box tells you almost nothing about what is inside.
Teleoperation is the workhorse of the robot-data market. Most of the demonstration data that trains manipulation policies today comes from a person driving a robot while the system records what the robot did. It is the closest thing the field has to a repeatable line for producing action data. Corpora like DROID were assembled this way across many labs, and open tooling such as Hugging Face LeRobot made publishing a teleoperation dataset cheap. Cheap to produce is not the same as valuable to train on, and that difference is the whole trade.
How teleoperation hours get made
Before you can judge a listing, you need to know how the hours were captured, because the method leaves a fingerprint on the data. Three styles dominate the market, and each ships a different kind of signal at a different cost.
Leader-follower rigs put a small replica arm in the operator's hands. The real robot mirrors it, and the recording captures action and proprioception on the exact arm a buyer may want to train. VR teleoperation is different: the operator wears a headset, sees through the robot's cameras, and drives it from inside the scene. Exoskeleton and wearable rigs read the human's joint angles directly, at high frequency, which is rich but has to be retargeted before another robot can use it. None of these is strictly better. They are better for different buyers.
Throughput drives the supply side. A single operator can capture a few hundred usable demonstrations a day on a good rig, fewer on a hard task, and that rate sets a floor under any price. It also explains why so much cheap teleoperation data exists, and why so little of it is worth buying twice.
| Capture method | What the data looks like | Market value driver |
|---|---|---|
| Leader-follower | Action and joint traces on the target robot, low latency, tightly matched | High transfer to the same embodiment, but locked to that robot |
| VR, camera-in-the-loop | Operator intent through robot cameras, natural viewpoints, some control lag | Scene diversity, though buyers check for teleop latency baked into actions |
| Exoskeleton or wearable | Direct human joint angles at high frequency, contact-rich potential | Rich kinematics, value depends on clean retargeting to a robot |
| Hand-tracking add-on | Fine finger motion and grasp detail layered on any of the above | Raises price for dexterous tasks, thin without force data |
Read the table as a map of tradeoffs, not a ranking. A leader-follower recording of your exact robot can be worth far more per hour than a generic VR session. A VR session across a hundred kitchens can be worth more for pretraining than a narrow leader-follower set. Price follows fit, and fit is specific to the buyer.
Why two datasets of the same size are not worth the same
Buy by the hour and you will overpay. A recorded hour is padded with idle time, reaching between objects, botched attempts nobody labeled, and long stretches where the operator was thinking. Strip that away, and the useful demonstration inside a nominal hour can be a fraction of it. Two vendors can each sell you one hour and hand you very different amounts of the thing a model learns from.
Then there is what the capture left out. Many teleoperation setups record video and joint positions but no force or torque, so contact-rich skills arrive without the signal that makes them learnable. Clocks drift, and a camera stream a few frames out of sync with the joint stream quietly corrupts every action label. Most fleets throw away the failures, which are exactly the recoveries a robust policy needs. A dataset like RH20T is interesting precisely because it was designed as structured, multimodal teleoperation rather than a pile of clips, and that design is what a serious buyer pays for.
None of this shows up in a preview. A thirty-second highlight looks the same whether the dataset behind it is clean or a mess, which is why the preview is the wrong thing to judge.
A teleoperation dataset is priced by the hour and valued by the demonstration. The two numbers are rarely the same.
What a buyer checks, and what a seller has to prove
Careful buyers converge on the same short list. Can I inspect a random slice before I commit, not a curated highlight reel? Does a manifest of hashes match the bytes that actually arrive? Was this captured on an embodiment close enough to mine to transfer? And is there a consent and provenance record I can defend after my model ships? If a seller cannot answer those, the listing is unpriced, whatever the sticker says.
There is almost no market machinery to fall back on here. No rating agency grades trajectories. No escrow releases payment only after the data proves out in training. No standard warranty covers a dataset that underdelivers. Until that scaffolding exists, the checks are the buyer's own job, and the discount for skipping them lands on the honest seller.
The seller's task is the mirror image. Raw footage becomes an asset when it ships with clean segmentation, documented rigs and capture rates, a schema a training pipeline already reads, and a consent basis a lawyer can sign. Industry trackers such as The Robot Report and ongoing coverage at IEEE Spectrum keep showing the same demand for manipulation data, so sellers who solve trust will not lack buyers. The winners are not those with the most hours. They are the ones whose hours are legible.
The provenance wedge
Human teleoperation is a person's movement, captured and sold. That makes it personal data, and it drags a legal question into every transaction. From August 2, 2026, providers of high-risk AI systems in Europe face data-governance duties under the EU AI Act that reach into the training set: where the data came from, how it was collected, how it was annotated. For a teleoperation seller that is not a tax. It is a way to charge more. A dataset with per-session consent, a documented jurisdiction, and an audit trail from operator to sample is procurement-ready in a way a folder of anonymous clips is not.
The practical read for both sides is the same. A buyer should shop for legible demonstrations, not raw hours: a slice to test, an embodiment to reuse, and a title to defend. A seller should build those properties in from the first recording, because they are cheap to design in and expensive to add later. Teleoperation will stay the workhorse of this market for years. The teleoperation data that earns a premium will be the data whose story a buyer can check.