Language models got the internet — a few trillion words of free training data, already lying around. Robots got nothing comparable. Every demonstration of a physical task has to be collected by someone, on a real machine, one at a time. That gap is now a market, and the money attached to it is getting strange.
XDOF — the robotics data startup that emerged from stealth barely three months ago — is in late-stage talks for a Series B at a valuation of roughly $1.2 billion, TechCrunch reported on September 4, with 8VC expected to lead. The round size isn’t confirmed and terms could still change. What makes the number notable isn’t the size alone: it would come maybe a quarter after a $70 million Series A in June, and it values a company with about 20 customers on reported annualized revenue approaching $50 million. That’s a revenue multiple in the twenties, applied to a company selling data collection.
What XDOF actually sells
The company’s core business is producing the demonstrations robots need: teleoperation recordings where humans drive real machines, sensor-based captures, and the annotation systems to make them trainable. Its customers include frontier AI labs — the same labs building the foundation models everyone expects to power the next generation of humanoids. Rivals circling the same demand include Scale AI, Mecka AI and Micro1.
Two things separate XDOF from being another data lab with contractors. The first is lineage: the company grew out of GELLO, the low-cost teleoperation rig that let researchers collect manipulation data on cheap hardware instead of six-figure arms.
The second is the open dataset. With UC Berkeley and Amazon’s FAR group, XDOF released ABC — 130,000-plus episodes, 3,553 hours, 195 bimanual tasks, published on Hugging Face with a simulation evaluation set (400 hours of MuJoCo) so researchers can benchmark progress without robot time. A company whose product is proprietary data is giving away 3,500 hours of it. The logic is the same one that built every platform business: the open layer sets the standard, the paid layer ships the quality.
Why a $1.2B multiple might — or might not — hold
The bull case is simple. If physical AI follows anything like the language-model curve, whoever holds the demonstration pipeline sits on the resource everyone else needs, and the crowdsourced-dataset plays already surveying this territory show labs know it. XDOF’s reported $50M annualized revenue three months out of stealth suggests demand is real, not narrative. LG and Nvidia’s Seoul data centre and Skild’s one-video learning model are attacks on the same bottleneck from other directions — compute-heavy synthesis and algorithmic efficiency respectively.
The bear case is that data collection has a ceiling nobody wants to name: teleoperation scales with human hours, and human hours don’t compound. If simulation-plus-learning approaches (Skild’s, notably) cut the demonstration requirement per task, the moat XDOF is being valued on shrinks with it. TechCrunch could not confirm the total capital being raised, and a valuation agreed in a private negotiation is a claim, not a fact — no shares have traded at that price.
What’s not in dispute is the direction of the market: robotics investment topped $18.8 billion in the first half of 2026, and the data layer — the least glamorous part of the stack — is suddenly where billion-dollar terms get discussed. Three months ago XDOF was a stealth startup with a good teleop rig. The pricing of this round, whenever it lands, will be the clearest read yet on how much of physical AI’s future investors think gets built on manufactured data.