XDOF Raises $70M to Build Robot Training Data Infrastructure as AI Labs Race to Catch Up
A new startup is positioning itself as the data backbone for physical AI, offering collection tools and annotation systems that frontier labs cannot easily build themselves.
Last verified:
The Data Bottleneck in Physical AI
According to TechCrunch AI, XDOF, a startup founded by UC Berkeley researchers Philippe Wu, Fred Shentu, and Nemo Jin, is emerging from stealth with $70 million in funding to solve a critical infrastructure gap: the absence of large-scale, high-quality training data for robots. The company secured backing from Thrive Capital, Spark Capital, Andreessen Horowitz (a16z), Lux Capital, and WndrCo. With approximately 60 employees, XDOF is already contracted with 20 customers spanning frontier AI labs, though the startup has not disclosed their identities.
The emergence of XDOF reflects a broader industry realization: as AI research pivots toward physical robotics, the infrastructure bottleneck has shifted from compute and model architectures to data itself. TechCrunch reports that while large language models were trained on vast repositories of publicly available text, robotics demands fundamentally different data—video footage of physical manipulation, interaction dynamics, and environmental response—that barely exists at scale.
Why Physical Data Collection Differs from Text
The robotics data problem is not merely one of volume. According to TechCrunch’s reporting, existing sources such as YouTube videos and footage collected by gig workers lack the fidelity and consistency needed for training robust physical AI systems. Low-fidelity data cannot reliably ground training signals to real-world robotic hardware, creating a gap that generic data sources cannot fill. TechCrunch notes that XDOF is addressing this by building not just collection pipelines but also data cleaning, tooling, and annotation systems—creating what the company describes as a self-reinforcing feedback loop for robot trainers.
The founders’ motivation stems from personal experience. Wu, working as a doctoral researcher at UC Berkeley, encountered what he calls a “chicken-and-egg problem”: answering how to train foundation models for robotics first required collecting data at scale. Wu and Shentu developed GELLO, a low-cost teleoperation system enabling human operators to control robotic arms and generate training data. That project became influential enough that other researchers began adopting the approach, validating the underlying bottleneck.
The Timing Pressure in Robotics
XDOF’s formation in October 2024 and today’s public emergence coincide with intensifying robotics efforts across the AI industry. TechCrunch notes that OpenAI recently relaunched a robotics program it had shuttered in 2021—a symbolic re-entry into physical AI. According to Wu, quoted by TechCrunch, frontier labs view the robotics race with the same urgency that surrounded language models. “All of the top labs are trying to pursue robotics,” Wu said. “You don’t want to be in this type of situation where you pursue this technology too late, and everyone is in this boat where physical AI is the next frontier.”
Why This Matters
The emergence of XDOF signals a maturing robotics AI ecosystem in which specialized infrastructure companies—not model builders alone—become essential. Teams pursuing vision-based or embodied foundation models will now face a build-versus-buy decision on data infrastructure, with XDOF positioned as a managed-services alternative. If XDOF’s $70M funding validates this market, expect similar infrastructure startups to emerge around simulation, simulation-to-reality transfer, and robot hardware abstraction. The critical question for both investors and AI labs: whether XDOF’s data pipelines can sustain competitive advantage as collection costs normalize and larger labs build internal capabilities.
Frequently Asked Questions
What does XDOF do?
XDOF builds data collection pipelines, tooling, and annotation systems specifically for robot training. The company helps AI labs and robotics companies gather, clean, and label high-quality physical interaction data at scale.
Why is robot training data a bottleneck?
Unlike language models trained on publicly available text, robots require data capturing physical interaction. Existing sources like YouTube videos and gig-worker footage are low-fidelity and difficult to align with real-world robotic systems.
Who is funding XDOF and what is the valuation?
XDOF raised $70 million from Thrive Capital, Spark Capital, a16z, Lux Capital, and WndrCo. The post-money valuation has not been disclosed.
Is XDOF working with OpenAI or other frontier labs?
XDOF's CEO confirmed the startup is working with 20 customers including several frontier AI labs, but declined to name them publicly.