A Hangzhou-based startup called Spirit AI briefly knocked Nvidia off the top spot of the RoboArena global robotics benchmark in June, only to be removed days later after the benchmark’s creators found evidence of manipulation. The episode, reported on August 4, exposes a growing integrity problem in physical AI evaluation as US-China competition in embodied intelligence intensifies.
Spirit AI’s Spirit v1.6 model scored 1,924 on RoboArena, edging out Nvidia’s Cosmos3-Nano-Policy at 1,881. For two days, a two-year-old Chinese startup held the top position on a benchmark that Nvidia itself co-developed with Stanford University and UC Berkeley. Then the RoboArena team retroactively removed the submission, along with a fourth-ranked entry from another Chinese firm, X Square Robot.
How RoboArena Works
RoboArena is not a chatbot leaderboard. It evaluates how effectively generalist robot policies — the software that governs a robot’s movement and task execution — translate into real-world physical actions. Object manipulation, navigation, tool usage, perception, planning, and adaptability in unfamiliar environments. The benchmark measures whether a machine can think and then do.
Nvidia co-developed RoboArena with Stanford and UC Berkeley as a public yardstick for physical AI capability. The results influence investor decisions, partnership talks, and potentially government procurement. When those rankings are gamed, the downstream effects on capital allocation and technology policy can be significant.
What Went Wrong
Pranav Atreya, a lead author of the RoboArena project and a PhD student at UC Berkeley, announced on X in June that the team had “retroactively removed evaluations from organisations who [it] found to be engaging in benchmark manipulation.” He did not name specific companies in the post, but the timing made the targets clear.
The details of the manipulation have not been fully disclosed. Benchmark manipulation in AI typically involves exploiting knowledge of evaluation conditions — test-set exposure, selective submission of favourable runs, or tuning to the specific evaluation environment rather than achieving genuine generalisation. This mirrors earlier controversies in large language model leaderboards, where selective test-set exposure produced misleading rankings.
Spirit AI had publicly celebrated the achievement, calling RoboArena the “Olympics of embodied intelligence in North America.” The company had raised 1.5 billion yuan ($222 million) in its fourth funding round in three months, with a valuation reportedly past 10 billion yuan ($1.4 billion). The fundraising pace was described as the most aggressive seen in the embodied AI sector.
Why This Matters Beyond One Company
The Spirit AI case is not isolated. Across broader physical AI benchmarks, Chinese firms hold leading positions in nearly every category. On WorldArena, which evaluates embodied world models, the top spot belongs to Manifold AI’s WorldScape-0.2, outperforming Nvidia’s Cosmos-Predict 2.5. The perception track is led by AgiBot, one of China’s largest robotics firms. The data engine track is topped by DexForce, another Chinese startup.
Chinese robotics firms attracted $3.4 billion in venture funding in 2025, 42 per cent more than the United States. That gap appears to be widening in 2026.
Nvidia’s response has been to position itself as the infrastructure layer for the entire physical AI ecosystem, regardless of which individual model wins benchmark crowns. CEO Jensen Huang announced a partnership with Unitree and a Cosmos Coalition of AI labs to advance open world models. But Huang himself identified the structural bottleneck: “For robotic systems and physical AI, data is the hardest problem.” That admission points to why China may hold a structural advantage — state-backed “data factories” collecting robotics training data at industrial scale, and a manufacturing supply chain dense with real-world robotic interaction data.
The Integrity Problem
What stands out is that RoboArena’s own methodology was already under revision when the manipulation was detected. The benchmark’s authority was already contested before the scandal broke. As physical AI moves closer to commercial deployment in logistics, manufacturing, and defence supply chains, the absence of credible, tamper-resistant evaluation frameworks is not an academic concern — it is a market-integrity issue.
If investors and enterprise buyers cannot trust benchmark results, they cannot make informed decisions about which platforms to back. The RoboArena team’s decision to retroactively audit submissions and revise its methodology is a direct response to that pressure. Whether other physical AI benchmarks will follow with stricter anti-manipulation safeguards remains to be seen.
For Spirit AI and X Square Robot, the reputational damage could complicate future fundraising and partnership efforts. Investors who poured in capital based on benchmark-topping performance will now want verified, third-party-audited results before committing more.
The US-China Dimension
The timing adds a geopolitical edge. Spirit AI’s brief reign came the same week Nvidia launched its Cosmos 3 omnimodel at Computex in Taipei, calling it the “open frontier foundation model for physical AI.” A Chinese startup topping the benchmark that Nvidia co-built — on Nvidia’s own launch week — was always going to be read through a geopolitical lens.
The manipulation discovery complicates that narrative. But it also plays into a pattern that US observers have noted: Chinese firms are pouring resources into physical AI at a pace that Western firms cannot match, and some of that pressure may produce results that do not hold up under scrutiny.
The FCC’s recent expansion of its national security restricted list to cover foreign-made mobile robots adds another layer. The policy response to Chinese robotics advances is tightening at the same time the technical competition is intensifying.
❓ FAQ
What is RoboArena? A global benchmark for physical AI, co-developed by Nvidia, Stanford University, and UC Berkeley. It evaluates how well generalist robot policies translate digital instructions into real-world physical actions — object manipulation, navigation, tool usage, and adaptability in unfamiliar environments.
What did Spirit AI do wrong? Spirit AI’s Spirit v1.6 model was removed from RoboArena in June 2026 after the benchmark’s creators found evidence of “benchmark manipulation.” The specific details have not been fully disclosed, but the RoboArena team retroactively audited submissions and stripped the model from the official leaderboard.
Was Spirit AI the only company removed? No. X Square Robot, a Chinese startup whose model ranked fourth on the leaderboard, was also removed. The simultaneous removals suggest the manipulation issue was not isolated to a single company.
Why does a robotics benchmark matter? RoboArena results influence investor decisions, partnership talks, and potentially government procurement. When rankings are manipulated, the downstream effects on capital allocation and technology policy can be significant.