Perfect motors were handed to the best minds. They scored 16.8%.
A multi-institution research group published a benchmark this month, HumanCLAW, that removes the most common excuse in robotics. In simulation, the low-level control is made perfect: grasps succeed, joints obey, nothing slips. The frontier models are then asked to do the part everyone assumes they are good at, running long-horizon household tasks, the multi-step sequences of cooking, tidying, fetching and staging that make up an actual day in an actual home. With flawless motors handed to them, the best frontier model completed 16.8 percent of tasks. Human teleoperators, on the same tasks, score essentially perfectly.
Sit with the shape of that result. The benchmark deliberately deletes the hard hardware problems, the ones the humanoid wave is spending billions to solve, and the ceiling barely moves. Whatever is between a demonstration and a dependable machine in a home, it is not in the wrists.
The missing layer is not motor skill
This is the second published result this month with the same lesson at different scales. The clapperboard traces showed frontier models struggling to act dependably at a one-verb task on a real rig. HumanCLAW shows the failure at the other end of the horizon: even with the one-verb layer solved by decree, the long arc of a household task collapses. What fails is not the grasp. It is everything around the grasp: remembering what has been done, holding the goal across interruptions, knowing what the house permits, deciding what must never be attempted, recovering when step four quietly invalidated step two.
That surrounding layer has a name in our stack: the mind. Memory, permissions, per-home certification, and supervision that is disclosed rather than hidden. None of it emerges from scale, because none of it is in the training distribution. A model can be taught to fold a shirt by ten thousand examples. No number of examples teaches it that this particular home stages medication at eight, that the back bedroom is off limits, that the daughter in Chicago must be told when something changes. That knowledge is installed, permissioned, and audited per household, or it does not exist.
Why we cheer this result
It would be easy to read a 16.8 percent as pessimism about the field. We read it the other way. The motor layer is commoditizing in public, week by week, and every improvement there makes a body cheaper for us to certify. The benchmark simply prices the layer that does not commoditize. When the substrate is handed to everyone for free and the task still fails five times out of six, the residual is the company. Our whole architecture is a bet on that residual: supervised from day one so the family gets a working machine while autonomy earns its way in skill by skill, every action attested, every certification judged on outcomes rather than curves. The gap between a demo and a dependable day is the product. This month, the field published its width.