The loss went down. The robot still failed.
A research group published a study this week, The Obsessed Encoder, on how self-supervised vision systems fail. It shows that a faint, predictable pattern in the training images, a watermark at five percent opacity, can come to dominate a representation a thousand dimensions wide, while every training metric stays healthy. The safeguards designed to prevent exactly this collapse were switched on. The authors reproduced it in three published systems, and released the code so anyone can check.
One result should stop every roboticist mid-scroll. In the world-model case, the model's prediction loss sinks below the clean baseline, the curve a team would celebrate, while the planner built on top of it succeeds at chance. The metric got better. The robot got worse. Nothing in the training dashboard could tell the difference.
A home is made of predictable patterns
The planted watermark in the study is not exotic. It is a clock on the wall. A camera's timestamp overlay. The same wallpaper behind every frame, the chyron on a television that is always on. A home is saturated with faint, repeating, perfectly predictable features, which is precisely the diet the study shows these encoders will fixate on, at the expense of everything a household actually needs the machine to see.
The bar has to be the outcome
We read this study as outside confirmation of a rule we already build by: no measurement taken inside the model is evidence that the robot works. In a lab, a misleading eval wastes a week. In a kitchen, it is a dropped plate, or a hand that closes on the wrong thing. That is why cereal-os certifies skills the only place certification means anything, at the level of outcomes: rehearsed in the digital twin across seeded trials, judged against the skill's own declared postconditions, installed only when every trial passes, with the certification signed into the attested chain. A training curve cannot lie its way through that gate, because the gate never looks at the training curve.
This is also why the household layer cannot be optional. If the models' internal signals can be confidently wrong, then the layer above them, the one that verifies outcomes, holds the permissions, and answers to the family, is not overhead. It is the product. Better models make that layer more valuable, not less, for the same reason better base models made coding harnesses more valuable: the stronger the capability, the more the question becomes who checked it, and against what.
The loss going down was never the point. The plate reaching the table, provably, was.