
Episode #67
Eighty Three Hertz: Robot-Use Agents and Who Holds the Endpoint
Phillip Isola's 767-word essay "Robot-Use Agents" makes a claim that needs no new hardware to come true: if a language model can competently drive a robot, then every internet-connected machine becomes a potential tool for a cloud model, with no new sensors, no new actuators and no new onboard GPUs, and a car becomes as controllable as a factory arm. He is careful about it, he hedges almost every consequence, and he concedes the capability argument himself in a subordinate clause, because his point is not that any one robot gets better but that intelligence diffuses through the entire installed base in a year rather than a decade. We verified his evidence and two things fell out. The hardest number in the whole story is inside the demo he cites in his own favour: Anthropic's own evaluation states that real-time control needs roughly 83 Hz while current non-reasoning inference runs at 0.2 to 0.4 Hz, a gap of two orders of magnitude, and the capability figures were obtained by pausing the simulator because otherwise the models act far too slowly for a physics loop. And the three demonstrations he footnotes as tantalizing evidence are three incommensurable things: a controlled evaluation whose code is unreleased, a launch page containing no measurements whatsoever, and an independent third party with an open harness and 120 published trials whose own description tag claims an interleaved blinded comparison that its own limitations section says never happened. Then the question his essay does not ask, which is ours and not his: if the intelligence layer of every physical machine is an endpoint, that layer is rented, and a robot whose mind is a cloud endpoint needs a permanent operator by construction. That is not a prediction, it is the architecture, and Google's own documentation already says the embodied reasoning 1.6 model will be shut down at the end of August. The counter-position is on-device open-weight robot models, which is a messier case than it sounds: GR00T gives three different licence answers in three NVIDIA-controlled documents and will not run until you accept a click-through for a gated backbone, pi-zero's weights ship with no licence file at all, and the only model that is exactly what it says is OpenVLA. But SmolVLA, at 450 million parameters under Apache-2.0, beats a 3.3-billion-parameter model and a 7-billion-parameter one on LIBERO, which means the thing you can download and keep is not the compromise option. Fry argues the diffusion speed is the genuinely revolutionary part and that delay has a cost paid by people who are never counted. Bob argues it arrives rented. Both are right, and the episode ends without a verdict.

