
Mechanical Dreams
The State-Prediction Separation Hypothesis
In this episode: • Introduction to State-Prediction Separation: Professor Norris and Linda introduce the paper and the core problem of standard Transformers conflating next-token prediction and future state representation. • The SPS Mechanism: Linda explains how the authors introduce a dummy predict token to separate the input state stream from the prediction stream. • Experimental Results and Baselines: The hosts discuss the massive data efficiency gains and how clever ablations like Delayed State prove that separation, not just compute, is the key. • Gradient Analysis and Inference: A deep dive into how gradients are routed during training and why the inference overhead is practically negligible. • Conclusion and Implications: Professor Norris is convinced by the arguments, and the hosts wrap up by discussing what this means for the future of LLM architectures.

