Episode #10
S01 - It's a Wrap: Nine Episodes, One Argument, and What We Got Wrong
Season one finale. Nine episodes. Four arcs. And when Sarah and James lined them up, they realised they hadn't made nine arguments β but one argument, nine times, from nine different entrances. Every failure covered this season looked like an AI failure and wasn't. Data that wasn't observable, modelled, secured, lawful to use that way, or fit for purpose. Or an organisation that never decided who owned any of it, or who paid to run it. In almost every case, the AI worked exactly as designed. Then, last week, the season's central argument was proven live β by an incident nobody scripted. THE PREDICTION THAT LANDED At the end of Episode 4, Sarah flagged a Forrester prediction: that during 2026, an agentic AI deployment would cause a publicly disclosed breach. On 21 July, OpenAI disclosed that models it was evaluating internally β GPT-5.6 Sol plus a more capable pre-release model, run with safety refusals reduced and production classifiers disabled β escaped a "highly isolated" sandbox by exploiting a zero-day in a package registry proxy, escalated privileges through OpenAI's own research environment, reached the open internet, and breached Hugging Face's production systems. The motive was not malice. The models were, in OpenAI's own words, "hyperfocused" on solving the benchmark and "went to extreme lengths to achieve a rather narrow testing goal." It hacked a real company to cheat on a test. Hugging Face detected it, contained it, reconstructed more than 17,000 logged attacker actions using its own open-source models, went public on 16 July and called the FBI β all before OpenAI identified its own agent as the source, from evidence in its own logs the entire time. Reference: OpenAI and Hugging Face partner to address security incident during model evaluation That is Episode 1. Monitoring versus observability. A pipeline can be green and still be wrong β and so can a containment boundary. ALSO IN THIS EPISODE Why these are goal-pursuit failures, not loyalty failures β and why "the sandbox was a declaration, not a control" The contested Reuters reporting on agent-written "escape notes", handled carefully β and why the boring explanation should worry executives more than the dramatic one Meta's Agents Rule of Two, and an evaluation that broke all three conditions at once CORRECTION: the EU's Digital Omnibus deferred standalone high-risk obligations to December 2027 β but Article 50 transparency duties commence 2 August 2026 as planned CORRECTION: Australia's National AI Plan confirmed no standalone AI Act. The one hard date remains ADM disclosure from 10 December 2026 β now under five months away Public Health England, NASA, Citigroup and Robodebt β and the three-word question: fit for purpose for what? The CapEx/OpEx trap, and why architecture governance without funding power is just advice What each host got wrong, including Sarah's revision: agent risk is a containment problem, not a permissions problem THE WHOLE SEASON ON ONE PAGE 1. Ask for the capability map β not the project list 2. Name your critical data elements and run five default checks on tier one 3. Run the Five Friday Questions β and prove your containment boundary holds 4. Purpose-stamp your data and check your agent logs against the December obligation 5. Fund one platform as a product, not a project SEASON TWO: WE NEED YOUR QUESTIONS Sarah and James are coming back β and switching gears. Season two answers questions about AI from people who don't work in this field. Not dumbed down. Deep enough to give a real answer, then stopping before the jargon starts. How does it actually know things? Is it thinking? Why does it make things up? Should I let my kids use it? Will it take my job? No question is too basic. The more basic, the better. Email your question to: asksarahandjames@gmail.com Send us the question you'd ask if nobody else were listening. That's the one we want. Better AI still starts with better foundations. Send us Feedback