
Episode #173
The Stack β October 02, 2026
Daily IT Briefing AI & Machine Learning Cloudflare open-sources small "decision models" for agent routing. Clef and Clef-flash are classifier-style models that produce typed, probabilistic outputs rather than free text β designed for the cheap, high-volume decisions inside agent workflows (which tool to call, whether to escalate) where a frontier LLM is overkill. They're hosted on Workers AI, released under Apache 2.0 on Hugging Face, and ship with a vision encoder and 64k context. Cloudflare also launched an RL fine-tuning service for them. The framing is explicitly competitive with TypeSafe's Jev, and the benchmark comparisons come from Cloudflare itself β worth treating as vendor-reported. AWS follows with its own open decision model. Strands Decider 2B is built on a Qwen3.5-2B base, sorts between pre-decided options, and returns confidence scores. It began as a side project by AWS distinguished engineer Marc Brooker and briefly topped the Jevbench leaderboard for its size class. TypeSafe's CEO downplayed the rivalry, noting how hard it is to make these models genuinely useful rather than just fast. This is now the third entrant in the decision-model niche in roughly a week, alongside OpenAI's Decisions API β the category is consolidating fast. A new paper proposes treating context as an editable file. Context Language Models (CLMs) let the model itself edit its own context, with the authors reporting zero-shot gains over existing context-management strategies β 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus β plus an RL training method and a serving optimization. It's a single preprint; the headline numbers need replication before they mean much. Imperfect-information AI gets a real advance. Researchers have cracked Stratego, a game where each player can't see the other's pieces. The trick was adding a second neural network dedicated to inferring hidden piece identities, letting the system play at a high level despite incomplete information. This matters more than another board-game win: hidden-information decision-making is much closer to real-world deployment than chess or Go. OpenAI dismissed three safety researchers. The company says an investigation found they "mishandled sensitive information outside established company procedures," reportedly sharing confidential material with an outside AI safety organization. No names, no organization, no details on what was shared. The context is a rough stretch for OpenAI's safety reputation β reporting that executives brushed aside employee warnings, security incidents involving its agents, and the earlier scrapping of the GPT-6.1 Astra launch. Note the sourcing here is anonymous and the social-media claims about the individuals involved are unconfirmed. A study quantifies AI writing tells β and finds labs drifting in opposite directions. Analyzing 10,000 pre-ChatGPT articles rewritten by frontier models, researchers found roughly 13,000 phrases at least twice as common in AI output as human writing. Some are extreme: one model overuses "this matters" 116x relative to human baseline, another favors corrective framings like "not simply X" at 100x+. Em-dash usage has collapsed across labs. The interesting finding is divergence β Claude-family models are trending toward human word distributions while GPT-family models drift further away, and the researchers argue labs have limited control over these patterns at scale. Industry BMW to cut 20% of management roles, citing AI. The company says it will push AI across vehicle development, procurement, sales, marketing, and after-sales, and will eliminate a fifth of divisions and management positions over the coming months, with comparable cuts at lower levels. Earlier reporting pointed to roughly 8,000 office staff reductions in Germany; BMW had about 155,000 employees at the end of 2025. The stated drivers are cost pressure, Chinese competition, and US import tariffs. Treat the direct AI-to-headcount causal link as the company's framing β this is a restructuring with AI as its public rationale. Photon raises $4.5M seed for messaging-native AI agents. The startup lets developers build agents that operate over iMessage, WhatsApp, Telegram, SMS/RCS, email, and voice. Co-led by Gradient and A* with participation from Vercel and others. Claims: 40,000+ developer sign-ups, 10x revenue growth in four months, under 3% churn. Its open source version still accounts for 98% of usage β a notable ratio for a company selling a hosted product. The thesis is that agents replace apps; the company held an "app funeral" in San Francisco to promote it. Infrastructure & Hardware Google is testing TPUs in orbit. A Falcon 9 placed a Google prototype satellite β built with Planet Labs β carrying four tensor processing units with about 1 kW of combined compute, roughly one standard data center server. The chips handle simple queries but run continuously for only ~15 minutes before thermal limits bite; in space, cooling means heat dissipation only, no convection. That's enough for query processing, nowhere near an orbital data center. The real goal is characterizing radiation tolerance and thermal behavior. Google plans further satellites to test inter-satellite communication, eventually with dozens of TPUs. Ground testing included simulated launch integrity, proton-beam exposure, and cooling validation. The RAM shortage is expected to run through 2028. Memory industry executives say prices for 2027-contracted memory are substantially above 2026 levels, and buyers should plan for elevated costs well into the decade. The cause is AI data center demand crowding out supply for other markets. If you're budgeting hardware refreshes, this is the number to plan around. Security Critical Zimbra vulnerability under active exploitation. The flaw can be triggered by a simple malicious email and allows remote OS command injection β a path to stealing mail and potentially full server takeover. Any organization running Zimbra should treat patching as urgent, not scheduled. Two US federal agencies breached in separate incidents within the past month. Together they exposed a large volume of sensitive data. Details on scope and attribution remain thin, but the incidents will renew scrutiny of federal network defenses. Dev World turbopuffer v3 demotes its vector index. The storage rewrite moves the ANN vector index from primary to just another secondary index. The argument: vector-primary layouts caused storage and write amplification and blocked vectorized query execution. The company frames this as the end of the "vector database" era for its own architecture β a notable position from a vendor whose product was that architecture. All CI passes; performance work continues. A sharp dispute over Postgres full-text search benchmarks. ParadeDB published a detailed rebuttal to PlanetScale's TIN extension, reporting that after optimization (per-term fieldnorm arrays, a MAXSCORE Blockmax path) it closed the performance gap without changing document identifiers, and pushing back on TIN's ctid-based architecture claims. More pointedly, ParadeDB flags what it calls a benchmark-configuration problem: TIN's "dense-term elision" shortcut, which ParadeDB says produces non-exact BM25 rankings on the benchmark's synthetic queries. Both parties have commercial incentives here β read the methodology, not the summaries. Improvements are being upstreamed to Tantivy. Undocumented SDR capabilities found in ESP32 chips. Multiple independent projects have demonstrated raw IQ baseband capture on ESP32 hardware β roughly 2.2β2.7 GHz at up to 80 MS/s. One team found phase-coherent capture is now possible; a separate project streams IQ continuously via an FPGA; another uses an ESP32-C5 as a 5.8 GHz FPV video receiver. Most setups export snapshots only, useful for spectrum analysis, though one variant can stream over Gigabit Ethernet. This is a genuinely surprising capability for a chip family nobody bought for radio work. Other releases worth noting: papero-extract, a CPU-only, no-ML PDF parser producing Markdown/JSON/Word/Excel with reading order, tables, formulas, and bounding boxes β runs in-browser, in Python, or as an API, MIT-licensed, with the author candid that ML tools still win on irregular layouts. Bez, an early-stage project generating a web rendering engine from specs and tests using the three shipping browsers plus WPT as oracles; three-browser majority voting agrees on ~95% of WPT keys and already surfaced a real Firefox length-rounding compat bug. And Pi Durable, an experimental framework for crash-resilient long-running agents that runs anywhere JavaScript does, with pluggable storage and execution environments, shipping alongside the Pi 1.0 coding agent. RacketCon 2026 runs this weekend (Oct 3β4), streaming online. Talks cover a new FFI layer, Typed Racket inference, the Pille low-level language, the BrandX OOP library, immutable dataframes, effect handlers, and a keynote from Pat Hanrahan. Figma's remote MCP server only accepts whitelisted clients β and Pi isn't on the list. A useful reminder that "open" MCP endpoints can still gate access at the client level. One more item, flagged as opinion: a widely-shared post argues Red Hat is being phased out inside IBM, citing brand changes, staff departures, and a shift toward marketing over Linux engineering. This is editorial commentary, not verified reporting β the "imploding" framing is one person's read.

