Episode #13
Maximizing intelligence per watt | Avanika Narayan & Jon Saad-Falcon, Stanford AI CS Ph.D.s
The next phase of AI adoption will depend on making intelligence more efficient, not just more capable. In this episode, Jaya is joined by Stanford Ph.D. researchers Jon Saad-Falcon and Avanika Narayan to discuss how local and hybrid AI can help bring these tools to billions more people. Much of today’s LLM inference runs through frontier models in centralized data centers. But Jon and Avanika argue that many tasks can already be handled by smaller, open-weight models, including models running on local hardware. Their concept of “intelligence per watt” measures how much task performance a given model-and-hardware configuration delivers relative to the power it consumes. They also discuss OpenJarvis and Minions, two projects that explore how AI systems can use local models when possible and turn to cloud models when more capability is needed. From there, the conversation turns to what this shift could mean for OpenAI, Anthropic, and other frontier labs. As open-weight models improve and inference becomes cheaper and more distributed, Jon and Avanika argue that low switching costs and limited lock-in could put pressure on the labs’ margins. They also explain why enterprises should be careful about handing proprietary data and decision-making processes over to model providers. They close with advice for Ph.D. students and academic researchers: rather than trying to reproduce the frontier labs’ work with fewer resources, choose problems where original ideas matter more than scale. What we covered: 00:00 – Cold open 00:52 – Welcome and guest introductions 02:03 – Why intelligence efficiency matters 03:22 – Defining intelligence per watt 05:46 – Overview of OpenJARVIS and Minions Projects 08:00 – Industry adoption of intelligence per watt metrics and hybrid execution strategies 09:10 – Why local capabilities will shift demand away from centralized cloud infrastructure toward fragmented, specialized hardware 12:35 – Technical drivers for local AI 12:56 – How tokenmaxxing complements efficiency research 14:36 – Why proprietary AI providers face revenue risks as open-weight models compress in size and rival frontier capabilities 17:15 – The real moats in AI today 18:23 – Advice for enterprise CIOs 22:15 – Hot takes 23:48 – The importance of US public policy and aligning public benefit with AI advancements 24:45 – Advice for PhD students and researchers 25:56 – Closing thoughts Links: Intelligence per Watt: Measuring Intelligence Efficiency of Local AI - https://arxiv.org/abs/2511.07885 OpenJarvis: Personal AI, On Personal Devices - https://arxiv.org/abs/2605.17172 Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models - https://arxiv.org/abs/2502.15964


