
Embedded AI Podcast
E24 Practical Local LLMs: Hardware, Tools, and Trade-offs
We explore the resurgence of self-hosted LLMs and why they're worth considering despite the convenience of frontier models. From data privacy concerns to cost management and consistent performance, there are solid reasons to run your own models. We break down the hardware requirements (spoiler: it's all about memory bandwidth), compare different runtime environments from beginner-friendly LM Studio to production-ready VLLM, and discuss practical model choices. Luca shares his experience running dual RTX 3090s in his basement, while Ryan questions whether his 2004 Athlon is still viable (it might be). We also touch on the reality that local models lag frontier models by about six months—which means they're roughly where Claude Opus 4A was in February, and that's perfectly usable for many tasks. Key Topics: [02:30] Why self-host? Data privacy, cost control, and avoiding unpredictable changes from hosted providers [08:45] Memory bandwidth is the real bottleneck—why GPUs still win and what that means for your hardware choices [15:20] Hardware options: Mac Minis vs. discrete GPUs, and why second-hand RTX 3090s offer the best price-performance [22:10] Available models: Quen, Gemini, DeepSeek, and Mistral—what works well locally and the six-month lag behind frontier models [28:40] Runtime environments compared: llama.cpp for flexibility, LM Studio for ease of use, Olama for Docker-like experience, and VLLM for production [35:15] Practical considerations: context windows, model size vs. available VRAM, and why speed matters more than you think [40:30] Announcement: UddleSpec SDD framework now available, and the Agile Embedded Podcast is now Pragmatic Embedded Notable Quotes: "You need memory bandwidth. All of the memory. Every single token filters through all of those neural network layers—potentially 8 billion parameters per token." — Luca "The marketing behind these tools is very much geared toward 'don't worry what's happening behind the curtain.' But there are more efficient ways of solving the same problem without incurring the same cost." — Ryan "I really feel like the most rational choice is a second-hand 3090 still. They have 24 gigabytes of VRAM, they're nice and fast, and they have about half the bandwidth of a 5090—which is still pretty good and very much usable." — Luca Resources Mentioned: llama.cpp - Command-line LLM runtime environment with broad hardware support and REST API LM Studio - User-friendly GUI for running local LLMs with built-in model downloader and recommendations Olama - Docker-like interface for managing and running local LLMs, good for experimentation VLLM - Production-grade LLM runtime with excellent parallel request handling and memory efficiency Hugging Face - Repository with hundreds of thousands of models, quantizations, and variants available for download UddleSpec - Luca's new SDD framework addressing real-world workflow needs—now available on GitHub Pragmatic Embedded Podcast - Sister podcast (formerly Agile Embedded) covering embedded systems development Agile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topics

