
Weaviate Podcast
In-Context Retrieval with Siddharth Gollapudi - Weaviate Podcast #146!
Siddharth Gollapudi, a researcher at UC Berkeley, joins the Weaviate Podcast to discuss in-context retrieval and his paper "Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale." The idea is simple and radical: instead of encoding documents into embeddings and building a nearest neighbor index, put the entire corpus into an LLM's context window and use attention itself as the retriever. The conversation opens by separating this from generative retrieval, which memorizes documents into model parameters and needs gradient updates every time the corpus changes, and from listwise re-ranking, which Siddharth frames as an easier subset of first-stage retrieval. Lost in the middle, he argues, turned out to be much less of a problem with today's long-context models.The discussion then dives into BlockSearch, a 0.6B parameter in-context retriever trained on MS MARCO (cleaned with RLHN) plus a mix of BEIR datasets. Reaching 500 documents was relatively easy; scaling to 2,500, 5,000, and 10,000 documents, or roughly a million tokens, is where things break. Siddharth explains why random document codes beat positional codes (they force the model to actually read the documents), and how an on-policy loss that corrects the model's own rollouts keeps it disciplined when hard negatives pull it off course midway through generating a code.From there, the conversation moves to attention dilution: as the corpus grows, the softmax denominator swamps the relevant document's score. Two fixes help, length-dependent temperature scaling and sparse attention that drops irrelevant documents before attention runs, bringing BlockSearch roughly level with dense retrieval at a million tokens. This opens a big question for vector databases: is there a sublinear, ANN-style version of attention that can be trained end to end, in the spirit of ReFrag and ColBERT's late interaction?The episode closes on timelines for LLM-based re-rankers, the latency trade-offs of high-latency search, "No More Free Lunch" and whether long-context LLMs can eat the database, recursive language models, and Siddharth's excitement for unsupervised notions of relevance that could help models make genuinely new discoveries, such as proving open theorems. Chapters0:00 Welcome Siddharth!1:21 In-Context Retrieval8:16 Drowning in Documents and Reranker Scaling13:05 BlockSearch, 0.6B In-Context Retriever28:21 Vector Databases for LLM Inference40:46 Timelines for In-Context Retrievers46:54 Will LLMs eat Databases?51:38 Exciting Directions for AI

