
Episode #5
Ep5: Multimodal AI Agents: Benchmarking, Adapting, and Adversarial attacks
<p>In this episode, we dive into the multimodal AI agents, starting with the recent release of runner H and diving into groundbreaking research, including:</p><p>04:15 VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks by Jing Yu Koh et. al</p><p>19:18 AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations by Gaurav Verma et. al.</p><p>32:32 Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast by Xiangming Gu et. al.</p>


