
Episode #19
Your Tests Pass. Can Your AI Agent Finish the Job? | Francesco Bonacci, Cua
AI agents can click, type, and operate software—but how do you know they actually finished the job? Cua co-founder and CEO Francesco Bonacci joins The Merge to discuss computer-use agents, agent evaluation, and the human judgment behind reliable automation. Hendrik Krack and Francesco explore what it takes to give AI agents computers they can work in—and how to check the results. From testing computer controls to evaluating real tasks in KiCad, the conversation examines the gap between an action succeeding and an assignment being complete. Francesco also shares his experiments with bot-driven development through Slack, including a bot CTO and chief of staff, and explains why he stays close to his coding agents to steer their work. In this episode: • Computer-use agents and the infrastructure behind them • Why working tools don’t guarantee successful tasks • Agent evaluation, Cua-Bench, and circuit-design tasks in KiCad • Catching regressions while shipping quickly • Isolated environments and protecting sensitive data • Delegating to bots while retaining human judgment • Why visibility matters: how repeated attempts led to five identical customer emails What task would you trust an AI agent to handle today—and what would you still check yourself? Tell us in the comments. Explore Cua: https://cua.ai/ Cua on GitHub: https://github.com/trycua/cua Learn about computer use: https://cua.ai/docs/concepts/what-is-computer-use Explore CodeRabbit: https://www.coderabbit.ai/ The Merge is CodeRabbit’s podcast about the people building the future of software. Subscribe for more conversations with founders, engineers, and open-source builders. #AIAgents #ComputerUse #TheMerge






