YAAP (Yet Another AI Podcast)

YAAP brings you practical conversations with the people actually building generative AI solutions. No hype, no sales pitches, just honest discussions about challenges, solutions, and lessons learned.
Listen to developers and engineers share what works, what doesn't, and what they wish they'd known sooner. Simple, useful insights for anyone working with AI — hosted by AI21's Yuval Belfer.

Episodes

Aug 18, 2026

16 min

Most agent-browser tools bolt on through MCP. Corey J. Gallon, Managing Director at Rexmore, ripped that out. He built Chrome Agent, which talks to Chrome directly over the DevTools protocol (CDP). That means his agent can share his actual browser session: watch it work, take over mid-task, even beat CAPTCHAs along the way.
Yuval and Corey get into why Playwright MCP fell short, what it actually means to hand an agent your exact browser context, and the scraping and outreach workflows this unlocked.
Then the twist. OpenAI tried to ban Corey for "cyber abuse." No explanation given. He asked them, politely, what he was even accused of, and they backed down within days.

Aug 18, 2026

16 min

Aug 4, 2026

21 min

Geoffrey Huntley hasn't hand-written code in over two years. He's the guy behind the Ralph Wiggum technique, he thinks Git is end of life, and he's rebuilding source control, operating systems, and programming languages from scratch just to see what survives. His new bar for code: explainable, not readable. Also why he ditched Anthropic and OpenAI for a $300-a-year Chinese model.

Aug 4, 2026

21 min

Jul 21, 2026

21 min

Model routing sounded solved: cheap stuff to the small model, hard stuff to the big one. Turns out that's how you blow your budget. Tomás Hernando Kofman, founder of Not Diamond, on why naive routing breaks your KV cache, why routing at scale is really a reinforcement learning problem, and why cheaper models mean you spend more, not less.

Jul 21, 2026

21 min

May 14, 2026

22 min

The video course had a good run. Ryan Keenan from DeepLearning.ai thinks it's over.
In this episode, recorded live at AI Dev 26 in San Francisco, Yuval sits down with Ryan to talk about what replaces the current online courses - and why the answer isn't another format: it's a conversation. They dig into why Jupyter Notebooks are out, why React is in, how Andrew Ng's voice clone became a teaching assistant, and what it actually takes to make learning feel personal at scale. The one-size-fits-all era of online education is ending. What comes next is wilder than you'd think.

May 14, 2026

22 min

May 7, 2026

26 min

Everything you need to know about harness engineering, in less than 30 minutes.
In this episode, Yuval sits down with Mike Chambers from AWS to unpack harness engineering - the term that's quietly taken over the AI developer conversation. They dig into what a harness actually is (spoiler: it's everything outside the LLM), how it differs from context engineering and scaffolding, and why getting agents into production has been so hard. Along the way: MCP's rocky debut, OpenClaw's calendar-clearing chaos, and why multi-tenanted agent architecture is harder than it sounds - but doesn't have to be.

May 7, 2026

26 min

Mar 18, 2026

10 min

Chunking is still one of the least-discussed but most decisive parts of RAG.
In this episode, we break down why no single chunk size works for all questions, how different queries benefit from different window sizes, and why fixed-window indexing quietly limits retrieval performance. We walk through a multi-window chunking approach, show how rank fusion ties it together, and explain why better agents can’t fix retrieval when the data is indexed the wrong way.

Mar 18, 2026

10 min

Feb 24, 2026

35 min

Chat was a great prototype. It’s a terrible product.
 
In this episode, Yuval sits down with CopilotKit co-founder Atai to unpack why most agentic apps stall at “chat + vibes”, and why the real bottleneck in production AI isn’t models or reasoning. It’s UI.
 
They break down what actually changed in the last year, why agents fundamentally break the request-response paradigm, and how a new generation of protocols is emerging to connect agents to real users. The conversation covers:
AG-UI, Model Context Protocol (and MCP Apps) and Agent-to-Agent (A2A) Protocols
The messy (but inevitable) transition from text-only chat to component-rich, voice-enabled, agentic applications.
If you’ve built an agent that works but users still bounce, this episode explains why, and what the new “glue layer” of AI UIs is starting to look like.

Feb 24, 2026

35 min

Feb 9, 2026

22 min

MCP standardized tool calling for agents but breaks down once agents start mutating state. In this episode, Yuval sits with Eran Gat from AI21 to dig into what happens when writing agents run in parallel, why shared environments fall apart, and how workspace isolation becomes a missing execution layer. Using real coding workloads and benchmarks, we walk through the architectural trade-offs behind making concurrent agents actually work.

Feb 9, 2026

22 min

Jan 15, 2026

56 min

Leaderboards reward “best average score.”
Real users reward “answer fast, don’t hallucinate, don’t bankrupt me.”
 
In this special deep dive episode, AI21’s CTO Barak Lenz walks through four gaps between what models can do and what real AI systems deliver: validation, contextualization (pick the right approach per input), latency (parallelize and stop early), and decomposition (making those choices continuously inside long workflows).
Less “best model.” More “best execution.”

Jan 15, 2026

56 min

Jan 13, 2026

30 min

Running multiple agents can improve quality. Doing it right is the hard part.
This time we look at the Agent Swarm Fallacy: the idea that throwing more agents at a problem automatically makes systems better. Yuval sits with Or Dagan, AI21 CPO, to explore why this breaks in practice, what happens when agents act instead of just think, and how test-time compute, structured execution, and smart decision points offer a solution.

Jan 13, 2026

30 min

© 2025 AI21

Podcast Powered By Podbean

Version: 20241125