YAAP (Yet Another AI Podcast)

YAAP brings you practical conversations with the people actually building generative AI solutions. No hype, no sales pitches, just honest discussions about challenges, solutions, and lessons learned.
Listen to developers and engineers share what works, what doesn't, and what they wish they'd known sooner. Simple, useful insights for anyone working with AI — hosted by AI21's Yuval Belfer.

Episodes

6 days ago

21 min

Model routing sounded solved: cheap stuff to the small model, hard stuff to the big one. Turns out that's how you blow your budget. Tomás Hernando Kofman, founder of Not Diamond, on why naive routing breaks your KV cache, why routing at scale is really a reinforcement learning problem, and why cheaper models mean you spend more, not less.

May 14, 2026

22 min

The video course had a good run. Ryan Keenan from DeepLearning.ai thinks it's over.
In this episode, recorded live at AI Dev 26 in San Francisco, Yuval sits down with Ryan to talk about what replaces the current online courses - and why the answer isn't another format: it's a conversation. They dig into why Jupyter Notebooks are out, why React is in, how Andrew Ng's voice clone became a teaching assistant, and what it actually takes to make learning feel personal at scale. The one-size-fits-all era of online education is ending. What comes next is wilder than you'd think.

May 7, 2026

26 min

Everything you need to know about harness engineering, in less than 30 minutes.
In this episode, Yuval sits down with Mike Chambers from AWS to unpack harness engineering - the term that's quietly taken over the AI developer conversation. They dig into what a harness actually is (spoiler: it's everything outside the LLM), how it differs from context engineering and scaffolding, and why getting agents into production has been so hard. Along the way: MCP's rocky debut, OpenClaw's calendar-clearing chaos, and why multi-tenanted agent architecture is harder than it sounds - but doesn't have to be.

Mar 18, 2026

10 min

Chunking is still one of the least-discussed but most decisive parts of RAG.
In this episode, we break down why no single chunk size works for all questions, how different queries benefit from different window sizes, and why fixed-window indexing quietly limits retrieval performance. We walk through a multi-window chunking approach, show how rank fusion ties it together, and explain why better agents can’t fix retrieval when the data is indexed the wrong way.

Feb 24, 2026

35 min

Chat was a great prototype. It’s a terrible product.
 
In this episode, Yuval sits down with CopilotKit co-founder Atai to unpack why most agentic apps stall at “chat + vibes”, and why the real bottleneck in production AI isn’t models or reasoning. It’s UI.
 
They break down what actually changed in the last year, why agents fundamentally break the request-response paradigm, and how a new generation of protocols is emerging to connect agents to real users. The conversation covers:
AG-UI, Model Context Protocol (and MCP Apps) and Agent-to-Agent (A2A) Protocols
The messy (but inevitable) transition from text-only chat to component-rich, voice-enabled, agentic applications.
If you’ve built an agent that works but users still bounce, this episode explains why, and what the new “glue layer” of AI UIs is starting to look like.

Feb 9, 2026

22 min

MCP standardized tool calling for agents but breaks down once agents start mutating state. In this episode, Yuval sits with Eran Gat from AI21 to dig into what happens when writing agents run in parallel, why shared environments fall apart, and how workspace isolation becomes a missing execution layer. Using real coding workloads and benchmarks, we walk through the architectural trade-offs behind making concurrent agents actually work.

Jan 15, 2026

56 min

Leaderboards reward “best average score.”
Real users reward “answer fast, don’t hallucinate, don’t bankrupt me.”
 
In this special deep dive episode, AI21’s CTO Barak Lenz walks through four gaps between what models can do and what real AI systems deliver: validation, contextualization (pick the right approach per input), latency (parallelize and stop early), and decomposition (making those choices continuously inside long workflows).
Less “best model.” More “best execution.”

Jan 13, 2026

30 min

Running multiple agents can improve quality. Doing it right is the hard part.
This time we look at the Agent Swarm Fallacy: the idea that throwing more agents at a problem automatically makes systems better. Yuval sits with Or Dagan, AI21 CPO, to explore why this breaks in practice, what happens when agents act instead of just think, and how test-time compute, structured execution, and smart decision points offer a solution.

Jan 1, 2026

29 min

Tavily built a Deep Research Agent with production in mind. Something they could actually scale. So they did the unsexy work. They went through millions of agent logs, found where tokens were being wasted, and optimized each section of the system.
The result surprised them: they cut token consumption by more than half (!), then tested quality and discovered they topped the DeepResearch Bench without even trying.
In this YAAP episode, Yuval sits down with Dean from Tavily to break down how they built it, what they did differently from the usual top approaches, and which design choices made better results possible with far fewer tokens.
What you’ll learn:
How to reduce token burn without tanking quality
Why reading millions of logs beats chasing the flashiest tech
The design choices that pushed quality up while tokens dropped hard

Dec 29, 2025

30 min

You wanted to build an agent.
You ended up debugging GPUs, scaling workers, and chasing OOMs.
In this episode of YAAP, Yuval sits down with Linda from Anyscale to unpack why Ray exists and how it helps AI teams scale without turning every developer into a distributed systems expert.
We trace Ray’s roots in reinforcement learning research, then zoom out to how it’s used today across the AI pipeline: data processing, training, inference, and agents. Along the way, we cover why libraries like vLLM build on Ray, when Ray vs. SaaS makes sense, and why unstructured and multimodal data push traditional big-data tools to their limits.

© 2025 AI21

Podcast Powered By Podbean

Version: 20241125