<?xml version="1.0" encoding="utf-8"?><rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" version="2.0">
  <channel>
    <title>Weekly AI Paper Spotlight</title>
    <link>https://dietrich.ai</link>
    <description><![CDATA[<p>Stay up to speed with the fast-moving world of artificial intelligence. Each Monday, two AI hosts break down the most relevant research paper of the week, highlight the key insights, explain why it matters right now and translate cutting-edge developments into practical understanding for listeners of all backgrounds.</p><p>Brought to you by <a href="https://dietrich.ai">https://dietrich.ai</a></p>]]></description>
    <language>en-us</language>
    <lastBuildDate>Mon, 14 Sep 2026 02:00:00 +0000</lastBuildDate>
    <atom:link href="http://65.21.60.15:8080/podcast.xml" rel="self" type="application/rss+xml"/>
    <itunes:author>dietrich.ai</itunes:author>
    <itunes:summary><![CDATA[<p>Stay up to speed with the fast-moving world of artificial intelligence. Each Monday, two AI hosts break down the most relevant research paper of the week, highlight the key insights, explain why it matters right now and translate cutting-edge developments into practical understanding for listeners of all backgrounds.</p><p>Brought to you by <a href="https://dietrich.ai">https://dietrich.ai</a></p>]]></itunes:summary>
    <itunes:subtitle>Weekly AI Paper Spotlight</itunes:subtitle>
    <itunes:image href="http://65.21.60.15:8080/podcast-cover.webp"/>
    <itunes:explicit>false</itunes:explicit>
    <itunes:owner>
      <itunes:name>dietrich.ai</itunes:name>
      <itunes:email>tilman@dietrich.ai</itunes:email>
    </itunes:owner>
    <itunes:category text="Technology">
      <itunes:category text="Tech News"/>
    </itunes:category>
    <item>
      <title>38/2026 - Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills</title>
      <link>https://huggingface.co/papers/2609.02749</link>
      <guid isPermaLink="false">2026-w38</guid>
      <pubDate>Mon, 14 Sep 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>AI research agents can write code and run experiments, but they waste time rediscovering how tools actually work in practice. This episode explores Repo-To-Skill, a system that distills GitHub repositories into verified, agent-readable &quot;skills&quot;—compact packages of operational knowledge covering setup traps, evaluation gotchas, and failure modes. The resulting library contains over 5,000 skills from 1,000 ML repos, giving practitioners a way to bootstrap agents with real-world know-how instead of trial and error.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2609.02749.pdf">https://arxiv.org/pdf/2609.02749.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W38/podcast_pro_de3aba44149a4d52abf5305788c6458d.mp3" length="8362605" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>AI research agents can write code and run experiments, but they waste time rediscovering how tools actually work in practice. This episode explores Repo-To-Skill, a system that distills GitHub repositories into verified, agent-readable &quot;skills&quot;—compact...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>AI research agents can write code and run experiments, but they waste time rediscovering how tools actually work in practice. This episode explores Repo-To-Skill, a system that distills GitHub repositories into verified, agent-readable &quot;skills&quot;—compact packages of operational knowledge covering setup traps, evaluation gotchas, and failure modes. The resulting library contains over 5,000 skills from 1,000 ML repos, giving practitioners a way to bootstrap agents with real-world know-how instead of trial and error.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2609.02749.pdf">https://arxiv.org/pdf/2609.02749.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:58</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>37/2026 - VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning</title>
      <link>https://huggingface.co/papers/2608.26105</link>
      <guid isPermaLink="false">2026-w37</guid>
      <pubDate>Mon, 07 Sep 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>What if AI models could reason by manipulating images and videos directly, instead of just describing their thinking in words? This episode explores VBVR-Pro, a new benchmark suite offering 300 procedurally generated visual reasoning tasks spanning perception, spatial reasoning, and abstraction. Models trained on it show consistent gains across seven external benchmarks, suggesting broad visual reasoning skills transfer to real-world scenarios.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2608.26105.pdf">https://arxiv.org/pdf/2608.26105.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W37/podcast_pro_abdbc881ee4b4944a79be5e4208b25fc.mp3" length="9361005" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>What if AI models could reason by manipulating images and videos directly, instead of just describing their thinking in words? This episode explores VBVR-Pro, a new benchmark suite offering 300 procedurally generated visual reasoning tasks spanning per...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>What if AI models could reason by manipulating images and videos directly, instead of just describing their thinking in words? This episode explores VBVR-Pro, a new benchmark suite offering 300 procedurally generated visual reasoning tasks spanning perception, spatial reasoning, and abstraction. Models trained on it show consistent gains across seven external benchmarks, suggesting broad visual reasoning skills transfer to real-world scenarios.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2608.26105.pdf">https://arxiv.org/pdf/2608.26105.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:48</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>36/2026 - StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling</title>
      <link>https://huggingface.co/papers/2608.15089</link>
      <guid isPermaLink="false">2026-w36</guid>
      <pubDate>Mon, 31 Aug 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Long-horizon AI agents often fail not because the model lacks capability, but because execution drifts—forgetting phases, skipping checks, or stopping too early. StateM introduces &quot;harness scaling,&quot; a runtime layer using human-readable runbooks that enforce state transitions and recovery without changing the underlying model. The approach hits 95.3% accuracy on Terminal-Bench 2.1 for just $15 in compute, suggesting practitioners may get more from better scaffolding than bigger models.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2608.15089.pdf">https://arxiv.org/pdf/2608.15089.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W36/podcast_pro_02a33e79b8464cad881e4d51f8c07aba.mp3" length="9321165" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Long-horizon AI agents often fail not because the model lacks capability, but because execution drifts—forgetting phases, skipping checks, or stopping too early. StateM introduces &quot;harness scaling,&quot; a runtime layer using human-readable runbooks that en...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Long-horizon AI agents often fail not because the model lacks capability, but because execution drifts—forgetting phases, skipping checks, or stopping too early. StateM introduces &quot;harness scaling,&quot; a runtime layer using human-readable runbooks that enforce state transitions and recovery without changing the underlying model. The approach hits 95.3% accuracy on Terminal-Bench 2.1 for just $15 in compute, suggesting practitioners may get more from better scaffolding than bigger models.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2608.15089.pdf">https://arxiv.org/pdf/2608.15089.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:46</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>35/2026 - BDH-CQ: In-Context Learning with Recurrent Latent Reasoning</title>
      <link>https://huggingface.co/papers/2608.09888</link>
      <guid isPermaLink="false">2026-w35</guid>
      <pubDate>Mon, 24 Aug 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Modern AI reasoning systems pay a steep computational tax by &quot;thinking out loud&quot; token by token. BDH-CQ takes a different approach, combining in-context learning with recurrent latent reasoning—processing examples through an evolving memory and solving queries in a hidden internal workspace rather than generating intermediate text. A 150M-parameter model achieves nearly 30% on ARC-AGI-1 at less than a tenth of a cent per task. For practitioners building systems requiring frequent lightweight reasoning calls, this cost-accuracy tradeoff opens new possibilities.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2608.09888.pdf">https://arxiv.org/pdf/2608.09888.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W35/podcast_pro_a9f0269e582e4874a3fd713e99664977.mp3" length="8181165" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Modern AI reasoning systems pay a steep computational tax by &quot;thinking out loud&quot; token by token. BDH-CQ takes a different approach, combining in-context learning with recurrent latent reasoning—processing examples through an evolving memory and solving...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Modern AI reasoning systems pay a steep computational tax by &quot;thinking out loud&quot; token by token. BDH-CQ takes a different approach, combining in-context learning with recurrent latent reasoning—processing examples through an evolving memory and solving queries in a hidden internal workspace rather than generating intermediate text. A 150M-parameter model achieves nearly 30% on ARC-AGI-1 at less than a tenth of a cent per task. For practitioners building systems requiring frequent lightweight reasoning calls, this cost-accuracy tradeoff opens new possibilities.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2608.09888.pdf">https://arxiv.org/pdf/2608.09888.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:49</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>34/2026 - Recursive Synthesis for Long-Horizon Terminal Tasks</title>
      <link>https://huggingface.co/papers/2608.05466</link>
      <guid isPermaLink="false">2026-w34</guid>
      <pubDate>Mon, 17 Aug 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Training terminal agents on long, multi-step workflows is notoriously expensive when tasks must be hand-crafted. This episode explores RST, a recursive method that grows verified seed tasks into harder descendants by extending reference solutions first, then aligning verifiers and instructions to match. The approach scaled a few hundred seeds to over 37,000 tasks at roughly five cents each. For practitioners building command-line agents, it offers a cost-effective path to diverse, verifiable training data.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2608.05466.pdf">https://arxiv.org/pdf/2608.05466.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W34/podcast_pro_40193aa4d6ae49a5a89fcd29fd629f99.mp3" length="9444525" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Training terminal agents on long, multi-step workflows is notoriously expensive when tasks must be hand-crafted. This episode explores RST, a recursive method that grows verified seed tasks into harder descendants by extending reference solutions first...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Training terminal agents on long, multi-step workflows is notoriously expensive when tasks must be hand-crafted. This episode explores RST, a recursive method that grows verified seed tasks into harder descendants by extending reference solutions first, then aligning verifiers and instructions to match. The approach scaled a few hundred seeds to over 37,000 tasks at roughly five cents each. For practitioners building command-line agents, it offers a cost-effective path to diverse, verifiable training data.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2608.05466.pdf">https://arxiv.org/pdf/2608.05466.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:52</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>33/2026 - Kimi K3: Open Frontier Intelligence</title>
      <link>https://huggingface.co/papers/2607.24653</link>
      <guid isPermaLink="false">2026-w33</guid>
      <pubDate>Mon, 10 Aug 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Open models have grown better at reasoning, but many still lag behind closed systems in raw foundational power. This episode explores Kimi K3, a 2.8 trillion parameter mixture-of-experts model that activates only 104 billion parameters per token, achieving 2.5x better scaling efficiency than its predecessor. The hosts unpack its three-axis architecture and million-token context training, discussing what it means for practitioners building agents that need to process massive codebases or documents.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2607.24653.pdf">https://arxiv.org/pdf/2607.24653.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W33/podcast_pro_3cf86c61fed840009186354c78915812.mp3" length="8386605" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Open models have grown better at reasoning, but many still lag behind closed systems in raw foundational power. This episode explores Kimi K3, a 2.8 trillion parameter mixture-of-experts model that activates only 104 billion parameters per token, achie...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Open models have grown better at reasoning, but many still lag behind closed systems in raw foundational power. This episode explores Kimi K3, a 2.8 trillion parameter mixture-of-experts model that activates only 104 billion parameters per token, achieving 2.5x better scaling efficiency than its predecessor. The hosts unpack its three-axis architecture and million-token context training, discussing what it means for practitioners building agents that need to process massive codebases or documents.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2607.24653.pdf">https://arxiv.org/pdf/2607.24653.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:59</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>32/2026 - ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU</title>
      <link>https://huggingface.co/papers/2607.19191</link>
      <guid isPermaLink="false">2026-w32</guid>
      <pubDate>Mon, 03 Aug 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>AI video generation is evolving from passive clips to interactive worlds users can actually control. This episode explores ABot-World-0, which achieves 720p interactive video at up to 16 fps on a single RTX 5090, using a full-stack approach that combines curated game recordings, simulation data, and an adaptive data collection system called WorldExplorer. For ML practitioners, this points toward locally-runnable world models for game prototyping, embodied AI training, and long-horizon video stability research.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2607.19191.pdf">https://arxiv.org/pdf/2607.19191.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W32/podcast_pro_ad2b2437fc27477e9839f578df001074.mp3" length="9450765" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>AI video generation is evolving from passive clips to interactive worlds users can actually control. This episode explores ABot-World-0, which achieves 720p interactive video at up to 16 fps on a single RTX 5090, using a full-stack approach that combin...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>AI video generation is evolving from passive clips to interactive worlds users can actually control. This episode explores ABot-World-0, which achieves 720p interactive video at up to 16 fps on a single RTX 5090, using a full-stack approach that combines curated game recordings, simulation data, and an adaptive data collection system called WorldExplorer. For ML practitioners, this points toward locally-runnable world models for game prototyping, embodied AI training, and long-horizon video stability research.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2607.19191.pdf">https://arxiv.org/pdf/2607.19191.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:52</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>31/2026 - Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable</title>
      <link>https://huggingface.co/papers/2607.13285</link>
      <guid isPermaLink="false">2026-w31</guid>
      <pubDate>Mon, 27 Jul 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>When AI agents evolve, the hardest part isn&#x27;t writing the code change—it&#x27;s finding where behavior actually lives in the codebase. This episode explores the Harness Handbook, a behavior-first documentation approach that maps agent actions to source code across three levels of abstraction, from system architecture down to specific code locations. For teams building and maintaining agent systems, this offers a practical framework for making harness code readable, navigable, and safely editable.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2607.13285.pdf">https://arxiv.org/pdf/2607.13285.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W31/podcast_pro_baca1be76cee45c59eb96f7ad842e681.mp3" length="9756525" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>When AI agents evolve, the hardest part isn't writing the code change—it's finding where behavior actually lives in the codebase. This episode explores the Harness Handbook, a behavior-first documentation approach that maps agent actions to source code...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>When AI agents evolve, the hardest part isn&#x27;t writing the code change—it&#x27;s finding where behavior actually lives in the codebase. This episode explores the Harness Handbook, a behavior-first documentation approach that maps agent actions to source code across three levels of abstraction, from system architecture down to specific code locations. For teams building and maintaining agent systems, this offers a practical framework for making harness code readable, navigable, and safely editable.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2607.13285.pdf">https://arxiv.org/pdf/2607.13285.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:08:08</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>30/2026 - The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning</title>
      <link>https://huggingface.co/papers/2606.29526</link>
      <guid isPermaLink="false">2026-w30</guid>
      <pubDate>Mon, 20 Jul 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>When training LLMs with reinforcement learning, the policy you optimize might not be the one you actually deploy—a mismatch that can silently derail performance. This episode explores how differences between training and inference engines cause instability, and introduces MIPI, a principle that judges updates by whether the deployed model actually improves. Essential listening for anyone building RL pipelines or reasoning systems who wants to avoid optimizing a mirage.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.29526.pdf">https://arxiv.org/pdf/2606.29526.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W30/podcast_pro_36e7f3a2ba644be1a0d93d575be414fe.mp3" length="8369805" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>When training LLMs with reinforcement learning, the policy you optimize might not be the one you actually deploy—a mismatch that can silently derail performance. This episode explores how differences between training and inference engines cause instabi...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>When training LLMs with reinforcement learning, the policy you optimize might not be the one you actually deploy—a mismatch that can silently derail performance. This episode explores how differences between training and inference engines cause instability, and introduces MIPI, a principle that judges updates by whether the deployed model actually improves. Essential listening for anyone building RL pipelines or reasoning systems who wants to avoid optimizing a mirage.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.29526.pdf">https://arxiv.org/pdf/2606.29526.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:58</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>29/2026 - Orca: The World is in Your Mind</title>
      <link>https://huggingface.co/papers/2606.30534</link>
      <guid isPermaLink="false">2026-w29</guid>
      <pubDate>Mon, 13 Jul 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Most AI systems predict tokens, frames, or actions separately—but what if they should predict how the world itself changes? This episode explores Orca, which learns a unified &quot;world latent&quot; representation from video, then uses lightweight readout modules for text, images, and robot control. The frozen backbone generalizes across all three tasks, and scaling curves suggest the approach hasn&#x27;t saturated yet. A compelling framework for practitioners building multimodal systems that need grounded world understanding.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.30534.pdf">https://arxiv.org/pdf/2606.30534.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W29/podcast_pro_11752784c1f241bc9c9250b34b0fb81a.mp3" length="9069165" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Most AI systems predict tokens, frames, or actions separately—but what if they should predict how the world itself changes? This episode explores Orca, which learns a unified &quot;world latent&quot; representation from video, then uses lightweight readout modul...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Most AI systems predict tokens, frames, or actions separately—but what if they should predict how the world itself changes? This episode explores Orca, which learns a unified &quot;world latent&quot; representation from video, then uses lightweight readout modules for text, images, and robot control. The frozen backbone generalizes across all three tasks, and scaling curves suggest the approach hasn&#x27;t saturated yet. A compelling framework for practitioners building multimodal systems that need grounded world understanding.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.30534.pdf">https://arxiv.org/pdf/2606.30534.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:33</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>28/2026 - MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision</title>
      <link>https://huggingface.co/papers/2606.17162</link>
      <guid isPermaLink="false">2026-w28</guid>
      <pubDate>Mon, 06 Jul 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>AI slide generators can produce polished decks, but they rarely remember your preferences across sessions or handle edits without breaking everything. MemSlides tackles this with a hierarchical memory system—separating long-term user profiles, session-specific working memory, and tool execution lessons—plus scoped local revision that targets only affected elements instead of regenerating entire decks. A useful framework for anyone building agents that need to personalize outputs while handling iterative feedback gracefully.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.17162.pdf">https://arxiv.org/pdf/2606.17162.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W28/podcast_pro_a39743a8ade9499286341531acc0fb7c.mp3" length="8188365" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>AI slide generators can produce polished decks, but they rarely remember your preferences across sessions or handle edits without breaking everything. MemSlides tackles this with a hierarchical memory system—separating long-term user profiles, session-...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>AI slide generators can produce polished decks, but they rarely remember your preferences across sessions or handle edits without breaking everything. MemSlides tackles this with a hierarchical memory system—separating long-term user profiles, session-specific working memory, and tool execution lessons—plus scoped local revision that targets only affected elements instead of regenerating entire decks. A useful framework for anyone building agents that need to personalize outputs while handling iterative feedback gracefully.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.17162.pdf">https://arxiv.org/pdf/2606.17162.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:49</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>27/2026 - Looped World Models</title>
      <link>https://huggingface.co/papers/2606.18208</link>
      <guid isPermaLink="false">2026-w27</guid>
      <pubDate>Mon, 29 Jun 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>World models help agents plan ahead, but small prediction errors compound over long horizons—and deeper models mean heavier compute costs. This episode explores LoopWM from FaceMind Research, which reuses a single transformer block in a loop to refine predictions rather than stacking separate layers, achieving up to 100x parameter efficiency. The approach also introduces adaptive early-exit gates that allocate more computation to complex transitions. Essential listening for anyone building model-based agents or simulation-heavy planning systems.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.18208.pdf">https://arxiv.org/pdf/2606.18208.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W27/podcast_pro_684809a6845c46bcb180e905e4af1ddc.mp3" length="7288365" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>World models help agents plan ahead, but small prediction errors compound over long horizons—and deeper models mean heavier compute costs. This episode explores LoopWM from FaceMind Research, which reuses a single transformer block in a loop to refine...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>World models help agents plan ahead, but small prediction errors compound over long horizons—and deeper models mean heavier compute costs. This episode explores LoopWM from FaceMind Research, which reuses a single transformer block in a loop to refine predictions rather than stacking separate layers, achieving up to 100x parameter efficiency. The approach also introduces adaptive early-exit gates that allocate more computation to complex transitions. Essential listening for anyone building model-based agents or simulation-heavy planning systems.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.18208.pdf">https://arxiv.org/pdf/2606.18208.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:04</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>26/2026 - ABot-Earth 0.5: Generative 3D Earth Model</title>
      <link>https://huggingface.co/papers/2606.09967</link>
      <guid isPermaLink="false">2026-w26</guid>
      <pubDate>Mon, 22 Jun 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Creating detailed 3D maps of the world remains expensive and slow, but ABot-Earth 0.5 proposes generating realistic environments directly from satellite imagery. The model outputs 3D Gaussian Splatting representations—millions of soft spatial blobs rather than clean polygons—enabling it to handle messy real-world scenes like trees and irregular rooftops. For ML practitioners working on embodied AI, drone navigation, or urban simulation, this offers a potential path to scalable, realistic training environments without costly aerial scanning.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.09967.pdf">https://arxiv.org/pdf/2606.09967.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W26/podcast_pro_decb5bc6412c4654a759e2a643354cc9.mp3" length="8302605" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Creating detailed 3D maps of the world remains expensive and slow, but ABot-Earth 0.5 proposes generating realistic environments directly from satellite imagery. The model outputs 3D Gaussian Splatting representations—millions of soft spatial blobs rat...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Creating detailed 3D maps of the world remains expensive and slow, but ABot-Earth 0.5 proposes generating realistic environments directly from satellite imagery. The model outputs 3D Gaussian Splatting representations—millions of soft spatial blobs rather than clean polygons—enabling it to handle messy real-world scenes like trees and irregular rooftops. For ML practitioners working on embodied AI, drone navigation, or urban simulation, this offers a potential path to scalable, realistic training environments without costly aerial scanning.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.09967.pdf">https://arxiv.org/pdf/2606.09967.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:55</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>25/2026 - On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters</title>
      <link>https://huggingface.co/papers/2606.02437</link>
      <guid isPermaLink="false">2026-w25</guid>
      <pubDate>Mon, 15 Jun 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>As AI assistants grow more capable, they still struggle with continuity—remembering user preferences and adapting behavior across sessions. This episode explores a new framework organizing parameter-efficient fine-tuning into three interdependent scales: scaling up base models so small adapters gain leverage, scaling down adapters for cheap reliable updates, and scaling out to support millions of persistent personal models. The discussion unpacks why reinforcement learning is &quot;prior-limited&quot; and how trillion-parameter mixture-of-experts architectures introduce tricky training-serving mismatches.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.02437.pdf">https://arxiv.org/pdf/2606.02437.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W25/podcast_pro_46944661be494e47a70c3c5892af7b08.mp3" length="7861005" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>As AI assistants grow more capable, they still struggle with continuity—remembering user preferences and adapting behavior across sessions. This episode explores a new framework organizing parameter-efficient fine-tuning into three interdependent scale...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>As AI assistants grow more capable, they still struggle with continuity—remembering user preferences and adapting behavior across sessions. This episode explores a new framework organizing parameter-efficient fine-tuning into three interdependent scales: scaling up base models so small adapters gain leverage, scaling down adapters for cheap reliable updates, and scaling out to support millions of persistent personal models. The discussion unpacks why reinforcement learning is &quot;prior-limited&quot; and how trillion-parameter mixture-of-experts architectures introduce tricky training-serving mismatches.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2606.02437.pdf">https://arxiv.org/pdf/2606.02437.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:33</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>24/2026 - Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players</title>
      <link>https://huggingface.co/papers/2605.28816</link>
      <guid isPermaLink="false">2026-w24</guid>
      <pubDate>Mon, 08 Jun 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>When multiple agents act simultaneously in a shared environment, video-based world models often collapse into inconsistency or computational chaos. This episode explores Gamma-World, which tackles synchronized cross-perspective consistency through Simplex Rotary Agent Encoding—a clever way to mark agent identities without privileging any particular ordering. The hosts break down why traditional &quot;stitching&quot; approaches fail and how symmetric geometric embeddings let models scale from two to four agents without retraining, offering practical insights for anyone building multi-agent simulation systems.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2605.28816.pdf">https://arxiv.org/pdf/2605.28816.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W24/podcast_pro_086bae0f5dd74c15affdb969463c1e02.mp3" length="8393805" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>When multiple agents act simultaneously in a shared environment, video-based world models often collapse into inconsistency or computational chaos. This episode explores Gamma-World, which tackles synchronized cross-perspective consistency through Simp...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>When multiple agents act simultaneously in a shared environment, video-based world models often collapse into inconsistency or computational chaos. This episode explores Gamma-World, which tackles synchronized cross-perspective consistency through Simplex Rotary Agent Encoding—a clever way to mark agent identities without privileging any particular ordering. The hosts break down why traditional &quot;stitching&quot; approaches fail and how symmetric geometric embeddings let models scale from two to four agents without retraining, offering practical insights for anyone building multi-agent simulation systems.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2605.28816.pdf">https://arxiv.org/pdf/2605.28816.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:00</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>23/2026 - CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence</title>
      <link>https://huggingface.co/papers/2605.12882</link>
      <guid isPermaLink="false">2026-w23</guid>
      <pubDate>Mon, 01 Jun 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Document QA models can produce correct-sounding answers while citing the wrong evidence—a hidden failure mode that current benchmarks miss entirely. CiteVQA tackles this with 1,897 questions across 711 multi-page PDFs, requiring models to provide element-level bounding-box citations for tables, figures, and text. The hosts unpack the clever masking ablation technique used to automatically identify crucial evidence at scale. Essential listening for teams building trustworthy AI in legal, financial, or medical document workflows.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2605.12882.pdf">https://arxiv.org/pdf/2605.12882.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W23/podcast_pro_e8da614a1ed04b52b4834f27127bd834.mp3" length="8569005" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Document QA models can produce correct-sounding answers while citing the wrong evidence—a hidden failure mode that current benchmarks miss entirely. CiteVQA tackles this with 1,897 questions across 711 multi-page PDFs, requiring models to provide eleme...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Document QA models can produce correct-sounding answers while citing the wrong evidence—a hidden failure mode that current benchmarks miss entirely. CiteVQA tackles this with 1,897 questions across 711 multi-page PDFs, requiring models to provide element-level bounding-box citations for tables, figures, and text. The hosts unpack the clever masking ablation technique used to automatically identify crucial evidence at scale. Essential listening for teams building trustworthy AI in legal, financial, or medical document workflows.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2605.12882.pdf">https://arxiv.org/pdf/2605.12882.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:08</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>22/2026 - MinT: Managed Infrastructure for Training and Serving Millions of LLMs</title>
      <link>https://huggingface.co/papers/2605.13779</link>
      <guid isPermaLink="false">2026-w22</guid>
      <pubDate>Mon, 25 May 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>As AI teams produce countless model variants through continuous post-training, copying full checkpoints for each one quickly becomes unmanageable. This episode explores MinT, a system that treats LoRA adapters—not full models—as the unit of behavior, keeping expensive base models resident in memory while swapping lightweight adapter revisions for training, serving, and rollback. Essential listening for platform engineers building infrastructure that needs to scale beyond a handful of model versions.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2605.13779.pdf">https://arxiv.org/pdf/2605.13779.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W22/podcast_pro_2e87ed392adc4b71b8d0df0260c8ba89.mp3" length="11056845" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>As AI teams produce countless model variants through continuous post-training, copying full checkpoints for each one quickly becomes unmanageable. This episode explores MinT, a system that treats LoRA adapters—not full models—as the unit of behavior, k...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>As AI teams produce countless model variants through continuous post-training, copying full checkpoints for each one quickly becomes unmanageable. This episode explores MinT, a system that treats LoRA adapters—not full models—as the unit of behavior, keeping expensive base models resident in memory while swapping lightweight adapter revisions for training, serving, and rollback. Essential listening for platform engineers building infrastructure that needs to scale beyond a handful of model versions.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2605.13779.pdf">https://arxiv.org/pdf/2605.13779.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:09:13</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>21/2026 - MolmoAct2: Action Reasoning Models for Real-world Deployment</title>
      <link>https://huggingface.co/papers/2605.02881</link>
      <guid isPermaLink="false">2026-w21</guid>
      <pubDate>Mon, 18 May 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>This week we spotlight MolmoAct2: Action Reasoning Models for Real-world Deployment out of Ai2. What&#x27;s the paper saying is broken with today&#x27;s vision-language-action models, in practical terms? And four, even after fine-tuning, success rates can still fall short of dependable deployment. They advance their prior system along five axes that line up with deployment needs. New open robot datasets on low-to-medium cost platforms.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2605.02881.pdf">https://arxiv.org/pdf/2605.02881.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W21/podcast_pro_3e0fd9f1d3d948d788ca3c1bc72e6861.mp3" length="9840525" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>This week we spotlight MolmoAct2: Action Reasoning Models for Real-world Deployment out of Ai2. What's the paper saying is broken with today's vision-language-action models, in practical terms? And four, even after fine-tuning, success rates can still...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>This week we spotlight MolmoAct2: Action Reasoning Models for Real-world Deployment out of Ai2. What&#x27;s the paper saying is broken with today&#x27;s vision-language-action models, in practical terms? And four, even after fine-tuning, success rates can still fall short of dependable deployment. They advance their prior system along five axes that line up with deployment needs. New open robot datasets on low-to-medium cost platforms.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2605.02881.pdf">https://arxiv.org/pdf/2605.02881.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:08:12</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>20/2026 - Recursive Multi-Agent Systems</title>
      <link>https://huggingface.co/papers/2604.25917</link>
      <guid isPermaLink="false">2026-w20</guid>
      <pubDate>Mon, 11 May 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Multi-agent AI systems are powerful but often grind to a halt under the weight of text-based communication between agents. This episode explores &quot;Recursive Multi-Agent Systems,&quot; which introduces RecursiveLink—a module that lets agents collaborate through continuous hidden representations instead of verbose text exchanges. The approach enables heterogeneous models to share internal states directly while keeping the base models frozen. For practitioners scaling multi-agent workflows, this offers a path to deeper collaboration without exploding latency and token costs.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2604.25917.pdf">https://arxiv.org/pdf/2604.25917.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W20/podcast_pro_019157c9a9244a47ab2c581a00e00fbf.mp3" length="10021965" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Multi-agent AI systems are powerful but often grind to a halt under the weight of text-based communication between agents. This episode explores &quot;Recursive Multi-Agent Systems,&quot; which introduces RecursiveLink—a module that lets agents collaborate throu...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Multi-agent AI systems are powerful but often grind to a halt under the weight of text-based communication between agents. This episode explores &quot;Recursive Multi-Agent Systems,&quot; which introduces RecursiveLink—a module that lets agents collaborate through continuous hidden representations instead of verbose text exchanges. The approach enables heterogeneous models to share internal states directly while keeping the base models frozen. For practitioners scaling multi-agent workflows, this offers a path to deeper collaboration without exploding latency and token costs.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2604.25917.pdf">https://arxiv.org/pdf/2604.25917.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:08:21</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>19/2026 - Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items</title>
      <link>https://huggingface.co/papers/2604.19748</link>
      <guid isPermaLink="false">2026-w19</guid>
      <pubDate>Mon, 04 May 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Virtual try-on systems often fail when faced with real-world messiness—cluttered backgrounds, awkward poses, and multiple garments at once. Tstars-Tryon 1.0 reframes the problem as general image editing, enabling coherent multi-item outfits across eight fashion categories while introducing a benchmark that evaluates identity, garment fidelity, and physical logic together. Already deployed to millions of Taobao users, this work offers ML practitioners a roadmap for bridging the gap between research demos and production-ready systems.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2604.19748.pdf">https://arxiv.org/pdf/2604.19748.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W19/podcast_pro_e734c19337c742dd9a01cc01507a6ae0.mp3" length="8131725" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Virtual try-on systems often fail when faced with real-world messiness—cluttered backgrounds, awkward poses, and multiple garments at once. Tstars-Tryon 1.0 reframes the problem as general image editing, enabling coherent multi-item outfits across eigh...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Virtual try-on systems often fail when faced with real-world messiness—cluttered backgrounds, awkward poses, and multiple garments at once. Tstars-Tryon 1.0 reframes the problem as general image editing, enabling coherent multi-item outfits across eight fashion categories while introducing a benchmark that evaluates identity, garment fidelity, and physical logic together. Already deployed to millions of Taobao users, this work offers ML practitioners a roadmap for bridging the gap between research demos and production-ready systems.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2604.19748.pdf">https://arxiv.org/pdf/2604.19748.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:46</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>18/2026 - WildDet3D: Scaling Promptable 3D Detection in the Wild</title>
      <link>https://huggingface.co/papers/2604.08626</link>
      <guid isPermaLink="false">2026-w18</guid>
      <pubDate>Mon, 27 Apr 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Detecting 3D objects from a single image typically breaks down outside narrow lab datasets with limited categories. WildDet3D introduces a unified model that accepts text, clicks, or 2D boxes as prompts and gracefully incorporates depth when available—all trained on a new million-image dataset spanning over 13,000 categories. For ML practitioners building robotics, AR, or mobile vision apps, this points toward truly general-purpose 3D perception.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2604.08626.pdf">https://arxiv.org/pdf/2604.08626.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W18/podcast_pro_c79105e3eb814dcf84eea54ec41fbc8d.mp3" length="9220365" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Detecting 3D objects from a single image typically breaks down outside narrow lab datasets with limited categories. WildDet3D introduces a unified model that accepts text, clicks, or 2D boxes as prompts and gracefully incorporates depth when available—...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Detecting 3D objects from a single image typically breaks down outside narrow lab datasets with limited categories. WildDet3D introduces a unified model that accepts text, clicks, or 2D boxes as prompts and gracefully incorporates depth when available—all trained on a new million-image dataset spanning over 13,000 categories. For ML practitioners building robotics, AR, or mobile vision apps, this points toward truly general-purpose 3D perception.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2604.08626.pdf">https://arxiv.org/pdf/2604.08626.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:41</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>17/2026 - Adam's Law: Textual Frequency Law on Large Language Models</title>
      <link>https://huggingface.co/papers/2604.02176</link>
      <guid isPermaLink="false">2026-w17</guid>
      <pubDate>Mon, 20 Apr 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>When prompts can be phrased multiple ways, which wording helps LLMs perform best? This episode explores &quot;Adam&#x27;s Law,&quot; which proposes that choosing higher-frequency expressions—more common words and phrasings—leads to more reliable model outputs. The researchers introduce Textual Frequency Law for selecting paraphrases and a distillation method that makes frequency estimates model-aware. For practitioners, it offers a simple tuning knob: rephrase prompts and training examples toward everyday language for better results without bigger models.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2604.02176.pdf">https://arxiv.org/pdf/2604.02176.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W17/podcast_pro_28766e094d8a4a82b26304ef7271e113.mp3" length="8961165" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>When prompts can be phrased multiple ways, which wording helps LLMs perform best? This episode explores &quot;Adam's Law,&quot; which proposes that choosing higher-frequency expressions—more common words and phrasings—leads to more reliable model outputs. The re...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>When prompts can be phrased multiple ways, which wording helps LLMs perform best? This episode explores &quot;Adam&#x27;s Law,&quot; which proposes that choosing higher-frequency expressions—more common words and phrasings—leads to more reliable model outputs. The researchers introduce Textual Frequency Law for selecting paraphrases and a distillation method that makes frequency estimates model-aware. For practitioners, it offers a simple tuning knob: rephrase prompts and training examples toward everyday language for better results without bigger models.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2604.02176.pdf">https://arxiv.org/pdf/2604.02176.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:28</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>16/2026 - DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models</title>
      <link>https://huggingface.co/papers/2603.26164</link>
      <guid isPermaLink="false">2026-w16</guid>
      <pubDate>Mon, 13 Apr 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Training large language models with smarter data strategies shouldn&#x27;t require weeks of engineering or rebuilding your entire pipeline. This episode explores DataFlex, a unified framework that treats data decisions—sample selection, domain mixing, and reweighting—as something you optimize alongside model weights rather than a one-time preprocessing step. For practitioners, the appeal is drop-in configs that work with existing workflows while supporting distributed training at scale.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.26164.pdf">https://arxiv.org/pdf/2603.26164.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W16/podcast_pro_ded0864e3a1f4ab1835847cc82bc4ebe.mp3" length="8287725" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Training large language models with smarter data strategies shouldn't require weeks of engineering or rebuilding your entire pipeline. This episode explores DataFlex, a unified framework that treats data decisions—sample selection, domain mixing, and r...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Training large language models with smarter data strategies shouldn&#x27;t require weeks of engineering or rebuilding your entire pipeline. This episode explores DataFlex, a unified framework that treats data decisions—sample selection, domain mixing, and reweighting—as something you optimize alongside model weights rather than a one-time preprocessing step. For practitioners, the appeal is drop-in configs that work with existing workflows while supporting distributed training at scale.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.26164.pdf">https://arxiv.org/pdf/2603.26164.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:54</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>15/2026 - MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding</title>
      <link>https://huggingface.co/papers/2603.22458</link>
      <guid isPermaLink="false">2026-w15</guid>
      <pubDate>Mon, 06 Apr 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Traditional OCR systems decode documents token by token, causing latency to scale with length and early errors to cascade through entire pages. This episode explores MinerU-Diffusion, which reframes document parsing as &quot;inverse rendering&quot; using masked diffusion to fill in and revise tokens in parallel rather than left-to-right. The block-wise decoding approach offers faster processing while reducing hallucination—critical for anyone building extraction pipelines in finance, legal, or scientific domains where accuracy matters.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.22458.pdf">https://arxiv.org/pdf/2603.22458.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W15/podcast_pro_eaf169c7a33e4db5bddc1213bb767824.mp3" length="7453005" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Traditional OCR systems decode documents token by token, causing latency to scale with length and early errors to cascade through entire pages. This episode explores MinerU-Diffusion, which reframes document parsing as &quot;inverse rendering&quot; using masked...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Traditional OCR systems decode documents token by token, causing latency to scale with length and early errors to cascade through entire pages. This episode explores MinerU-Diffusion, which reframes document parsing as &quot;inverse rendering&quot; using masked diffusion to fill in and revise tokens in parallel rather than left-to-right. The block-wise decoding approach offers faster processing while reducing hallucination—critical for anyone building extraction pipelines in finance, legal, or scientific domains where accuracy matters.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.22458.pdf">https://arxiv.org/pdf/2603.22458.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:13</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>14/2026 - AI Can Learn Scientific Taste</title>
      <link>https://huggingface.co/papers/2603.14473</link>
      <guid isPermaLink="false">2026-w14</guid>
      <pubDate>Mon, 30 Mar 2026 04:00:00 +0200</pubDate>
      <description><![CDATA[<p>Can AI develop the judgment to know which research directions are worth pursuing? This episode explores how the OpenMOSS team tackles &quot;scientific taste&quot; through Reinforcement Learning from Community Feedback, training models on 700K citation-based paper comparisons to predict long-term impact. Their Scientific Judge achieves mid-80s accuracy at picking higher-impact work. For ML practitioners, this points toward AI systems that filter signal from noise rather than just scaling idea generation.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.14473.pdf">https://arxiv.org/pdf/2603.14473.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W14/podcast_pro_ca5fa935623e4edda83f7a59fe905718.mp3" length="8368365" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Can AI develop the judgment to know which research directions are worth pursuing? This episode explores how the OpenMOSS team tackles &quot;scientific taste&quot; through Reinforcement Learning from Community Feedback, training models on 700K citation-based pape...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Can AI develop the judgment to know which research directions are worth pursuing? This episode explores how the OpenMOSS team tackles &quot;scientific taste&quot; through Reinforcement Learning from Community Feedback, training models on 700K citation-based paper comparisons to predict long-term impact. Their Scientific Judge achieves mid-80s accuracy at picking higher-impact work. For ML practitioners, this points toward AI systems that filter signal from noise rather than just scaling idea generation.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.14473.pdf">https://arxiv.org/pdf/2603.14473.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:58</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>13/2026 - Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning</title>
      <link>https://huggingface.co/papers/2603.04597</link>
      <guid isPermaLink="false">2026-w13</guid>
      <pubDate>Mon, 23 Mar 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>When language models only get binary success/fail feedback, they stumble through trial and error—especially when rewards are sparse. This episode explores GOLF, a framework that combines natural language critiques with insights from multiple failed attempts to bootstrap smarter exploration. The key innovation is adaptive refinement injection, which scaffolds training when models get stuck in low-reward regions. Essential listening for practitioners working on RL fine-tuning who want to move beyond simple reward signals.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.04597.pdf">https://arxiv.org/pdf/2603.04597.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W13/podcast_pro_1ce31bdc3c0f4e4f886de919121b9ea8.mp3" length="9047565" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>When language models only get binary success/fail feedback, they stumble through trial and error—especially when rewards are sparse. This episode explores GOLF, a framework that combines natural language critiques with insights from multiple failed att...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>When language models only get binary success/fail feedback, they stumble through trial and error—especially when rewards are sparse. This episode explores GOLF, a framework that combines natural language critiques with insights from multiple failed attempts to bootstrap smarter exploration. The key innovation is adaptive refinement injection, which scaffolds training when models get stuck in low-reward regions. Essential listening for practitioners working on RL fine-tuning who want to move beyond simple reward signals.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.04597.pdf">https://arxiv.org/pdf/2603.04597.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:32</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>12/2026 - Heterogeneous Agent Collaborative Reinforcement Learning</title>
      <link>https://huggingface.co/papers/2603.02604</link>
      <guid isPermaLink="false">2026-w12</guid>
      <pubDate>Mon, 16 Mar 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Training multiple RL models on the same task means paying the steep cost of rollout generation and verification over and over. This episode breaks down HACRL, a new paradigm where heterogeneous agents—different sizes, architectures, even tokenizers—share verified rollouts during training while still deploying independently. The key is HACPO&#x27;s capability-aware mechanisms that handle distribution shift across models. A practical win for teams running parallel experiments or maintaining model families.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.02604.pdf">https://arxiv.org/pdf/2603.02604.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W12/podcast_pro_27f57c85a60b4ee5bf0c820090b3fccd.mp3" length="8226765" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Training multiple RL models on the same task means paying the steep cost of rollout generation and verification over and over. This episode breaks down HACRL, a new paradigm where heterogeneous agents—different sizes, architectures, even tokenizers—sha...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Training multiple RL models on the same task means paying the steep cost of rollout generation and verification over and over. This episode breaks down HACRL, a new paradigm where heterogeneous agents—different sizes, architectures, even tokenizers—share verified rollouts during training while still deploying independently. The key is HACPO&#x27;s capability-aware mechanisms that handle distribution shift across models. A practical win for teams running parallel experiments or maintaining model families.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2603.02604.pdf">https://arxiv.org/pdf/2603.02604.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:51</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>11/2026 - A Very Big Video Reasoning Suite</title>
      <link>https://huggingface.co/papers/2602.20159</link>
      <guid isPermaLink="false">2026-w11</guid>
      <pubDate>Mon, 09 Mar 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Video models can generate stunning visuals, but can they actually reason? This episode explores the VBVR suite, which introduces a massive dataset of over one million video clips spanning 200 reasoning tasks—roughly a thousand times larger than previous benchmarks. The key innovation is deterministic task generators paired with rule-based scoring that correlates above 0.9 with human judgment, eliminating subjective model-as-judge evaluation. Essential listening for teams building or evaluating video generation systems.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2602.20159.pdf">https://arxiv.org/pdf/2602.20159.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W11/podcast_pro_ed415e9c35964731b5bc814dd6a9e1b8.mp3" length="8936205" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Video models can generate stunning visuals, but can they actually reason? This episode explores the VBVR suite, which introduces a massive dataset of over one million video clips spanning 200 reasoning tasks—roughly a thousand times larger than previou...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Video models can generate stunning visuals, but can they actually reason? This episode explores the VBVR suite, which introduces a massive dataset of over one million video clips spanning 200 reasoning tasks—roughly a thousand times larger than previous benchmarks. The key innovation is deterministic task generators paired with rule-based scoring that correlates above 0.9 with human judgment, eliminating subjective model-as-judge evaluation. Essential listening for teams building or evaluating video generation systems.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2602.20159.pdf">https://arxiv.org/pdf/2602.20159.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:27</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>10/2026 - Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs</title>
      <link>https://huggingface.co/papers/2602.10388</link>
      <guid isPermaLink="false">2026-w10</guid>
      <pubDate>Mon, 02 Mar 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Scaling synthetic data often fails because surface-level diversity doesn&#x27;t translate to meaningful model improvements. This episode explores Feature Activation Coverage, a metric that measures diversity where it matters—inside the model&#x27;s learned feature space using sparse autoencoders. The authors&#x27; FAC Synthesis method targets missing capabilities rather than adding volume, generating data that fills genuine coverage gaps. A practical framework for teams looking to do more with less synthetic data.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2602.10388.pdf">https://arxiv.org/pdf/2602.10388.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W10/podcast_pro_10b139d135ea4e5fb0641990b7eadd46.mp3" length="7666125" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Scaling synthetic data often fails because surface-level diversity doesn't translate to meaningful model improvements. This episode explores Feature Activation Coverage, a metric that measures diversity where it matters—inside the model's learned featu...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Scaling synthetic data often fails because surface-level diversity doesn&#x27;t translate to meaningful model improvements. This episode explores Feature Activation Coverage, a metric that measures diversity where it matters—inside the model&#x27;s learned feature space using sparse autoencoders. The authors&#x27; FAC Synthesis method targets missing capabilities rather than adding volume, generating data that fills genuine coverage gaps. A practical framework for teams looking to do more with less synthetic data.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2602.10388.pdf">https://arxiv.org/pdf/2602.10388.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:23</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>09/2026 - OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration</title>
      <link>https://huggingface.co/papers/2602.05400</link>
      <guid isPermaLink="false">2026-w09</guid>
      <pubDate>Mon, 23 Feb 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>As high-quality training data grows scarce, the focus shifts from &quot;more tokens&quot; to &quot;better tokens&quot;—but most data selection methods ignore how modern optimizers actually update models. OPUS scores training samples in the optimizer&#x27;s own coordinate space, using a clever &quot;ghost&quot; computation trick to keep costs manageable. The approach shows gains across benchmarks without teaching to the test, offering practitioners a principled way to squeeze more value from limited data budgets.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2602.05400.pdf">https://arxiv.org/pdf/2602.05400.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W09/podcast_pro_770254fb3ee14e85944d4e8d97c796c3.mp3" length="7726125" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>As high-quality training data grows scarce, the focus shifts from &quot;more tokens&quot; to &quot;better tokens&quot;—but most data selection methods ignore how modern optimizers actually update models. OPUS scores training samples in the optimizer's own coordinate space...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>As high-quality training data grows scarce, the focus shifts from &quot;more tokens&quot; to &quot;better tokens&quot;—but most data selection methods ignore how modern optimizers actually update models. OPUS scores training samples in the optimizer&#x27;s own coordinate space, using a clever &quot;ghost&quot; computation trick to keep costs manageable. The approach shows gains across benchmarks without teaching to the test, offering practitioners a principled way to squeeze more value from limited data budgets.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2602.05400.pdf">https://arxiv.org/pdf/2602.05400.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:26</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>08/2026 - Green-VLA: Staged Vision-Language-Action Model for Generalist Robots</title>
      <link>https://huggingface.co/papers/2602.00919</link>
      <guid isPermaLink="false">2026-w08</guid>
      <pubDate>Mon, 16 Feb 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Getting robots from impressive demos to dependable real-world deployment remains a stubborn challenge. This episode explores Green-VLA, a staged vision-language-action model that tackles chaotic multi-robot data through semantic action slots with masking, and breaks through behavior cloning&#x27;s plateau using reinforcement learning alignment. The team deploys on a high-DOF humanoid with bimanual dexterous manipulation—offering ML practitioners a practical curriculum for training generalist robot policies that actually transfer.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2602.00919.pdf">https://arxiv.org/pdf/2602.00919.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W08/podcast_pro_5f414294f2f94c89b991b53ecaa09196.mp3" length="8675565" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Getting robots from impressive demos to dependable real-world deployment remains a stubborn challenge. This episode explores Green-VLA, a staged vision-language-action model that tackles chaotic multi-robot data through semantic action slots with maski...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Getting robots from impressive demos to dependable real-world deployment remains a stubborn challenge. This episode explores Green-VLA, a staged vision-language-action model that tackles chaotic multi-robot data through semantic action slots with masking, and breaks through behavior cloning&#x27;s plateau using reinforcement learning alignment. The team deploys on a high-DOF humanoid with bimanual dexterous manipulation—offering ML practitioners a practical curriculum for training generalist robot policies that actually transfer.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2602.00919.pdf">https://arxiv.org/pdf/2602.00919.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:14</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>07/2026 - Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs</title>
      <link>https://huggingface.co/papers/2601.17058</link>
      <guid isPermaLink="false">2026-w07</guid>
      <pubDate>Mon, 09 Feb 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Data preparation remains a stubborn bottleneck, with messy, inconsistent datasets draining resources and breaking pipelines. This episode explores a new survey examining how LLMs are shifting data prep from brittle rule-based systems to prompt-driven, context-aware workflows across three core tasks: cleaning, integration, and enrichment. For ML practitioners tired of regex nightmares and hand-tuned pipelines, the paper offers a practical taxonomy for where LLMs actually deliver production-ready results.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2601.17058.pdf">https://arxiv.org/pdf/2601.17058.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W07/podcast_pro_319a7181f833493bbd763e4f104ff1f8.mp3" length="8537325" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Data preparation remains a stubborn bottleneck, with messy, inconsistent datasets draining resources and breaking pipelines. This episode explores a new survey examining how LLMs are shifting data prep from brittle rule-based systems to prompt-driven,...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Data preparation remains a stubborn bottleneck, with messy, inconsistent datasets draining resources and breaking pipelines. This episode explores a new survey examining how LLMs are shifting data prep from brittle rule-based systems to prompt-driven, context-aware workflows across three core tasks: cleaning, integration, and enrichment. For ML practitioners tired of regex nightmares and hand-tuned pipelines, the paper offers a practical taxonomy for where LLMs actually deliver production-ready results.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2601.17058.pdf">https://arxiv.org/pdf/2601.17058.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:07</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>06/2026 - Agentic Reasoning for Large Language Models</title>
      <link>https://huggingface.co/papers/2601.12538</link>
      <guid isPermaLink="false">2026-w06</guid>
      <pubDate>Mon, 02 Feb 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Large language models excel at isolated puzzles but often struggle when tasks require multiple steps, real-time adaptation, or tool use. This survey introduces &quot;agentic reasoning&quot; as a framework where reasoning becomes the control center for a perceive-plan-act-verify loop, organizing approaches across three layers: foundational single-agent reasoning, self-evolving adaptation through feedback, and multi-agent collaboration. For practitioners, it clarifies when to invest in prompt engineering versus post-training optimization.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2601.12538.pdf">https://arxiv.org/pdf/2601.12538.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W06/podcast_pro_7ad1301ab2474d38bf7c19ab712bd9e5.mp3" length="7712205" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Large language models excel at isolated puzzles but often struggle when tasks require multiple steps, real-time adaptation, or tool use. This survey introduces &quot;agentic reasoning&quot; as a framework where reasoning becomes the control center for a perceive...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Large language models excel at isolated puzzles but often struggle when tasks require multiple steps, real-time adaptation, or tool use. This survey introduces &quot;agentic reasoning&quot; as a framework where reasoning becomes the control center for a perceive-plan-act-verify loop, organizing approaches across three layers: foundational single-agent reasoning, self-evolving adaptation through feedback, and multi-agent collaboration. For practitioners, it clarifies when to invest in prompt engineering versus post-training optimization.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2601.12538.pdf">https://arxiv.org/pdf/2601.12538.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:26</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>05/2026 - Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning</title>
      <link>https://huggingface.co/papers/2601.06943</link>
      <guid isPermaLink="false">2026-w05</guid>
      <pubDate>Mon, 26 Jan 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>When a video contains only subtle clues—a sign, a logo, a location detail—but the real answer lives somewhere on the web, how should AI systems bridge that gap? This episode explores VideoDR, a new benchmark that tests whether models can watch videos, extract the right visual anchors, and verify answers through web search. The surprising finding: agentic approaches don&#x27;t automatically beat structured workflows, especially when models lose track of their initial clues during long search chains.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2601.06943.pdf">https://arxiv.org/pdf/2601.06943.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W05/podcast_pro_df2eca3e1845487e9c59e4ecabadabb8.mp3" length="8370765" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>When a video contains only subtle clues—a sign, a logo, a location detail—but the real answer lives somewhere on the web, how should AI systems bridge that gap? This episode explores VideoDR, a new benchmark that tests whether models can watch videos,...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>When a video contains only subtle clues—a sign, a logo, a location detail—but the real answer lives somewhere on the web, how should AI systems bridge that gap? This episode explores VideoDR, a new benchmark that tests whether models can watch videos, extract the right visual anchors, and verify answers through web search. The surprising finding: agentic approaches don&#x27;t automatically beat structured workflows, especially when models lose track of their initial clues during long search chains.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2601.06943.pdf">https://arxiv.org/pdf/2601.06943.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:58</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>04/2026 - GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization</title>
      <link>https://huggingface.co/papers/2601.05242</link>
      <guid isPermaLink="false">2026-w04</guid>
      <pubDate>Mon, 19 Jan 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>When training language models with multiple rewards simultaneously, standard GRPO normalization can collapse distinct outcomes into identical gradient signals—making it impossible for models to distinguish &quot;good&quot; from &quot;great.&quot; This episode breaks down GDPO, a new approach that normalizes each reward separately before combining them, preserving the learning signal across dimensions. Essential listening for anyone building RL-tuned assistants, agents, or tutors that need to satisfy multiple objectives without training instability.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2601.05242.pdf">https://arxiv.org/pdf/2601.05242.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W04/podcast_pro_e6611164926c477d81c47471e6a65d91.mp3" length="8837805" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>When training language models with multiple rewards simultaneously, standard GRPO normalization can collapse distinct outcomes into identical gradient signals—making it impossible for models to distinguish &quot;good&quot; from &quot;great.&quot; This episode breaks down...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>When training language models with multiple rewards simultaneously, standard GRPO normalization can collapse distinct outcomes into identical gradient signals—making it impossible for models to distinguish &quot;good&quot; from &quot;great.&quot; This episode breaks down GDPO, a new approach that normalizes each reward separately before combining them, preserving the learning signal across dimensions. Essential listening for anyone building RL-tuned assistants, agents, or tutors that need to satisfy multiple objectives without training instability.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2601.05242.pdf">https://arxiv.org/pdf/2601.05242.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:22</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>03/2026 - mHC: Manifold-Constrained Hyper-Connections</title>
      <link>https://huggingface.co/papers/2512.24880</link>
      <guid isPermaLink="false">2026-w03</guid>
      <pubDate>Mon, 12 Jan 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Residual connections keep deep networks stable, but Hyper-Connections—which boost performance by adding parallel &quot;lanes&quot;—can break that stability and cause training to blow up at scale. This episode explores mHC, which constrains residual mixing to doubly stochastic matrices using Sinkhorn-Knopp normalization, keeping signals from exploding across layers. For practitioners training large models, this could mean the difference between smooth runs and mysterious loss spikes.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.24880.pdf">https://arxiv.org/pdf/2512.24880.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W03/podcast_pro_89ed3aa2284845bf84451795cc603b1e.mp3" length="8897805" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Residual connections keep deep networks stable, but Hyper-Connections—which boost performance by adding parallel &quot;lanes&quot;—can break that stability and cause training to blow up at scale. This episode explores mHC, which constrains residual mixing to dou...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Residual connections keep deep networks stable, but Hyper-Connections—which boost performance by adding parallel &quot;lanes&quot;—can break that stability and cause training to blow up at scale. This episode explores mHC, which constrains residual mixing to doubly stochastic matrices using Sinkhorn-Knopp normalization, keeping signals from exploding across layers. For practitioners training large models, this could mean the difference between smooth runs and mysterious loss spikes.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.24880.pdf">https://arxiv.org/pdf/2512.24880.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:07:25</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>02/2026 - DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI</title>
      <link>https://huggingface.co/papers/2512.16676</link>
      <guid isPermaLink="false">2026-w02</guid>
      <pubDate>Mon, 05 Jan 2026 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Modern LLM development depends on complex data pipelines, yet most teams still cobble them together with ad-hoc scripts—breaking reproducibility and limiting iteration. This episode explores DataFlow, a framework that formalizes data preparation through nearly 200 reusable operators and an agent that converts natural-language specs into executable pipelines. For practitioners tired of reinventing data workflows, this offers a path toward composable, shareable data manufacturing.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.16676.pdf">https://arxiv.org/pdf/2512.16676.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W02/podcast_pro_d646133e00284b30ba41a889c0fe6b8f.mp3" length="9858765" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Modern LLM development depends on complex data pipelines, yet most teams still cobble them together with ad-hoc scripts—breaking reproducibility and limiting iteration. This episode explores DataFlow, a framework that formalizes data preparation throug...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Modern LLM development depends on complex data pipelines, yet most teams still cobble them together with ad-hoc scripts—breaking reproducibility and limiting iteration. This episode explores DataFlow, a framework that formalizes data preparation through nearly 200 reusable operators and an agent that converts natural-language specs into executable pipelines. For practitioners tired of reinventing data workflows, this offers a path toward composable, shareable data manufacturing.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.16676.pdf">https://arxiv.org/pdf/2512.16676.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:08:13</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>01/2026 - Kling-Omni Technical Report</title>
      <link>https://huggingface.co/papers/2512.16776</link>
      <guid isPermaLink="false">2026-w01</guid>
      <pubDate>Mon, 29 Dec 2025 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Video generation today is fragmented across dozens of specialized tools, but Kuaishou&#x27;s Kling-Omni proposes a unified framework that handles text-to-video, editing, and visual reasoning in one system. The key innovation is their Multimodal Visual Language paradigm, which lets users combine text prompts with image and video references to specify exactly what they want. For ML practitioners, this points toward a future where generalist models replace brittle multi-stage pipelines.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.16776.pdf">https://arxiv.org/pdf/2512.16776.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2026-W01/podcast_pro_ac7e02b0655f4b73968ff433ed6b4a92.mp3" length="14824845" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Video generation today is fragmented across dozens of specialized tools, but Kuaishou's Kling-Omni proposes a unified framework that handles text-to-video, editing, and visual reasoning in one system. The key innovation is their Multimodal Visual Langu...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Video generation today is fragmented across dozens of specialized tools, but Kuaishou&#x27;s Kling-Omni proposes a unified framework that handles text-to-video, editing, and visual reasoning in one system. The key innovation is their Multimodal Visual Language paradigm, which lets users combine text prompts with image and video references to specify exactly what they want. For ML practitioners, this points toward a future where generalist models replace brittle multi-stage pipelines.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.16776.pdf">https://arxiv.org/pdf/2512.16776.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:12:21</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>52/2025 - DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI</title>
      <link>https://huggingface.co/papers/2512.16676</link>
      <guid isPermaLink="false">2025-w52</guid>
      <pubDate>Mon, 22 Dec 2025 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Data pipelines for LLM training are often a mess of ad-hoc scripts that are hard to reproduce or trust. This episode explores DataFlow, a framework that treats data preparation as composable operators with explicit dependencies, plus an LLM-powered agent that can turn natural language requests into executable pipelines—even synthesizing missing operators on the fly. Essential listening for anyone building training data at scale.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.16676.pdf">https://arxiv.org/pdf/2512.16676.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2025-W52/podcast_pro_dff5a91d88d54f45bf1c78cc37602dc7.mp3" length="7884525" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Data pipelines for LLM training are often a mess of ad-hoc scripts that are hard to reproduce or trust. This episode explores DataFlow, a framework that treats data preparation as composable operators with explicit dependencies, plus an LLM-powered age...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Data pipelines for LLM training are often a mess of ad-hoc scripts that are hard to reproduce or trust. This episode explores DataFlow, a framework that treats data preparation as composable operators with explicit dependencies, plus an LLM-powered agent that can turn natural language requests into executable pipelines—even synthesizing missing operators on the fly. Essential listening for anyone building training data at scale.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.16676.pdf">https://arxiv.org/pdf/2512.16676.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:06:34</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
    <item>
      <title>51/2025 - Kling-Omni Technical Report</title>
      <link>https://huggingface.co/papers/2512.16776</link>
      <guid isPermaLink="false">2025-w51</guid>
      <pubDate>Mon, 15 Dec 2025 04:00:00 +0100</pubDate>
      <description><![CDATA[<p>Video AI pipelines often require stitching together multiple tools to go from creative intent to polished output. Kling-Omni introduces a unified approach using multimodal visual language (MVL), where text, reference images, and video snippets become a single representation for generation and editing. The system&#x27;s Prompt Enhancer uses reasoning-guided transformation trained with reinforcement learning to improve identity preservation and physical plausibility. ML practitioners working on generative video will find insights on bridging the gap between messy human instructions and coherent cinematic results.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.16776.pdf">https://arxiv.org/pdf/2512.16776.pdf</a></p>]]></description>
      <enclosure url="http://65.21.60.15:8080/2025-W51/podcast_pro_ce6c85f7339b4df3afbcc08585ccb50f.mp3" length="9695565" type="audio/mpeg"/>
      <itunes:author>dietrich.ai</itunes:author>
      <itunes:subtitle>Video AI pipelines often require stitching together multiple tools to go from creative intent to polished output. Kling-Omni introduces a unified approach using multimodal visual language (MVL), where text, reference images, and video snippets become a...</itunes:subtitle>
      <itunes:summary><![CDATA[<p>Video AI pipelines often require stitching together multiple tools to go from creative intent to polished output. Kling-Omni introduces a unified approach using multimodal visual language (MVL), where text, reference images, and video snippets become a single representation for generation and editing. The system&#x27;s Prompt Enhancer uses reasoning-guided transformation trained with reinforcement learning to improve identity preservation and physical plausibility. ML practitioners working on generative video will find insights on bridging the gap between messy human instructions and coherent cinematic results.</p><p>Link to the paper discussed in this episode: <a href="https://arxiv.org/pdf/2512.16776.pdf">https://arxiv.org/pdf/2512.16776.pdf</a></p>]]></itunes:summary>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:duration>00:08:05</itunes:duration>
      <itunes:explicit>false</itunes:explicit>
    </item>
  </channel>
</rss>