Playbook · 8 min read

How to Use AI for Podcast Production & Clips (2026 Guide)

A founder's guide to using AI tools like Descript, Riverside, and NotebookLM to produce podcasts, auto-generate clips, and ship episodes faster.

We record The Inference Podcast in person in Birmingham. Every episode is a single, uninterrupted conversation with a founder, 60 to 120 minutes of raw footage that has to turn into a full episode, chapters, show notes, a transcript, and a dozen short-form clips for LinkedIn, YouTube Shorts, X, and Instagram.

Two years ago that took a producer a week per episode. In 2026 it's mostly AI, and a founder-operator can ship it in an afternoon. Here's the stack we use and how each piece fits together.

The 2026 AI podcast stack

Think of production as four jobs. Pick one tool per job and don't over-engineer it.

  1. Recording & separation: Riverside, Descript Rooms, or a local Zoom H6 for in-person.
  2. Editing & cleanup: Descript for transcript-driven cuts and filler-word removal.
  3. Clip generation: Opus Clip or Descript for scored, captioned short-form.
  4. Prep & show notes: NotebookLM for guest research; Claude or GPT for summaries and titles.

How to use AI to make podcast clips

The single highest-leverage AI workflow in podcasting is clip generation. Here's the process that works for us.

  1. Upload the full episode to Opus Clip or Descript. Both accept a raw MP4 up to a few hours long.
  2. Let the model score the moments. The tool ranks segments by hook strength, emotional beat, and quotability. Ignore the score at first, scan the transcript and mark the moments you know landed in the room.
  3. Trim to 45–75 seconds. Shorter for TikTok and Shorts, up to 90 seconds for LinkedIn. Always cut on a complete thought.
  4. Auto-caption, then fix names. AI captions are ~97% accurate but they mangle founder names, company names, and acronyms every time. Fix them by hand, it's the one non-negotiable step.
  5. Frame vertical, add a bold hook overlay. The first 1.5 seconds decides everything. A one-line question overlay ("Why is vertical AI eating horizontal?") outperforms speaker labels 3-to-1 in our data.
  6. Batch export, schedule with Buffer or Metricool. One episode → 8–12 clips → 4 weeks of content.

Using NotebookLM for guest prep

NotebookLM is not a production tool: it's a research tool. Before every recording we build a notebook with:

  • The guest's LinkedIn profile (exported to PDF).
  • Their last 3–5 podcast appearances (transcripts).
  • Their company's blog posts and product docs.
  • Any published papers, decks, or investor letters.

Then ask it three questions: "What contradictions appear across these sources?", "Which questions has this guest never been asked?", and "What would a founder-to-founder audience find non-obvious here?" The output is your prep doc.

Automating show notes and chapters

Feed the transcript to Claude or GPT with a strict prompt: return chapter markers with timestamps, a 150-word episode summary, five pull-quotes, and a title A/B pair optimised for YouTube CTR. Paste output into YouTube Studio, done.

One tip: give the model your show's tone-of-voice as three example episodes' show notes. Otherwise every summary reads like a LinkedIn thought-leader post.

What AI still can't do

AI won't decide what your show is about. It won't tell you which guests to book. It won't feel the room when a founder is about to say something they've never said before. Those are the parts that make the show worth listening to, and they're still 100% human.

Use AI to compress the 20 hours of production work into 2, and spend the 18 you get back on booking better guests and asking sharper questions.

Watch the show that uses this exact workflow

Every episode of The Inference Podcast is edited, clipped, and shipped with the stack above.

Watch the latest episode