# How to Clip a Two-Person Podcast Into Shorts

> Split screen, one face or the wide shot? How to clip a two-person podcast into Shorts: a layout table, 1080p-vs-4K crop math and captions for two voices.

HTML version: https://everpop.app/blog/how-to-clip-a-two-person-podcast-into-shorts

Published: 2026-10-01
Updated: 2026-10-01
By Everpop

**To clip a two-person podcast into Shorts, choose the layout per moment. Crop to one speaker when one person carries the answer, cutting to the other for a reaction that matters. Stack both in a top-and-bottom split screen when the exchange is the point. Stay wide only when neither crop holds it, and caption whoever is off screen.**

Clipping a two-person episode is a geometry problem: a vertical frame holds one face, and the centre of a wide two-shot holds whatever sat between them. A microphone, a lamp, a bookshelf.

## Which speaker should you show in a two-person podcast Short?

Show whoever is carrying the point, such as the guest giving the answer, and cut to the other only on a real turn change or a reaction that changes the meaning: a laugh, a wince, a disagreement. A listener's "right" or "mm-hm" is not a turn; cutting to every one makes a clip twitch.

When a clip is a one-line question and a long answer, open on the asker if the question is the hook; if the answer's first line is stronger, start on the guest and caption the question as tagged audio. Choosing the moment itself is covered in [turn a long interview into Shorts](/blog/turn-a-long-interview-into-shorts).

## Split screen or single speaker: which layout works for podcast Shorts?

Match the layout to the moment: one person's story suits a single-speaker crop; an argument full of interruptions suits a split. The last two columns are arithmetic for the widest full-height window; zooming in enlarges more.

| Layout | Use it when | Trade-off | Share of 16:9 frame kept | From 1080p |
|---|---|---|---|---|
| Single-speaker crop (9:16) | One person carries the answer | Reactions need cuts; off-screen voice needs a tag | 31.6% of width | Enlarged 1.78x |
| Stacked split, camera per person | Back-and-forth, or a telling reaction | Half the screen per face | 63.3% of each width | Reduced 0.89x |
| Stacked split, one wide two-shot | The same, one camera | Needs a gap between them | Half the width, 79% of height, each | Enlarged 1.125x |
| Square (1:1) or 4:5 two-shot | The two sit close together | Screen partly empty | 56.25% or 45% of width | 1.0x or 1.25x |
| Wide 16:9 inside 9:16 | A shared look no crop holds | Picture is 31.6% of the height | All | Reduced 0.56x |

Of the layouts that fill the screen, only a stacked split from a camera per person needs no enlargement from 1080p. For single-speaker crops, record in 4K: from 3840x2160 the crop is a 1215x2160 window, larger than the 1080x1920 output. Crop mechanics are in [how to reframe horizontal video to vertical](/blog/how-to-reframe-horizontal-video-to-vertical).

Square and 4:5 clips qualify for longer Shorts too: YouTube's change to Shorts "up to 3 minutes long" covers videos "square or taller in aspect ratio" ([YouTube Blog, 2024](https://blog.youtube/news-and-events/tall-updates-coming-to-shorts/)).

## How do you make a stacked split-screen podcast Short?

A stacked split screen is a 9:16 frame divided into two 9:8 panels: 1080x960 each on a 1080x1920 timeline, with the seam at mid-height. Any editor with tracks can build one.

1. Put each camera on its own track.
2. Scale and crop one camera into each panel.
3. Put whoever speaks most on top, so a caption band at the seam sits under them: [DCMP's Captioning Key](https://dcmp.org/learn/603-captioning-key---speaker-identification) identifies an on-screen speaker "by placing the caption under the speaker."
4. Frame the bottom speaker's face high in its panel: most of that panel falls in the bottom margin of Google's safe-zone template, where the player's text and buttons can cover it ([where to put captions on a Short](/blog/where-to-put-captions-so-youtube-ui-doesnt-cover-them)).
5. Cut back to a single-speaker crop for a long answer; a silent face in half the screen adds nothing.

YouTube announced a test of recomposition tools to "adjust the layout, zoom, and crop", with "Split screen effects" ([YouTube Blog, 2023](https://blog.youtube/news-and-events/6-new-youtube-shorts-creation-tools/)), then Auto layout on Android, "which automatically tracks the main subject" ([YouTube Blog, 2024](https://blog.youtube/news-and-events/6-new-youtube-shorts-tools/)). The announcement does not say whether it follows whoever is talking, so check its choice on a two-shot before posting.

## How do you caption a Short when two people are talking?

Captions for two voices should say who is speaking whenever the picture does not. The [W3C](https://www.w3.org/WAI/WCAG22/Understanding/captions-prerecorded.html) says captions "identify who is speaking"; [Section 508](https://www.section508.gov/create/captions-transcripts/) says: "Clearly communicate who is speaking when it is not immediately obvious from the video."

- **Single-speaker crop:** tag the off-screen voice. [DCMP](https://dcmp.org/learn/603-captioning-key---speaker-identification) puts a known speaker's name in parentheses, "on a line of its own, separate from the captions": (Maya), then "What went wrong at launch?" below.
- **Stacked split:** a caption at the seam reads as the top speaker's, so tag the bottom speaker's lines.
- **Cross-talk:** multi-speaker speech recognition "faces significant challenges in transcribing overlapped speech", states a 2025 [arXiv paper](https://arxiv.org/abs/2506.05796). Start the clip where one voice takes the floor; if the overlap is the moment, use the split and check every captioned word by hand.

## How should you record a podcast so it clips well to vertical?

Record one matched camera per person at eye level, keep each face inside the centre third of its frame with headroom above, and give each voice its own microphone track, so every crop is already inside the shot and every overlap can be mixed.

- **Why the centre third:** a 9:16 window keeps 31.6% of the width, a little less than a third, so leave a margin.
- **Why headroom:** a split from one wide shot keeps about 79% of the height, so a head framed tight to the top leaves the crop no room without cutting the crown or the chin.
- **Recording remotely?** Save each person's camera as its own file, if your call software allows, not one [gallery view](/blog/how-to-clip-a-zoom-recording-into-shorts).
- **Only one wide camera?** Record in 4K with a clear gap between the two, so each half-frame holds one whole person.

## Do you need your guest's permission to post podcast clips?

Ask for it: get your guest's agreement to short clips, naming the platforms, when they agree to the episode, ideally before you record. A guest who agreed to a long conversation may not expect one sentence of it on a platform nobody mentioned, so send them the finished clips before anything posts. Good practice, not legal advice.

## Can Everpop clip a two-person podcast?

Yes: Everpop, an AI clipping tool that turns your uploaded episode into captioned vertical clips and publishes them review-first, frames each shot on one person and does not build a split screen. It burns in word-by-word captions, and any clip exports to Final Cut Pro, Premiere Pro or DaVinci Resolve as FCPXML, EDL and SRT, so you can finish a stacked split or a speaker tag there ([pricing](https://everpop.app/pricing)).

How it frames two people depends on the file you upload ([podcast to Shorts](https://everpop.app/podcast-to-shorts)). From an edited multicam export, each shot longer than a second or two is framed on the person it holds, up to eight shots per clip, though in a faster-cut clip some shots keep the framing of the shot before. From a single wide camera, it holds the frame on one person, usually the most prominent face, and does not yet follow who is talking: upload the multicam edit instead, or re-render those clips on any plan with Framing set to "Keep the whole picture".

Upload the episode or link a Google Drive folder; one episode counts as one video, with files up to 32 GB on Pro (8 GB on Starter, 4 GB on Free). For a file over 32 GB, the page suggests re-exporting at 1080p or 720p, which gives up the 4K headroom a single-speaker crop needs, or trimming "to the part you want", which keeps it.

Clips come in 9:16, 1:1, 16:9 or 4:5 and wait for your approval unless you switch on auto-post to YouTube, Instagram and Facebook (Pro and Scale); TikTok gets an inbox draft, and going public stays your tap ([pricing](https://everpop.app/pricing)). Agencies can give each client a workspace folder for its YouTube channels ([Everpop for agencies](https://everpop.app/ai-clipping-for-agencies)); on Pro or Scale, [an AI agent can upload and clip](/blog/can-chatgpt-or-claude-clip-your-videos) through the REST API or MCP server, and publishing still needs the user's explicit confirmation ([Everpop for AI agents](https://everpop.app/agents)).

## Claims table

| Claim | Source |
|---|---|
| Shorts "up to 3 minutes long", "square or taller" | [YouTube Blog, Oct 2024](https://blog.youtube/news-and-events/tall-updates-coming-to-shorts/) |
| Announced: a test of tools to "adjust the layout, zoom, and crop"; "Split screen effects". Auto layout "automatically tracks the main subject" | [YouTube Blog, Aug 2023](https://blog.youtube/news-and-events/6-new-youtube-shorts-creation-tools/); [YouTube Blog, Jul 2024](https://blog.youtube/news-and-events/6-new-youtube-shorts-tools/) |
| "identify who is speaking" | [W3C](https://www.w3.org/WAI/WCAG22/Understanding/captions-prerecorded.html) |
| "Clearly communicate who is speaking"; role "if their name is not known" | [Section508.gov](https://www.section508.gov/create/captions-transcripts/) |
| "under the speaker"; name in parentheses "on a line of its own" | [DCMP](https://dcmp.org/learn/603-captioning-key---speaker-identification) |
| "significant challenges in transcribing overlapped speech" (2025) | [arXiv 2506.05796](https://arxiv.org/abs/2506.05796) |
| Bottom panel mostly in the template's bottom margin | [where to put captions on a Short](/blog/where-to-put-captions-so-youtube-ui-doesnt-cover-them) |
| 9:16: 81/256 ≈ 31.6%, 1.78x; 4K 1215x2160. 9:8: 81/128 ≈ 63.3%, 0.89x. Halves: 960x853, ≈79%, 1.125x. 1:1: 56.25%, 1.0x. 4:5: 45%, 1.25x. Letterbox: 607.5/1920 ≈ 31.6%, 0.56x | Arithmetic |
| Aspect ratios; word-by-word captions; FCPXML/EDL/SRT; auto-post (Pro, Scale); TikTok inbox draft | [everpop.app/pricing](https://everpop.app/pricing) |
| One episode = one video; files up to 32 GB on Pro (8 GB Starter, 4 GB Free); upload or Drive; over 32 GB, re-export or trim; multicam shots framed on the person each holds; one wide camera held on one person, not yet following who talks; "Keep the whole picture" | [everpop.app/podcast-to-shorts](https://everpop.app/podcast-to-shorts) |
| A folder per client; auto-post off until switched on | [everpop.app/ai-clipping-for-agencies](https://everpop.app/ai-clipping-for-agencies) |
| Agents on Pro or Scale; publishing needs explicit confirmation | [everpop.app/agents](https://everpop.app/agents) |

## Frequently asked questions

### Should a two-person podcast Short show both people?

Only when the exchange is the point: a disagreement, a fast back-and-forth, or a reaction that changes the meaning. Then use a stacked split screen, two 9:8 panels in a 9:16 frame. When one person carries the answer, crop to that speaker and cut to the other only for a reaction that matters.

### What size is each half of a vertical split screen?

Each panel is 9:8, which is 1080x960 on a 1080x1920 frame, with the seam at exactly mid-height. A 9:8 window cropped at full height keeps about 63.3% of a 16:9 camera's width, so a camera per person fits with room to spare.

### What resolution should you record a two-person podcast in for Shorts?

Record in 4K if you plan full-screen single-speaker crops or have only one wide camera; 1080p per person is enough for a stacked split. From 1080p, a full-screen 9:16 crop is enlarged about 1.78x and a split cut from one wide shot 1.125x, while a split from a camera per person is scaled down to about 0.89x.

### How do you caption a podcast clip where two people talk?

Say who is speaking whenever the picture does not; the W3C, Section 508 and DCMP caption guides all ask for speaker identification. In a single-speaker crop, tag the off-screen voice with a name in parentheses on its own line, such as (Maya), or with a role when the name is not known, as Section 508 allows. In a stacked split, a caption at the seam reads as the top speaker's, so tag the bottom speaker's lines.

### How long can a podcast Short be?

YouTube announced in October 2024 that Shorts could run up to 3 minutes, a change that applies to videos that are square or taller, so 9:16, 4:5 and 1:1 clips all qualify. If a clip needs the whole ceiling to make sense, try a later starting point first.

### What do you do when both people in a podcast clip talk at once?

Start the clip where one voice takes the floor: multi-speaker speech recognition "faces significant challenges in transcribing overlapped speech" (arXiv 2506.05796). If the overlap is the moment, use a stacked split so both faces stay visible, and check every captioned word by hand.

More articles: https://everpop.app/blog
