Two or three people on a podcast, one vertical frame: the layouts that hold attention

A podcast with more than one person does not fit inside a 9:16 frame, and the most common mistake is trying to make it fit. There are four layouts that work, each for a different kind of moment, and choosing between them decides more retention than captions or hooks.

Two or three people on a podcast, one vertical frame: the layouts that hold attention

Two or three people on a podcast, one vertical frame: the layouts that hold attention

The fact: a podcast recorded in 16:9 with two people side by side will not fit in a 9:16 frame without losing someone. Four layouts solve it: full screen on the speaker, split screen, full frame with a blurred background and reaction cutaways. Each fits a different kind of moment, and choosing wrong costs more retention than any badly styled caption.

The most common mistake is not picking the wrong layout. It is picking one and using it for the whole clip.

Why podcasts are the hardest format to go vertical

Gameplay has a webcam and a screen, and the solution became standard. Vlogs are shot with one person in the middle. Podcasts are the worst case: two or three people spread horizontally, each with a mic in front of their mouth, and the content lives precisely in the exchange between them.

Crop a 9:16 frame out of a 16:9 video at the same height and you keep about a third of the original width (608 out of 1920 pixels, in a 1080p video). Fitting both faces at once without splitting the screen almost always means tiny faces. A small face in a vertical feed is the fastest signal of "not worth stopping", the problem we covered in skip rate.

So the right question is not "how do I fit everyone", it is "who needs to be on screen this second".

The four layouts and when to use each

1. Full screen on the speaker

The crop closes in on whoever is talking and switches when the floor changes hands.

Use it when: one person tells a story, explains an idea or gives a long opinion. It is the default layout for most podcast clips, because most good stretches have a protagonist.

Watch out for: switching on every "yeah", "right", "wow". A listener's interjection is not a change of speaker.

2. Split screen

One face per half, stacked: one on top, one below.

Use it when: both people are in the moment at once. A quick back-and-forth, a joke with a comeback, an interview where the question matters as much as the answer.

Watch out for: long single-speaker stretches. The silent half becomes dead space and the speaker gets half the size they could have.

3. Full frame with a blurred background

The whole original frame in a horizontal band in the middle, with an enlarged, blurred copy filling the rest.

Use it when: the moment depends on the set or on three people reacting together. The group laugh, someone walking into the studio, an object being shown.

Watch out for: making it the default. Faces get small and the video looks recycled from horizontal, which is exactly what it is. It works for two or three seconds, not a minute.

4. Reaction cutaway

A second or two on the listener's face, in the middle of the other person's story.

Use it when: the reaction is the point. The guest's face at the awkward question. The host trying not to laugh.

Watch out for: reactions with no reason. If nothing happened on the listener's face, cutting to it just breaks the rhythm.

Three people: the rules change

With three people, a three-band split almost never works on a phone. Every face is too small and the eye does not know where to go.

What works better:

  • Full screen on the speaker as the base.
  • Two-way split when two of the three are going at it and the third is just listening.
  • Full frame only for the group moment, and only for a few seconds.

In practice, a three-person table clip is a two-person clip with a third person who shows up when they have something to say.

Where faces and captions go

Vertical video has regions the app interface covers. For Reels, Meta recommends keeping 14% of the top, 35% of the bottom and 6% of each side free of text, logos or key elements. The bottom is the worst part: that is where the account name, caption and buttons sit.

For podcasts, that means:

Full screen: the speaker's eyes around the upper third, captions just below the chin, still above the bottom 35%.

Split screen: captions go in the seam between the two faces, near the center of the screen. It is the only spot that covers nobody and stays clear of the interface. Captions under the lower face land on top of the account name and description.

Full frame: the band with the original frame sits in the middle, and captions can go right under it.

One tip that fixes a lot: a different caption color for each person. In a split screen, viewers can tell who is talking without hunting for the moving mouth. How caption style affects retention is in caption style and clip retention.

Switching mistakes that wear people out

Camera switching is what separates amateur podcast clips from professional ones. Four common mistakes:

Switching mid-word. The switch should land at the start of the new speaker's sentence, not half a second early and not in the middle.

Switching on sound instead of speech. Laughs, "mm-hm" and coughs are not a change of speaker.

Switching too fast. Two switches inside a second look like an editing error even when the conversation really was that quick. If it is that quick, it is a split-screen moment.

Never switching. The opposite also tires people: a full minute on one person while the other asks questions off-camera leaves the viewer without half the conversation.

General pacing is covered in hard cuts and pacing, and it counts double here.

A decision checklist for every clip

Before framing, answer three questions:

  1. How many people do I need to understand this stretch? One: full screen. Two at once: split. The group: full frame, briefly.
  2. Is there a reaction that tells the story? If yes, mark the exact second and insert the cutaway.
  3. Does the first image have a big face? If the clip opens on a split with small faces, consider opening full screen and splitting later.

A good podcast clip almost always uses two or three layouts in sequence, not one.

Where a tool does the boring part

Doing this by hand means setting a keyframe every time the floor changes hands. In a two-hour episode with twenty clips, that is hundreds of keyframes.

In Cut.Pro, vertical reframing follows the face and the transcript identifies who is speaking, so the camera switch follows the speaker switch without you marking anything. In the editor's templates you can choose which layouts the AI may use (full screen on the speaker, a two-person split, full frame with a blurred background, among others), and it picks among the ones you allowed based on each scene. You review and fix the odd switch that landed in the wrong place, which is much faster than building it from scratch.

If you are choosing a tool for podcasts, the comparison is in the best AI podcast clipping tool. For volume, see how to turn a podcast into clips every week.

The short version

  • A multi-person podcast does not fit whole in 9:16. The question is who needs to be on screen right now.
  • Full screen on the speaker is the default; split screen is for back-and-forth; full frame is for group moments; reaction cutaways are for when a face tells the story.
  • With three people, treat it as two plus one who joins when they speak.
  • In a split, captions go in the seam between faces. Meta asks for 14% top, 35% bottom and 6% sides free of text on Reels.
  • Switch at the start of the new speaker's sentence, never on an interjection.

Viewers do not notice good framing. They notice bad framing, and they leave before they know why.

Sources: Meta, Instagram Reels ad specs (safe zone)

Share

Keep reading

More insights and tutorials to help you grow as a content creator.