Ad format
Footage the viewer thinks somebody else found first
Podcast Clip is Lollipop’s format for an ad that reads as a moment lifted out of an episode. A host and a guest sit at the same mic’d desk in one studio, covered by several fixed cameras that only ever cut between angles on the same sitting. The product is never presented. The guest tells a story, and the host’s questions pull the specifics out of it.
When is a fake podcast clip the right call?
Its definition names the trigger precisely: use it when the product has a counterintuitive mechanism, a surprising number, or a before and after that earns a wait-what, the kind of claim a host would interrupt to question. Avoid it when the value is purely visual or aesthetic, since the authority comes from two people talking rather than from anything you can show them.
The shape follows from that. It opens on the arresting claim mid-sentence rather than a welcome, and ends on a reaction rather than a summary. Neither person speaks more than two or three sentences before the other cuts in.
Make one of these for your own product.
Make a Podcast Clip adEvery new account starts with 4,000 credits.
What has to be in frame for the set to hold together?
Every angle has to contain the shared set. In a single, that means the speaker’s microphone plus a slice of desk and the other person’s shoulder or the backdrop behind them. A single with no mic and no desk is a portrait, not a podcast shot.
Mic geometry is where generated footage usually gives itself away. Booms are clamped at each occupant’s outer edge and swing inward, mirroring each other, because two arms reaching toward the middle of a real table would collide. Short desk stands are the alternative. One style for both people, never one boom and one stand, and never a mic parked over a mouth.
The seating decides every eyeline in the ad
The Podcast Clip format seats a host and a guest at one shared desk with a large microphone in front of each of them, caught mid-conversation. The host holds the left seat and the guest the right for the whole ad, the host looks screen right and the guest screen left in every shot, and nobody ever looks at the lens.
Flip one single and the two of them appear to be staring the same way, at which point the conversation stops existing. It is a rule you can check on a still frame, which is exactly why it is written as a rule rather than as a note about coverage.
Who writes the captions that run over it?
A Podcast Clip carries burned-in captions the whole way through, and they are generated by Lollipop from the audio rather than written into the script. Authored typography is reserved for a genuine graphic on a proof insert, such as a number or a stat card.
Captions come in 10 animated presets, including karaoke highlighting. Which one an ad wears is a platform choice rather than something this format decides, which is why the scene stage leaves the authored overlay field empty and lets the caption system work off the spoken track.
How long does one of these run?
This format defaults to 45 seconds, with allowed lengths of 30, 45, 60 and 90.
Scripts are written to about three spoken words per second, so a format’s length is also its word budget. Two voices share that budget here, backchannel included, which is why the longer stops exist at all.
How does it differ from the other conversational formats?
Podcast and interview work is one of six families in the catalogue. Every format in it films people talking to each other, and what separates them is the situation, the geometry of the room, and how far the edit is allowed to wander from a face.
| Format | Where the people are | Where they look | What the edit may cut to |
|---|---|---|---|
| Podcast Clip | One studio, both people at the same desk | Across the table, never at the lens | Short proof inserts over unbroken audio |
| FaceTime Call | Two different rooms, one per caller | Straight down each phone lens | Nothing, apart from a shared photo inset |
| Comedy Skit | One lived-in room, everybody in it | At each other or the product | Nothing at all |
| Street Interview | One real pavement, one afternoon | At the host, who alone may look at us | A product passed inside the framing |
The limits, before you commit to it
- Two voices is two pipelines. An actor who speaks runs the whole actor pipeline, a generated portrait, a preview video, a cloned voice and a video-model element, while an actor who is only seen stops at the portrait. This format needs both people talking, so it is a two-speaker build by definition.
- The set is generated once and inherited everywhere. A base image showing one person alone at a mic, or two separate desks, is not a shot you discard. Every angle is generated against it, so the mistake propagates and the ad reads as two solo recordings cut together.
- You cannot crop your way out of trouble. Tightening a single until only the face is left removes the mic and the desk, and with them the only evidence that this was ever a podcast.
- Beautiful products get nothing out of it. Two people describing how something looks is a weaker argument than showing it, and this format has almost no room to show it.
Neighbours worth considering instead
Moving between formats changes what the ad is allowed to argue. Each of the 21 formats in Lollipop ships its own scripting agent, its own scene agent and its own visual rules, and generation is bound by them.
When the credibility should come from a crowd rather than from one expert at a desk, the Street Interview format stops a run of strangers on one real pavement. When the two people should sound like friends rather than professionals, the FaceTime Call format puts them in two different rooms on a video call. And when the sell would land better as a story than as testimony, a Comedy Skit builds a scene the product is load-bearing in.
Start a Podcast Clip ad now.
Make a Podcast Clip adDescribe the product in a chat. Lollipop writes the script, casts the actor and renders every shot, and you can open the canvas to change any of it.
