Ad format
A silent ad in which the captions are the whole script
This is a silent montage. One performer moves through their day to music, using the product inside whatever they are already doing, and every word the ad needs arrives as an on-screen caption. Nothing is spoken anywhere in it. There is no voiceover, nobody talks to camera, and no close-up of a face appears in the whole ad.
What kind of product survives with nothing spoken?
Use it when the benefit reads on sight. Energy, ease, time back, comfort, confidence: things a viewer can watch happen to a body rather than be told about. The other half of the test is the copy. The hook, the point and the CTA each have to land in one short line sitting over a real action.
Avoid it when the argument needs sentences to make. There is nowhere for a sentence to go. A caption that has to explain a mechanism is a caption nobody finishes reading, and the format gives you no voice to fall back on.
Make one of these for your own product.
Make a Text & Music adEvery new account starts with 4,000 credits.
If nobody speaks, where do the words go?
Into the captions, which are the only copy the ad has. They are written at the script stage, before a single shot is described, and then placed verbatim on the scene they were mapped to. The scene agent may not invent one, expand one or rephrase one to fit a shot.
Placement is fixed too. Lines sit in the upper third, away from the performer’s face, high-contrast against whatever is behind them, and they stay up long enough to read. A scene with no caption mapped to it simply runs clean.
That is unusual. Whether an ad carries authored on-screen text is decided by the format rather than by the user, and most of the catalogue either forbids overlays or rations them to a sparse line. Here they are load-bearing.
Why is there no close-up of a face?
Because the ad is built to look like footage somebody shot of their own day, and that footage is wide. The consequence is a strange one: with no face close-up anywhere, the performer is recognised across shots by silhouette, posture, hair shape, body proportion and wardrobe instead.
So their reference portrait is full body rather than head and shoulders, standing, wardrobe visible, both hands in frame. The other rule that falls out of it: the product never appears without a hand on it. No bottle standing alone, no detached macro.
Framing has a ceiling rather than a target. Most of the ad is wide enough to see the room, a chest-up shot is as tight as the performer ever gets, and the only genuinely tight shot the format permits is hands plus product doing a detail action. Which means the room is doing real work, and has to be specific, lived-in and readable to the edges.
How long is one, and how fast does it cut?
Its default length is 20 seconds, and its allowed lengths are 15, 20, 30 and 45.
There is no words-per-second budget to spend, because there are no words to speak. The one timing rule that survives is reading speed: a line the eye cannot finish while it is on screen is a line that never landed, so the shot under a caption is sized to the caption.
What else in the catalogue is silent?
One other format, and it is the closer comparison. The Story Slideshow format cuts every 1 to 3 seconds through product-in-use footage in plot order, with no presenter, no voiceover and no face. The difference is who is in the frame: that one has nobody in it, this one is built around a person.
Within the UGC family the separation is mechanical rather than a matter of degree. UGC Text and Music has no speech at all, while UGC Voiceover never shows a mouth speaking and Yapper speaks only in shot.
| Format | What is spoken | Who is in it | Where the words live |
|---|---|---|---|
| UGC Text & Music | Nothing, anywhere in the ad | One performer, never in close-up | Captions, and they are the entire script |
| Story Slideshow | Nothing, anywhere in the ad | Nobody. Product-in-use footage only | Captions, and they are the entire script |
| UGC Voiceover | A voice, laid over the picture | A performer, never seen speaking | Spoken. On-screen text essentially off |
| Yapper | Straight down the lens, the whole way | One person, in every single frame | Spoken. Text off, bar a closing card |
The limits, and they are strict ones
- No argument can be made in sentences. Everything the ad asserts has to survive being cut down to a line the viewer reads in passing, over a picture that is already carrying the situation.
- There is no reaction shot to lean on. With no face close-up permitted, the emotional beat has to be readable from a body across a room. Continuity is riskier here for the same reason, since recognition rides on silhouette and wardrobe rather than on a face the model can lock to.
- The captions cannot be improvised later. They are the script, placed verbatim, so a line that does not fit its shot is fixed by resizing the shot rather than by trimming the words.
- No product beauty shot. A hand has to be on it, always, which rules out the detached macro that most product ads open on.
If silence is the wrong choice
Two neighbours solve the same brief differently. When the product is the whole subject and there is no reason for a person to be on screen, the Story Slideshow format runs product-in-use footage in plot order with nobody in it and keeps the caption-led shape. And when the claim really does need saying out loud by somebody the viewer can look at, Yapper keeps one person on camera in every frame with no cutaways at all. Both are worth watching before you commit, because every ad format has a real example ad you can watch before you pick it.
At the opposite end of the same family, where the words are spoken by somebody the viewer is meant to believe rather than read off the screen, the crew-lit expert talking head states each claim to camera.
Start a UGC Text & Music ad now.
Make a Text & Music adDescribe the product in a chat. Lollipop writes the script, casts the actor and renders every shot, and you can open the canvas to change any of it.
