Ad format
The format that stops pretending to be anything but an ad
A VSL is Lollipop’s video sales letter format. Nobody presents it. One unbroken narrator voiceover runs over a fast-cut mix of anatomical renders, graphs and stat cards, lived-in room footage and product macros, and the numbers and the offer appear on screen as well as in the voice.
When should you write a VSL instead of something native?
Use it when you can openly argue. The format wants a pain the viewer feels right now, a mechanism you can draw, real numbers, and an offer with a guarantee to close on. It walks a Problem, Agitation, Solution and CTA corridor, and it never pretends to be a vlog.
Its definition names the cost of that honesty. Avoid a VSL when the brand needs to feel native and undefended, because this format never pretends to be anything but an ad. That is the whole trade.
Where does the picture come from if nobody is on camera?
A VSL keeps no room. Each shot lives in one of four named spaces, and each space has its own locked light and camera recipe: an anatomical render space against a dark void, a flat graphic field for graphs and stat cards, a real lived-in room shot rough, and a product macro surface. Cohesion comes from repeating a recipe, not from returning to a location.
The plates themselves carry no people, no faces and no baked-in text. Callouts arrive as overlays afterwards rather than being painted into the environment, which is what keeps a number editable instead of burned into a render.
Two kinds of actor, and they are opposites
The narrator is never on camera. She still gets a real identity portrait, because her voice is cloned from it, but that portrait is never seen by a viewer and she is never written into a shot.
A VSL can also carry a second actor who is the exact inverse: seen and never heard. That is the recurring figure who carries two or more of the pain beats, the same commuter or sufferer in every one of them, while the narrator speaks every word in the ad. No speaker on camera is not the same thing as no person on camera, and collapsing the two is how an ad ends up with two different faces in the same seat.
What stops a VSL inventing a statistic?
Specificity is this format’s engine, so fabricated specificity is its worst failure. The scripting agent is instructed never to manufacture a number, a study, a percentage, a discount or a guarantee, and to build the argument on a mechanism it can substantiate when the brief supplies no hard figures. Where there is no real offer, the close is the action alone.
Read that as a rule rather than a promise. It is an instruction to a language model, and nothing in the pipeline checks a claim against a source. If the numbers matter, they have to come from you and you still have to read the script.
How long is a VSL?
The default length is 60 seconds, and the allowed lengths are 30, 45, 60 and 90.
Scripts are written to about three spoken words per second, so a format’s length is also its word budget: at the 60-second default that is roughly a hundred and eighty spoken words. The cutting moves ahead of the sentence, so one line often plays across several shots while the voice never stops.
How it compares with the other footage-driven formats
Footage-driven is one of six families in the catalogue. The formats in it separate on who carries the words, and that decision reaches all the way down to whether the ad needs a voice clone at all.
| Format | Who carries the words | Where the picture comes from | Authored on-screen text |
|---|---|---|---|
| VSL (Video Sales Letter) | An unseen narrator, arguing | Renders, graphs, room footage, product macros | Mandatory on the hook and on the CTA |
| Story Slideshow | Nobody speaks; captions carry it | A fast run of product-in-use shots | The captions are the script |
| Product Mashup | Nobody speaks; the pattern carries it | Phone clips locked to one template, one thing swapped each cut | Sparse, pre-written, placed rather than authored |
| Voiceover Product Demo | An unseen narrator, explaining | Full-frame footage of the real product | On generated shots only |
| Expert Talking Head | A presenter, down the lens | High-end proof B-roll cut against them | Yes, and it is a signature of the format |
Where a VSL stops working
- It reads as an ad, by design. Every other choice in the format follows from that. If the placement rewards looking native, this is the wrong format and no amount of good footage will fix it.
- The argument is only as good as your inputs. With no numbers in the brief the script falls back on mechanism, which is honest but softer, and the close carries no offer at all.
- There is no continuing location or face to anchor to. A VSL is held together by a shared colour grade and one type system, so a lifestyle figure who recurs has to be declared explicitly or she is regenerated from scratch on every cut.
When a different faceless format fits better
Formats in Lollipop are rule sets rather than templates. Each of the 21 formats ships its own scripting agent, its own scene agent and its own visual rules. Switching format changes what the script is allowed to argue, not only what the shots look like.
The nearest relative here is also faceless and also narrator-led: the whiteboard explainer runs the same kind of argument, drawn by hand on one board instead of cut across four render spaces. Pick that when the reasoning needs to be followed step by step; pick a VSL when the proof is footage and numbers. If instead you want the undefended, native register a VSL deliberately gives up, the Yapper format is one person telling a friend about it, with no cutaways at all.
If you would rather choose by what you sell than by what you want to film, which formats suit an online store shipping physical products starts from the buyer and works back to the catalogue.
Start a VSL ad now.
Make a VSLDescribe the product in a chat. Lollipop writes the script, casts the actor and renders every shot, and you can open the canvas to change any of it.
