B-Roll & Visual Finder
Paste a passage of narration and get what goes on screen for every line — what to search, what to prompt, and what to do when nothing comes back.
Paste your narration and press Plan the Visuals.
One beat per line of narration, with search terms and a fallback.
Working through the narration…
This usually takes ten to twenty seconds.
Need the scene structure first?
Split the script into scenes →Scripts are written in abstractions. Stock libraries are indexed in nouns.
That sentence is the entire problem this tool exists for. Your script says "the economy slowed", and no stock library on earth has a clip of that. So you type it in, get nothing usable, and settle for a drone shot of a city skyline that means nothing to anyone watching.
The work is translating the idea into something a camera could physically have been pointed at: an empty shopping mall walkway, a shop shutter coming down, a half-full car park at midday. That is what happens here, line by line, and it is why the search terms look nothing like your script. A tool that handed your own wording back as a search phrase would be doing none of the work.
Why the search terms stay in English
Stock libraries and image generators are indexed in English almost regardless of where you are. A Spanish or Urdu search phrase typed into a stock site returns nothing, so handing you one would be handing you a dead end politely. On every language version of this page the search terms and the generation prompt stay in English; the plan around them — what the viewer should be looking at, the framing, the fallback — is written in your language. Your narration is never translated or reworded, only echoed back so you can see which line each beat covers.
Every beat has a fallback, and that is not padding
Some beats have no footage. Abstract arguments, specific historical claims, anything from the last few months — the search will come back empty however good the terms are. A simple motion graphic, a map, a text card or an adjacent literal object is almost always better than a vaguely related clip, which quietly tells the viewer you had nothing and hoped they would not notice.
Vary the framing, or it looks like a screensaver
Six identical slow aerials in a row is the visual signature of a faceless video nobody finished watching. The shot type on each beat is there to be read as a sequence — if three consecutive beats all say the same thing, change one yourself. Cutting between scales is most of what makes footage you did not shoot feel deliberate.
Where it sits in the workflow
After the script, alongside the shot list. Draft narration in the Script Generator, structure it in the Scene Splitter, then bring a section here when it is time to actually go and find footage. Once you know how much of the video is generated rather than stock, the production cost calculator will tell you what that choice costs per video.
Common questions
How is this different from the Scene Splitter?
The Scene Splitter takes a whole script and breaks it into timed scenes — it answers "how is this video structured". This takes a passage of narration and answers "what do I actually put on screen, and where do I get it". They compose: split the script first, then bring a section here when you are ready to go looking for footage.
Why are the search terms in English on a non-English page?
Because that is how stock libraries and image generators are indexed, almost regardless of where you are. A search phrase in another language typed into a stock site returns nothing, so giving you one would be a politely worded dead end. The terms and the generation prompt stay English; the plan around them — what should be on screen, the framing, the fallback — is written in your language, and your narration is never touched.
Why do the search terms look nothing like my script?
That is the whole job. Scripts are written in abstractions and stock libraries are indexed in nouns. No library has a clip of "the economy slowed" — it has clips of empty mall walkways, shop shutters coming down, half-full car parks. Turning the first into the second is the work, and a tool that handed your own wording back would be doing none of it.
What if a beat has no stock footage at all?
That happens constantly — abstract arguments, specific historical claims, anything recent — which is why every beat comes with a fallback rather than only a search. A simple motion graphic, a map, a text card or an adjacent literal object is usually better than a vaguely related clip that quietly tells the viewer you had nothing.
Does it use one of my daily AI generations?
Yes. The free allowance is shared across every AI tool on the site rather than given separately to each one, because each generation costs real money whichever tool spends it. Everything that runs in your browser stays unlimited.