Skip to content
SUPA
Tools

Video generation from a prompt

Describe a scene in words — and get a finished shot with motion: the camera glides, the light changes, fabric and water come alive. Want it more precise — give the first frame, or both ends of the scene, and the AI shoots exactly the transition you need.

Create a video from a description

A smooth flight over an autumn forest with a winding river, morning fog between the trees, moving forward, an aerial shot

Just a description — and the scene is shot

No source footage, no filming: under each clip is the request it came from. Everything you see plays exactly as it came back from the generation.

Macro: a pitcher of milk in a barista's hand, milk pours in a thin stream into a cup of espresso, a heart pattern gradually appears on the surface, warm morning light, shallow depth of field
A smooth flight over an autumn forest with a winding river, morning fog between the trees, moving forward, an aerial shot
A night street in the rain, neon signs reflected in the puddles, the camera slowly moves forward, a cinematic shot
A girl in a beige coat stands on a city street, the wind stirs her hair and the hem of the coat, she turns to the camera, soft evening light
Vertical shot: a potter's hands raise the walls of a clay bowl on a wheel, the clay glistens with water, warm side light, close-up
A wristwatch with no lettering on the dial slowly rotates on a dark stone pedestal, highlights glide over the steel case, studio light, an ad shot

Bring a ready frame to life

Above each clip is the frame it starts from: a photograph, a generated image or a finished layout. The prompt describes only the motion — the scene stays yours.

The camera glides forward through the room, patches of sunlight crawl slowly across the floor, the curtains sway gently
She turns her head to the camera and smiles, steam rises from the cup, the light changes softly
Olive oil pours into the frame from above in a thin stream, the salad leaves tremble slightly, the camera moves in a little

Set the beginning and the end

Two frames — and the AI works out what happened between them. That's how you get precise transitions: a bud opens, day turns to evening, summer to fall.

The bud slowly opens petal by petal, the motion continuous and smooth
Day fades into evening, the lights come on in the coffee shop, a garland lights up
The foliage slowly turns from green to gold, leaves begin to fall, the light grows warmer

What to know about the clips

Eight seconds is one shot of a scene. A full video is assembled from several such shots in the editor.

Landscape and portrait

16:9 for a website or a presentation, 9:16 for stories and Shorts. The proportions are chosen before the generation, so the shot is composed for the format from the start.

Three ways to set the scene

Prompt only — when the scene doesn't exist yet. A first frame — when you have a photo or a layout. First and last — when the ending matters.

Duration and quality

A clip runs four, six or eight seconds and is generated in 720p or 1080p. The longer and larger, the longer the wait and the higher the cost.

Then — into the edit

The finished clip stays in your project: trim it, put it on the timeline next to other scenes, add titles, music and a voice-over.

How to generate a video

From a description to a finished clip — four steps in the browser.

  1. 1

    Sign up for SUPA and open a video project

  2. 2

    In the AI section, pick video generation

  3. 3

    Describe the scene and the camera movement. If needed, upload a first frame — or the first and the last at once

  4. 4

    Choose the proportions and the duration and run the generation. The finished clip appears on the timeline

How to describe a scene to get the shot you need

Describe the motion, not just the picture

Video differs from a photo in that something happens in it. “The camera slowly moves forward”, “the wind stirs the hem of the coat”, “steam rises” — these are the words that turn a still into a scene.

One action per clip

Eight seconds fit one event. Three actions in a row the model will either compress into a blur or show only the first: better to shoot three short scenes and cut them together.

Name the type of shot

“Macro”, “aerial shot”, “ad shot”, “cinematic shot with shallow depth of field” set the lens, the pace and how carefully the camera behaves.

The first frame decides more than the prompt

If the look of the scene matters — a product, an interior, a person, brand colors — start from a picture. Then the prompt is only responsible for the motion, and the result stops being a lottery.

Two frames — for a precise transition

When the ending matters, set both ends. The frames must be from the same scene: same angle, same composition, and the only change is the one the transition is about.

Write about the light and the time of day

“Warm morning light”, “dusk”, “a rainy night” — light sets the mood better than any adjective, and it also keeps the shot from drifting in color.

Frequently asked questions about video generation

How long is a generated clip?

Four, six or eight seconds. All the examples on this page are eight-second clips.

How long does the generation take?

Usually from one to several minutes — noticeably longer than generating a picture. You can close the tab: the finished clip appears in your project.

Can I use my own photo as the first frame?

Yes, and it's the most reliable way to a predictable result: upload a shot of your product, interior or a finished layout and describe only the motion.

Does the clip have sound?

The model can generate sound along with the image. The examples on this page were made without it — they play muted and looped in the grid, and an extra audio track would serve no purpose.

Can I build a full video out of the clips?

Yes. Each clip is a scene, and the editing, titles, music and transitions are done in the SUPA editor.

Can I use the videos in ads?

Yes. As with pictures, check the frame for third-party logos and don't present a generated scene as documentary footage.