Text to Video AI: The Ultimate Prompt Guide (With Examples)

Published on: 27 August 2026

The difference between a mediocre AI video and a stunning one almost always comes down to the prompt. This guide breaks down the exact formula for writing text to video AI prompts that work — with 15+ real examples across every use case, so you can try these prompts on Websistant's text to video AI and see the results yourself.


How Text to Video AI Actually Works

AI video generation uses a class of model called a video diffusion model. These models are trained on millions of hours of footage and learn to associate patterns in language with patterns in video — motion types, lighting conditions, camera behaviours, visual styles, and temporal sequences.

When you write a text prompt, the model doesn't search a video database and return a match. It literally generates every frame from scratch — synthesising pixels based on what it learned during training. This means the quality of its output is entirely determined by how well your prompt matches the kinds of descriptions it was trained on.

The practical implication: vague prompts produce generic output. Specific, structured prompts that describe subject, motion, camera, and style consistently produce far better results. Websistant's text to video AI runs a 22-billion-parameter video diffusion model — one of the most capable available — which means it can handle nuanced, detailed prompts well. The more you give it, the more it has to work with.

Think of the AI as a cinematographer who has never seen your project brief. Your prompt is the director's shot description. If you'd hand a vague sticky note to a real cinematographer, expect a vague result. If you'd write a detailed shot card, you'll get something worth watching.


The 4-Part Prompt Formula

After testing hundreds of prompts on text to video AI models, one structure consistently produces the best results. We call it the SMCS formula: Subject + Motion + Camera + Style.

Every high-performing text to video prompt contains all four of these elements. You don't need to use them in a rigid order, but you do need all of them present.


Breaking Down the Four Elements

Element 1 — Subject Who or what is in the shot? Be specific — describe appearance, clothing, species, material, size, and context. "A woman" is weak. "A woman in her 40s wearing a navy blazer, seated at a glass desk" is strong.

Element 2 — Motion / Action What is the subject doing, and how? Describe the speed, direction, and quality of movement. "Walking" is weak. "Strides briskly, glancing over her shoulder" is strong. Motion is what makes video different from a photo.

Element 3 — Camera Behaviour Where is the camera? Is it moving? How fast? The camera description is the most underused element in beginner prompts — and the one that most dramatically improves results. "Static medium shot" vs "slow push-in from behind" produces completely different footage.

Element 4 — Visual Style & Mood What does the overall image look like? Reference film stocks, lighting conditions, colour grades, and genres. "Cinematic, anamorphic, warm golden tones, shallow depth of field" gives the model a full aesthetic brief.


15 Prompt Examples by Use Case

All 15 prompts below are ready to paste directly into Websistant's text to video AI generator. Recommended settings are listed alongside each prompt.


SOCIAL MEDIA

Prompt 01 — Lifestyle Reel (TikTok / Reels) "A young woman with curly hair pours an iced matcha latte in a sun-drenched kitchen. Close-up on the pour, ice swirling in slow motion. Camera tilts up slowly to her smiling face. Soft morning light, warm tones, aesthetic lifestyle, 4K quality." Settings: 576×1024 vertical · 5–8 seconds · 25 fps

Prompt 02 — Motivational Quote Background (Instagram Stories) "Abstract golden particles drift upward against a deep navy background. Slow, meditative drift. No subjects. Camera static. Elegant, minimal, luxury brand aesthetic, soft bokeh, loopable motion." Settings: 576×1024 vertical · 5 seconds · 24 fps

Prompt 03 — Cinematic Travel Clip (YouTube Shorts) "Aerial drone shot sweeping over turquoise Maldivian waters at golden hour. Camera glides forward slowly, revealing a row of overwater bungalows. Rich warm tones, cinematic colour grade, wide anamorphic format. Gentle ocean sounds." Settings: 1280×720 landscape · 8 seconds · 24 fps


E-COMMERCE & PRODUCT

Prompt 04 — Skincare Product Hero (Product Page / Ads) "A minimalist white glass serum bottle sits on a marble surface. Water droplets form on the glass and slide down slowly. Camera orbits the product in a slow 180-degree arc. Studio lighting, clean white background, pharmaceutical aesthetic, macro detail." Settings: 1024×576 widescreen · 5–8 seconds · 25 fps

Prompt 05 — Fashion Lifestyle Ad (Social Ads) "A tall woman in a flowing cream linen dress walks along a sunlit cobblestone street in southern Europe. Camera tracks alongside her at waist height, slow steady motion. Warm Mediterranean afternoon light, light wind moving fabric. Editorial fashion, film grain, Kodak 400 colour palette." Settings: 576×1024 vertical · 8 seconds · 24 fps

Prompt 06 — Food & Beverage (Restaurant / Delivery Ads) "A chef's hands slice through a perfectly seared wagyu steak on a dark slate board. Juices flow in slow motion. Camera starts in extreme close-up and slowly pulls back to reveal the full plating. Dark moody restaurant lighting, cinematic food photography style. Sizzling audio." Settings: 1024×576 widescreen · 8 seconds · 24 fps


BUSINESS & SAAS

Prompt 07 — Corporate Hero Video (Website Hero) "A diverse team of professionals collaborates around a large glass conference table in a modern open-plan office. People lean in, point at screens, and nod. Camera slowly pushes in from outside through floor-to-ceiling windows. Natural daylight, clean corporate aesthetic, optimistic mood. Ambient office hum." Settings: 1280×720 HD · 8–10 seconds · 25 fps

Prompt 08 — Tech / SaaS Abstract (App Demo / Pitch) "Glowing blue data nodes connect and pulse across a dark digital network map. Lines of light travel between nodes at speed. Camera slowly zooms out, revealing the network spans a globe. Dark background, electric blue and cyan, futuristic data visualisation aesthetic. Ambient electronic pulse." Settings: 1280×720 HD · 8 seconds · 30 fps


CREATIVE & CINEMATIC

Prompt 09 — Cinematic Establishing Shot (Film / Pitch) "A lone figure in a long dark coat stands at the edge of a cliff overlooking a storm-lit ocean. Camera orbits slowly around them from behind. Dramatic overcast light, crashing waves below, salt spray in the air. Cinematic, anamorphic widescreen, desaturated cool grade, orchestral score building." Settings: 1280×720 HD · 8–10 seconds · 24 fps

Prompt 10 — Abstract Art Loop (Music / Art) "Ink drops fall into water in extreme slow motion. Deep crimson ink disperses into midnight blue water, forming organic cloud shapes. Camera static, top-down macro view. Silent, meditative, high-contrast, art photography aesthetic." Settings: 640×640 square · 5 seconds · 24 fps


EDUCATION & EXPLAINER

Prompt 11 — Nature Explainer B-Roll (Educational Content) "A time-lapse of a flower blooming from tight bud to full bloom. Camera static, tight close-up on the petals. Soft diffused natural light. Botanical illustration aesthetic, clean white background. No audio." Settings: 1024×576 widescreen · 5 seconds · 24 fps

Prompt 12 — Architecture Walk-Through (Property / Real Estate) "First-person walk through a light-filled modern apartment. Floor-to-ceiling windows, white walls, Scandinavian furniture. Camera moves forward through the living room toward the view. Smooth, stabilised, late afternoon sunlight streaming in. Real estate photography aesthetic." Settings: 1280×720 HD · 8–10 seconds · 25 fps


ATMOSPHERE & BACKGROUND

Prompt 13 — Looping Website Background "Slow, gentle movement of deep space nebula clouds in purple and midnight blue. Distant stars shimmer softly. Camera barely moves. Abstract, cosmic, meditative. Loopable. No subjects. Silent." Settings: 1280×720 HD · 8 seconds · 24 fps

Prompt 14 — Seasonal Brand Moment "Snow falls gently in a quiet pine forest at dusk. A single warm light glows through a frosted window in a log cabin in the distance. Camera static wide shot. Blue-hour light, extreme quiet, festive but understated. Soft wind in the trees." Settings: 1280×720 HD · 8 seconds · 24 fps

Prompt 15 — Urban Night Scene (Brand / Lifestyle) "Rain-soaked city street at night. Neon signs reflect in puddles. A yellow taxi passes through frame, blurring light trails. Camera static on a tripod, slightly low angle. Long exposure feel, cinematic noir, Tokyo-style urban aesthetic. Distant traffic, rain on concrete." Settings: 1280×720 HD · 8 seconds · 24 fps


Camera Movement Reference List

Camera direction is the most underused element in beginner prompts. Adding even one specific movement term can transform your clip into something that feels genuinely cinematic.

Push & Pull - Slow push in — camera gradually moves closer to the subject. Creates intimacy and focus. - Slow pull back / zoom out — reveals wider context. Creates scale and distance. - Whip zoom in — fast, energetic push toward subject. High-impact, social-native feel.

Tracking & Following - Camera tracks alongside subject — moves laterally with the subject. Creates momentum. - Camera follows from behind — POV-style following shot. Immersive and personal. - Camera leads subject — camera moves backwards in front of a walking subject.

Rotation & Orbit - Slow orbit around subject — camera circles the subject. Great for product reveals. - 360-degree rotate — full circle around a central point. - Dutch tilt — camera tilted at an angle. Creates unease or drama.

Aerial & Elevated - Aerial drone shot — high overhead perspective looking down or forward. - Bird's eye view, static — directly overhead, locked off. Great for flat lays. - Crane shot sweeping up — starts low, rises to reveal wider scene.

Static & Locked - Static wide shot — camera completely locked off. Subject and environment move, camera does not. - Static close-up — tight on a detail, no camera movement. Forces attention on the subject.


Style & Mood Modifiers That Work

Cinematic looks - Cinematic, anamorphic widescreen, shallow depth of field - Film grain, Kodak 400 colour palette, warm highlights - Desaturated cool grade, muted shadows, high contrast - Golden hour, lens flare, atmospheric haze

Commercial & brand looks - Clean, minimal, white studio background - Editorial fashion, high-key lighting - Lifestyle photography, warm natural light, authentic - Luxury brand aesthetic, muted tones, precise composition

Stylised & creative looks - Anime-style, vibrant colour palette, fluid motion - Oil painting, impressionist brushwork, painterly - 3D render, CGI, Unreal Engine aesthetic - Retro VHS, scan lines, 1980s colour grade


Prompting for Audio

Websistant's AI video model supports native audio rendering — meaning you can describe ambient sounds directly in your prompt and the model will synthesise synchronised audio into the finished clip. Add audio descriptions naturally at the end of your prompt:

  • Gentle ocean waves breaking on shore
  • Soft piano score, distant and melancholic
  • City traffic hum, distant sirens
  • Rain on a tin roof, quiet and rhythmic
  • Orchestral score building to a crescendo
  • Electronic ambient pulse, low and steady
  • Sizzling sound, kitchen ambience
  • Birdsong, morning forest atmosphere
  • Wind through pine trees, remote and quiet

Keep audio descriptions to one or two cues per prompt. Don't mix conflicting environments — the model will attempt to blend them with inconsistent results.


7 Common Prompting Mistakes to Avoid

1. Describing intent, not imagery The model renders what a camera would see — not concepts or feelings. "Show the power of our brand" gives the AI nothing to render. "A bold red logo emerges from darkness, spotlit, camera slowly pushing in" gives it everything.

2. Packing multiple scenes into one prompt The model handles a single scene per clip best. Generate each shot separately and combine them in a video editor if you need a multi-shot sequence.

3. Omitting camera direction entirely Without a camera instruction, the model chooses arbitrarily. Add at minimum one camera direction — even "static wide shot" or "slow push in" — and quality improves immediately.

4. Using overloaded style terms "Realistic" and "high quality" are meaningless to the model. Use specific technical terms instead: "photorealistic", "anamorphic lens", or "documentary-style handheld".

5. Describing the wrong subject identity If your subject's appearance matters — skin tone, age, clothing, hair — describe it fully upfront. The model defaults to whatever is most common in its training data for underspecified subjects.

6. Requesting text or logos on screen Video diffusion models cannot reliably render legible text within a clip. Add titles and logos in post-production instead.

7. Expecting consistent characters across clips Each generation is independent. If you need consistent character identity across multiple shots, use Image to Video mode with the same portrait image as your reference.

Pro tip: Don't judge the model on your first prompt. Generate a 3-second test clip ($0.06), evaluate what the model understood and what it missed, refine the prompt, and regenerate. Most great results come on the second or third attempt.


Frequently Asked Questions

How long should a text to video AI prompt be? Aim for 40 to 100 words. Long enough to cover all four SMCS elements, short enough to stay coherent. Beyond 120 words you risk confusing the model with competing instructions.

What language should I write prompts in? English produces the most consistent results, as the majority of training data is in English. You can write in other languages, but expect more variability in output quality.

Can I reuse the same prompt multiple times? Yes. Each generation uses a different random seed, so the same prompt can produce noticeably different results on each run. If a prompt is producing broadly good results, run it 2–3 times and choose the best output.

How do I make the AI video match my brand colours? Describe your colour palette directly in the style section of your prompt. For example: "warm amber and deep navy palette", "muted terracotta tones", or "high-contrast black and white, silver accents".

Where can I test these prompts? All 15 prompts in this guide work directly with Websistant's text to video AI. Sign up free, claim your $5 starter credit, and paste any prompt into the Text to Video tab. A 3-second test clip costs just $0.06.


Ready to put these prompts to work? Generate your first AI video free on Websistant — $5 credit, no card required.

Try Our Free AI Website Builder now!

Need product image? Generate your AI product image here!

← Back to all posts