# How to write prompts for AI video generators

A practical guide to Seedance, Kling, Veo, Runway, Hailuo, H3 Max and Wan, with examples and ready-to-use instructions for ChatGPT.

Imagine a simple advertising shot: a coffee bag stands on a countertop, a cup steams beside it, and the camera slowly moves towards the product. Before describing the scene to a generator, decide what you are starting with. Do you already have a photo to animate? Should the model create the kitchen from text? Should it reproduce the packaging from a separate photograph? Does the video need sound?

These decisions shape the prompt: the instructions you give the generator. The model you choose matters too. Runway Gen-4 documentation recommends describing the desired result rather than adding prohibitions. Veo has a documented field for negative instructions, while MiniMax describes special camera commands for Hailuo. Copying the same elaborate template into every tool is therefore not a good approach. [Sources: Runway — Gen-4 guide](https://help.runwayml.com/hc/en-us/articles/39789879462419-Gen-4-Video-Prompting-Guide), [Google — Veo guide](https://cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide), [MiniMax — image-to-video](https://platform.minimax.io/docs/api-reference/video-generation-i2v).

The guide contains 19 chapters covering individual models and their variants. The examples suggest ways to write prompts; they are not the results of a comparative test. Instructions can specify the intended result precisely, but cannot guarantee that a generator will reproduce it without errors.

## Before choosing the instructions for your model

### Which language should you write in?

The example prompts in this guide are in English, so they can use the terminology found in documentation and descriptions of camera work. This does not mean that English always produces better results than Polish or other languages.

Google's documentation states that Veo fully supports English. For the other model families, treat English as this guide's working language, not as a choice proven best by research. [Source: Google — video generation in the Gemini API](https://ai.google.dev/gemini-api/docs/veo).

Keep three things separate:

| Element | How to describe it |
| --- | --- |
| Instructions for the generator | You can describe the scene, action and camera work in English. |
| A character's dialogue | Specify the exact wording and language. Check whether the model can generate speech in that language. |
| Text visible in the video | Keep the original wording. Important advertising copy or a price is best planned as a separate layer added during editing. |

The prompt language and the spoken language are not the same thing. The Kling 2.6 and 3.0 guides describe automatic translation of dialogue from unsupported languages into English. Simply putting a Polish sentence in the prompt therefore does not ensure Polish dialogue. [Sources: Kling 2.6 — audio](https://kling.ai/quickstart/klingai-video-26-audio-user-guide), [Kling 3.0 — model guide](https://kling.ai/quickstart/klingai-video-3-model-user-guide).

### First, define the role of your materials

**Text-to-video: video from text.** The model creates both the scene's appearance and what happens in it. Describe the setting, main subject, action, lighting and camera work.

**Image-to-video: animating a starting image.** The photo becomes the first frame, or *start frame*, of the video. It defines the opening composition. Focus the prompt mainly on movement and what should remain unchanged.

**Generation using reference materials.** A photo or video serves as a reference for a product, character, style or movement. It does not have to be the first frame. Specify which features the model should take from it.

**Editing an existing video.** You start with a finished recording. Identify the element to change, the desired result and what should be preserved.

Not every platform offers every mode supported by a model family. The model name and platform name are therefore two separate pieces of information. [Sources: Runway — image-to-video](https://help.runwayml.com/hc/en-us/articles/48324313115155), [Kling — Omni](https://kling.ai/quickstart/klingai-video-3-omni-model-user-guide).

### How to use the ready-made ChatGPT instructions

Blocks headed **“ChatGPT starter prompt”** are for setting up a conversation with the chatbot. Do not paste them directly into the video generator.

First, select the instructions for your model and send them to ChatGPT. The chat should check the listed materials, summarise the key principles and ask for any missing information. Then describe the shot and attach a photo or other materials. You will receive an English prompt for the generator, with settings such as duration, aspect ratio and sound listed separately.

If you provide the shot description and all the necessary information straight away, there is no reason to wait for another message. As you continue, the chat should also remember the chosen model and platform rather than asking about them with every photo.

These instructions assume access to web search. If the chat cannot read the documentation, it should clearly separate what it has verified from what still needs confirmation. [Source: OpenAI — ChatGPT Search](https://help.openai.com/en/articles/9237897-chatgpt-search).

**These instructions are only for preparing text. They do not ask the chat to generate images or videos.**

---

## 1. Seedance 2.5

### How to structure the prompt

A Seedance 2.5 prompt can be treated as a short set of production notes for a shot. Define the result, assign roles to the attached materials, then describe the action and camera. The guide also covers references, timing, sound and editing existing footage. [Source: fal — Seedance 2.5 guide](https://fal.ai/learn/devs/how-to-use-seedance-2-5).

A useful order is:

> Shot type → role of the materials → sequence of actions → camera movement → sound → elements to preserve.

This is a way to organise information, not a mandatory syntax. If three precise sentences are enough to animate a photo, there is no need to turn them into a full script.

#### Give each reference a specific role

If two photos show the same package from different sides, say so explicitly. For an additional interior photo, specify whether it should define the whole setting or just the direction of the light.

```text
The first two reference images show the same coffee package from
different angles. Use them as references for the package's shape,
colors, and label design.

Use the third reference image for the kitchen setting and the
direction of the window light.
```

The guide to the fal integration uses markers such as `@Image1`. Use them only if the selected interface supports them. Typing a filename or marker does not replace attaching the material to the correct field. [Source: fal — Seedance 2.5 prompting](https://fal.ai/learn/devs/seedance-2-5-prompting-guide).

#### Describe actions in a logical order

A vague “woman picks up a cup, dynamic advert” leaves many decisions to the model. If you need a specific sequence, describe it more precisely:

> A right hand enters from the right side of the frame, grips the cup by its handle and lifts it above the countertop. After a brief pause, it puts the cup back in the same place, releases the handle and leaves the frame.

Each action should follow from the previous one. The cup cannot sit on the countertop and be held in a hand at the same time. If it needs to return to its starting position, include putting it down and releasing the handle.

#### Use a timeline when order matters

The example below assumes a starting image showing the package and a cup with its handle accessible from the right. It also requires a mode that can produce a 12-second clip.

```text
One continuous shot based on the starting image. The coffee package
remains stationary throughout.

0–3 seconds: a right hand enters from the right edge of the frame
and gently grasps the cup by its handle.

3–8 seconds: the hand lifts the cup a few centimeters above the table
and holds it steady. The camera slowly moves forward toward the
package, keeping the front label fully in view.

8–12 seconds: the hand lowers the cup to its original position,
releases the handle, and leaves the frame. The camera gradually
slows to a stop, holding a close view of the package.

Audio: quiet room ambience and a soft tap as the cup is
set back on the table.
```

Timestamps describe the intended sequence. They do not guarantee frame-accurate execution of every action. Match the number of actions to the video's duration. [Source: fal — Seedance 2.5 guide](https://fal.ai/learn/devs/how-to-use-seedance-2-5).

### Language, sound and limitations

The generator instructions in this guide are in English. Write dialogue in the language in which it should be spoken, however, and assign it to a specific character. Check support for that language separately.

The documentation also describes optional notation for sound and subtitles: parentheses for music, angle brackets for effects, braces for dialogue and `〖〗` for subtitles. You do not need these in every prompt. A plain description linking a sound to a visible action is a clear starting point. [Source: fal — Seedance 2.5 guide](https://fal.ai/learn/devs/how-to-use-seedance-2-5).

“Preserve the label layout” is not the same as a separate `negative_prompt` field. Whether that field is available depends on the integration. When editing video, select one base recording and describe the scope of the change, rather than assigning several files the role of primary footage at once. [Source: Dreamina — Seedance 2.5 guide](https://dreamina.capcut.com/seedance/seedance-2-5-prompt).

#### ChatGPT starter prompt — Seedance 2.5

```text
Help me prepare prompts for Seedance 2.5. Respond in English and write the final generator prompts in English. Do not generate images or videos.

Start by reviewing the available documentation:
https://fal.ai/learn/devs/how-to-use-seedance-2-5
https://fal.ai/learn/devs/seedance-2-5-prompting-guide
https://dreamina.capcut.com/seedance/seedance-2-5-prompt

Briefly summarise the principles relevant to writing prompts. If you cannot read a source, state what you were unable to confirm. Once I name the platform, also check the features of its specific integration.

If I have not described the shot yet, ask for its description, platform, duration, aspect ratio and materials. Ask only for missing information. If I have already supplied the description and attachments, get to work. Do not invent an example scene before I share my own idea.

Identify whether the task is text-to-video, first-frame animation, reference-based generation or editing footage. Assign each attached file a specific role. Use markers supported by the platform.

Preserve my idea. Describe actions in a logical order, separate camera movement from object movement and define how the shot ends. When animating a photo, focus on changes without repeating everything in the image. Use timed stages only when needed.

Keep dialogue in the language I specify. If the model does not support it, explain the limitation rather than translating the text yourself. Assign sounds to specific events.

Return one ready-to-use prompt in a code block. Below it, list the settings available on the platform in English. Prepare a separate negative prompt only if the selected mode supports one. Mention only limitations relevant to this shot.

For subsequent shots, use the previously supplied settings unless I change them. Do not ask again for information you already have.
```

---

## 2. Seedance 2.0 and the name “Seedance 2 Pro”

### Identify the model first

The name shown on a platform is not always enough to identify the model version. Magnific lists Seedance 2.0, Seedance 2.0 Fast and Seedance 2.0 Mini. The sources cited in this guide do not describe separate prompting rules for the name “Seedance 2 Pro”, so there is no basis for assigning it its own syntax. First, match the name to a specific variant in that service. [Source: Magnific — Seedance 2](https://www.magnific.com/seedance-2).

### How to write for Seedance 2.0

Runway's documentation covers text-based generation, references, and first and last frames. The prompt should match the selected mode. [Source: Runway — Seedance 2.0](https://help.runwayml.com/hc/en-us/articles/50488490233363-Creating-with-Seedance-2-0).

For **text-to-video**, describe the scene's appearance first, then the movement:

```text
A simple product commercial in a bright kitchen. A matte coffee
package stands on a pale wooden table, with a white ceramic cup
slightly behind it.

The camera tracks slowly to the right, parallel to the edge of the
table. The package remains fully in frame, with its front label
visible. The camera movement creates gentle parallax between the
package and the kitchen background.

Steam rises slowly from the cup. Soft daylight enters from the left.
The package remains upright and stationary throughout the shot.
```

For **image-to-video**, the photo already defines the opening scene's appearance, so you can shorten the description:

```text
The camera tracks slowly to the right, parallel to the table edge.
The coffee package remains upright and stationary, with its front
label visible throughout. Steam rises gently from the cup.
The lighting remains consistent with the starting image.
```

With **references**, specify what should be taken from each file. A video may serve only as a camera-movement reference, not as a set-design reference. A product photo may define its appearance without defining its position in the new frame.

With **first and last frames**, describe a logical transition between the supplied images. If the last frame is wider than the first, a continuous move towards the product may contradict it. [Source: Runway — Seedance 2.0](https://help.runwayml.com/hc/en-us/articles/50488490233363-Creating-with-Seedance-2-0).

### Language and negative instructions

English remains the working language here. Do not automatically transfer Seedance 2.5's special audio notation to version 2.0, however. Start with a plain description and check the documentation for your integration.

Apply the same approach to negative instructions. Define the intended result first, and only write a separate *negative prompt*, describing unwanted elements, if the platform provides the appropriate field.

#### ChatGPT starter prompt — Seedance 2.0

```text
Help me write prompts for Seedance 2.0. Give explanations and final prompts in English. Do not generate videos.

First, check the available documentation:
https://seed.bytedance.com/en/seedance2_0
https://help.runwayml.com/hc/en-us/articles/50488490233363-Creating-with-Seedance-2-0
https://fal.ai/seedance-2.0

Summarise the key principles. Distinguish model features from platform capabilities. Do not attribute Seedance 2.5 features to version 2.0 without confirmation. If you cannot read a source, state the scope of the uncertainty.

If I have not described the shot yet, ask for the description, platform, duration, aspect ratio and materials. Ask only for missing information. Check the platform's documentation once you know its name.

Prepare one coherent prompt based on my description. For text-to-video, define the scene's appearance and action. For photo animation, focus on movement. Assign references specific roles. For first and last frames, describe a transition consistent with both images.

Write natural sentences. Distinguish camera movement from object movement. Do not add unrequested characters, cuts or effects. Preserve brand names, on-screen text and dialogue language. If the spoken language is unsupported, explain this rather than translating the dialogue yourself.

Return the prompt in a code block, with settings separately in English. Add a negative prompt only for a mode that supports one. Keep the agreed settings for subsequent shots until I change them.
```

#### ChatGPT starter prompt — a model labelled “Seedance 2 Pro” on the platform

```text
My platform labels the model “Seedance 2 Pro”. Help me identify which version that name refers to, then prepare prompts for it. Do not generate images or videos.

If I have not named the platform, ask for its name or a screenshot of the settings. Use that to check the integration's documentation. You can start with these sources:
https://www.magnific.com/seedance/seedance-2
https://www.magnific.com/ai/docs/video-ai-models
https://seed.bytedance.com/en/seedance2_0

Do not treat Pro, Standard, Fast and Mini as interchangeable names. If identification remains uncertain, explain what could not be established. In that case, use only confirmed Seedance 2.0 family principles and do not give unverified parameters.

Once you know the model, ask for any missing shot details: description, duration, aspect ratio and materials. If I have already provided them, prepare the prompt without another round of questions.

Write explanations and the generator prompt in English. Adapt the prompt to text-to-video, first-frame animation or references. Describe the action, camera movement and elements that should remain unchanged.

Preserve the dialogue language and on-screen text. Check support for them in the specific variant. Add a negative prompt only if an appropriate field exists. Return the final text in a code block, with settings underneath.

For subsequent shots, keep using the agreed model and platform until I request a change.
```

---

## 3. Seedance 2.0 Fast

### How to test the Fast variant

The documentation presents Seedance 2.0 Fast as a variant of the same family with its own settings. That does not imply a need for different sentence structure or special language. [Source: fal — Seedance 2.0](https://fal.ai/seedance-2.0).

Limit the number of variables in your first attempt. Test camera movement first, then add motion in the surroundings, and finally an interaction with an object. If a particular change makes the result worse, it will be easier to identify the cause. This is a testing method, not a claim that Fast is only suitable for simple scenes.

Example for a product photo:

```text
One continuous shot based on the starting image. The camera slowly
moves forward toward the coffee package. The package stays fixed
on the table, with its front label visible throughout.

A small amount of steam rises from the cup behind it.
The soft window light remains constant.
```

In the next attempt, you could add a hand lifting the cup. Do not also change the lighting, background and direction of camera movement at the same time, or it will be difficult to tell which change affected the result.

### Language and settings

The examples use English. Set duration, aspect ratio and audio generation separately, according to the platform's capabilities.

The API schema linked below has no separate `negative_prompt` field for the Fast mode described. Do not add that parameter based on instructions for another generator. [Source: fal — Seedance 2.0 Fast, text-to-video](https://fal.ai/models/bytedance/seedance-2.0/fast/text-to-video/api).

#### ChatGPT starter prompt — Seedance 2.0 Fast

```text
Help me prepare prompts for Seedance 2.0 Fast. Respond in English and write the final prompts in English. Do not generate videos.

First, check the available documentation:
https://fal.ai/seedance-2.0
https://fal.ai/models/bytedance/seedance-2.0/fast/text-to-video/api
https://help.runwayml.com/hc/en-us/articles/50488490233363-Creating-with-Seedance-2-0

Briefly explain the key principles. Check my platform's features once I name it. Do not assume Fast has identical settings to other variants. Explicitly state when a source is inaccessible.

If I have not described the shot, ask for its description, platform, duration, aspect ratio and materials. Do not ask again for known information. Once there is enough information, prepare the prompt immediately.

Preserve the main idea and describe events in a logical order. Separate camera movement from object movement. When animating a photo, focus on changes. Assign references specific roles. Do not expand the scene with extra actions or effects on your own.

Preserve brand names, on-screen text and dialogue language. Check speech support if the shot requires it. Prepare a negative prompt only if the chosen mode supports one.

Return one prompt in a code block. List duration, aspect ratio, resolution and audio settings separately, within the options the platform offers. Keep the previously agreed configuration for subsequent shots.
```

---

## 4. Seedance 2.0 Mini

### Start with a shot that is easy to assess

Seedance 2.0 Mini appears in the Runway and Magnific catalogues. However, the materials cited here do not provide an equally detailed, dedicated prompting guide for this variant. The advice below is therefore a suggested testing approach, not a description of additional Mini limitations. [Sources: Runway — model catalogue](https://docs.dev.runwayml.com/guides/models/), [Magnific — Seedance 2](https://www.magnific.com/seedance-2).

Start with one task: the camera moves sideways, the product stays still and the lighting remains constant. Leave opening the package, dialogue and changing the set for separate tests.

When comparing Mini with the full Seedance 2.0, use the same image and as close to the same prompt as possible. Otherwise, you are judging both the model and differences in the description at once.

```text
The camera tracks slowly to the right, parallel to the table edge.
The coffee package remains stationary and fully in frame.
The camera movement creates gentle parallax between the package
and the background.

The package's shape, label design, and soft daylight remain
consistent with the starting image.
```

English is the working language here. Do not automatically apply rules from older models named “Lite” to Mini. Check reference, audio and last-frame support for the specific integration.

#### ChatGPT starter prompt — Seedance 2.0 Mini

```text
Help me write prompts for Seedance 2.0 Mini. Give explanations and final prompts in English. Do not generate videos.

First, check the available information:
https://docs.dev.runwayml.com/guides/models/
https://www.magnific.com/seedance/seedance-2
https://help.runwayml.com/hc/en-us/articles/50488490233363-Creating-with-Seedance-2-0

Separate principles shared by Seedance 2.0 from information specific to Mini. Do not transfer Lite or 2.5 features to it. If the documentation does not resolve something, say so.

If I have not described the shot, ask for its description, platform, duration, aspect ratio and materials. Ask only for missing information. Once you know the platform, check the features of that integration.

Prepare one prompt that follows my idea. Clearly define the action, direction and speed of camera movement, and stationary elements. When animating a photo, do not repeat its entire visual description. Refer only to materials I have actually attached.

If there are too many actions for the clip's duration, point this out and suggest reducing their number. Preserve on-screen text and dialogue language, and check speech support separately.

Return the finished prompt in a code block and the settings in English. Provide a negative prompt and additional parameters only after confirming their availability. Keep the agreed settings for subsequent shots.
```

---

## 5. Kling 2.5 / Kling 2.5 Turbo

### Describe camera and object movement separately

In the fal integration described here, the full model name includes “2.5 Turbo”, with additional labels identifying the quality variant. Check the full name before choosing settings. [Source: fal — Kling 2.5 Turbo Standard](https://fal.ai/models/fal-ai/kling-video/v2.5-turbo/standard/image-to-video).

Kling's text-to-video guide suggests describing the main subject or character, action and surroundings, then adding camera and lighting details. In image-to-video, movement matters more because the photo already defines the elements' appearance. [Source: Kling — text-to-video guide](https://kling.ai/quickstart/text-to-video-prompt-guide).

If the photo shows a package, a cup and a curtain, you could write:

```text
The camera tracks slowly to the right, parallel to the table edge.
The coffee package remains upright and stationary, with its front
label visible throughout.

Steam rises gently from the cup behind the package. The curtain
in the background sways slightly in a light breeze.
```

Each sentence refers to a specific element, so a correction does not require rewriting the whole prompt. If there is no curtain in the photo, do not add that sentence just because it appears in the example.

Pay particular attention to the difference between “the package rotates” and “the camera moves in an arc around the stationary package”. In the first case the product changes position; in the second the camera moves. If the label must remain visible throughout, describe a small arc, not a full orbit around the product.

### Negative instructions and language

The examples use English. A separate *negative prompt* can contain a short list of errors relevant to the shot, provided the selected mode supports that field.

That field is available in the Kling 2.5 Turbo Standard API described here. For example:

```text
duplicate packages, distorted packaging, altered label design,
flickering lighting, sudden camera shake
```

The same schema describes `cfg_scale`. Leave it at the default value initially. To assess its effect, change it while keeping the prompt and input material the same. Do not assume that the maximum value is automatically the best setting. [Source: fal — Kling 2.5 Turbo Standard, API](https://fal.ai/models/fal-ai/kling-video/v2.5-turbo/standard/image-to-video/api).

#### ChatGPT starter prompt — Kling 2.5 Turbo

```text
Help me prepare prompts for Kling 2.5 Turbo. Respond in English and write generator prompts in English. Do not generate video.

First, check the available documentation:
https://kling.ai/quickstart/text-to-video-prompt-guide
https://kling.ai/quickstart/image-to-video-guide
https://fal.ai/models/fal-ai/kling-video/v2.5-turbo/standard/image-to-video/api

Briefly summarise the principles. Once I name the platform, check the available Standard or Pro variant and its settings. Use the documentation for that specific variant. Flag anything you cannot confirm.

If I have not described the shot, ask for its description, platform, duration, aspect ratio and materials. Ask only for missing information.

Prepare one prompt based on my description. For text-to-video, describe the main subject, action, surroundings and camera. For image-to-video, focus on animating visible elements.

Clearly separate camera movement from product movement. Define direction, speed and extent of movement. Do not add unrequested cuts or effects. For an orbit, adapt its extent to the part of the product that must remain visible.

If the integration supports a negative prompt, prepare a short list of unwanted changes separately. Do not exclude text or logos that should remain in the image. Preserve brand names and label wording.

Return the prompt in a code block, any negative prompt in a second block, and settings in English. Do not add undocumented markers or word weights. Use the agreed configuration for subsequent shots.
```

---

## 6. Kling 2.6 Pro

### Plan picture and sound as one scene

The Kling 2.6 guide describes native audio generation. It recommends defining the scene, action and what should be heard. For dialogue, specify who is speaking, the exact words and how they are delivered. [Source: Kling 2.6 — audio guide](https://kling.ai/quickstart/klingai-video-26-audio-user-guide).

A clear prompt can have separate visual and audio sections, but both should describe the same event. If the sound of a cup being set down should be heard, include that action in the visual description too.

```text
A barista stands behind a small café counter, holding a white cup.
The camera starts in a medium shot and slowly moves closer as the
barista places the cup beside the coffee package.

After setting down the cup, the barista looks toward the customer
and says in a relaxed tone, in English:
"Good morning. Your coffee is ready."

Audio: a soft tap as the cup touches the counter, followed by the
barista's voice. Quiet café ambience continues in the background.
```

The camera moves, not the “medium shot”. Shot size describes framing. The example also makes clear when the character puts down the cup, when they start speaking and whom they address.

With two speakers, use consistent labels such as “the barista” and “the customer”. “She” can be ambiguous when two women are in the frame. For the first attempt, plan one speaker at a time, with no overlapping voices.

### Check the dialogue language carefully

The documentation lists Chinese and English speech, with other languages translated into English. Do not therefore plan a Polish advert with native Polish dialogue without checking the current integration. If the language is unsupported, a separately recorded or generated voiceover is an alternative. [Source: Kling 2.6 — audio guide](https://kling.ai/quickstart/klingai-video-26-audio-user-guide).

An English scene description does not require English dialogue in ALL CAPS. Use normal spelling, keeping capitals where needed, such as in names and abbreviations. [Source: fal — Kling 2.6 Pro, text-to-video](https://fal.ai/models/fal-ai/kling-video/v2.6/pro/text-to-video/api).

If the platform can assign a saved voice to a character, use that voice's actual identifier. A placeholder such as `@VoiceName` does not create a voice and will not work without the appropriate setup.

#### ChatGPT starter prompt — Kling 2.6 Pro

```text
Help me write prompts for Kling 2.6 Pro. Give explanations and final prompts in English. Do not generate videos.

First, check the available documentation:
https://kling.ai/quickstart/klingai-video-26-audio-user-guide
https://fal.ai/models/fal-ai/kling-video/v2.6/pro/text-to-video/api

Summarise the principles for visuals, sound and spoken languages. Once I name the platform, check the specific mode's capabilities, including audio and negative-prompt support. Flag information you cannot confirm.

If I have not described the shot, ask for its description, platform, duration, aspect ratio, materials and any dialogue. Ask only for missing details. If my description is complete, prepare the prompt immediately.

Separate visuals and sound in the prompt, but keep them consistent. Assign every sound effect to a specific event. Assign every line of dialogue to a character, specifying tone and pace. Assess whether the dialogue and actions fit the clip's duration.

For photo animation, focus on movement. Distinguish framing from camera movement. Do not add characters or objects I have not requested.

Preserve the exact wording and language of the dialogue. If Polish speech is unsupported or unconfirmed, explain this and suggest a separate voiceover track. Do not translate dialogue on your own. Use only actual voice identifiers available on my platform.

Return the prompt in a code block, with settings in English, including the audio setting. Add a negative prompt only for a mode that supports one. Use the previously agreed details for subsequent shots.
```

---

## 7. Kling 3.0 / Kling 3.0 Pro

### Choose: one shot or several

The Kling 3.0 guide describes Multi-Shot mode, which lets you plan a sequence with changes in framing. It also lists element references linked to the starting image. [Source: Kling 3.0 — model guide](https://kling.ai/quickstart/klingai-video-3-model-user-guide).

Before writing, decide whether you need one continuous shot or a short sequence. For a smooth push-in without cuts, you can start with:

```text
One continuous shot.
```

Do not later add “cut to a close-up”, because that means cutting to another shot. Also check that the platform settings match your intended approach.

Example of a three-shot sequence:

```text
Shot 1 — 4 seconds:
A medium shot of a coffee package and a cup on a kitchen table.
A hand enters from the right, gently turns the cup until its handle
points to the right, releases it, and leaves the frame.

Shot 2 — 4 seconds:
Cut to a close-up of the same package in the unchanged setting.
The camera slowly moves closer, keeping the front label in view.
The cup remains in the position established in the first shot.

Shot 3 — 4 seconds:
Cut to a wider, static view of the same tabletop arrangement.
Keep the package in the lower half of the frame, with clear space
above it for a title to be added during editing.
```

This breakdown assumes 12 seconds and a mode that supports such a sequence. The hand finishes its action and leaves the frame before the first cut, while the cup stays in its new position. The transitions therefore do not leave the model to work out what happened to these elements.

### Keep character and product references consistent

If the platform lets you save a product or character as a reference element, use that same element across all shots. A long list of features does not replace correctly attaching the reference.

If a character already has a specific voice assigned, do not define a conflicting voice in the next shot. Kling's guide highlights such conflicts. [Source: Kling 3.0 — model guide](https://kling.ai/quickstart/klingai-video-3-model-user-guide).

### Language and the Pro variant

The examples use English. For generated speech, the guide lists Chinese, English, Japanese, Korean and Spanish. Polish is not on that list; other languages are described as being translated into English. [Source: Kling 3.0 — model guide](https://kling.ai/quickstart/klingai-video-3-model-user-guide).

The Pro label does not itself mean you need a different prompting style. Check which parameters and features it denotes at that provider. [Source: Artlist — Kling 3 family](https://help.artlist.io/hc/en-us/articles/37026869403933-Kling-3-Model-Family).

#### ChatGPT starter prompt — Kling 3.0 / Pro

```text
Help me prepare prompts for Kling 3.0 or Kling 3.0 Pro. Respond in English and write generator text in English. Do not generate images or videos.

First, check the available documentation:
https://kling.ai/quickstart/klingai-video-3-model-user-guide
https://help.artlist.io/hc/en-us/articles/37026869403933-Kling-3-Model-Family

Briefly explain the key principles. Once you know my platform, check the model variant, Multi-Shot, element references, audio and negative prompts. State what could not be confirmed.

If I have not described the shot, ask for its description, platform, duration, aspect ratio and materials. Ask only for missing information.

Use my description to determine whether one continuous shot or a sequence is needed. Do not combine a no-cuts instruction with cutting instructions. For a sequence, prepare separate descriptions and timings that match the mode. Maintain continuity of object positions and character actions between shots.

When animating a first frame, focus on movement. Use only references actually attached and existing element names.

Assign every line of dialogue to a character. Preserve the exact words and language and check support for them. Do not override a voice already assigned to a reference or translate dialogue unless asked.

Return the final prompt or individual shot descriptions in code blocks. List settings separately in English. Include additional fields only if the integration supports them. For subsequent shots, use the agreed model, platform and references.
```

---

## 8. Kling 3.0 Turbo

### Check the mode and input format first

The Kling 3.0 Turbo Pro text-to-video API schema describes either a single `prompt` or a `multi_prompt` structure with separate descriptions and durations for each shot. These are alternative ways to submit instructions, not two fields to fill in simultaneously. [Source: fal — Kling 3.0 Turbo Pro, API](https://fal.ai/models/fal-ai/kling-video/v3/turbo/pro/text-to-video/api).

In a graphical interface, choose the correct mode first. If the platform provides separate storyboard fields, simply typing “Shot 1, Shot 2” into the normal text field does not replace using them.

When reading API documentation, check the **Input section for the selected mode**. The page may also include data-type definitions used elsewhere. Finding `negative_prompt` or `generate_audio` further down the page does not necessarily mean that this mode accepts that parameter. [Source: fal — Kling 3.0 Turbo Pro, API](https://fal.ai/models/fal-ai/kling-video/v3/turbo/pro/text-to-video/api).

### Describe the direction, speed and extent of movement

The sources cited here do not describe a separate prompt syntax for Turbo. You can start with the same precise sentences used for Kling 3.0.

```text
One continuous product shot. The camera moves slowly to the right
along a small arc around the stationary coffee package, keeping
approximately the same distance from it.

The camera remains aimed at the package. Its front label stays
visible throughout, while the changing viewpoint creates gentle
parallax in the background.
```

This describes an arc movement. Do not use it when the goal is a straight sideways move. Specify the movement's speed in the prompt; the model variant's name does not determine it.

English remains the working language. Check dialogue, reference and additional-parameter support for the selected mode rather than transferring settings from standard Kling 3.0.

#### ChatGPT starter prompt — Kling 3.0 Turbo

```text
Help me write prompts for Kling 3.0 Turbo. Give explanations and final prompts in English. Do not generate videos.

First, check the available documentation:
https://fal.ai/models/fal-ai/kling-video/v3/turbo/pro/text-to-video/api
https://console.higgsfield.ai/models/kling-video/v3.0-turbo/text-to-video/playground
https://kling.ai/quickstart/klingai-video-3-model-user-guide

Summarise the principles relevant to this variant. Once the platform is known, check the specific mode's features and Input section. Do not assume that other models' parameters are available in Turbo. Explicitly flag uncertainty caused by missing documentation.

If I have not described the shot, ask for its description, platform, duration, aspect ratio and materials. Do not ask again for known information.

For a single shot, prepare one coherent prompt. For a sequence, write separate shots and durations in the platform's required format. Treat prompt and multi_prompt as alternative ways to submit a description, not as fields to fill in simultaneously.

Specify the camera movement's direction, speed and extent precisely. Distinguish lateral movement, an arc, camera rotation and a change in focal length. Describe product movement separately. Preserve my idea without additional cuts, characters or effects.

Keep the dialogue language and check speech support for the selected mode. Include audio, references and negative prompts only when available in that specific integration.

Return the final text in a code block, with settings in English. Use the previously agreed configuration for subsequent shots.
```

---

## 9. Kling O3 / Kling 3.0 Omni

### When editing, describe the change, not the entire scene again

Distinguish Kling O3 from standard Kling 3.0. The Omni guide covers references, elements, voice and storyboards, and also refers to O1 documentation for editing. [Source: Kling — 3.0 Omni guide](https://kling.ai/quickstart/klingai-video-3-omni-model-user-guide).

Creating a new scene requires a description of what should be created. When editing existing footage, the instruction for the change matters more:

> Replace the cup in the footage with the cup from the reference photo. Preserve the action, camera work and surroundings.

This is not the same as describing a new shot in Runway Gen-4. Here, the direct instruction “replace” has a specific target and source recording.

Example of editing footage in which a hand lifts a cup and puts it back on the countertop:

```text
Replace the cup in the input video with the cup shown in the reference
image. Use the image only as a reference for the replacement cup,
not for the surrounding scene.

Keep the original hand movement, camera movement, tabletop, lighting,
and timing. Match the replacement cup to the original cup's position
and orientation throughout the shot.

Keep the fingers naturally wrapped around the handle, with correct
overlap between the hand and the cup. When the cup rests on the table,
maintain natural contact with the surface and a matching contact shadow.
```

The last paragraph defines the relationship between the hand, cup handle and countertop. It does not demand identical occlusion for two potentially different shapes; it specifies the natural contact expected.

### Define the scope of the edit and each file's role

Choose one base recording. Specify the object to change, the relevant part of the video if needed, and the result. If you are only replacing the packaging, do not redescribe the kitchen, the character's clothing and the lighting.

A reference photo may define only the product. Make that explicit so the description does not suggest copying the entire background.

The editing guide covers replacing elements, changing surroundings and style, and using reference motion, among other things. It does not guarantee that every pixel outside the changed object will remain identical. [Source: Kling — O1 guide](https://kling.ai/quickstart/klingai-video-o1-user-guide).

Prepare the prompt in English. Check audio support separately for generation and editing: its presence in one mode does not confirm that it works in the other.

#### ChatGPT starter prompt — Kling O3

```text
Help me prepare prompts for Kling O3 / Kling 3.0 Omni. Respond in English and write the final generator text in English. Do not generate images or videos.

First, check the available documentation:
https://kling.ai/quickstart/klingai-video-3-omni-model-user-guide
https://kling.ai/quickstart/klingai-video-o1-user-guide
https://help.artlist.io/hc/en-us/articles/37026869403933-Kling-3-Model-Family

Briefly summarise the principles. Once I name the platform, check the capabilities of its specific mode. If sources are unavailable or leave an issue unresolved, state the scope of the uncertainty.

If I have not described the task, ask for its description, platform, materials and necessary video settings. Ask only for missing information.

Distinguish creating a new scene, working with references and editing video. For editing, identify the base footage, target object, scope of the change and elements to preserve. Do not rebuild the entire scene when I ask to replace one object.

Give every reference a specific role. Use only real attachments and markers supported by the platform. Account for object position, contact with hands and surfaces, occlusion between elements, shadows and motion continuity.

Describe preserving unchanged elements as a requirement, not a guarantee of pixel-for-pixel identity. Check audio and negative prompts for the selected mode, not just the whole model family. Preserve the dialogue language and verify its support.

Return one prompt in a code block, with settings in English. For later tasks, use the details already agreed unless I change the model, platform or workflow.
```

---

## 10. Veo 3.1

### Separate the visual description from the audio description

For Veo, English is a choice supported by Google's documentation, which states that it is fully supported. However, the language of the instructions should not be confused with support for every dialogue language. [Source: Google — Veo in the Gemini API](https://ai.google.dev/gemini-api/docs/veo).

For text-based generation, define the main subject, setting, action, camera and lighting. Describe sound separately, referring to the same scene.

```text
A close-up product shot of a coffee package beside a white cup
and saucer on a pale wooden table in a quiet kitchen.

The camera slowly moves forward toward the package. Steam rises
gently from the cup. Soft morning light enters from the left.

A hand enters from the right, places a teaspoon on the saucer,
and leaves the frame. The package remains stationary.

Audio: quiet kitchen ambience, the faint sound of a kettle in the
background, and a light clink as the metal teaspoon touches the
ceramic saucer.
```

In this example, the saucer and hand appear in the visual description, not only in the audio instructions. The sound is linked to the spoon touching the saucer.

When animating a photo, shorten the visual description. Google recommends a high-quality input image and avoiding unnecessary repetition of its contents. Use the spoon example only if it matches the photo and intended action. [Source: Google Cloud — video generation best practices](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/video/best-practice).

#### Dialogue

Specify the speaker, exact words, language and tone. The amount of dialogue must fit the clip's duration.

```text
The barista looks toward the customer and says warmly, in English:
"Take a moment. Your coffee is ready."
```

For Polish dialogue, check the selected integration's capabilities and run a test before planning a larger production. Stated support for English prompts does not confirm reliable Polish speech generation.

### Negative prompt

Google's guide recommends listing unwanted elements without constructions such as “no” and “don't”. Put that list in the separate field, if the platform provides one. [Source: Google Cloud — Veo prompting guide](https://cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide).

```text
duplicate products, distorted packaging, altered label design,
flickering exposure, sudden camera shake
```

Do not write the generic “text” if the product label must remain visible. The list should exclude specific errors, not elements the advert needs.

### References, first frames and last frames

A product photo used as a reference does not have to be the starting image. In reference mode, specify the features the model should retain. With first and last frames, describe the action connecting both images.

Google's documentation also covers video extension, but feature availability depends on the variant. Do not assume the full feature set is available in Lite. [Source: Google — Veo in the Gemini API](https://ai.google.dev/gemini-api/docs/veo).

#### ChatGPT starter prompt — Veo 3.1

```text
Help me write prompts for Veo 3.1. Respond in English and prepare final prompts in English. Do not generate video.

First, check the available documentation:
https://ai.google.dev/gemini-api/docs/veo
https://cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide
https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/video/best-practice

Briefly explain the principles for visuals, sound and references. Once I name the platform, check its modes and settings. Flag information you could not confirm.

If I have not described the shot, ask for its description, platform, duration, aspect ratio, materials and desired sound. Ask only for missing information. If the description is complete, prepare the prompt immediately.

For text-to-video, define appearance, action, camera and lighting. For image-to-video, focus on movement. Distinguish a reference photo, first frame and last frame and give them the correct roles.

Describe sound separately but keep it consistent with the visuals. If an action should be heard, include it in the scene's sequence. Assign each line of dialogue to a character, preserving its wording and language. Check language support rather than translating dialogue on your own.

If the integration provides a negative prompt, prepare a short list of unwanted elements without “no” or “don't”. Do not exclude labels, text or logos that should be preserved.

Return the prompt and any negative prompt in separate code blocks. List settings in English, respecting the available duration and mode. Use the previously agreed configuration for subsequent shots.
```

---

## 11. Veo 3.1 Fast

### Use the same principles, but check the variant's settings

Google's documentation does not specify a separate syntax for Fast. The starting point remains the scene description, camera movement, sound and, where applicable, a negative prompt. Check the available settings for the specific variant. [Source: Google — Veo in the Gemini API](https://ai.google.dev/gemini-api/docs/veo).

For the first test, choose one result that is easy to assess: for example, the camera should move closer while the product stays still.

```text
One continuous shot based on the starting image. The camera slowly
moves forward toward the coffee package, keeping its front label
fully in view. The package remains fixed on the table.
Steam rises gently from the cup.

Audio: quiet indoor ambience.
```

When switching from Fast to another variant to compare results, keep the prompt and materials the same. Do not treat the switch as upscaling an existing video, however. It creates a new generation, so the action may differ.

Write the instructions in English, following the language choice used for the Veo family. [Source: Google — Veo in the Gemini API](https://ai.google.dev/gemini-api/docs/veo).

#### ChatGPT starter prompt — Veo 3.1 Fast

```text
Help me prepare prompts for Veo 3.1 Fast. Write explanations and generator text in English. Do not generate videos.

First, check the available documentation:
https://ai.google.dev/gemini-api/docs/veo
https://cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide

Summarise the key principles. Once the platform is known, check the Fast variant's settings. Do not assign it limitations or a separate syntax that the documentation does not confirm. Flag gaps in the sources.

If I have not described the shot, ask for its description, platform, duration, aspect ratio, materials and desired sound. Ask only for missing details. If there is enough information, prepare the prompt immediately.

Preserve my idea. For photo animation, focus on the movement of the camera, objects and surroundings. Fit the number of actions to the clip's duration. Do not add cuts or new scene elements on your own.

Describe sound separately, consistently with the visuals. Assign dialogue to a specific character, preserve its words and language, and check speech support. Do not translate dialogue unless I ask.

Add a negative prompt only in a supported mode, as a list of unwanted elements without “no” or “don't”. Do not present switching model variants as a way to reproduce an earlier video identically at higher quality.

Return the prompt in a code block, with settings in English. Use the previously agreed details for subsequent shots.
```

---

## 12. Veo 3.1 Lite

### Match the idea to the available inputs

According to the Gemini API documentation linked below, Lite supports text-to-video, image-to-video, and first and last frames. It does not support the `referenceImages` input, extending an existing video or 4K generation. Check that this information still applies to your selected mode before production. [Source: Google — Veo in the Gemini API](https://ai.google.dev/gemini-api/docs/veo).

Listing several photo filenames in the prompt does not replace a missing reference input. To animate a specific product composition, prepare a suitable starting image and describe its movement.

The example below assumes a product photo taken with the camera slightly above the product's centre:

```text
The camera slowly lowers to the height of the package's front label,
keeping the same horizontal distance from the product. It adjusts
its tilt smoothly so that the package remains in frame, ending
with the label centered in the image.

The package and cup remain stationary. Steam rises gently from
the cup, and the lighting stays consistent with the starting image.
```

The description separates lowering the camera's position from changing where it points. A vague “the camera moves down towards the product” does not explain whether this means a downward move, a tilt or a diagonal push-in.

Attach the last frame to the appropriate field, if the platform provides one. “End frame: attached image” does not automatically assign a normal attachment to the last-frame function.

English remains the working language. Shorter scenes are convenient for testing, but are not in themselves a documented Lite requirement.

#### ChatGPT starter prompt — Veo 3.1 Lite

```text
Help me write prompts for Veo 3.1 Lite. Respond in English and write final prompts in English. Do not generate video.

First, check the available documentation:
https://ai.google.dev/gemini-api/docs/veo
https://cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide

Check information specific to Lite. Once the platform is known, also verify its integration's features. Do not assign Lite reference inputs, video extension or resolutions just because another Veo variant supports them. Flag unconfirmed information.

If I have not described the shot, ask for its description, platform, duration, aspect ratio and materials. Ask only for missing details.

Prepare one prompt from the description. Adapt it to text-based generation or image animation. Use first and last frames only if the selected mode supports them. Do not treat a photo as supplied to the model unless it is attached to the correct input.

Clearly describe the action and camera movement. Distinguish a change in camera position from rotation or tilt. Link sound to the visuals. Preserve the dialogue language and check support for it. Add a negative prompt only after confirming that the appropriate field exists.

Return the prompt in a code block and settings in English. If the idea requires an unavailable feature, explain the problem and suggest the closest workable approach. Keep the agreed settings for subsequent shots.
```

---

## 13. Runway Gen-4.5

### Focus on a precise description, not prompt length

Runway's guide stresses that there is no mandatory formula or ideal word count. Natural sentences help describe relationships between elements. The order in which features are listed is not a documented weighting system. [Source: Runway — text-to-video guide](https://help.runwayml.com/hc/en-us/articles/42460036199443-Text-to-Video-Prompting-Guide).

For **text-to-video**, describe appearance and movement. For **image-to-video**, the photo defines the initial composition, lighting and style; the prompt should focus on what happens next. [Source: Runway — Gen-4.5](https://help.runwayml.com/hc/en-us/articles/46974685288467-Creating-with-Gen-4-5).

Example for a product photograph:

```text
The camera moves slowly forward and slightly to the right along
a straight diagonal path. The coffee package stays fixed on the
table, with its front label visible throughout.

Steam rises gently from the cup. Near the end of the shot, the camera
gradually slows to a stop and holds a close-up of the package.
```

The description makes clear that the camera moves in a straight line rather than an arc. The ending is also defined: movement gradually stops, then the camera holds the frame. After the first attempt, you can change the distance travelled or the product's size in the final frame.

### You can describe successive stages

There is no need to reduce every scene to one action. Gen-4.5 documentation covers both natural sequences of actions and approximate timestamps. [Source: Runway — image-to-video guide](https://help.runwayml.com/hc/en-us/articles/48324313115155).

The example below assumes a cup in the foreground and a package behind it:

```text
The camera remains stationary. The cup in the foreground is in focus,
while the coffee package behind it is softly out of focus.

A hand enters from the right, grasps the cup by its handle, and lifts
it out of the frame. As the cup leaves, focus shifts smoothly to the
coffee package, bringing its front label into sharp focus.
```

The focus change is not combined with unspecified camera movement. First, it is clear what is in focus, then what leaves the foreground, and finally which object receives focus. This sequence must still fit the clip's duration.

### Check that the movement fits the photo

Motion blur, a leaning figure or the arrangement of objects can imply a direction of action. If the prompt asks for the opposite, image and instructions contradict each other. Runway highlights these visual cues. [Source: Runway — image-to-video guide](https://help.runwayml.com/hc/en-us/articles/48324313115155).

The examples use English and positive descriptions of the desired result. Do not add a separate negative prompt without confirming that the field is available. Plan sound separately rather than assuming the animation model supports it.

#### ChatGPT starter prompt — Runway Gen-4.5

```text
Help me prepare prompts for Runway Gen-4.5. Give explanations and final prompts in English. Do not generate videos.

First, check the available documentation:
https://help.runwayml.com/hc/en-us/articles/46974685288467-Creating-with-Gen-4-5
https://help.runwayml.com/hc/en-us/articles/42460036199443-Text-to-Video-Prompting-Guide
https://help.runwayml.com/hc/en-us/articles/48324313115155-Image-to-Video-Prompting-Guide

Briefly summarise the principles. Once the platform is known, check its supported modes and settings. Flag information you cannot confirm.

If I have not described the shot, ask for its description, platform, duration, aspect ratio and any photo. Ask only for missing details.

Prepare one prompt that follows my idea. For text-to-video, describe appearance and motion. For image-to-video, focus on animation, camera, direction and speed without repeating the photo's entire contents.

Use natural, precise sentences. Do not impose an artificial word limit or treat the order of descriptions as a weighting system. Describe stages and approximate timings only when they clarify the scene's progression. Check that the actions fit the clip's duration.

Distinguish camera movement, focus changes and subject movement. Check the instructions against the input image. Use positive descriptions of the desired result rather than long lists of prohibitions. Do not add unconfirmed negative-prompt or audio-generation features.

Return the prompt in a code block and settings in English. If you find a contradiction, briefly explain it and suggest a correction. Use the previously agreed settings for subsequent shots.
```

---

## 14. Runway Gen-4 Turbo

### Start with the initial image

In the workflow described by Runway's documentation, Gen-4 Turbo requires an input image. A text description alone cannot replace a missing first frame. [Source: Runway — model catalogue](https://docs.dev.runwayml.com/guides/models/).

The Gen-4 guide recommends starting with a simple instruction and gradually adding necessary details. Describe the movement of the subject, camera and surroundings. [Source: Runway — Gen-4 guide](https://help.runwayml.com/hc/en-us/articles/39789879462419-Gen-4-Video-Prompting-Guide).

```text
The camera slowly moves forward toward the coffee package.
The package and cup remain stationary. Steam rises gently from
the cup, while the package's front label stays visible throughout.
```

If the push-in works as intended, you can add a small sideways movement. Combining an orbit, zoom, focus change and product movement from the outset makes it harder to tell which instruction caused a problem.

### Describe the desired result instead of prohibitions

Runway explicitly advises against negative phrasing in Gen-4 prompts. Instead of “no camera movement”, use:

```text
The camera remains stationary.
```

Instead of “the package must not move”, write:

```text
The package remains stationary.
```

Both sentences clearly define the desired state. A long list such as “no blur, no deformation, no extra objects” is not a recommended substitute for a precise shot description. [Source: Runway — Gen-4 guide](https://help.runwayml.com/hc/en-us/articles/39789879462419-Gen-4-Video-Prompting-Guide).

English remains the working language here. If a generation fails, check the starting image and the consistency of the movement before making the prompt longer.

#### ChatGPT starter prompt — Runway Gen-4 Turbo

```text
Help me write prompts for Runway Gen-4 Turbo. Respond in English and write the final generator text in English. Do not generate images or videos.

First, check the available documentation:
https://help.runwayml.com/hc/en-us/articles/39789879462419-Gen-4-Video-Prompting-Guide
https://docs.dev.runwayml.com/guides/models/

Summarise the principles relevant to photo animation. Once I name the platform, check its settings. Flag information you could not confirm.

If I have not attached a starting image, ask for one. Also request any missing details: movement description, platform, duration and aspect ratio. Do not ask again for information I have already supplied.

Prepare one prompt based on the photo and description. Treat the image as the basis for the scene's appearance. Focus on object, camera and environmental movement rather than describing every detail of the photo again.

Use positive descriptions of the desired result. Instead of forbidding movement, say that the camera or object remains stationary. Do not add a negative prompt, word weights or polite introductions such as “please add”.

Preserve my idea and prepare one coherent shot. Do not add cuts or independent movements unless I request them. Check that the planned animation does not contradict the photo.

Return the prompt in a code block, with settings in English. Use the previously agreed configuration for subsequent photos.
```

---

## 15. MiniMax Hailuo 2.3 Fast

### Use the correct camera commands

MiniMax documents Hailuo 2.3 Fast as an image-to-video variant that takes a starting image and a text description of the animation. It also lists camera commands written in square brackets. [Source: MiniMax — image-to-video](https://platform.minimax.io/docs/api-reference/video-generation-i2v).

| Command | Intended movement |
| --- | --- |
| `[Push in]` | A push-in: the camera moves closer to the subject. |
| `[Truck right]` | The camera moves sideways to the right, changing its position. |
| `[Pan right]` | The camera turns to the right without moving sideways. |
| `[Tilt up]` | The camera tilts upwards, changing its viewing direction rather than rising. |
| `[Static shot]` | A stationary camera with fixed framing. |

The command specifies the type of movement. Use a plain sentence to clarify its speed and how the objects should behave:

```text
[Push in]
The camera moves slowly and smoothly toward the coffee package,
keeping its front label fully in view.

The package remains stationary on the table. Steam rises gently
from the cup. The shot ends on a close-up of the package.
```

Write the prompt in English and keep the command names exactly as documented. Do not translate the text inside the brackets.

### Do not combine movements unnecessarily

The documentation allows movements to be combined within one pair of brackets and recommends no more than three simultaneous commands. Start with one command in your first test. This makes it easier to judge whether the model is producing the right type of movement. [Source: MiniMax — image-to-video](https://platform.minimax.io/docs/api-reference/video-generation-i2v).

Also check the `prompt_optimizer` setting. The documentation describes automatic prompt rewriting and the option to disable it. When comparing your own precise descriptions, the optimiser’s behaviour can affect how you interpret the results. [Source: MiniMax — image-to-video](https://platform.minimax.io/docs/api-reference/video-generation-i2v).

### Do not assume features from other MiniMax models apply

In the documentation linked below, the separate first-and-last-frame mode applies to Hailuo-02. This does not confirm last-frame support in Hailuo 2.3 Fast. [Source: MiniMax — first and last frames](https://platform.minimax.io/docs/api-reference/video-generation-fl2v).

Likewise, the audio feature documented for H3 does not confirm its availability in Hailuo 2.3 Fast. For this workflow, plan sound separately unless the documentation for your specific integration says otherwise.

#### ChatGPT starter prompt — Hailuo 2.3 Fast

```text
Help me prepare prompts for MiniMax Hailuo 2.3 Fast. Give explanations in English and write generator prompts in English. Do not generate videos.

First check the available documentation:
https://platform.minimax.io/docs/api-reference/video-generation-i2v
https://platform.minimax.io/docs/api-reference/video-generation-fl2v
https://hailuoai.video/tools/hailuo-2-3-model

Briefly summarise the rules for animating a photo and the camera commands. Once you know the platform, check which features its integration provides. Flag anything you could not confirm.

If I have not attached a starting image, ask for it. Fill in the missing details: shot description, platform name, duration and aspect ratio. Do not ask again for information I have already provided.

Prepare one prompt focused on animating the photo. If the integration supports MiniMax commands, choose the appropriate command from the documentation and keep its original spelling in square brackets.

Distinguish between camera translation, panning, tilting and zooming. Specify the speed and range of movement. Do not combine movements unless my description calls for it. Refer only to an image that has actually been attached.

Check whether prompt_optimizer is available. Discuss this setting separately from the prompt and point out when enabling it means the description is automatically rewritten before generation.

Do not attribute Hailuo-02 or H3 features to this model. Include a last frame, audio and a negative prompt only after confirming support in the selected mode.

Return the prompt in a code block and the settings in English. Use the configuration already agreed upon for subsequent photos.
```

---

## 16. MiniMax H3

### Distinguish cloud access, open weights and 2K mode

MiniMax H3 generates video with stereo sound. It supports text, first and last frames, and reference materials, with an official duration range of 4–15 seconds. The public base version defaults to 768p at 24 fps. [Source: MiniMax — video generation](https://platform.minimax.io/docs/guides/video-generation).

H3’s open weights can be run in ComfyUI. Ready-made workflows include text-to-video, FL2VA for animating first and last frames, and Ref2VA for working with references. These workflows may use reference tags such as `<Picture 1>`, `<Video 1>` and `<Audio 1>`. Always check the syntax for your chosen integration. [Source: ComfyUI — MiniMax H3 workflow](https://docs.comfy.org/tutorials/video/minimax/minimax-h3-native).

Turbo LoRA can reduce the number of local generation steps, but it does not turn H3 into H3 Max or H3 Max Turbo. The H3-Regenerate-2K module has not been released for local installation; the developer offers this process through an API. A conventional upscaler is not equivalent. [Source: Hugging Face — MiniMax H3](https://huggingface.co/MiniMaxAI/MiniMax-H3).

### Licensing and hardware in Poland

**The H3 community licence excludes the European Union, including Poland.** Before downloading weights for local use, you must confirm that you have the necessary rights. The Professional licence offered by Comfy starts at USD 5,000 per month. Generating in a licensed cloud service does not give you a licence for your own installation. [Source: Comfy — MiniMax licence](https://comfy.org/minimax/license).

ComfyUI describes configurations that reduce total memory requirements to around 42.5 GB and allow the model to run with offloading even on an RTX 3060. This does not mean you need 42.5 GB of VRAM. The figures below are a cautious starting point, not a guaranteed minimum for every workflow. [Source: ComfyUI — H3 support](https://blog.comfy.org/p/minimax-h3-day-0-support-in-comfyui).

| Component | Starting point for testing | More frequent use |
| --- | --- | --- |
| Graphics card | NVIDIA with 12 GB VRAM, a memory-saving configuration and offloading | NVIDIA with 24 GB VRAM and suitable weights |
| RAM | 64 GB | 64–128 GB, depending on the workflow |
| Storage | NVMe SSD with 100–150 GB free | Additional headroom for models and working files |
| First test | Short clip, lower resolution, one generation | Move to 768p and longer shots only after measuring memory use |

Do not promise a single generation time. Getting the model to run and comfortably producing many tests a day are different requirements.

#### ChatGPT starter prompt — MiniMax H3

```text
Help me prepare prompts for MiniMax H3. Reply in English and write finished generator prompts in English. Do not generate videos.

First check the MiniMax H3 documentation for text, first and last frames, and reference materials. Distinguish between the official API, local ComfyUI workflows and features of the platform I use.

If I have not described the shot, ask for the platform, mode, duration, aspect ratio, materials and desired sound. Ask only for missing information. Assign a specific role to each reference material.

Describe camera movement separately from object movement. When there are several actions, define their order and how the shot ends. Do not carry over Hailuo commands or H3 Max settings without confirming support.

For a local installation in Poland, first check the licence, hardware, ComfyUI version, quantisation and offloading. Do not equate Turbo LoRA with H3 Max Turbo or promise a local Regenerate-2K process.

Return the prompt in a code block and explain settings and limitations in English.
```

#### Codex instruction — preparing a local test

```text
Prepare a local MiniMax H3 installation in ComfyUI on this computer.

First detect the operating system, GPU, VRAM, RAM, driver and free storage. Check the current MiniMax and ComfyUI documentation. Verify the licence terms for a user in Poland. If separate permission is required, tell me and do not download weights until the necessary rights have been confirmed.

Once confirmed, prepare an isolated installation, choose compatible PyTorch and library versions, download only the official files needed for one image-to-video workflow, configure quantisation and offloading, and run a short test measuring time and memory use.

Do not equate local H3 or Turbo LoRA with H3 Max. Do not use a paid API, change drivers or disable security measures without my consent.
```

---

## 17. MiniMax H3 Max

### Treat H3 Max as a cloud variant

H3 Max is a variant tuned for instruction following and faster generation. MiniMax offers 480p and 768p, while the fal integration also offers 1080p. The verified offerings are cloud services and APIs. Do not apply installation instructions for the regular H3’s open weights to H3 Max. [Sources: MiniMax — video generation](https://platform.minimax.io/docs/guides/video-generation), [fal — MiniMax H3 Max](https://fal.ai/minimax-h3-max).

Before generating, check the selected resolution, duration, input mode and automatic prompt expansion. If the platform does not document a separate negative prompt field, put your visual requirements in the main description.

#### ChatGPT starter prompt — MiniMax H3 Max

```text
Help me write prompts for MiniMax H3 Max. Reply in English and prepare finished prompts in English. Do not generate videos.

First check the current MiniMax documentation and the integration I use. Distinguish between H3, H3 Max, H3 Max Turbo and Hailuo. Do not assume their features and settings are identical.

If information is missing, ask for the platform, shot description, duration, resolution, aspect ratio, materials and sound. Prepare one coherent shot, describe the sequence of actions and separate camera movement from object movement.

Check how automatic prompt expansion works. Include a negative prompt, additional references and audio only after confirming support in the specific mode. Do not present H3 Max as a model with public weights for local installation.

Return one prompt in a code block and the settings in English.
```

---

## 18. MiniMax H3 Max Turbo

### Distinguish H3, H3 Max and H3 Max Turbo

According to fal, H3 Max and H3 Max Turbo are variants developed by fal from MiniMax H3. They should not be confused with Hailuo 2.3, nor assumed to share identical inputs and parameters. [Source: fal — MiniMax H3 Max](https://fal.ai/minimax-h3-max).

**H3 Max Turbo is one of the strongest candidates for low-cost advertising tests in terms of generation price.** That does not automatically make it the best model for quality or the cheapest on the entire market. Ultimately, compare the cost of an accepted shot, not just the cost of running the model.

| Model and generation service | Base rate | 5 seconds | 10 seconds | 30 seconds of footage |
| --- | ---: | ---: | ---: | ---: |
| H3 Max Turbo, 768p — fal promotion | USD 0.02/s | PLN 0.38 | PLN 0.76 | PLN 2.28 |
| H3 Max Turbo, 1080p — fal promotion | USD 0.04/s | PLN 0.76 | PLN 1.52 | PLN 4.56 |
| H3 Max Turbo, 768p — Magnific Premium+ annual plan | 200 credits / 5 s | PLN 0.47 | PLN 0.94 | PLN 2.83 |

Data as of **18 September 2026**. The fal promotion ends on 30 September 2026; the published pricing states that rates will then double. The Magnific figures assume that the entire annual allowance of 600,000 credits is used.

The description can combine visuals, camera movement and sound, provided the selected mode supports them. For a sequence of actions, specify what happens first, what follows and how the shot ends.

This example assumes a starting image showing a package and a cup on a countertop:

```text
One continuous shot based on the starting image. The camera is
initially stationary, and the coffee package stays fixed throughout.

A hand enters from the right and gently rotates the cup on the table
until its handle points toward the camera. The hand releases the cup
and leaves the frame.

The camera then slowly moves forward toward the package, keeping
the front label in view. It gradually slows to a stop and holds
a close-up for the end of the shot.

Audio: quiet indoor ambience and a soft scraping sound as the base
of the cup moves across the table.
```

The cup remains on the countertop as it turns, the hand finishes its action before the push-in begins, and the camera stops at the end. There is no need to describe every colour and object in the photo again.

English is the working language used in this guide. This is not a conclusion from a test comparing the effectiveness of different languages in H3 Max Turbo.

### Check automatic prompt expansion

The fal API schema describes the `prompt_expansion_mode` parameter, including `balanced`, `quality` and `disabled`. The last option turns expansion off. The response may also include the rewritten text. [Source: fal — H3 Max Turbo, text-to-video](https://fal.ai/models/minimax/h3-max-turbo/text-to-video/api).

For a broad idea, expansion may be part of your chosen workflow. For a precise script, it is worth comparing the result with rewriting disabled, if the platform lets you change that setting. Do not judge only the text you entered: also check whether it was expanded before generation.

Do not carry Hailuo’s bracketed commands over to H3 Max without confirming support. The basic H3 Max Turbo schema described here also does not list a separate `negative_prompt` field. [Source: fal — H3 Max Turbo, text-to-video](https://fal.ai/models/minimax/h3-max-turbo/text-to-video/api).

#### ChatGPT starter prompt — H3 Max Turbo

```text
Help me write prompts for MiniMax H3 Max Turbo. Reply in English and prepare finished prompts in English. Do not generate videos.

First check the available documentation:
https://fal.ai/minimax-h3-max
https://fal.ai/models/minimax/h3-max-turbo/text-to-video/api
https://platform.minimax.io/docs/guides/video-generation

Distinguish between H3, H3 Max, H3 Max Turbo and Hailuo. Summarise the rules for the chosen model. Once I name the platform, check its integration’s features. Flag information you could not confirm.

If I have not described the shot, ask for its description, platform, duration, aspect ratio, materials and desired sound. Ask only for missing information. When the description is complete, prepare the prompt immediately.

Describe the action and camera movement in natural sentences. For several actions, define their sequence and ending. When animating a first frame, focus on changes instead of repeating the entire photo description.

Assign sounds to specific events and dialogue to specific characters. Preserve the wording and language of speech and check speech support. Do not add actions, cuts or effects I have not requested.

Check prompt_expansion_mode and discuss available settings outside the prompt. Do not automatically carry over camera commands from Hailuo. Include a negative prompt or additional reference inputs only when the specific mode supports them.

Return one prompt in a code block and the settings in English. Separate documented recommendations from your own suggestions. Use the configuration already agreed upon for subsequent shots.
```

---

## 19. Wan 3.0

### Use the instructions for the correct version

The fal materials linked below describe generation from text, images and references. They also mention automatic prompt expansion and additional planning before generation. [Source: fal — Wan 3](https://fal.ai/wan-3).

This does not mean you can copy settings and extensive negative lists from local Wan 2.x configurations without checking. Use the documentation for the specific Wan 3.0 mode.

English is the working language used here. Adopting this convention does not mean the model always performs better in English than in Chinese or another language.

### Describe the scene and how it unfolds

In text-to-video, start with the location and the arrangement of objects. Then specify the movement and the final framing.

```text
A close-up product shot in a small, sunlit kitchen. A coffee package
stands beside a ceramic cup on the counter.

The package is slightly to the right of center in the opening frame.
The camera tracks slowly to the right, parallel to the counter,
with a fixed viewing direction. The move ends with the package near
the center of the frame and its front label fully visible.

Steam rises gently from the cup while the package remains stationary.

Audio: quiet kitchen ambience and the faint sound of a kettle
in the background.
```

The product’s starting position and the direction of travel are linked. The description does not simultaneously demand a fixed viewing direction and a product that stays centred throughout the movement.

Give each reference material a clear role. The Wan 3.0 schema describes referring to images and videos by their position in the input list. Use the notation for this integration, rather than automatically borrowing Seedance tags. [Source: fal — Wan 3.0, reference-to-video](https://fal.ai/models/alibaba/wan-3.0/reference-to-video/api).

```text
Use the first reference image for the coffee package's shape,
colors, and label design.

Use the first reference video only for the camera movement.
Place the package in a new, bright kitchen setting rather than
recreating the background from the reference video.
```

### Set parameters outside the prompt

In the integration described here, `enable_thinking` and `enable_prompt_expansion` are separate settings. Writing “think carefully” in the scene description is not equivalent to changing their values.

For a planned advertising shot, specify the duration and aspect ratio rather than leaving them to automatic settings. This makes successive tests easier to compare.

The basic text-to-video schema described here does not list a separate `negative_prompt` field. Put requirements for the product’s appearance in the main description, without adding an unsupported parameter. [Source: fal — Wan 3.0, text-to-video](https://fal.ai/models/alibaba/wan-3.0/text-to-video/api).

#### ChatGPT starter prompt — Wan 3.0

```text
Help me prepare prompts for Wan 3.0. Give explanations in English and finished prompts in English. Do not generate videos.

First check the available documentation:
https://fal.ai/wan-3
https://fal.ai/models/alibaba/wan-3.0/text-to-video/api
https://fal.ai/models/alibaba/wan-3.0/reference-to-video/api

Briefly summarise the rules for version 3.0. Once you know the platform, check the features of the chosen mode. Do not automatically carry over Wan 2.x parameters. Flag information you could not confirm.

If I have not described the scene, ask for its description, platform, duration, aspect ratio, materials and desired sound. Ask only for missing information. When the description is complete, prepare the prompt immediately.

In text-to-video, describe the scene’s appearance and movement. In image-to-video, focus on animation. Give references specific roles and use notation supported by the integration. Do not copy tags from other models.

Separate camera movement from object movement. Check that the starting position, movement direction and final framing are consistent. Describe successive actions in a logical order.

Describe sound separately and keep it consistent with the image. Assign dialogue to a character, preserve its wording and language, and check speech support.

Discuss enable_thinking and enable_prompt_expansion in English outside the prompt. Do not replace them with commands inside the scene description. Add a negative prompt only after confirming that this field is available.

Return the finished prompt in a code block and the settings in English. Use the configuration already agreed upon for subsequent shots.
```

---

## How to describe a shot in your next message to ChatGPT

You do not need to know the technical names for camera movements. Describe what should happen in everyday language. What matters is separating camera movement from the behaviour of objects and explaining the attachment’s role.

Example message after choosing a model:

```text
Platform: Higgsfield.
Model: the one we chose earlier in this conversation.
Duration: 10 seconds. If this mode does not support that duration, explain the available options.
Aspect ratio: vertical 9:16.

The attached photo should be the video’s first frame, not just a reference for the product’s appearance.

I want a calm advertising shot. The camera should slowly move to the right in a straight line while also moving slightly closer to the product. I do not mean an orbit around it or a zoom.

The product should stay in the same place throughout and not rotate. The front of the package should remain visible. At the end I need some empty space above the product, because I will add text later in Premiere.

No dialogue. I will add sound myself during editing. Preserve the surroundings and lighting from the photo. Do not add new people or objects.

The priorities are an unchanged product appearance and smooth camera movement. If the proposed framing conflicts with the photo, point out the problem instead of changing the concept yourself.
```

The *aspect ratio* is the relationship between image width and height. 9:16 denotes a vertical frame; it does not specify the number of pixels. Choose the resolution separately from the settings available on your platform.

In your message to ChatGPT, you can use prohibitions such as “the product must not rotate”. This is a task description, not a finished generator prompt. The chat should turn your intention into wording appropriate for the model, such as “The package remains stationary” in Runway.

## How to improve a prompt after an unsuccessful generation

Instead of simply writing “make it better”, identify the difference between your intention and the result. The most useful feedback names the object, the type of error and the point at which it appears.

Vague:

> It looks artificial.

Specific:

> The camera moved in an arc even though it was supposed to move sideways in a straight line. From the third second, the package becomes wider. The steam movement looks good: do not change that instruction.

Correct one category of problem at a time. This makes it easier to judge whether the prompt change actually helped.

| Problem | What to check or change |
| --- | --- |
| The product rotates instead of the camera moving | Describe the stationary product and the intended camera movement separately. |
| Unwanted cuts appear | Check Multi-Shot mode and wording that suggests moving to another shot. |
| The product’s appearance changes | Check the input material’s quality, the role of references and the range of motion. Simply making the prompt longer will not address the cause. |
| A character performs actions in the wrong order | Put the actions in sequence and specify the state of the scene after each one. |
| The wrong person speaks the dialogue | Use consistent character descriptions and assign each line separately. |
| The result departs from a precise description | Check whether the platform automatically expands or rewrites the prompt, and account for this when comparing tests. |

Gradually refining instructions follows the approach described by Runway. In the MiniMax and Wan integrations discussed here, settings that transform the prompt before generation may also matter. [Sources: Runway — Gen-4 guide](https://help.runwayml.com/hc/en-us/articles/39789879462419-Gen-4-Video-Prompting-Guide), [MiniMax — image-to-video](https://platform.minimax.io/docs/api-reference/video-generation-i2v), [fal — Wan 3.0, text-to-video](https://fal.ai/models/alibaba/wan-3.0/text-to-video/api).

### What does the documentation leave unresolved?

This guide cites the most detailed resources for Runway, Veo, Kling and Seedance 2.5. It does not present an equally extensive, dedicated prompting manual for Seedance 2.0 Mini. The name “Seedance 2 Pro” needs to be identified on the specific platform. Nor does it cite comparable studies that would establish one best prompt language for all the models discussed.

Rather than memorising a universal recipe such as “always 150 words, always English, always a long negative list”, record the configuration of successful tests: model, platform, mode, materials, prompt and settings. For your next shot, you will have a reference point instead of a broad rule that may not suit the tool.

---

*Editorial note: the language, syntax, consistency of the instructions and logic of the example scenes have been revised. Information about model features and the references has been retained from the supplied article. This revision does not include a fresh review of the documentation or tests of the generators.*
