第 4 章(共 9 章)
Video Generation: Text-to-Video and Image-to-Video
Video generation happens on the same canvas as everything else: you create a Video Node, give it an input, and pick a model to run. What you feed the node decides the mode — a written description for text-to-video, or an existing image for image-to-video.
Choosing a video model
Different models excel at different requirements. Before you write a prompt, decide what the shot actually needs — movement, texture, realism, references, or camera language — and pick the model that is strongest there.
| Model | What it does well | Versions named |
|---|---|---|
| Hailuo | Excellent for action scenes. | Hailuo 2.3 |
| Kling | Strong in detail and texture; features audio-visual synchronization, and supports multimodal references. | Kling 3.0 / 2.6 (audio-visual sync), Kling O1 (multimodal references) |
| Vidu | Supports multiple references with balanced action performance. | Vidu Q3 |
| Veo | Balanced performance with a high sense of realism. | Veo 3.1 |
| Sora | Rich cinematic camera language, though it currently does not support real-person references. | — |
Text-to-video (T2V)
Text-to-video starts from a description alone, with no source image. Three steps:
- Create a Video Node.
- Input your description.
- Select a model to generate.
Image-to-video (I2V)
Image-to-video starts from artwork you already have. There are three ways to feed images into the node, each suited to a different kind of shot.
- Single image — add a Video Node from an existing Image Node to use that image as the first frame.
- Dual image (start and end frames) — best for defining a clear start and end point, controlling movement trajectories, and improving consistency.
- Multi-reference video — connect multiple keyframes to control the complex evolution of a scene.
What comes next
During the creative process you may encounter practical hurdles: awkward angles, unsatisfactory details, low efficiency in storyboarding, or blurry video. Instead of rewriting prompts and re-running models, the next part covers the "one-click" utility tools built for exactly those cases.
Frequently Asked Questions
What is the difference between text-to-video and image-to-video?
Text-to-video starts from a written description only: you create a Video Node, input your description, and select a model to generate. Image-to-video starts from an image you already have — you add a Video Node from an existing Image Node so that image becomes the first frame of the clip.
Which video model should I use for action scenes?
Hailuo is called out for action scenes, specifically Hailuo 2.3. Vidu Q3 is another option, since it supports multiple references with balanced action performance. If you care more about detail and texture, Kling is the stronger pick, and Veo 3.1 leans toward balanced performance with a high sense of realism.
How do I control the start and end of a video?
Use the dual image approach and supply both a start frame and an end frame. It is best for defining a clear start and end point, controlling movement trajectories, and improving consistency. If the scene needs to change in more complex ways, connect multiple keyframes instead to build a multi-reference video.
Can Sora use a real person as a reference?
No. Sora currently does not support real-person references. It is highlighted instead for its rich cinematic camera language. If your shot depends on references, other models support them: Kling O1 supports multimodal references, and Vidu Q3 supports multiple references.