Vidu

Vidu

Vidu is a reference-driven video generation family with five specialized variants. It excels at maintaining character consistency across clips using multiple image references and video reference inputs.

The Vidu family is built around reference-guided generation. Q1 generates video from multiple reference images, combining visual cues into a coherent clip. Q3 is the general-purpose model supporting text-to-video and image-to-video across five aspect ratios at 1080p. Q2 offers subject reference with 1-7 input images for strong identity locking. Q2 Image Reference takes multi-image reference sets for broader visual guidance. Q2 Pro Video Reference operates in V2V mode, using an existing video plus reference images to produce a new clip that merges both inputs. This layered approach makes Vidu the strongest choice when character consistency across a series of clips is the top priority.

Illustrative sample of a Vidu still showing a reference-locked character holding consistent identity across a generated scene on the Astorie canvas
Illustrative sample — representative output, not a verbatim model render

Vidu Variants

VariantDescription
Vidu Q1Reference-to-video from multiple images, combining visual cues.
Vidu Q3General T2V and I2V across 5 aspect ratios at 1080p.
Vidu Q2 Subject RefSubject reference with 1-7 input images for strong identity locking.
Vidu Q2 Image RefMulti-image reference sets for broader visual guidance.
Vidu Q2 Pro Video RefV2V with reference images, merges existing video with new visual direction.

Capabilities

Text-to-Video
Image-to-Video
Video-to-Video
Reference Images
End Frame
Storyboard
Audio-Driven

Supported Aspect Ratios

16:99:164:33:41:1

Best For

  • Multi-clip series with consistent character identity
  • Brand campaigns requiring the same spokesperson across videos
  • Reference-heavy workflows with multiple visual inputs
  • Video remixing that blends source footage with new style references
  • Product showcase series maintaining visual continuity

Strengths

  • Strongest multi-reference input system with up to 7 subject images
  • Five specialized variants for different reference-driven workflows
  • Q3 provides a solid general-purpose baseline at 1080p
  • V2V Pro variant merges video and image references in a single pass
  • Five aspect ratio options on the general model

Limitations

  • Learning curve to choose the right variant for each use case
  • Q1 requires multiple images, which increases setup time
  • General quality on Q3 may trail top-tier single-purpose models

Tips & Best Practices

For Q2 Subject Ref, use 3-5 images showing the character from different angles for best identity lock.
Start with Q3 for general clips, then switch to Q1 or Q2 when you need reference consistency.
For V2V Pro, pair a simple source video with strong style references for the cleanest merge.
Keep reference images consistent in lighting and color palette across your image set.

Use Vidu on Astorie

Connect Vidu with other AI models on Astorie's infinite canvas. No GPU required — start free.

Get Started Free

Related Features

How-To Guides

Related Reading

Related Video Models

Back to All Video Models