Wan v2.6 Reference-to-Video
Wan v2.6 Reference-to-Video is Alibaba Cloud's quality-first reference-to-video model, extracting identity from short clips and rendering new scenes with high-fidelity appearance, voice, and motion preservation at up to 1080p. Your use is subject to Alibaba Cloud's Terms & Privacy Policies.
View API reference- Price
- $0.10, Per secondLowest available configuration
import { experimental_generateVideo as generateVideo } from 'ai';
const result = await generateVideo({ model: 'alibaba/wan-v2.6-r2v', prompt: 'A serene mountain lake at sunrise.'});Copy link to headingPlayground
Try out Wan v2.6 Reference-to-Video by Alibaba Cloud. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Your generated video will appear here.
Copy link to headingProviders
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|
Getting started
Generate videos with Wan v2.6 Reference-to-Video using the experimental_generateVideo function from AI SDK 6 or later. AI Gateway handles routing and polls until the video is ready.
Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the video generation quickstart.
import { experimental_generateVideo as generateVideo } from 'ai';import fs from 'node:fs';import 'dotenv/config';
async function main() { const result = await generateVideo({ model: 'alibaba/wan-v2.6-r2v', prompt: 'character1 and character2 have a friendly conversation in a cozy cafe', inputReferences: [ 'https://example.com/cat.png', 'https://example.com/dog.png', ], });
// Save the generated video fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');}
main().catch(console.error);Top-level parameters
Pass references through inputReferences and control the output with the top-level resolution and duration parameters. Wan uses resolution, not aspectRatio.
import { experimental_generateVideo as generateVideo } from 'ai';import fs from 'node:fs';import 'dotenv/config';
async function main() { const result = await generateVideo({ model: 'alibaba/wan-v2.6-r2v', prompt: 'character1 and character2 have a friendly conversation in a cozy cafe', inputReferences: [ 'https://example.com/cat.png', 'https://example.com/dog.png', ], resolution: '1920x1080', duration: 4, });
// Save the generated video fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');}
main().catch(console.error);| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | No | Text description of the video to generate. Max 1500 characters. |
duration | number | No | Video length in seconds. 2-10 seconds. |
resolution | string | No | Resolution ('1280x720', '1920x1080'). |
aspectRatio | string | No | Aspect ratio ('16:9', '9:16', '1:1', '4:3', '3:4'). |
generateAudio | boolean | No | Generate synchronized audio with the video. |
inputReferences | Array<string> | No | Reference images and videos, mapped in order onto character1, character2, and so on. URLs only — a non-URL reference is skipped with a warning, so upload local files to Vercel Blob first. See the Input limits table for supported counts and formats. |
Input limits
| Input | Formats | Sources | Max count | Max size | Limits |
|---|---|---|---|---|---|
| Text | — | — | — | — | Up to 1500 characters |
| Image | jpeg, jpg, png, bmp, webp | url | 5 | 20 MB | ≥240px · ≤8000px |
| Video | mp4, mov | url | 3 | 100 MB | 1-30s |
Provider options
Load every Wan reference-to-video option under providerOptions.alibaba. References themselves go in the top-level inputReferences, where the first entry maps to character1, the second to character2, and so on.
import { experimental_generateVideo as generateVideo } from 'ai';import fs from 'node:fs';import 'dotenv/config';
async function main() { const result = await generateVideo({ model: 'alibaba/wan-v2.6-r2v', prompt: 'character1 and character2 have a friendly conversation in a cozy cafe', inputReferences: [ 'https://example.com/cat.png', 'https://example.com/dog.png', ], resolution: '1920x1080', duration: 4, generateAudio: true, providerOptions: { alibaba: { negativePrompt: 'blurry, low quality', shotType: 'single', watermark: false, pollIntervalMs: 5000, pollTimeoutMs: 600000, }, }, });
// Save the generated video fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');}
main().catch(console.error);Pass Wan-specific options under providerOptions.alibaba in your generateVideo call. References come from the top-level inputReferences; referenceUrls is the legacy fallback.
| Parameter | Type | Required | Description |
|---|---|---|---|
referenceUrls | string[] | No | Array of URLs to reference images or videos. The first URL maps to character1, the second to character2, and so on. Legacy alternative to the top-level inputReferences, used only when inputReferences is omitted. See the Input limits table for supported counts and formats. |
negativePrompt | string | No | What to avoid in the video. Max 500 characters. |
shotType | 'single' | 'multi' | No | 'single' for a continuous shot. 'multi' for multiple camera angles. |
audio | boolean | No | Generate audio with the video. Provider-side alias for the top-level generateAudio, which wins when both are set. v2.6 only — v2.7 always generates audio and ignores it with a warning. |
watermark | boolean | No | Add watermark to the video. Defaults to false. |
pollIntervalMs | number | No | How often to check task status. Defaults to 5000. |
pollTimeoutMs | number | No | Maximum wait time. Defaults to 600000 (10 minutes). |
Reference-to-video vs image-to-video
Reference-to-video uses the top-level inputReferences to show the model what your characters look like, then generates a brand-new scene from your prompt. The reference media never becomes the video content; reference each one in the prompt with character1, character2, and so on (first entry maps to character1).
Image-to-video instead animates the actual image you pass in frameImages or prompt.image. The image you provide becomes the video content, and you add motion to that exact scene.
Reference media with Vercel Blob
References must be URLs. If you have local files, upload them to Vercel Blob first, then pass the returned URLs in inputReferences.
import { experimental_generateVideo as generateVideo } from 'ai';import fs from 'node:fs';import 'dotenv/config';import { put } from '@vercel/blob';
async function main() { const catImage = fs.readFileSync('./cat.png'); const { url: catUrl } = await put('cat.png', catImage, { access: 'public' });
const dogImage = fs.readFileSync('./dog.png'); const { url: dogUrl } = await put('dog.png', dogImage, { access: 'public' });
const result = await generateVideo({ model: 'alibaba/wan-v2.6-r2v', prompt: 'character1 and character2 play together in a sunny garden', inputReferences: [catUrl, dogUrl], resolution: '1280x720', duration: 4, providerOptions: { alibaba: { shotType: 'single', }, }, });
// Save the generated video fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');}
main().catch(console.error);Copy link to headingAbout Wan v2.6 Reference-to-Video
Alibaba Cloud positioned Wan v2.6 Reference-to-Video as China's first reference-to-video generation model. The core idea is simple: feed the model a short video clip of a person, and it extracts enough about their face, body, clothing, voice, and movement style to convincingly place them into an entirely new scene described by text.
What sets the standard R2V apart within the Wan lineup is its emphasis on reconstruction depth. The model devotes additional inference time to faithfully reproducing fine-grained identity signals, the specific way light falls across facial features, subtle mannerisms in how a subject moves, the particular resonance of a voice. For final-delivery video where a client or audience will scrutinize whether the generated character truly matches the reference, this level of fidelity matters.
The reference pipeline accepts multiple characters from reference images or videos. On AI Gateway, pass reference URLs in order and name them character1, character2, and so on in the prompt (see the reference-to-video docs). You can mix images and videos within provider limits (up to five references total). Each video reference can be 2 to 30 seconds. Output duration is 2 to 10 seconds at 720p or 1080p, with five aspect ratio options covering landscape, portrait, square, and intermediate formats.
Copy link to headingWhat To Consider When Choosing a Provider
- Configuration: Because R2V allocates more compute to identity reconstruction than the Flash variant, generation times are longer. Budget wall-clock time accordingly when planning render queues for final-delivery assets.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
- Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.
Copy link to headingWhen to Use Wan v2.6 Reference-to-Video
Best for
- Final-delivery marketing assets: Generated characters that must pass close visual inspection against the reference subject
- Brand campaign consistency: Voice and appearance preservation across a series of scenes rendered over time
- Virtual production pipelines: Directors who need confident identity transfer before approving a generated sequence
- Multi-subject compositions: Scenes where several reference identities appear together in one generated clip
Consider alternatives when
- Speed over peak fidelity: Wan-v2.6-r2v-flash provides faster turnaround during early creative exploration
- Photograph source material: Wan-v2.6-i2v animates from still images when the source is a photograph rather than a video clip
- Text-only video generation: Wan-v2.6-t2v is the appropriate model when no reference subject is involved
Copy link to headingConclusion
Wan v2.6 Reference-to-Video is purpose-built for cases where identity fidelity matters more than generation speed. It trades wall-clock time for meticulous reconstruction of appearance, voice, and motion from reference material, making it the right choice when the output needs to survive close comparison to the original subject.
Copy link to headingFrequently Asked Questions
Why does R2V take longer to generate than the Flash variant?
The standard R2V model spends additional compute on identity reconstruction, fine-grained facial detail, voice matching, and movement pattern extraction all receive more processing time. This is a deliberate design choice favoring output quality over speed.
What happens if my reference clip is very short, like 2 seconds?
The model can work with clips as short as 2 seconds, but shorter references provide less identity data. For the highest fidelity, longer clips within the 2-30 second range give the extraction pipeline more material to work with, particularly for voice and movement characteristics.
How should I tag multiple subjects in a single prompt?
Use
character1,character2, and so on in the prompt text, in the same order as your reference URLs. Each name maps to the corresponding reference image or video.Does R2V preserve clothing and accessories from the reference?
Yes. The identity extraction covers visual appearance broadly, including clothing, accessories, hairstyle, and body proportions, not just facial features. The generated output aims to maintain the full visual signature of the reference subject.
What aspect ratios work best for social media delivery?
For vertical social content, 9:16 is the standard choice. For feed posts, 1:1 provides square framing. The model also supports 16:9, 4:3, and 3:4 for other distribution contexts.
Can R2V be used for product or object identity transfer, or only people?
The reference extraction is designed broadly enough to capture objects and animals in addition to people, though the pipeline is most heavily optimized for human subjects where facial and vocal identity are the primary signals.