Skip to content
Dashboard

Wan v2.7 Reference-to-Video

Wan v2.7 Reference-to-Video is the Wan 2.7 reference-to-video model from Alibaba Cloud, generating new scenes from reference images and videos with combined subject and voice referencing at 720p or 1080p. Your use is subject to Alibaba Cloud's Terms & Privacy Policies.

View API reference
Price
$0.10, Per second
Lowest available configuration
import { experimental_generateVideo as generateVideo } from 'ai';
const result = await generateVideo({
model: 'alibaba/wan-v2.7-r2v',
prompt: 'A serene mountain lake at sunrise.'
});
Read docs

Copy link to headingPlayground

Try out Wan v2.7 Reference-to-Video by Alibaba Cloud. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

alibaba logo
Images
Add up to 5 images
Videos
Add up to 3 videos
Prompt(optional)

Duration8s
2s10s
Resolution
Aspect ratio
Videos to generate
alibaba logo

Your generated video will appear here.

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Checking availability for your team
Provider
Input
Output
Capabilities
ZDR
No Training
Free AI Gateway Credit
Release Date
$0.10/sec+1 more
04/07/2026

Getting started

Generate videos with Wan v2.7 Reference-to-Video using the experimental_generateVideo function from AI SDK 6 or later. AI Gateway handles routing and polls until the video is ready.

Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the video generation quickstart.

index.ts
import { experimental_generateVideo as generateVideo } from 'ai';
import fs from 'node:fs';
import 'dotenv/config';
async function main() {
const result = await generateVideo({
model: 'alibaba/wan-v2.7-r2v',
prompt: 'Image 1 and Image 2 have a friendly conversation in a cozy cafe',
inputReferences: [
'https://example.com/cat.png',
'https://example.com/dog.png',
],
});
// Save the generated video
fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');
}
main().catch(console.error);

Top-level parameters

Pass references through inputReferences and size the output with resolution and duration. v2.7 also accepts an aspect ratio, through providerOptions.alibaba.ratio.

wan-v2.7-reference-to-video-top-level.ts
import { experimental_generateVideo as generateVideo } from 'ai';
import fs from 'node:fs';
import 'dotenv/config';
async function main() {
const result = await generateVideo({
model: 'alibaba/wan-v2.7-r2v',
prompt: 'Image 1 and Image 2 have a friendly conversation in a cozy cafe',
inputReferences: [
'https://example.com/cat.png',
'https://example.com/dog.png',
],
resolution: '1920x1080',
duration: 4,
});
// Save the generated video
fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');
}
main().catch(console.error);
ParameterTypeRequiredDescription
promptstringNoText description of the video to generate. Max 5000 characters.
durationnumberNoVideo length in seconds. 2-10 seconds.
resolutionstringNoResolution ('1280x720', '1920x1080').
aspectRatiostringNoAspect ratio ('16:9', '9:16', '1:1', '4:3', '3:4').
generateAudiobooleanNoGenerate synchronized audio with the video.
inputReferencesArray<string | { data: string; mediaType: string }>NoReference images and videos, mapped onto the request automatically. Pass videos as { data, mediaType: "video/mp4" }. See the Input limits table for counts and formats.
frameImagesArray<{ image: string; frameType: 'first_frame' }>NoOpening frame of the clip, as a single first_frame entry. Replaces prompt.image and wins when both are set. URLs only. Wan does not interpolate to an ending image, so a last_frame entry is ignored with a warning.

Input limits

InputFormatsSourcesMax countMax sizeLimits
Text————Up to 5000 characters
Imagejpeg, jpg, png, bmp, webpurl, base64520 MB≥240px · ≤8000px · aspect 1:8–8:1
Videomp4, movurl3100 MB1-30s · ≥240px · ≤4096px
Audiowav, mp3url—15 MB1-10s
Up to 5 reference inputs total across images and videos.

Provider options

Set the ratio and, when the automatic mapping is not what you want, list the media explicitly. media overrides inputReferences entirely.

wan-v2.7-reference-to-video-provider-options.ts
import { experimental_generateVideo as generateVideo } from 'ai';
import fs from 'node:fs';
import 'dotenv/config';
async function main() {
const result = await generateVideo({
model: 'alibaba/wan-v2.7-r2v',
prompt: 'Image 1 walks past Video 1 as the camera pans',
resolution: '1920x1080',
duration: 4,
providerOptions: {
alibaba: {
media: [
{ type: 'reference_image', url: 'https://example.com/cat.png' },
{ type: 'reference_video', url: 'https://example.com/street.mp4' },
],
ratio: '16:9',
negativePrompt: 'blurry, low quality',
},
},
});
// Save the generated video
fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');
}
main().catch(console.error);

Pass Wan-specific options under providerOptions.alibaba in your generateVideo call. References come from the top-level inputReferences unless media overrides them.

ParameterTypeRequiredDescription
mediaArray<{ type: 'reference_image' | 'reference_video' | 'first_frame'; url: string; referenceVoice?: string }>NoExplicit media list, overriding the mapping from the top-level inputReferences and frameImages. Images and videos are numbered separately in array order, so a prompt refers to them as Image 1, Video 1, and so on. referenceVoice attaches a voice reference to one item.
ratio'16:9' | '9:16' | '1:1' | '4:3' | '3:4'NoAspect ratio of the generated video. v2.7 text-to-video and reference-to-video only.
negativePromptstringNoWhat to avoid in the video. Max 500 characters.
promptExtendbooleanNoEnhance prompt for better quality. Defaults to true.
watermarkbooleanNoAdd watermark to the video. Defaults to false.
pollIntervalMsnumberNoHow often to check task status. Defaults to 5000.
pollTimeoutMsnumberNoMaximum wait time. Defaults to 600000 (10 minutes).

Reference-to-video vs image-to-video

Reference-to-video uses the top-level inputReferences to show the model what your characters look like, then generates a brand-new scene from your prompt. The reference media never becomes the video content; reference each one in the prompt with character1, character2, and so on (first entry maps to character1).

Image-to-video instead animates the actual image you pass in frameImages or prompt.image. The image you provide becomes the video content, and you add motion to that exact scene.

Copy link to headingMore models by Alibaba Cloud

Model
Context
Latency
Throughput
Input
Output
Cache
Search
Capabilities
Providers
ZDR
No Training
Free AI Gateway Credit
Release Date
1M2.3 s89 tps
$0.15/M
$0.47/M
Read$0.02/M
—
+2
alibaba logo
09/17/2026
991K2.8 s77 tps
$0.15/M
$0.47/M
Read$0.02/M
Write$0.20/M
—
+4
alibaba logo
08/26/2026
1M0.4 s1628 tps
$0.02/M
$0.40/M
Read$0.01/M
Write$0.63/M
—
+2
alibaba logo
cerebras logo
deepinfra logo
+4
08/14/2026
1M1.5 s89 tps
$2/M
$6/M
Read$0.25/M
Write$2.50/M
—
+3
alibaba logo
fireworks logo
08/02/2026
991K2.0 s103 tps
$0.03/M+2 more
$0.13/M+2 more
Read$0.006/M
Write$0.04/M
—
+4
alibaba logo
07/28/2026
1M2.5 s58 tps
$0.40/M+1 more
$1.60/M+1 more
Read$0.08/M
Write$0.50/M
—
+3
alibaba logo
06/02/2026

Copy link to headingAbout Wan v2.7 Reference-to-Video

Wan v2.7 Reference-to-Video is the reference-to-video member of Alibaba Cloud's Wan 2.7 release. You supply reference images, reference videos, or both, and Wan v2.7 Reference-to-Video places the referenced subjects into an entirely new scene described by an optional text prompt. Up to three reference videos can be attached to a single generation.

Combined subject and voice referencing is the notable addition in the 2.7 generation. Wan v2.7 Reference-to-Video lets you bind a subject's visual identity and vocal identity together from your reference material, keeping both consistent in the generated output. The reference pipeline also supports multi-subject work, so several referenced identities can interact in one scene, and multi-shot workflows keep those identities stable across scene cuts.

Output runs 2 to 10 seconds, with a default of 5 seconds, at 720p or 1080p. Five aspect ratio options cover landscape, portrait, square, and intermediate formats: 16:9, 9:16, 1:1, 4:3, and 3:4. The 2.7 generation also brings smoother, more coherent motion than the 2.6 line, which helps generated subjects hold up under close comparison to their references.

Copy link to headingWhat To Consider When Choosing a Provider

  • Configuration: Reference quality drives output quality. Clear, well-lit reference material gives the extraction pipeline more identity signal to work with, so curate your references before scaling up a render queue.
  • Configuration: Pricing is per second of generated video and varies by resolution, so a 1080p clip costs more than the same clip at 720p. Run a few test prompts in the AI Gateway playground to calibrate cost and generation time before full integration.
  • Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

Copy link to headingWhen to Use Wan v2.7 Reference-to-Video

Best for

  • Character-Driven Video Series: Subjects that must look and sound consistent across many generated scenes
  • Brand and Spokesperson Content: Referenced people or mascots placed into new settings without a reshoot
  • Multi-Subject Compositions: Scenes where several referenced identities interact in a single generated clip
  • Multi-Shot Identity Stability: Sequences that keep referenced subjects consistent across automatic scene cuts

Consider alternatives when

  • Text-Only Generation: Wan-v2.7-t2v handles pure text-to-video when no reference subject is involved
  • Clips Beyond 10 Seconds: Wan-v2.7-t2v extends output duration to 15 seconds
  • Previous-Generation Pipelines: Wan-v2.6-r2v remains available for workflows already tuned to the earlier release

Wan v2.7 Reference-to-Video is the model to reach for when generated video must stay faithful to a real subject. Combined subject and voice referencing, multi-subject support, and smoother motion than the 2.6 line make Wan v2.7 Reference-to-Video a strong default for identity-sensitive video work on AI Gateway.

Copy link to headingFrequently Asked Questions

  • What reference inputs does Wan v2.7 Reference-to-Video accept?

    Wan v2.7 Reference-to-Video accepts reference images, reference videos, or both together, plus an optional text prompt describing the scene to generate. Up to three reference videos can be attached to a single generation.

  • What durations and resolutions does Wan v2.7 Reference-to-Video support?

    Clips run 2 to 10 seconds, with a default of 5 seconds, at 720p or 1080p. Five aspect ratios are available: 16:9, 9:16, 1:1, 4:3, and 3:4.

  • How does Wan v2.7 Reference-to-Video differ from the previous Wan reference-to-video model?

    Wan v2.7 Reference-to-Video adds combined subject and voice referencing, so appearance and vocal identity bind together from your reference material. Motion is also smoother and more coherent than in the 2.6 generation.

  • Can Wan v2.7 Reference-to-Video keep a subject's voice consistent across generations?

    Yes. Combined subject and voice referencing carries a referenced subject's vocal identity into the generated output alongside their visual appearance, keeping both consistent in new scenes.

  • How do I use Wan v2.7 Reference-to-Video through AI Gateway?

    Call Wan v2.7 Reference-to-Video with generateVideo from the AI SDK, passing your reference images or videos and an optional prompt. You can also try prompts first in the playground on this page.

  • Does AI Gateway offer Zero Data Retention for Wan v2.7 Reference-to-Video?

    Yes, Zero Data Retention is available for this model. Zero Data Retention is offered on a per-provider basis. See https://vercel.com/docs/ai-gateway/capabilities/zdr for details.