Skip to content
Dashboard

Wan v2.6 Image-to-Video

Wan v2.6 Image-to-Video is Alibaba Cloud's image-to-video model that animates still images into high-fidelity video clips up to 1080p and 15 seconds, with optional audio and precise motion control from text guidance. Your use is subject to Alibaba Cloud's Terms & Privacy Policies.

View API reference
Price
$0.10, Per second
Lowest available configuration
import { experimental_generateVideo as generateVideo } from 'ai';
const result = await generateVideo({
model: 'alibaba/wan-v2.6-i2v',
prompt: 'A serene mountain lake at sunrise.'
});
Read docs

Copy link to headingPlayground

Try out Wan v2.6 Image-to-Video by Alibaba Cloud. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

alibaba logo
Start frame (required)
Prompt(optional)

Duration8s
2s15s
Resolution
Videos to generate
alibaba logo

Your generated video will appear here.

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Checking availability for your team
Provider
Input
Output
Capabilities
ZDR
No Training
Free AI Gateway Credit
Release Date
$0.10/sec+1 more
12/16/2025

Getting started

Generate videos with Wan v2.6 Image-to-Video using the experimental_generateVideo function from AI SDK 6 or later. AI Gateway handles routing and polls until the video is ready.

Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the video generation quickstart.

index.ts
import { experimental_generateVideo as generateVideo } from 'ai';
import fs from 'node:fs';
import 'dotenv/config';
async function main() {
const result = await generateVideo({
model: 'alibaba/wan-v2.6-i2v',
prompt: {
image: 'https://example.com/cat.png',
text: 'The cat waves hello and smiles',
},
});
// Save the generated video
fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');
}
main().catch(console.error);

Top-level parameters

Control the output with the top-level resolution and duration parameters. Wan uses resolution, not aspectRatio.

wan-image-to-video-top-level.ts
import { experimental_generateVideo as generateVideo } from 'ai';
import fs from 'node:fs';
import 'dotenv/config';
async function main() {
const result = await generateVideo({
model: 'alibaba/wan-v2.6-i2v',
prompt: {
image: 'https://example.com/cat.png',
text: 'The cat waves hello and smiles',
},
resolution: '1280x720',
duration: 5,
});
// Save the generated video
fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');
}
main().catch(console.error);
ParameterTypeRequiredDescription
prompt.imagestringYesURL of the image to animate.
prompt.textstringNoDescription of the motion or animation. Max 1500 characters.
durationnumberNoVideo length in seconds. 2-15 seconds.
resolutionstringNoResolution ('1280x720', '1920x1080').
aspectRatiostringNoAspect ratio ('16:9', '9:16', '1:1', '4:3', '3:4').
generateAudiobooleanNoGenerate synchronized audio with the video.
frameImagesArray<{ image: string; frameType: 'first_frame' }>NoOpening frame of the clip, as a single first_frame entry. Replaces prompt.image and wins when both are set. URLs only. Wan does not interpolate to an ending image, so a last_frame entry is ignored with a warning.

Input limits

InputFormatsSourcesMax countMax sizeLimits
Text————Up to 1500 characters
Imagejpeg, jpg, png, bmp, webpurl120 MB≥240px · ≤8000px
Audiowav, mp3url—15 MB3-30s

Provider options

Load the Wan options under providerOptions.alibaba. audioUrl (an external audio track for lip-sync) is a separate workflow and is documented in the table below.

wan-image-to-video-provider-options.ts
import { experimental_generateVideo as generateVideo } from 'ai';
import fs from 'node:fs';
import 'dotenv/config';
async function main() {
const result = await generateVideo({
model: 'alibaba/wan-v2.6-i2v',
prompt: {
image: 'https://example.com/cat.png',
text: 'The cat waves hello and smiles',
},
duration: 5,
generateAudio: true,
providerOptions: {
alibaba: {
negativePrompt: 'blurry, low quality',
watermark: false,
pollIntervalMs: 5000,
pollTimeoutMs: 600000,
},
},
});
// Save the generated video
fs.writeFileSync('output.mp4', result.videos[0].uint8Array);
console.log('Video saved to output.mp4');
}
main().catch(console.error);

Pass Wan-specific options under providerOptions.alibaba in your generateVideo call.

ParameterTypeRequiredDescription
negativePromptstringNoWhat to avoid in the video. Max 500 characters.
audioUrlstringNoURL to audio file for audio-video sync — see the Input limits table for supported formats, duration, and size.
audiobooleanNoGenerate audio with the video. Provider-side alias for the top-level generateAudio, which wins when both are set. v2.6 only — v2.7 always generates audio and ignores it with a warning.
watermarkbooleanNoAdd watermark to the video. Defaults to false.
pollIntervalMsnumberNoHow often to check task status. Defaults to 5000.
pollTimeoutMsnumberNoMaximum wait time. Defaults to 600000 (10 minutes).

Reference-to-video vs image-to-video

Reference-to-video uses the top-level inputReferences to show the model what your characters look like, then generates a brand-new scene from your prompt. The reference media never becomes the video content; reference each one in the prompt with character1, character2, and so on (first entry maps to character1).

Image-to-video instead animates the actual image you pass in frameImages or prompt.image. The image you provide becomes the video content, and you add motion to that exact scene.

Copy link to headingMore models by Alibaba Cloud

Model
Context
Latency
Throughput
Input
Output
Cache
Search
Capabilities
Providers
ZDR
No Training
Free AI Gateway Credit
Release Date
1M2.3 s89 tps
$0.15/M
$0.47/M
Read$0.02/M
—
+2
alibaba logo
09/17/2026
991K2.8 s77 tps
$0.15/M
$0.47/M
Read$0.02/M
Write$0.20/M
—
+4
alibaba logo
08/26/2026
1M0.4 s1628 tps
$0.02/M
$0.40/M
Read$0.01/M
Write$0.63/M
—
+2
alibaba logo
cerebras logo
deepinfra logo
+4
08/14/2026
1M1.5 s89 tps
$2/M
$6/M
Read$0.25/M
Write$2.50/M
—
+3
alibaba logo
fireworks logo
08/02/2026
991K2.0 s103 tps
$0.03/M+2 more
$0.13/M+2 more
Read$0.006/M
Write$0.04/M
—
+4
alibaba logo
07/28/2026
1M2.5 s58 tps
$0.40/M+1 more
$1.60/M+1 more
Read$0.08/M
Write$0.50/M
—
+3
alibaba logo
06/02/2026

Copy link to headingAbout Wan v2.6 Image-to-Video

Wan v2.6 Image-to-Video belongs to Alibaba Cloud's 2.6-generation video lineup, the quality-focused image-animation model in the series. The model accepts a source image, between 360px and 2000px on either dimension, up to 100MB, paired with a descriptive text prompt that guides the motion, timing, and scene direction. From those inputs it produces video at 480p, 720p, or 1080p in five aspect ratios.

The 2.6 generation introduced substantial upgrades over its predecessors, including improved temporal consistency (less flicker between frames), sharper fine detail retention from the source image, and better instruction-following when the text prompt specifies particular motions or environmental conditions. Audio integration is optionally available, allowing ambient sound or scene audio to accompany the animated output.

Unlike the R2V models, I2V doesn't attempt to extract a character identity for reuse across multiple shots; it animates the submitted image as a single continuous visual source. Multi-shot mode is disabled by default, which keeps the output as a single unbroken clip, appropriate for product showcases, still-life animations, and short narrative scenes anchored to one visual composition.

Copy link to headingWhat To Consider When Choosing a Provider

  • Configuration: For rapid iteration and draft-quality previews, consider wan-v2.6-i2v-flash, which trades some visual fidelity for significantly faster generation times.
  • Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

Copy link to headingWhen to Use Wan v2.6 Image-to-Video

Best for

  • Animated product advertising: Turning product photography into short video ads without a source video reference
  • Bringing static art to life: Animating concept art, illustrations, or architectural renders with guided motion
  • Single-hero-image social clips: Portrait or landscape outputs driven by a motion-direction prompt
  • High-fidelity 1080p animation: Workflows where visual quality outweighs generation speed

Consider alternatives when

  • Speed over peak quality: Use wan-v2.6-i2v-flash for faster turnaround at 720p or 1080p during iteration
  • Text-only video generation: Wan-v2.6-t2v generates video from a text description without a source image
  • Cross-scene character consistency: Wan-v2.6-r2v preserves a character's identity across multiple generated scenes

Wan v2.6 Image-to-Video transforms static images into polished animated video clips with precise text-guided motion control, supporting the full resolution range up to 1080p and durations up to 15 seconds. It is the quality-first choice within the Wan I2V lineup for teams where output fidelity outweighs generation speed.

Copy link to headingFrequently Asked Questions

  • What image formats and sizes are accepted?

    The model accepts images between 360px and 2000px on each dimension, with a maximum file size of 100MB.

  • Can I control the direction of motion with a text prompt?

    Yes. A text prompt accompanying the image guides the generated motion, camera direction, and scene atmosphere.

  • Does Wan v2.6 Image-to-Video support audio output?

    Audio integration is optionally available, ambient sounds can accompany the generated video when enabled.

  • How long can generated clips be?

    Clips can be 5, 10, or 15 seconds long, making this the longest-output option in the I2V variants.

  • What aspect ratios are supported?

    The model supports 16:9, 9:16, 1:1, 4:3, and 3:4, the widest aspect ratio selection in the Wan 2.6 I2V lineup.

  • When should I choose I2V Flash instead?

    If you are prototyping or need fast feedback on motion ideas, the flash variant offers faster generation. For final-quality deliverables at 1080p, the standard I2V model is recommended.