Skip to main content
WaveSpeed

Wan 2.1 Image-to-Video 720p

Wan 2.1 Image-to-Video 720p is an image-to-video generation model listed under WaveSpeed (alias wavespeedai/wan-2.1-i2v-720p) that animates a still input image into a 720p video clip.

From
$0.250 / second
across 1 provider

API Pricing

ProviderPrice /sec
$0.250/sec

Prices updated daily. Last check: Sep 25, 2026

Model Details

General

Creator
WaveSpeed
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No
Aliases
wavespeedai/wan-2.1-i2v-720p

Strengths & Limitations

Strengths

  • Image-to-video conditioning gives precise control over the first frame, composition, and subject appearance
  • 720p output resolution targets more visual detail than lower-resolution variants in the Wan 2.1 line
  • Accepts a text prompt alongside the input image to direct motion and camera behavior
  • Available through a hosted API alias (wavespeedai/wan-2.1-i2v-720p), avoiding local GPU setup for video inference
  • Part of the Wan 2.1 family, so text-to-video and lower-resolution i2v siblings can be swapped in for different quality/cost points
  • Suited to batch generation of short clips from existing image libraries and asset catalogs

Limitations

  • Output is capped at 720p, below the 1080p and higher resolutions offered by some other video models
  • Requires an input image, so it cannot generate a scene from a text prompt alone
  • Video generation is compute-heavy, so latency per clip is far higher than for text models and unsuitable for interactive use
  • Clip length, frame rate, and aspect ratio limits are provider-defined and can vary between hosts
  • We do not track audio generation, upscaling, or video-editing capabilities for this entry

Key Features

•Image-to-video (i2v) generation from a single input frame
•720p output resolution configuration
•Text prompt guidance for motion and camera direction
•Preservation of the input image's subject, style, and framing in the generated clip
•Temporal consistency across generated frames
•Hosted API access via the wavespeedai/wan-2.1-i2v-720p alias
•Part of the Wan 2.1 model family, with alternative resolution and text-to-video variants

About Wan 2.1 Image-to-Video 720p

Wan 2.1 Image-to-Video 720p is a video generation entry in our catalog, listed under WaveSpeed with the provider alias wavespeedai/wan-2.1-i2v-720p. It belongs to the Wan 2.1 family of video models and specifically covers the image-to-video (i2v) configuration at 720p output resolution, as distinct from text-to-video variants or lower-resolution i2v endpoints in the same family. As an image-to-video model, its input is a starting image, typically combined with a text prompt that describes the motion, camera movement, or scene evolution the user wants. The model's job is to preserve the identity, composition, and style of the supplied frame while generating temporally coherent motion across the clip. Because it is a generation model rather than a chat model, the relevant quality dimensions are output resolution (720p here), motion realism, frame-to-frame consistency, prompt adherence for the requested motion, and generation latency per clip rather than context window or token throughput. In practice this variant is used for animating product shots, illustrations, concept art, character stills, and stock photography into short motion clips for social media, ads, previsualization, and design mockups. Compared with 480p siblings in the Wan 2.1 line, the 720p configuration targets higher output detail, which generally comes with longer generation times and higher per-clip cost; compared with text-to-video models, it trades open-ended scene invention for tighter control over the exact starting frame. Detailed specifications such as maximum clip length, frame rate, and supported aspect ratios are set by the serving provider, so check the provider's documentation alongside the pricing table on this page.

Common Use Cases

Wan 2.1 Image-to-Video 720p fits workloads where the starting visual already exists and the goal is to add motion: animating product photography for e-commerce listings and paid social, turning illustrations or concept art into short looping clips, adding parallax or camera drift to character and environment stills, previsualizing shots from storyboard frames, and generating motion variants of a single hero image for A/B testing ad creative. The 720p setting is a reasonable middle ground for web and mobile delivery where full HD masters are not required. Because generation is asynchronous and compute-intensive, it is best used in batch or queue-based pipelines rather than in real-time interactive products, and teams that need higher-resolution masters or text-only scene generation should compare it against other resolution tiers and text-to-video models in the Wan 2.1 family.

Frequently Asked Questions

How much does Wan 2.1 Image-to-Video 720p cost?

Pricing for video generation models varies by provider and by billing model — some hosts charge per generated clip, others per second of output video or per compute-time unit, and rates change over time. Check the pricing table on this page for current provider rates rather than relying on a fixed figure.

What is Wan 2.1 Image-to-Video 720p best used for?

It is best used for turning an existing still image into a short 720p video clip with prompt-directed motion — animating product shots, illustrations, character stills, and stock imagery for social media, advertising, and previsualization work.

How does the 720p variant differ from lower-resolution Wan 2.1 image-to-video endpoints?

The 720p configuration produces higher-resolution output than 480p-class variants in the same family, which generally means more visible detail but longer generation times and higher cost per clip. If your delivery target is small-format social video, a lower-resolution sibling may be sufficient.

Can it generate video from text alone?

No — this entry is the image-to-video configuration, so it requires a starting image as input. The text prompt is used to guide motion and camera behavior. For prompt-only generation, use a text-to-video model instead.

How long are the clips it produces?

Maximum clip length, frame rate, and supported aspect ratios are set by the serving provider and are not fields we track for this entry. Consult the provider's API documentation for the exact limits on the endpoint you plan to use.