Percify
All modelsOpen the playground
  1. Playground
  2. /
  3. Best AI Video Models (2026)

Text-to-video, image-to-video, motion control and lip-sync

Best AI Video Models (2026)

Compare and run the best AI video generation models in one place. Generate in your browser or over a simple API — pay per run, credits refunded on failed runs.

72 models ready to run

72 models
ElevenLabs Dubbing
ElevenLabs Dubbing

ElevenLabs Dubbing automatically translates and dubs video/audio content into different languages while preserving the original speakers' voices. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

HeyGen Video Translate
HeyGen Video Translate

HeyGen Video Translate: AI video translation into 70+ languages and 175+ dialects with no voice actors or dubbing. Fast, accurate, easy to use at $0.0375/sec. Ready-to-use REST API, no coldstarts, affordable pricing.

MoCha Character Swap
MoCha Character Swap

Swap the character in a video for someone else. Give it the clip and a reference image, and it replaces the performer while keeping the original motion, timing and audio.

P-Video Replace
P-Video Replace

Replace or insert a subject into an existing video. Give it the clip and up to three reference images of who should appear, and it preserves the original motion, timing and audio.

Realtime Video (Image to Video)

Animate a still image into video in seconds rather than minutes. Fast enough to iterate on.

Seedance 2.5 Image-to-Video
Seedance 2.5 Image-to-Video

Seedance 2.5 (Image-to-Video) generates Hollywood-grade cinematic videos from reference images and text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability. Built on Seed's unified multimodal architecture, it preserves the input image's subject and composition while adding expressive, physically accurate motion.

Seedance 2.5 Image-to-Video Turbo
Seedance 2.5 Image-to-Video Turbo

Seedance 2.5 (Image-to-Video Turbo) generates cinematic 720p/1080p videos from reference images and text prompts —a faster, more affordable high-resolution tier with native audio-visual synchronization, director-level control, and exceptional motion stability. Built on Seed's unified multimodal architecture.

Seedance 2.5 Text-to-Video
Seedance 2.5 Text-to-Video

Seedance 2.5 (Text-to-Video) generates Hollywood-grade cinematic videos from text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability. Built on Seed's unified multimodal architecture, it leads on instruction adherence, motion quality, and visual aesthetics. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Seedance 2.5 Text-to-Video Turbo
Seedance 2.5 Text-to-Video Turbo

Seedance 2.5 (Text-to-Video Turbo) generates cinematic videos from text prompts at 720p and 1080p with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability — optimized for turbo output. Built on Seed's unified multimodal architecture. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan 2.7 Video Edit

Edit a video by describing the change. Takes up to three reference images to guide who or what appears, and can keep the original audio. Outputs up to 10 seconds.

Wan 3.0 Image-to-Video

Alibaba WAN 3.0 Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan 3.0 Reference-to-Video

Alibaba WAN 3.0 Reference-to-Video combines reference images, videos, and audio with prompts to create coherent videos with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan 3.0 Text-to-Video

Alibaba WAN 3.0 Text-to-Video generates videos from text prompts with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Realtime Video (Text to Video)

Describe a shot and get video back in seconds rather than minutes.

InfiniteTalk Fast
InfiniteTalk Fast

Fast lip-sync — animate a portrait to speak any audio in seconds.

InfiniteTalk
InfiniteTalk

Lip-sync any portrait to any audio for a natural talking-avatar video.

P-Video Turbo
P-Video Turbo

P-Video — fast, efficient image-to-video generation.

Seedance V1.5 Pro I2V
Seedance V1.5 Pro I2V

Seedance 1.5 Pro — high-fidelity image-to-video.

Sora 2 I2V
Sora 2 I2V

Sora 2 — turn an image into a coherent, high-quality video.

Sora 2 Pro I2V
Sora 2 Pro I2V

Sora 2 Pro — premium image-to-video with longer, sharper results.

Kling v2.6 Standard Motion Control
Kling v2.6 Standard Motion Control

Kling 2.6 — animate a subject with motion control for cinematic video.

Kling v3 Motion Control
Kling v3 Motion Control

Kling 3.0 motion control: transfer motion from a reference video to any character image with improved consistency and quality.

Grok Imagine Video Edit
Grok Imagine Video Edit

Generate videos using xAI's Grok Imagine Video model

Wan 2.2 Animate Replace
Wan 2.2 Animate Replace

Use Wan 2.2 Animate to replace a character in a video scene

Seedance 2.0 Image-to-Video
Seedance 2.0 Image-to-Video

Seedance 2 — image-to-video with smooth, dynamic motion.

Seedance 2.0 Text-to-Video
Seedance 2.0 Text-to-Video

Seedance 2 — generate video from a text prompt.

Seedance 2.0 Fast Image-to-Video
Seedance 2.0 Fast Image-to-Video

Seedance 2 Fast — quick image-to-video generation.

Seedance 2.0 Fast Text-to-Video
Seedance 2.0 Fast Text-to-Video

Seedance 2 Fast — quick text-to-video generation.

Sora 2 (Text-to-Video)
Sora 2 (Text-to-Video)

OpenAI Sora 2 is a state-of-the-art text-to-video model with realistic visuals, accurate physics, synchronized audio, and strong steerability. Ready-to-use REST inference API, best performance, no coldstarts, affordable

Kling v3 Standard (Text-to-Video)
Kling v3 Standard (Text-to-Video)

Kling 3.0 Standard delivers high-quality text-to-video generation with smooth motion, cinematic visuals, accurate prompt adherence, and native audio for ready-to-share clips.

Seedance V1.5 Pro (Text-to-Video)
Seedance V1.5 Pro (Text-to-Video)

Seedance 1.5 Pro (Text-to-Video) generates cinematic, live-action–leaning clips from text with strong prompt adherence, expressive motion, and stable aesthetics. It supports 4–12s duration control (including Smart Durati

Veo 3.1 Lite (Text-to-Video)
Veo 3.1 Lite (Text-to-Video)

Google Veo 3.1 Lite generates high-fidelity videos with native audio from text prompts, optimized for cost efficiency.

Veo 3 Fast (Text-to-Video)
Veo 3 Fast (Text-to-Video)

Google Veo 3 Fast creates text-to-video with synchronized audio, delivering faster, more cost-effective results than standard Veo 3; commercial use allowed and pricing starts at $0.25/second. Ready-to-use REST inference

Hailuo 02 Standard (Text-to-Video)
Hailuo 02 Standard (Text-to-Video)

Hailuo 02 is a text-to-video model, fine-tuned to output responsive 768P videos even for complex physics-driven scenes. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Wan 2.5 (Text-to-Video)
Wan 2.5 (Text-to-Video)

Alibaba WAN 2.5 makes 480p-1080p text/image-to-video with synced audio and is faster, more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

PixVerse V5 (Text-to-Video)
PixVerse V5 (Text-to-Video)

PixVerse V5 Text-to-Video generates smooth, natural 5s videos from text prompts in seconds, with 720p output available ($0.20 per 5s). Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

P-Video (Text-to-Video)
P-Video (Text-to-Video)

p-video model running.

LTX-2 Fast (Text-to-Video)
LTX-2 Fast (Text-to-Video)

LTX-2 Fast is a production-grade text-to-video engine that creates synchronized audio and 1080p video from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Grok Imagine (Text-to-Video)
Grok Imagine (Text-to-Video)

Generate videos from text descriptions using xAI's Grok Imagine Video model. Create high-quality videos with customizable duration, aspect ratio, and resolution.

Wan 2.2 Ultra-Fast 480p (Text-to-Video)
Wan 2.2 Ultra-Fast 480p (Text-to-Video)

Wan 2.2 t2v 480p Ultra-Fast generates unlimited AI videos from text prompts at 480p with ultra-fast inference. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Gemini Omni Flash Video
Gemini Omni Flash Video

Gemini Omni Flash Text-to-Video creates short videos with synchronized audio from a text prompt.

Hailuo 2.3
Hailuo 2.3

Hailuo 2.3 is a text-to-video model creating physics-aware 768p videos with 2.5× efficiency and 85% complex instruction response rate. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Hailuo 2.3 (Image to Video)
Hailuo 2.3 (Image to Video)

Hailuo 2.3 Standard is an image-to-video model producing physics-aware 768p output with a 2.5x efficiency improvement. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Hailuo 2.3 Pro
Hailuo 2.3 Pro

Hailuo 2.3 Pro is a text-to-video model delivering 1080p videos with 2.5x efficiency and 85% complex-instruction accuracy. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Kling V3 Turbo Pro
Kling V3 Turbo Pro

Kling V3 Turbo Pro generates high quality 1080p videos from text prompts, with support for single prompts and multi-shot storyboards.

Kling V3 Turbo Pro (Image to Video)
Kling V3 Turbo Pro (Image to Video)

Kling V3 Turbo Pro generates high quality 1080p videos from a first-frame image, with optional text prompts and multi-shot storyboards.

Kling V3 Turbo Standard
Kling V3 Turbo Standard

Kling V3 Turbo Standard generates fast, affordable 720p videos from text prompts, with support for single prompts and multi-shot storyboards.

Kling V3 Turbo Standard (Image to Video)
Kling V3 Turbo Standard (Image to Video)

Kling V3 Turbo Standard generates fast, affordable 720p videos from a first-frame image, with optional text prompts and multi-shot storyboards.

Kling V3.0 Pro
Kling V3.0 Pro

Kling 3.0 Pro delivers top-tier text-to-video generation with smooth motion, cinematic visuals, accurate prompt adherence, and native audio for ready-to-share clips.

Kling Video O3 Standard
Kling Video O3 Standard

Kling Omni Video O3 (Standard) is Kuaishou's advanced unified multi-modal video model with MVL (Multi-modal Visual Language) technology. Text-to-Video mode generates cinematic videos from text prompts with subject consistency, natural physics simulation, and precise semantic understanding. Supports audio generation. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.

Leonardo Motion 2.0
Leonardo Motion 2.0

Leonardo Motion 2.0 delivers upgraded image-to-video generation, producing more realistic, detailed videos than its predecessor. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

LTX-2 Pro
LTX-2 Pro

LTX-2 Pro is a text-to-video engine that generates synchronized audio and 1080P video from text prompts for production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

LTX-2 Pro (Image to Video)
LTX-2 Pro (Image to Video)

LTX-2 is an AI creative engine for production workflows, generating synchronized audio and 1080p video output (cost $0.06/s). Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Lucy Edit Pro (Video Edit)
Lucy Edit Pro (Video Edit)

Lucy Edit Pro is a state-of-the-art video editing model that produces studio-quality results in minutes, not weeks. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Luma Ray 3.2
Luma Ray 3.2

Luma Ray 3.2 Text-to-Video generates cinematic videos from text prompts with controllable aspect ratio, resolution, duration, and optional reference images.

Luma Ray 3.2 (Image to Video)
Luma Ray 3.2 (Image to Video)

Luma Ray 3.2 Image-to-Video animates a source image into cinematic video guided by a text prompt, with controllable aspect ratio, resolution, duration, and optional reference images.

Ovi (Video + Audio)
Ovi (Video + Audio)

Ovi is a veo-3-like model that converts text or text+image prompts into synchronized video with audio. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Pika 2.2
Pika 2.2

Pika v2.2 is a text-to-video model that creates high-quality videos from text prompts, supporting multiple video sizes and advanced prompt optimization. Ready-to-use REST API, no coldstarts, affordable pricing.

Pika 2.2 (Image to Video)
Pika 2.2 (Image to Video)

Pika V2.2 Image-to-Video converts images into high-quality videos in various sizes with prompt optimization for precise results. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Pixverse V6
Pixverse V6

PixVerse V6 generates high-quality videos from text prompts with flexible duration (1-15s), multiple resolutions up to 1080p, and optional audio generation. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Pixverse V6 (Image to Video)
Pixverse V6 (Image to Video)

PixVerse V6 generates high-quality videos from images with flexible duration (1-15s), multiple resolutions up to 1080p, and optional audio generation. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Seedance 2.0 Mini
Seedance 2.0 Mini

Seedance 2.0 Mini is 's faster, lower-cost tier of Seedance 2.0 for cinematic multi-shot video — narrative sequences, AI camera control (zoom/pan/tracking), and consistent characters across scenes, from text or image prompts. 480p-4k, 4-15s, aspect ratios 16:9 / 4:3 / 1:1 / 3:4 / 9:16. Priced at 50% of standard Seedance 2.0.

Seedance 2.0 Mini (Image to Video)
Seedance 2.0 Mini (Image to Video)

Seedance 2.0 Mini is 's faster, lower-cost tier of Seedance 2.0 for cinematic multi-shot video — narrative sequences, AI camera control (zoom/pan/tracking), and consistent characters across scenes, from text or image prompts. 480p-4k, 4-15s, aspect ratios 16:9 / 4:3 / 1:1 / 3:4 / 9:16. Priced at 50% of standard Seedance 2.0.

SkyReels V4 (Image to Video)
SkyReels V4 (Image to Video)

SkyReels V4 Image to Video generates videos from image references and text prompts using the SkyReels V4 image2video workflow.

Sora 2 Pro
Sora 2 Pro

openai/sora2

Veo 3.1
Veo 3.1

Google Veo 3.1 converts text prompts into videos with synchronized audio at native 1080p for high-quality outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Veo 3.1 (Image to Video)
Veo 3.1 (Image to Video)

Google Veo 3.1 is an Image-to-Video model that converts images into high-quality videos with native 1080P output for enhanced detail and creative flexibility. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Veo 3.1 Fast
Veo 3.1 Fast

Google Veo 3.1 Fast creates text-to-video with native 1080p and synchronized audio, delivering high-quality videos for creators. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Veo 3.1 Fast (Image to Video)
Veo 3.1 Fast (Image to Video)

Google Veo 3.1 Fast is an Image-to-Video model with native 1080p output for high-detail videos from images and fast performance. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Wan 2.7
Wan 2.7

Alibaba WAN 2.7 Text-to-Video turns plain prompts into coherent, cinematic clips with crisp detail, stable motion, and strong instruction-following—great for ads, explainers, and social posts. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan 2.7 (Image to Video)
Wan 2.7 (Image to Video)

Alibaba WAN 2.7 converts images into videos (720p/1080p) with optional audio, supporting first and last frame control. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan 2.7 Pro (Image to Video)
Wan 2.7 Pro (Image to Video)

Alibaba WAN 2.7 Pro converts images into ultra-high-resolution videos (1080p/2K/4K) with cinematic detail and smooth motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Frequently asked

What is the best AI video model in 2026?

It depends on your goal: Sora and Veo lead on cinematic realism, Kling on motion control, and Seedance and PixVerse on speed and cost. Every model below runs in the same playground, so you can compare them side by side on your own prompt.

Can I use these AI video models via API?

Yes — every model is callable over a simple REST API with a single key, billed per run in credits. See the developer docs to get started.

How much does AI video generation cost?

Each model shows its exact per-run credit cost before you run it. Cheaper models start in the single digits of credits per clip; premium cinematic models cost more. Failed runs are refunded automatically.