Turn text into native 4K video up to 15 seconds long. Kuaishou's Kling V3 delivers physics-aware motion and built-in audio in portrait, square, or landscape.
Models
All Models
Animate any image into native 4K video at up to 60fps. Cinematic motion physics, multi-shot scenes, and built-in lip-sync from Kuaishou's Kling V3.
Meshy Multi Image to 3D by Meshy. Upload 1-4 photos from different angles to generate a textured 3D mesh, with PBR maps, pose modes, and polycount control.
SAM 3.1 Video by Meta tracks and segments objects across video frames using a text prompt. Returns up to 16 isolated mask tracks per video.
SAM 3.1 by Meta segments any image into isolated object masks. Guide detection with a text prompt or up to 10 bounding boxes. Outputs one PNG mask per object.
ERNIE Image Turbo by Baidu is a fast distilled variant of ERNIE Image with the same bilingual text-in-image rendering, built for speed-sensitive workflows.
ERNIE Image by Baidu is an 8B text-to-image model built for accurate text rendering in images. Ideal for posters, infographics, signage, and UI mockups.
Phota Enhance by PhotaLabs upscales and restores photos with identity-preserving AI. No prompt needed: drop in an image, get back a sharper, higher-resolution result.
Edit photos with text instructions. Upload up to 10 references and Phota transforms the scene, background, or lighting while preserving subject identity.
Phota by PhotaLabs generates photorealistic images of people from text: portraits, lifestyle, fashion, and group shots at 1K or 4K in six aspect ratios.
LTX-2.3 Pro Audio to Video by Lightricks generates video synchronized to your audio. Voice cadence shapes pacing, music drives motion. Up to 20s, 1080p.
ElevenLabs Speech to Speech by ElevenLabs re-voices any recording into 21 preset voices or your own custom cloned voices, preserving the words, timing, and emotional delivery.
Minimax Music Cover by MiniMax transforms any song into a new genre or style, preserving the original melody while reimagining vocals, instruments, and arrangement.
Dub video or audio into 30 languages, preserving each speaker's voice via cloning. Auto-detects source language and up to 10 speakers; optional background audio removal.
Pixverse V6 by PixVerse: animate any image into a cinematic clip up to 15 seconds. Choose style, resolution, and optionally add audio.
PixVerse V6 by PixVerse. Text-to-video in five artistic styles, up to 15 seconds at 1080p, with optional native audio and multi-clip storytelling.
Seedance 2.0 Fast by ByteDance. Speed-optimized text, image, and video-to-video model with multimodal references, optional audio, and up to 15 seconds at 720p.
Seedance 2.0 by ByteDance: text, image, and video to video generation with up to 1080p output, synced audio, and multi-reference support.
Ideogram V3 Layerize Text by Ideogram splits flat graphics into a clean base image and editable text layers, ready to localize or restyle.
Generate images with a native alpha channel using Ideogram V3. Four speed tiers, 15 aspect ratios, negative prompt support. No background removal needed.
Veo 3.1 Lite by Google. Text-to-video and image-to-video with native audio generation. Choose 720p or 1080p, landscape or portrait, and 4 to 8 seconds.
Upscale real footage to 1K, 2K, or 4K while preserving every detail exactly as shot. Ideal for professional video that needs size, not reinterpretation.
Magnific Video Upscaler Creative by Magnific. Upscale video to 1K, 2K, or 4K with creativity, sharpening, smart grain, FPS boost, and Vivid or Natural color mode.
ReconViaGen 0.5 turns 1 to 8 photos of an object into a textured 3D mesh, with controls for mesh detail, texture resolution, and multi-view blending.
Sync-3 Lipsync by Sync Labs syncs any audio to a speaker's mouth in video. Built for dubbing, voice-over, and ADR with five duration-mismatch modes.
JoyAI Image Edit by JD Open Source: edit any photo with a plain-language instruction. Control guidance, inference steps, and negative prompts for precise results.
P-Image Upscale by Pruna AI. Fast upscaling up to 8 MP via target megapixel or side-factor mode, with optional detail and realism enhancement.
Wan 2.7 Video Edit by Alibaba rewrites video content from text instructions: swap backgrounds, shift lighting, apply styles, or restyle using a reference image.
Wan 2.7 by Alibaba. Text-to-video at up to 1080p, 2-15 seconds, five aspect ratios, optional synced audio, and built-in prompt expansion.
Wan 2.7 Image Pro by Alibaba generates up to 4K images from text and edits or fuses up to 9 reference images, with optional thinking mode for deeper prompt reasoning.
Wan 2.7 Image by Alibaba: text-to-image and multi-reference editing at 1K or 2K, with image set mode and thinking mode for up to four outputs per run.
Wan 2.7 I2V by Alibaba animates images into 720p or 1080p video, 2-15s. Supports first/last frame bracketing, clip continuation, and optional audio input.
Minimax Speech 2.8 Turbo by MiniMax. Fast TTS with 17 voices, 10 emotions, and 40+ language boosts, built for low-latency apps.
Minimax Speech 2.8 HD by MiniMax. Premium TTS with 17 voices, 10 emotions, 40+ languages, natural interjections, and precise control over speed, pitch, and volume.
Hy Wu Edit by Tencent. Transfer outfits, swap faces, and blend textures using up to 3 reference photos, with no fine-tuning required.
Google Lyria 3 Clip by Google generates 30-second music clips from text prompts or reference images, with control over genre, tempo, mood, and song structure.
Grok Edit Video by xAI: edit any clip with a text prompt. Restyle scenes, swap objects, change environments, all while preserving what you don't touch.
Grok Extend Video by xAI continues your clip from its last frame, generating 2 to 10 seconds of new AI footage guided by a text prompt.
Flux.2 LoRA for stylized 3D costume sets. Renders outfit collections and equipment on invisible figures with detailed textures and studio lighting.
Tada 3B by Hume AI clones any voice from a short audio reference and synthesizes multilingual speech across 10 languages with no transcript hallucinations.
Tada 1B by Academia / Open Source clones any voice from a short audio sample and synthesizes multilingual speech across 10 languages with speed control.
HeyGen Avatar 4 by HeyGen: turn a face photo into a talking video with 80+ preset voices, resolutions up to 1080p, and stable or expressive motion.
HeyGen Video Translate Speed dubs your video into 170+ languages with AI lip sync, optimized for fast, high-volume translation at scale.
HeyGen Video Translate Precision by HeyGen. Translate video speech into 170+ languages with high-fidelity lip sync, voice cloning, and multi-speaker support.
HeyGen Video Agent by HeyGen turns a text prompt into a polished presenter video, handling scripting, avatar selection, scenes, and narration pacing automatically.
Pixelcut Background Removal by Pixelcut: isolate subjects from any image with clean edge detection. Export as full RGBA composite or alpha-only mask.
Tripo P1 Multi View by Tripo AI. Generate high-fidelity 3D meshes with PBR maps from up to four reference angles using native 3D diffusion.
Physic Edit applies physics-aware transformations to images: flood cities, melt armor, shatter glass, or freeze scenes with accurate refraction and deformation.
A Flux 2 LoRA for 3D low-poly environments suited to platformers and adventure games, with geometric shapes, vibrant colors, and soft lighting.