Flash VSR by Tsinghua/Shanghai AI Lab upscales video 2x to 4x using one-step diffusion, delivering sharp detail restoration with real-time streaming speed.
Models
All Models
Pixverse Swap by PixVerse replaces people or backgrounds in any video using a reference image, with Person and Background modes up to 720p.
Hunyuan 3D 3.0 Pro (Sketch) by Tencent converts hand-drawn sketches into textured 3D meshes with adjustable face count and optional PBR materials.
Hunyuan 3D Pro 3.0 by Tencent converts a single image into a detailed 3D mesh with up to 1.5M faces, PBR textures, and a choice of triangle or quad topology.
Hunyuan 3D 3.0 Pro (Multiview) by Tencent generates 3D assets from up to 4 reference views, with PBR materials, face count control, and tri or quad topology.
Qwen Edit Multi-Angle by Alibaba edits images using camera controls: rotate, zoom, vertical tilt, and wide-angle, with optional text prompts.
Flux Kontext by Black Forest Labs. Edit images with text instructions, preserving characters and style across local or global edits. Supports up to 10 reference images.
GPT Image 1.5 by OpenAI. Instruction-driven image editing with precise reference fidelity, transparent background support, and aspect ratio control.
Seedream 4.5 by ByteDance edits images from natural language instructions, with up to 10 reference images, 4K output, and strong subject preservation.
P-Image Edit by Pruna AI is a fast, instruction-driven image editor. Edit, composite, or transform images using text and up to 10 reference images.
Meshy Remesh by Meshy rebuilds your 3D model's geometry with clean triangle or quad topology and a target polycount you control, from 100 to 300,000 polygons.
Meshy Text-to-3D by Meshy generates textured, game-ready 3D assets from a text prompt, running on Meshy 6.
Flux.2 [klein] 4B by Black Forest Labs. A 4B-parameter model distilled for sub-second image generation and editing, with LoRA support.
FLUX.2 [klein] 4B Base by Black Forest Labs. Compact, undistilled 4B image model built for LoRA compatibility, fine-tuning, and precise prompt control.
Flux 2 (Turbo) Edit by Black Forest Labs: fast, instruction-based image editing. Describe the change, get a transformed image in seconds.
FLUX.2 [klein] 9B Base by Black Forest Labs. Undistilled 9B foundation model for text-to-image and image-to-image, built for LoRA fine-tuning and high output diversity.
FLUX 2 (Max) by Black Forest Labs. Top-tier image editing with up to 8 reference images, up to 4 MP output, and precise consistency across colors, faces, and objects.
FLUX 2 (Flex) by Black Forest Labs. Precision image editing with up to 10 reference images, up to 4 MP output, and strong typography and detail control.
FLUX 2 (Pro) Edit by Black Forest Labs. Instruction-based image editor with up to 8 reference images, 4 MP output, and flexible aspect ratios.
LongCat Image by Meituan edits photos via plain-language instructions, no masks needed. Supports 15 edit types and accurate text rendering in Chinese and English.
Sparc3D (Portrait) by Hitem3D converts 1-4 face photos into detailed 3D head models, up to 1536³ Pro resolution, as mesh-only or fully textured assets.
Minimax Hailuo 2.3 (Fast) by MiniMax. Text-to-video and image-to-video with optimized speed. Choose 6 or 10-second clips at 768p, or 6 seconds at 1080p.
Minimax Hailuo 2.3 by MiniMax generates text and image-to-video at 768p or 1080p, with strong motion coherence, temporal consistency, and anime style support.
A Flux Kontext LoRA that ages clean building images into worn, crumbling ruins, preserving the original structure, layout, and perspective.
Turns a character image into a 9-emotion expression sheet, preserving their look and art style.
Clarity Crystal Upscaler by Clarity AI enhances images with an adjustable scale factor and a creativity dial, optimized for portraits, faces, and product photos.
Flux Kontext LoRA that transforms 3D blockout shapes into photorealistic renders, preserving the original geometry and spatial layout.
Veo 3.1 (Fast) by Google. Fast text-to-video and image-to-video with native audio, first/last frame anchoring, and subject-consistent generation from reference images.
Generate a four-view character turnaround sheet from any character image, showing every side on a clean background.
Flux Kontext LoRA that converts photos of buildings into isometric 3D models placed on square tiles, ideal for game assets and top-down scene design.
Hunyuan Image 3 by Tencent. An 80B MoE text-to-image model with exceptional multilingual text rendering, complex scene understanding, and photorealistic output.
SeedVR2 - Image Upscale by Bytedance scales images up to 4K using one-step diffusion super-resolution, with factor or target resolution modes and a noise slider.
SeedVR2 - Video Upscale by ByteDance upscales videos up to 4K using one-step diffusion, restoring fine detail and temporal consistency with minimal hallucination.
Omni Human 1.5 by ByteDance animates a single photo into a talking avatar, syncing lip movements, expressions, and natural gestures to your audio.
Wan 2.5 I2V by Alibaba animates any image into a 720p or 1080p video clip of 5 or 10 seconds, with optional audio synchronization.
Wan 2.5 T2V by Alibaba generates 720p or 1080p videos from text, in 16:9 or 9:16, with synchronized audio support, in 5 or 10 second clips.
Wan 2.2 Animate (Replace) by Alibaba swaps a person in a video with your reference character image, preserving the original motion and audio.
Wan 2.2 Animate (Move) by Alibaba brings any character image to life by transferring body motion and expressions from a reference video.
Wan 2.2 Reframe by Alibaba converts any video to a new aspect ratio, keeping the subject intelligently framed. Outputs in 16:9, 1:1, or 9:16 at up to 720p.
Lucy Edit (Pro) by Decart edits videos via text: swap outfits, change objects, or replace scenes while preserving motion and identity. Up to 720p.
Lucy Edit (Dev) by Decart AI edits videos from text prompts, swapping outfits, characters, or scenes while preserving original motion and composition.
Expand any video's borders left, right, up, or down with Wan 2.2 Outpainting by Alibaba. New content is generated to match your scene seamlessly.
Sync Lipsync 2 (Pro) by Sync Labs syncs mouth movements to any audio track with studio-grade detail preservation, up to 4K resolution.
Luma Modify Video by Luma Labs rewrites existing footage with a text prompt across nine Adhere, Flex, and Reimagine strength levels.
ElevenLabs 3 (Alpha) is an expressive text-to-speech model for 70+ languages, using inline tags like [whispers] or [excited] to steer emotion, emphasis, and delivery.
Pixverse Lipsync by PixVerse syncs any audio track to a speaker's mouth movements in your video, with built-in TTS across 14 voices.