Qwen Image 2512 by Alibaba generates and transforms images up to 2048px with photorealistic human rendering, native text in images, and support for up to 6 LoRA models.
Models
All Models
Seedance 1.5 Pro by ByteDance generates text-to-video and image-to-video clips up to 12 seconds in 1080p, with native audio and multilingual lip-sync.
MM Audio 2 Text-To-Audio (SFX) by Academia / Open Source generates realistic sound effects from text, with clips up to 30 seconds and a guidance strength control.
MM Audio 2 by MMAudio adds synchronized, high-quality audio to silent videos using a text prompt to guide the generated soundscape.
Kling V2.6 Motion Control by Kuaishou transfers motion from a reference video to a character image, replicating body movement, facial expressions, and lip sync.
Qwen Image Layered by Alibaba converts any image into 2 to 8 editable RGBA layers. Each layer is independently movable, resizable, and recolorable.
Flux LoRA for storybook-style isometric backgrounds. Generates warm, detailed fantasy scenes: villages, forests, medieval towns, and architectural environments.
Extend any 16:9 video with Google's Veo 3.1, continuing your clip with visual and audio consistency. Prompt-guided to keep your story flowing.
SAM3D Objects by Meta reconstructs textured 3D meshes from a single photo, isolating multiple objects at once via text prompts, points, boxes, or masks.
SAM3D Human Body by Meta converts a single photo into a full-body 3D mesh, covering body, hands, and feet. Supports multiple characters per image.
SAM3D Align by Meta places a body mesh and object mesh into a shared, spatially consistent 3D scene using your original image as reference.
Wan 2.6 T2V by Alibaba generates up to 1080p video in 5, 10, or 15 seconds with native audio sync and multi-shot storytelling from a text prompt.
Wan 2.6 I2V by Alibaba animates images into 720p or 1080p video clips of 5, 10, or 15 seconds, with native audio lip-sync and multi-shot storytelling.
Magnific Upscaler Precision by Magnific adds high-fidelity detail as it upscales images 2x to 16x, with controls for sharpness, grain, and ultra detail.
Magnific Creative by Freepik is a generative upscaler that invents new detail while enlarging 2x to 16x, best on an already-good image that needs more density and finish.
Photoroom Uncrop restores subjects clipped at the image edge, seamlessly reconstructing missing parts for clean, full-frame product shots.
Erase text, logos, and watermarks from any image with Photoroom's engine. Choose artificial overlays, natural scene text, or both.
Photoroom Relighting by Photoroom adjusts light sources, exposure, and mood on product images, with an option to lock brand color accuracy.
Expand any image's background to a new canvas size with AI-generated fill. Original content stays intact, new areas blend seamlessly.
Photoroom Background Replacer by Photoroom replaces image backgrounds using a text prompt or reference image, preserving the original subject.
Photoroom Background Removal by Photoroom. Clean, precise subject cutouts from any photo, with optional shadow effects and HD mode for high-resolution images.
Kling O1 Reference Video by Kuaishou generates new scenes guided by a source video, preserving its motion and style while compositing characters and images.
Kling O1 Video Editing by Kuaishou modifies existing footage via text prompts, with optional reference images and multi-angle Elements for subject swaps.
Kling O1 Reference Images by Kuaishou generates video from reference images, keeping characters and objects visually consistent. Up to 3 elements, 3-10 second clips.
Creatify Aurora by Creatify animates a single portrait into a lifelike speaking avatar with audio-driven lipsync, expressive gestures, and full-body motion.
Sync Lipsync React-1 by Sync Labs syncs audio to video with emotion presets and three animation modes: lips only, face expressions, or full head movement.
FLUX 1.1 (Pro) by Black Forest Labs is a fast text-to-image model that tops prompt-adherence and aesthetics benchmarks, generating photorealistic images in seconds.
FLUX 1.1 (Pro Ultra) by Black Forest Labs generates images at up to 4MP across 11 aspect ratios, with a Raw mode for natural, candid photography looks.
Kling AI Avatar 2 (Pro) by Kuaishou turns a single image and an audio track into a lifelike talking avatar video with precise lip sync.
Kling 2.6 I2V (Pro) by Kuaishou animates a still into 1080p video with its strongest motion yet, stable character identity, and optional native audio synced in one pass.
Kling 2.6 T2V (Pro) by Kuaishou Technology generates 1080p video clips of 5 or 10 seconds with optional native audio, including voice, sound effects, and ambience.
Kling O1 I2V by Kuaishou animates a start image into video with optional end-frame control for precise transitions. Choose 5 or 10 second clips.
Voxel Crafter 1.0 generates blocky 3D voxel models from text or images, with grid size controls (16 to 256) for width, height, and depth.
Grid Maker arranges multiple images into a clean grid layout. Control columns, rows, padding, background color, and cell aspect ratio.
Texture Converter turns a flat image into a surface material, with sliders for how raised, shiny, polished, and angular it looks, plus an option to invert the relief.
Extract individual frames from any video as PNG, JPEG, or WebP images. Choose a frame interval or pull every frame at once.
Assemble up to 1,000 images into an MP4 or GIF with control over frame rate, compression, looping, and optional audio.
Scenario Gemini Upscale by Google uses multimodal reasoning to sharpen and enhance images up to 4K, with an optional prompt and creativity dial.
A Flux.1 LoRA that renders any landscape or environment in a bold-line cartoon illustration style, with clean outlines and smooth cel shading.
Flux.1 LoRA for glossy, candy-colored surreal landscapes. Generates dripping liquids, swirling textures, and vibrant floral scenes.
Flux.1 LoRA for stylized 3D cartoon characters with vibrant colors and exaggerated proportions, built for games and animation assets.
Flux LoRA that renders complete costume sets as flat-lay collections, from medieval armor to sci-fi gear, in muted, semi-realistic tones.
Z-Image Turbo by Tongyi-MAI is a distilled open-source image model built for near-instant generation with strong photorealism and accurate text rendering in images.
Bria Remove Background by Bria isolates subjects with clean cutouts, preserving fine details like hair. Outputs a transparent PNG or a fully opaque image.
Image Slicer by Scenario splits any image into a custom grid of up to 6x6 sections, outputting each tile as its own file.
Pixel Snapper cleans up pixel art by snapping every pixel to a consistent grid and reducing colors to a strict palette, removing blurry AI artifacts.