Recraft V4.1 by Recraft AI generates design-ready images with accurate in-image text. Six aspect ratios, palette steering, and background color control.
Models
All Models
P-Image Try-On by Pruna AI dresses a person photo in a full outfit using up to 11 garment images, from tops and accessories to outerwear.
Meshy UV Unwrap by Meshy auto-generates clean UV maps for any GLB mesh. Upload your model and get a texture-ready, UV-unwrapped GLB back in seconds.
Kling Video to Audio by Kuaishou adds synchronized sound effects and background music to silent video clips of 3 to 20 seconds.
Generate seamless textures and patterns that repeat perfectly in any direction. Prompt from scratch or convert an existing image into a tile, at up to native 2K.
Luma's flagship image model combining reasoning with generation. Up to 9 reference images, web search grounding, 9 aspect ratios, and crisp text rendering.
Luma Uni-1 by Luma Labs generates and edits images with a reasoning model, reliable text rendering, web search grounding, and up to 9 reference images.
Luma Ray 3.2 Edit by Luma Labs. Restyle any video while preserving its motion. Nine edit strengths, HDR output, and face, pose, depth, and trajectory controls.
Luma Ray 3.2 Reframe by Luma Labs converts any video to a new aspect ratio via AI outpainting, filling the added canvas with prompt-guided content.
Luma Ray 3.2 by Luma Labs. Cinematic text or image to video up to 10s and 1080p, with HDR, seamless looping, start/end keyframes, and six aspect ratios.
Add styled captions to any video: Whisper transcription, translation into 18 languages, 7 presets or your own style, burned in or as an SRT file.
Generate seamless 360 equirectangular panoramas from a text prompt, powered by FLUX.1. Choose from 21 styles with automatic seam and pole correction.
Generate seamless 360° skyboxes from a text prompt. Outputs equirectangular panoramas and cubemap formats, engine-ready for real-time scenes.
Sonilo V1.1 by Sonilo analyzes your video's pacing, motion, and mood to generate an original licensed soundtrack automatically.
Sonilo V1.1 by Sonilo. Text to original instrumental music, 1 to 600 seconds. Describe genre, mood, and instruments, then generate up to 3 variations per prompt.
MAI Image 2.5 Edit by Microsoft. Instruction-based editing: swap backgrounds, update text, change styles, or fix lighting. Up to 4 outputs per run.
MAI Image 2.5 by Microsoft generates photorealistic images with reliable in-image text. Ideal for posters, ads, and key art across 11 aspect ratios.
Audio Extract by Scenario. Pull the original audio track from any video as MP3, WAV, or AAC. Optional broadcast-safe loudness normalization.
ElevenLabs Voice Isolator by ElevenLabs strips background noise, music, and ambient sounds from audio or video, returning a clean isolated vocal track.
ElevenLabs Voice Changer by ElevenLabs transforms any voice recording, preserving words, timing, and emotion, using 21 preset voices or your own cloned voice.
Split any image into separate, editable layers. Each object is isolated with a transparent background and the gaps filled by AI, ready to move or restyle.
Extract subjects from any video as separate layers plus a clean background plate, ready to edit, reposition, or reuse each element independently.
Generate seamless, tileable textures from a text prompt. Optional seam erasing ensures perfect 2D tiling. Add reference images for style guidance.
Rodin Gen-2.5 by Deemos Technology converts 1 to 5 images into production-ready 3D models with five quality tiers, quad or triangle topology, and PBR textures.
Rodin Gen-2.5 by Deemos Technology. Generate production-ready 3D meshes from text. Five quality tiers, quad or triangle topology up to 500K faces.
Transcribe audio or video into text or SRT subtitles using Whisper. Supports auto language detection, English translation, and voice activity filtering.
Split a video into ordered segments at precise cut points, preserving audio and exporting each clip as MP4, MOV, WebM, or GIF.
Split any audio file into precise segments by timestamp. Outputs N+1 clips from N cut points, exported as MP3, WAV, OGG, or M4A.
Resize and reframe any image to exact target dimensions up to 4K, preserving art style, subjects, on-image text, brand elements, and color palette.
P-Video Replace by Pruna AI swaps up to four identities into an existing video, preserving the original background, motion, and audio. Outputs up to 1080p.
P-Video Animate by Pruna AI transfers motion from a source video onto a still character image, with no rigging needed. Outputs at up to 1080p with audio.
Ideogram's generative background remover isolates subjects on a transparent PNG, keeping hair, fur, glass, and fine edges clean. One image in, compositing-ready output.
Grok Imagine Image Quality by xAI generates and edits 2K images with precise text and multilingual typography across 16 aspect ratios, now including cinematic 21:9 and 5:2.
Foley Control adds synchronized sound effects and ambience to any video, guided by text, a negative prompt, or a short reference audio clip.
Upscale images to 2x, 4x, 8x, or 16x with a diffusion engine that adds photorealistic detail. A creativity slider controls fidelity vs. texture enhancement. Up to 8K.
Extract ControlNet-ready detection maps from any image. Ten preprocessors in one tool: Canny edges, depth, pose, normals, lines, segmentation, and more.
Uthana Character Rigging by Uthana automatically rigs any uploaded 3D humanoid model for animation, outputting a production-ready GLB or FBX file.
Auto Subtitles by Scenario transcribes and burns subtitles into any video, with full control over font, color, border style, segment length, and language.
Happy Horse Video Edit by Alibaba transforms existing clips with text instructions, swapping style, characters, or scenes. Up to 15s input, 720P or 1080P output.
Sparc3D 2.1 by Hitem3D turns 1-4 photos into a watertight 3D mesh at up to 1536 Pro resolution, with optional PBR texturing and up to 2M faces.
Sparc3D 2.1 Portrait by Hitem3D converts 1 to 4 portrait photos into a detailed 3D face model with up to 2M faces and PBR textures.
Turn text into native 4K video up to 15 seconds long. Kuaishou's Kling V3 delivers physics-aware motion and built-in audio in portrait, square, or landscape.
Animate any image into native 4K video at up to 60fps. Cinematic motion physics, multi-shot scenes, and built-in lip-sync from Kuaishou's Kling V3.
SAM 3.1 Video by Meta tracks and segments objects across video frames using a text prompt. Returns up to 16 isolated mask tracks per video.
SAM 3.1 by Meta segments any image into isolated object masks. Guide detection with a text prompt or up to 10 bounding boxes. Outputs one PNG mask per object.
ERNIE Image Turbo by Baidu is a fast distilled variant of ERNIE Image with the same bilingual text-in-image rendering, built for speed-sensitive workflows.
ERNIE Image by Baidu is an 8B text-to-image model built for accurate text rendering in images. Ideal for posters, infographics, signage, and UI mockups.