Models
All Models
Generate editable SVG vector images that follow a reusable Recraft style ID or attached reference images. Built for logos, icons, mascots, and other vector assets.
Generate high-quality editable SVG vector images that follow a reusable Recraft style ID or attached reference images. Built for polished logos, icon systems, and brand assets.
Generate high-resolution raster images that follow a reusable Recraft style ID or attached reference images. Built for print-ready campaign assets and high-detail brand work.
Generate 1K raster images that follow a reusable Recraft style ID or attached reference images. Built for consistent campaign, brand, and product visuals.
Gemini 3.5 Transcribe converts speech to text with speaker diarization, word-level timestamps, 85+ language auto-detection, and custom vocabulary biasing.
Gemini Omni 1.1 Flash continues an existing video, appending 3 to 10 seconds of new native-audio footage that matches the original motion, at 360p through 4K.
Gemini Omni 1.1 Flash edits an existing video from natural-language instructions, preserving motion while changing the look, with optional reference images and native audio.
Gemini Omni 1.1 Flash generates subject-consistent video with native audio from reference images and/or a short reference video, at 360p through 4K.
Gemini Omni 1.1 Flash generates video with native audio from a text prompt or an input image, at 360p through 4K, with optional last-frame interpolation.
Lightricks' higher-fidelity audio-driven video model: feed it a music track and it generates a matching visual performance in sync, up to 1080p.
Lightricks' higher-fidelity video model, natively multi-shot: cut across several connected scenes in one prompt with synced audio, up to 1080p.
Lightricks' audio-driven video model: feed it a music or speech track and it generates a matching visual performance in sync, up to 4K and 20 seconds.
Lightricks' fast video model, natively multi-shot: one prompt can cut across several connected scenes with synced audio, up to 4K and 20 seconds, plus image-to-video.
Turn text prompts into cinematic 5-15s videos at 480p or 768p. Fal's MiniMax H3 with stronger prompt adherence and richer aesthetics for fast video generation.
Animate a still image into a cinematic 5-15s video at 480p or 768p. Fal's post-trained MiniMax H3 with stronger prompt adherence, plus optional end-frame control.
Generate a standalone character motion clip from a text prompt: Retarget the clip onto your own rigged character in 2-10 second clips.
Meshy 7 text-to-3D: higher-fidelity geometry than Meshy 6, optional Ultra mode for finer surface detail, and 2K/4K/8K PBR texturing in one generation.
Krea 2 Turbo Style is Krea's dedicated style-transfer model: feed it 1 to 3 reference images and it applies their palette, texture, and composition to any new prompt, fast.
Krea 2 Turbo is Krea's fastest general-purpose image model: quick, versatile, and equally capable at photoreal and illustrated looks, with resolution presets and optional prompt expansion.
Krea 2 Medium Turbo is the fast version of Krea 2 Medium: same illustration and anime strengths, optional style references, and creativity control, tuned for rapid iteration and quick concepting.
Krea 2 Large is Krea's photorealism specialist, built for raw aesthetics like natural grain, motion blur, and true dynamic range. Style reference images and a creativity dial push it toward a specific photographic mood.
Krea 2 Medium is Krea's illustration and anime specialist, expressive and painterly with striking visual variety. Add style reference images to steer the look, and tune creativity from literal to wildly imaginative.
Isolates and flattens one material from a photo of a real surface into a clean, seamlessly tiling PBR set: base color, normal, roughness, metalness, and height maps.
Generates a seamlessly tiling PBR material from text prompt, with img2img variation and inpainting. Outputs base color, normal, roughness, metalness, and height maps, up to 8K.
Predicts a full seamless PBR map set, base color, normal, roughness, metalness, and height, from a single photo or texture. Reads the source image without rewriting it.
FLUX.3 super-resolution from Black Forest Labs via Fal. Upscale clips up to 20s to 1080p, 2K, or 4K, with Precise (source-faithful) or Creative (detail enhancement) mode.
Automatic dubbing that translates speech in audio or video into 31 languages while preserving each speaker's voice, tone, and timing. Source auto-detects; keyterms keep names intact.
Qwen Image 3.0 Pro is the high quality variant of Qwen Image 3.0, great for ads, game art, including dense text layouts and multilingual headlines, with sharper detail up to 2K.
Alibaba's Qwen Image 3.0 makes visuals for films, ads, and game art, plus dense text layouts, posters, UI, and multilingual headlines across 12+ languages.
Splits a 3D mesh into clean, assembly-ready parts. Six character templates with ball or dovetail print joints, or granularity-based splitting for any object.
Quantizes a textured 3D mesh into 1 to 8 flat color regions for multi-filament printing. Set the palette size and Multicolor snaps every surface to the nearest color, keeping the model watertight and ready to slice.
Continue an existing clip into a seamless new shot with synchronized native audio, up to 20 more seconds at 1080p. Source under 15 seconds.
Pin up to 10 keyframe images at exact frame positions and get a video that follows them, with synchronized native audio, up to 20 seconds at 1080p.
Give a first and last frame and get a smooth video that transitions between them, with synchronized native audio, up to 20 seconds at 1080p.
Animate a still image into a cinematic clip with synchronized native audio, up to 20 seconds at 1080p. Your image sets the opening frame.
Turn a text prompt into a cinematic video with synchronized native audio, up to 20 seconds at 1080p. Handles dialogue, sound effects and music in one pass.
Reference-to-video from xAI: guide people, objects, and clothing with 1 to 3 reference images tagged @image1 in your prompt, without locking the first frame. Native audio, up to 15 seconds at 720p.
Generate commercial-use sound effects from a text prompt. Describe any sound, choose an exact length from 1 to 180 seconds, and export as AAC, MP3, WAV, or FLAC.
Generate synchronized, royalty-free sound effects for any video and get back the audio track. Auto-detects scenes, or steer with a prompt or per-range segments.
Score any video with an original, commercially licensed soundtrack matched to its pacing and mood. Optionally keep the original speech while replacing the music.
Add synchronized, royalty-free sound effects to a video and get back the finished clip with audio mixed in, plus a separate audio track.
Fast multilingual text-to-speech from Alibaba Qwen Audio 3.0 (Flash): 45 preset voices across 11 languages, with natural delivery for voiceover, narration, ads, and game dialogue.
Studio-quality product shots without the studio. Pixelcut cuts your product out of any photo and sets it on a solid color, transparent, or custom background, with margins, shadows, and an optional watermark.
Microsoft's highest-fidelity instruction-guided image editor. Swap text, recolor, restyle, replace objects, and change lighting from a plain prompt while preserving structure.
Microsoft's flagship text-to-image model for hero imagery with standout legible typography. Photoreal and stylized, 8 aspect ratios, up to 4 variations per prompt.
Turn one product photo into a 3D Gaussian splat. TripoSplat reconstructs an object's full shape and color from a single image, output as a .spz preview & a downloadable .ply.
Turn a 360 degree panorama into an explorable 3D Gaussian splat scene you can move through. Trajectory planning adds coverage; tune splat density and detail.