Reimagine any track in a new style: describe the target genre, instruments, and vocals, then set how closely the result follows the original. A reference audio can guide the feel.
Models
All Models
Regenerate one section of a track while everything outside stays untouched: rework an intro, swap a chorus, restyle a passage, or fix a flubbed line, duration preserved exactly.
Create a full song from a style prompt and lyrics. Script it with [Verse]/[Chorus] tags, set BPM and key, sing in 50+ languages, and render tracks from 10 seconds to 10 minutes.
Remove backgrounds from images using the transparent-background package. Fast, open-source background removal powered by InSPyReNet.
Create a full song from a style prompt and lyrics. Use [Verse]/[Chorus], set BPM and key, choose 50+ languages, up to 10 min. Open-source, half the cost of ACE-Step 1.5 "Quality".
Regenerate any section while keeping the rest untouched. Rework an intro, swap an instrumental, or restyle a passage. Half the cost of ACE-Step 1.5 "Quality".
Reimagine a track in a new style. Set the genre and vocals, control how much of the original remains, and optionally add reference audio. Half the cost of ACE-Step 1.5 "Quality".
Subject-consistent 720p video with native audio from 1 to 3 reference images and an optional prompt. Your characters, products, or place stays the same across the clip.
Restyle an existing video with a plain-language instruction. Motion is preserved, look changes: season, wardrobe, palette, weather, style. Native audio comes with it.
Generate expressive speech with a prompt-defined voice (age, accent, mood, character...) plus optional audio/image references and controls for speed, pitch, or volume.
Cartwheel Video to Motion converts any video into 3D character animation with markerless mocap, retargeted onto your image or 3D mesh, exported as GLB or FBX.
Play any video backwards in one click, picture and sound reversed together. No prompt needed: drop in a clip and get a clean rewind.
Cartwheel Text to Motion turns a text prompt into 3D character animation, applied to your own image or 3D mesh, exported as GLB or FBX for Blender, Maya, Unreal, or Roblox.
Replace characters in a video with any mascot or custom character, guided by reference images and a prompt. Powered by Cartwheel.
Cartwheel Character Rigging by Cartwheel turns a character image or an existing 3D mesh into an animation-ready 3D model with a built-in skeleton, no manual rigging needed.
Meshy Animation by Meshy rigs and animates 3D characters in one step, applying any motion from Meshy's library of 500+ game-ready clips.
Sync-3 by Sync Labs animates a still portrait to lip-sync any audio track. Works on real photos, illustrations, anime, and 3D characters.
Telestyle V2 by Tele-AI transfers style, materials, and lighting from any reference image onto your content, while keeping composition fully intact.
YVO3D Retexture by YVO3D applies AI-generated textures to any existing GLB model, guided by a reference image for color, material, and style, from 1K up to Ultima 8K.
YVO3D Image to 3D by YVO3D converts a single photo into a textured GLB, or 2-4 multi-view photos into a precise mesh, across four quality tiers.
Happy Horse 1.1 R2V by Alibaba generates videos from up to 9 reference images, keeping characters consistent, with native audio and multilingual lip-sync.
Alibaba's video model that natively synthesizes audio alongside visuals, with multilingual lip-sync, from a text prompt or first-frame image.
ElevenLabs Music v2 by ElevenLabs generates studio-quality music from text: any genre, vocal or instrumental, 3 to 180 seconds, exported as MP3 or Opus.
ElevenLabs Music Advanced v2 by ElevenLabs. Compose full songs section by section, with per-section styles, lyrics, duration, and a context-adherence dial.
Rodin Hyper3D Gen-2.5 Fast by Deemos Technology. Turn 1-5 images into a game-ready 3D model. Choose triangle or quad mesh, PBR or shaded textures, GLB or FBX.
Rodin Hyper3D Gen-2.5 by Deemos: text-to-3D fast lane. Choose triangle or quad mesh, PBR or shaded materials, GLB or FBX. Built for rapid prototyping and iteration.
Uthana Text to Motion 3.0 by Uthana. Describe any motion in text and get a rigged 3D animation, exported as GLB or FBX at 24, 30, or 60 fps.
Riverflow 2.5 Pro by Sourceful. Agentic image model with multi-step reasoning, custom scoring, up to 4K, 10 reference images, and custom font support.
Riverflow 2.5 Fast by Sourceful. Speed-optimized image model for marketing and design, with crisp text, custom fonts, transparent backgrounds, and up to 2K resolution.
Recraft AI's premium utility model for predictable commercial imagery. Flat lighting, tight compositions, six aspect ratios, and custom palette controls.
Recraft V4.1 Utility by Recraft AI. Clean, flat-lit image generation built for product shots, mockups, and graphics that need consistent, predictable results.
Recraft V4.1 SVG by Recraft AI generates real editable SVGs from prompts. Logos, icons, badges, and mascots with clean geometry and accurate text rendering.
The premium vector model for elaborate, multi-element work: detailed crests and mascots, complex isometric scenes, and infographics with many correctly labeled callouts.
Recraft V4.1 Pro by Recraft AI. Premium raster generation with design-forward compositions, photorealism, and reliably legible in-image text for demanding creative work.
Recraft V4.1 by Recraft AI generates design-ready images with accurate in-image text. Six aspect ratios, palette steering, and background color control.
P-Image Try-On by Pruna AI dresses a person photo in a full outfit using up to 11 garment images, from tops and accessories to outerwear.
Meshy UV Unwrap by Meshy auto-generates clean UV maps for any GLB mesh. Upload your model and get a texture-ready, UV-unwrapped GLB back in seconds.
Kling Video to Audio by Kuaishou adds synchronized sound effects and background music to silent video clips of 3 to 20 seconds.
Luma's flagship image model combining reasoning with generation. Up to 9 reference images, web search grounding, 9 aspect ratios, and crisp text rendering.
Luma Uni-1 by Luma Labs generates and edits images with a reasoning model, reliable text rendering, web search grounding, and up to 9 reference images.
Luma Ray 3.2 Edit by Luma Labs. Restyle any video while preserving its motion. Nine edit strengths, HDR output, and face, pose, depth, and trajectory controls.
Luma Ray 3.2 Reframe by Luma Labs converts any video to a new aspect ratio via AI outpainting, filling the added canvas with prompt-guided content.
Luma Ray 3.2 by Luma Labs. Cinematic text or image to video up to 10s and 1080p, with HDR, seamless looping, start/end keyframes, and six aspect ratios.
Add styled captions to any video: Whisper transcription, translation into 18 languages, 7 presets or your own style, burned in or as an SRT file.
Generate seamless 360 equirectangular panoramas from a text prompt, powered by FLUX.1. Choose from 21 styles with automatic seam and pole correction.
Generate seamless 360° skyboxes from a text prompt. Outputs equirectangular panoramas and cubemap formats, engine-ready for real-time scenes.
Sonilo V1.1 by Sonilo analyzes your video's pacing, motion, and mood to generate an original licensed soundtrack automatically.
Sonilo V1.1 by Sonilo. Text to original instrumental music, 1 to 600 seconds. Describe genre, mood, and instruments, then generate up to 3 variations per prompt.
MAI Image 2.5 Edit by Microsoft. Instruction-based editing: swap backgrounds, update text, change styles, or fix lighting. Up to 4 outputs per run.
MAI Image 2.5 by Microsoft generates photorealistic images with reliable in-image text. Ideal for posters, ads, and key art across 11 aspect ratios.