Fast multilingual text-to-speech from Alibaba Qwen Audio 3.0 (Flash): 45 preset voices across 11 languages, with natural delivery for voiceover, narration, ads, and game dialogue.
Models
All Models
Studio-quality product shots without the studio. Pixelcut cuts your product out of any photo and sets it on a solid color, transparent, or custom background, with margins, shadows, and an optional watermark.
Microsoft's highest-fidelity instruction-guided image editor. Swap text, recolor, restyle, replace objects, and change lighting from a plain prompt while preserving structure.
Microsoft's flagship text-to-image model for hero imagery with standout legible typography. Photoreal and stylized, 8 aspect ratios, up to 4 variations per prompt.
Turn one product photo into a 3D Gaussian splat. TripoSplat reconstructs an object's full shape and color from a single image, output as a .spz preview & a downloadable .ply.
Turn a 360 degree panorama into an explorable 3D Gaussian splat scene you can move through. Trajectory planning adds coverage; tune splat density and detail.
Turn overlapping photos or a video walkthrough of a scene into an explorable 3D Gaussian splat. No camera rig needed, up to 64 images or one clip.
Turn one photo of a place into a seamless 360 degree equirectangular skybox. Pick a high-fidelity or fast backend, and optionally describe the unseen surroundings.
Turn a prompt, image, 360 panorama, or video into a navigable 3D Gaussian-splat world. Marble 1.1 Plus adds dynamic world sizing and the highest fidelity in the family.
Turn a prompt, image, 360 panorama, or video into a navigable 3D Gaussian-splat world. Marble 1.1 is the balanced default: strong quality with coherent, reliable geometry.
Turn a prompt, image, 360 panorama, or video into a navigable 3D Gaussian-splat world in about a minute. Marble 1.0 Draft is the fastest tier: iterate here, then reproduce the same world at higher quality on 1.1 or 1.1 Plus.
Split any 3D model into clean, semantic parts and regenerate PBR materials in one pass. Adjustable split strength (2 to 12), 2K or 4K textures, guided by a reference image.
Turn a 3D mesh into semantic, editable parts. Tripo Segmentation v2 offers adjustable detail levels (Simple, Balanced, Detailed) and optional reference-image guidance
Automatically split a 3D mesh into editable parts. Tripo Segmentation v1 uses geometry analysis to separate characters, props, and hard-surface models.
Add one new instrument layer onto an existing track, matched to its key, tempo, and groove. Pick the stem and the section, and build arrangements one layer at a time.
Turn a partial track into a full arrangement: feed a sung vocal or a played riff and generate drums, bass, keys, and more around it. Vocal2BGM on the ACE-Step 1.5 edit engine.
Isolate any single stem from a finished track: vocals, drums, bass, guitar, keys, strings, and more. Full-quality ACE-Step 1.5 edit engine, 12 selectable stems, up to 4 variations/run.
Generate a perfectly synced soundtrack for any video. Sonilo reads pacing, mood, and action, composing music that fits while preserving speech and replacing only music.
Turns one image (or up to 4 views) into a clean, game-ready 3D mesh: organized topology, target polycount control, PBR textures, and optional A-pose or T-pose.
Veed Lipsync v2 by Veed IO re-syncs any talking video to a new audio track: swap the script, the voice, or the language, and the mouth follows. Works on real and stylized faces.
Turn 3D and game renders into photoreal video with LTX 2.3. Optional first-frame anchoring, Strong V2 intensity, Detail Refine, and synced audio at 480p or 720p.
Reimagine any track in a new style: describe the target genre, instruments, and vocals, then set how closely the result follows the original. A reference audio can guide the feel.
Regenerate one section of a track while everything outside stays untouched: rework an intro, swap a chorus, restyle a passage, or fix a flubbed line, duration preserved exactly.
Create a full song from a style prompt and lyrics. Script it with [Verse]/[Chorus] tags, set BPM and key, sing in 50+ languages, and render tracks from 10 seconds to 10 minutes.
Cut out any subject in one click. Open-source InSPyReNet gives clean edges on hair, fur, and glass, output as a transparent PNG, white, green screen, or blurred background.
Create a full song from a style prompt and lyrics. Use [Verse]/[Chorus], set BPM and key, choose 50+ languages, up to 10 min. Open-source, half the cost of ACE-Step 1.5 "Quality".
Regenerate any section while keeping the rest untouched. Rework an intro, swap an instrumental, or restyle a passage. Half the cost of ACE-Step 1.5 "Quality".
Reimagine a track in a new style. Set the genre and vocals, control how much of the original remains, and optionally add reference audio. Half the cost of ACE-Step 1.5 "Quality".
Subject-consistent 720p video with native audio from 1 to 3 reference images and an optional prompt. Your characters, products, or place stays the same across the clip.
Restyle an existing video with a plain-language instruction. Motion is preserved, look changes: season, wardrobe, palette, weather, style. Native audio comes with it.
Generate expressive speech with a prompt-defined voice (age, accent, mood, character...) plus optional audio/image references and controls for speed, pitch, or volume.
Cartwheel Video to Motion converts any video into 3D character animation with markerless mocap, retargeted onto your image or 3D mesh, exported as GLB or FBX.
Play any video backwards in one click, picture and sound reversed together. No prompt needed: drop in a clip and get a clean rewind.
Cartwheel Text to Motion turns a text prompt into 3D character animation, applied to your own image or 3D mesh, exported as GLB or FBX for Blender, Maya, Unreal, or Roblox.
Replace characters in a video with any mascot or custom character, guided by reference images and a prompt. Powered by Cartwheel.
Cartwheel Character Rigging by Cartwheel turns a character image or an existing 3D mesh into an animation-ready 3D model with a built-in skeleton, no manual rigging needed.
Meshy Animation by Meshy rigs and animates 3D characters in one step, applying any motion from Meshy's library of 500+ game-ready clips.
Sync-3 by Sync Labs animates a still portrait to lip-sync any audio track. Works on real photos, illustrations, anime, and 3D characters.
Telestyle V2 by Tele-AI transfers style, materials, and lighting from any reference image onto your content, while keeping composition fully intact.
YVO3D Retexture by YVO3D applies AI-generated textures to any existing GLB model, guided by a reference image for color, material, and style, from 1K up to Ultima 8K.
YVO3D Image to 3D by YVO3D converts a single photo into a textured GLB, or 2-4 multi-view photos into a precise mesh, across four quality tiers.
Happy Horse 1.1 R2V by Alibaba generates videos from up to 9 reference images, keeping characters consistent, with native audio and multilingual lip-sync.
Alibaba's video model that natively synthesizes audio alongside visuals, with multilingual lip-sync, from a text prompt or first-frame image.
ElevenLabs Music v2 by ElevenLabs generates studio-quality music from text: any genre, vocal or instrumental, 3 to 180 seconds, exported as MP3 or Opus.
ElevenLabs Music Advanced v2 by ElevenLabs. Compose full songs section by section, with per-section styles, lyrics, duration, and a context-adherence dial.
Rodin Hyper3D Gen-2.5 Fast by Deemos Technology. Turn 1-5 images into a game-ready 3D model. Choose triangle or quad mesh, PBR or shaded textures, GLB or FBX.
Rodin Hyper3D Gen-2.5 by Deemos: text-to-3D fast lane. Choose triangle or quad mesh, PBR or shaded materials, GLB or FBX. Built for rapid prototyping and iteration.
Uthana Text to Motion 3.0 by Uthana. Describe any motion in text and get a rigged 3D animation, exported as GLB or FBX at 24, 30, or 60 fps.