Turn a 360 degree panorama into an explorable 3D Gaussian splat scene you can move through. Trajectory planning adds coverage; tune splat density and detail.
Models
All Models
Turn overlapping photos or a video walkthrough of a scene into an explorable 3D Gaussian splat. No camera rig needed, up to 64 images or one clip.
Turn one photo of a place into a seamless 360 degree equirectangular skybox. Pick a high-fidelity or fast backend, and optionally describe the unseen surroundings.
Turn a prompt, image, 360 panorama, or video into a navigable 3D Gaussian-splat world. Marble 1.1 Plus adds dynamic world sizing and the highest fidelity in the family.
Turn a prompt, image, 360 panorama, or video into a navigable 3D Gaussian-splat world. Marble 1.1 is the balanced default: strong quality with coherent, reliable geometry.
Turn a prompt, image, 360 panorama, or video into a navigable 3D Gaussian-splat world in about a minute. Marble 1.0 Draft is the fastest tier: iterate here, then reproduce the same world at higher quality on 1.1 or 1.1 Plus.
Split any 3D model into clean, semantic parts and regenerate PBR materials in one pass. Adjustable split strength (2 to 12), 2K or 4K textures, guided by a reference image.
Turn a 3D mesh into semantic, editable parts. Tripo Segmentation v2 offers adjustable detail levels (Simple, Balanced, Detailed) and optional reference-image guidance
Automatically split a 3D mesh into editable parts. Tripo Segmentation v1 uses geometry analysis to separate characters, props, and hard-surface models.
Add one new instrument layer onto an existing track, matched to its key, tempo, and groove. Pick the stem and the section, and build arrangements one layer at a time.
Turn a partial track into a full arrangement: feed a sung vocal or a played riff and generate drums, bass, keys, and more around it. Vocal2BGM on the ACE-Step 1.5 edit engine.
Isolate any single stem from a finished track: vocals, drums, bass, guitar, keys, strings, and more. Full-quality ACE-Step 1.5 edit engine, 12 selectable stems, up to 4 variations/run.
Generate a perfectly synced soundtrack for any video. Sonilo reads pacing, mood, and action, composing music that fits while preserving speech and replacing only music.
Turns a text prompt or a single image into a clean, game-ready 3D mesh: organized topology, target polycount control, PBR textures, and optional A-pose or T-pose.
Veed Lipsync v2 by Veed IO re-syncs any talking video to a new audio track: swap the script, the voice, or the language, and the mouth follows. Works on real and stylized faces.
Turn 3D and game renders into photoreal video with LTX 2.3. Optional first-frame anchoring, Strong V2 intensity, Detail Refine, and synced audio at 480p or 720p.
Reimagine any track in a new style: describe the target genre, instruments, and vocals, then set how closely the result follows the original. A reference audio can guide the feel.
Regenerate one section of a track while everything outside stays untouched: rework an intro, swap a chorus, restyle a passage, or fix a flubbed line, duration preserved exactly.
Create a full song from a style prompt and lyrics. Script it with [Verse]/[Chorus] tags, set BPM and key, sing in 50+ languages, and render tracks from 10 seconds to 10 minutes.
Cut out any subject in one click. Open-source InSPyReNet gives clean edges on hair, fur, and glass, output as a transparent PNG, white, green screen, or blurred background.
Create a full song from a style prompt and lyrics. Use [Verse]/[Chorus], set BPM and key, choose 50+ languages, up to 10 min. Open-source, half the cost of ACE-Step 1.5 "Quality".
Regenerate any section while keeping the rest untouched. Rework an intro, swap an instrumental, or restyle a passage. Half the cost of ACE-Step 1.5 "Quality".
Reimagine a track in a new style. Set the genre and vocals, control how much of the original remains, and optionally add reference audio. Half the cost of ACE-Step 1.5 "Quality".
Subject-consistent 720p video with native audio from 1 to 3 reference images and an optional prompt. Your characters, products, or place stays the same across the clip.
Restyle an existing video with a plain-language instruction. Motion is preserved, look changes: season, wardrobe, palette, weather, style. Native audio comes with it.
Generate expressive speech with a prompt-defined voice (age, accent, mood, character...) plus optional audio/image references and controls for speed, pitch, or volume.
Cartwheel Video to Motion converts any video into 3D character animation with markerless mocap, retargeted onto your image or 3D mesh, exported as GLB or FBX.
Play any video backwards in one click, picture and sound reversed together. No prompt needed: drop in a clip and get a clean rewind.
Cartwheel Text to Motion turns a text prompt into 3D character animation, applied to your own image or 3D mesh, exported as GLB or FBX for Blender, Maya, Unreal, or Roblox.
Replace characters in a video with any mascot or custom character, guided by reference images and a prompt. Powered by Cartwheel.
Cartwheel Character Rigging by Cartwheel turns a character image or an existing 3D mesh into an animation-ready 3D model with a built-in skeleton, no manual rigging needed.
Meshy Animation by Meshy rigs and animates 3D characters in one step, applying any motion from Meshy's library of 500+ game-ready clips.
Sync-3 by Sync Labs animates a still portrait to lip-sync any audio track. Works on real photos, illustrations, anime, and 3D characters.
Telestyle V2 by Tele-AI transfers style, materials, and lighting from any reference image onto your content, while keeping composition fully intact.
YVO3D Retexture by YVO3D applies AI-generated textures to any existing GLB model, guided by a reference image for color, material, and style, from 1K up to Ultima 8K.
YVO3D Image to 3D by YVO3D converts a single photo into a textured GLB, or 2-4 multi-view photos into a precise mesh, across four quality tiers.
Happy Horse 1.1 R2V by Alibaba generates videos from up to 9 reference images, keeping characters consistent, with native audio and multilingual lip-sync.
Alibaba's video model that natively synthesizes audio alongside visuals, with multilingual lip-sync, from a text prompt or first-frame image.
ElevenLabs Music v2 by ElevenLabs generates studio-quality music from text: any genre, vocal or instrumental, 3 to 180 seconds, exported as MP3 or Opus.
ElevenLabs Music Advanced v2 by ElevenLabs. Compose full songs section by section, with per-section styles, lyrics, duration, and a context-adherence dial.
Rodin Hyper3D Gen-2.5 Fast by Deemos Technology. Turn 1-5 images into a game-ready 3D model. Choose triangle or quad mesh, PBR or shaded textures, GLB or FBX.
Rodin Hyper3D Gen-2.5 by Deemos: text-to-3D fast lane. Choose triangle or quad mesh, PBR or shaded materials, GLB or FBX. Built for rapid prototyping and iteration.
Uthana Text to Motion 3.0 by Uthana. Describe any motion in text and get a rigged 3D animation, exported as GLB or FBX at 24, 30, or 60 fps.
Riverflow 2.5 Pro by Sourceful. Agentic image model with multi-step reasoning, custom scoring, up to 4K, 10 reference images, and custom font support.
Riverflow 2.5 Fast by Sourceful. Speed-optimized image model for marketing and design, with crisp text, custom fonts, transparent backgrounds, and up to 2K resolution.
Recraft AI's premium utility model for predictable commercial imagery. Flat lighting, tight compositions, six aspect ratios, and custom palette controls.
Recraft V4.1 Utility by Recraft AI. Clean, flat-lit image generation built for product shots, mockups, and graphics that need consistent, predictable results.
Recraft V4.1 SVG by Recraft AI generates real editable SVGs from prompts. Logos, icons, badges, and mascots with clean geometry and accurate text rendering.
The premium vector model for elaborate, multi-element work: detailed crests and mascots, complex isometric scenes, and infographics with many correctly labeled callouts.
Recraft V4.1 Pro by Recraft AI. Premium raster generation with design-forward compositions, photorealism, and reliably legible in-image text for demanding creative work.