Minimax Music 2.5 by MiniMax generates full songs from lyrics and a style prompt, with 14 arrangement tags, vocal and instrumental modes, up to 44.1kHz / 256kbps.
Models
All Models
Restyle and transform any video with a text prompt using LTX-2 19B by Lightricks. IC-LoRA structure control, first-frame conditioning, up to 4K output.
LTX-2 19b by Lightricks. Animate a sequence of keyframe images into smooth video with synced audio, up to 2160p and 20 seconds.
Vidu Q1 by Shengshu Technology. Generate 5-second 1080p videos from text in general or anime style, with control over motion intensity and aspect ratio.
LTX-2 19B Fast by Lightricks generates text-to-video and image-to-video with synchronized audio, optimized for rapid iteration at 8 inference steps.
LTX-2 19b Extend Video by Lightricks adds up to 20 seconds of continuation to any video, with synchronized audio and optional camera motion control.
The standard design-centric model for high-quality raster image generation with professional artistic flair.
Pika 2.2 T2V by Pika Labs generates cinematic video from text prompts at 720p or 1080p, in 5 or 10 second clips, across landscape, square, or portrait formats.
Pika 2.2 Scenes by Pika Labs makes video from up to 10 reference images in one scene. Choose 5 or 10 seconds, 720p or 1080p, and precise or creative blending.
Pika 2.2 I2V by Pika Labs animates a still image into a 5 or 10-second video at up to 1080p, with strong character and style consistency.
Pika 2.2 Frames by Pika Labs generates video transitions between 2 to 5 keyframe images. Up to 25 seconds, per-transition prompts, 720p or 1080p output.
Minimax Music 2.0 by MiniMax generates full songs with vocals and instruments from a text prompt and lyrics, producing tracks up to 5 minutes long.
LTX-2 Retake by Lightricks re-renders any segment of a video, replacing visuals, audio, or both, without touching the rest of the clip.
Gemini 2.5 Flash Image by Google edits and generates images from text instructions, with up to 6 reference images and aspect ratios from ultrawide to portrait.
GPT Image 1 by OpenAI. Edit or generate images from text instructions, with up to 10 reference images and controls for fidelity, quality, and background.
Seedream 4.0 by ByteDance edits or generates images up to 4K using text prompts, with support for up to 14 reference images and sequence output.
Runway Gen 4 by Runway generates and edits images from up to 3 reference images and a text instruction, keeping characters, objects, and style consistent across results.
Meshy Image-to-3D by Meshy converts one photo or up to four multiview images into a textured, PBR-ready 3D mesh with full geometry coverage.
Meshy Retexture by Meshy applies new textures to existing 3D models from a text prompt or reference image, with optional full PBR map generation.
Minimax Speech 2.6 (HD) by MiniMax. Studio-quality text-to-speech in 40+ languages, with 17 built-in voices, emotion control, and adjustable speed, pitch, and volume.
Minimax Speech 2.6 (Turbo) by MiniMax. Fast, low-latency text-to-speech in 40+ languages with 17 voices, emotion control, and pitch, speed, and volume adjustment.
BiRefNet v2 removes video backgrounds frame by frame, preserving fine details like hair and transparent edges. Exports with true transparency via WebM or ProRes.
Seedance 1 (Pro Fast) by ByteDance generates 1080p cinematic video from text or a first-frame image, at 3x the speed of the standard Pro model.
LTX-2 (Fast) by Lightricks generates text and image-to-video clips up to 4K and 20 seconds, optimized for rapid iteration and quick preview workflows.
LTX-2 (Pro) by Lightricks generates video from text or a first frame at up to 4K, with synchronized audio produced in the same pass. Choose 6, 8, or 10 second clips.
Kling 2.5 I2V (Standard) by Kuaishou turns a still image into a 720p video clip. Choose 5 or 10 seconds and tune the guidance dial for creativity or faithfulness.
SORA 2 (Pro) by OpenAI. Cinematic video from text or a first-frame image. Up to 20 seconds, 1080p, with synced audio. Landscape or portrait.
Rodin Gen-2 by Deemos Technology generates 3D models from text or images, with quad and triangle mesh quality options and PBR or shaded materials.
Kling 2.5 T2V (Pro) by Kuaishou. Text-to-1080p video at Turbo speed. Choose 5 or 10 seconds in 16:9, 9:16, or 1:1. Strong prompt adherence and dynamic motion.
Kling 2.5 I2V (Pro) by Kuaishou turns a still image into a smooth 1080p video clip, with optional last-frame anchoring for precise start-to-end control.
Google Lyria 2 generates high-fidelity instrumental music from text prompts, producing up to 30-second clips at 48kHz stereo across a wide range of genres.
Kling 2.1 (Pro) by Kuaishou turns images into 1080p video with optional first and last frame anchoring for controlled scene transitions.
Luma Video Reframe by Luma Labs resizes any video across 7 aspect ratios by outpainting beyond the original frame, preserving your subject.
Wan 2.2 T2V by Alibaba is an open-source text-to-video model with a Mixture-of-Experts architecture. Generate 480p or 720p clips with motion speed and seed controls.
Wan 2.2 I2V by Alibaba. Open-source image-to-video with first and last frame conditioning, 480p or 720p output, and a motion speed control.
Runway Gen4 Turbo by Runway ML is a fast image-to-video model. Animate a first frame into clips of 2 to 10 seconds across six aspect ratios, from 21:9 to 9:16.
PartCrafter by PKU turns a single image into up to 16 separate, semantically distinct 3D meshes in one pass, no segmentation required.
Minimax Video 02 by MiniMax generates realistic video from text or images, with natural motion, physics accuracy, and first and last frame anchoring.
Minimax Image 01 by MiniMax generates photorealistic images with advanced lighting, natural skin rendering, and optional character reference support.
Luma Photon by Luma AI generates images with strong prompt adherence across seven aspect ratios, with reference controls for content, style, and character.
Luma Photon Flash by Luma Labs is a fast text-to-image model built for rapid concept iteration, with reference image, style, and character guidance controls.
Ideogram 3 (Quality) by Ideogram generates images with precise text rendering and complex layouts. Supports inpainting, up to 4 style references, and 60+ style presets.
Ideogram 3 (Balanced) by Ideogram generates images and handles inpainting with strong text rendering, 60+ style presets, and up to 4 style reference images.
Ideogram 3 (Turbo) by Ideogram. Fast text-to-image with inpainting, 15 aspect ratios, 60+ style presets, up to 4 style references, and multilingual Magic Prompt.
Rodin Gen-1 (HighPack) by Deemos Technology generates detailed 3D models from text or images, with 4K textures and up to 500k polygon geometry.
Rodin Gen-1 by Deemos Technology generates textured 3D models from text or images, with PBR materials, multi-view input, pose control, and quality settings.
Kling 2.1 (Master) by Kuaishou generates 1080p video from text or a first frame, with advanced 3D motion, cinematic camera work, and refined facial detail.