Retro Diffusion Tile by Retro-Diffusion generates seamless pixel art tilesets and game assets in six modes, from full tilesets to placeable objects.
Models
All Models
Retro Diffusion Plus by Retro-Diffusion generates crisp, grid-aligned pixel art in 18 styles, from isometric assets to Minecraft textures.
Retro Diffusion Animation by Retro-Diffusion generates pixel art character walk cycles, idle states, small sprites, and VFX effects as animated GIFs or spritesheets.
FLUX 2 (Dev) by Black Forest Labs is an open-weight image model built for fine-tuning. Stack up to 6 LoRAs, add up to 5 reference images, and generate up to 2048x2048.
Meshy Rigging by Meshy adds a skeleton and skin weights to humanoid GLB characters automatically, making them animation-ready in one step.
Hunyuan 3D Part by Tencent segments an existing GLB mesh into clean, organized sub-parts. Especially effective on mechanical and hard-surface models, ready for editing.
Gemini 3.0 Pro by Google. Instruction-based image generation and editing with advanced reasoning, multi-image fusion, and outputs up to 4K resolution.
Flash VSR by Tsinghua/Shanghai AI Lab upscales video 2x to 4x using one-step diffusion, delivering sharp detail restoration with real-time streaming speed.
Pixverse Swap by PixVerse replaces people or backgrounds in any video using a reference image, with Person and Background modes up to 720p.
Hunyuan 3D 3.0 Pro (Sketch) by Tencent converts hand-drawn sketches into textured 3D meshes with adjustable face count and optional PBR materials.
Hunyuan 3D Pro 3.0 by Tencent converts a single image into a detailed 3D mesh with up to 1.5M faces, PBR textures, and a choice of triangle or quad topology.
Hunyuan 3D 3.0 Pro (Multiview) by Tencent generates 3D assets from up to 4 reference views, with PBR materials, face count control, and tri or quad topology.
Qwen Edit Multi-Angle by Alibaba edits images using camera controls: rotate, zoom, vertical tilt, and wide-angle, with optional text prompts.
Gemini 2.5 Flash Image by Google edits and generates images from text instructions, with up to 6 reference images and aspect ratios from ultrawide to portrait.
Flux Kontext by Black Forest Labs. Edit images with text instructions, preserving characters and style across local or global edits. Supports up to 10 reference images.
GPT Image 1.5 by OpenAI. Instruction-driven image editing with precise reference fidelity, transparent background support, and aspect ratio control.
Seedream 4.5 by ByteDance edits images from natural language instructions, with up to 10 reference images, 4K output, and strong subject preservation.
P-Image Edit by Pruna AI is a fast, instruction-driven image editor. Edit, composite, or transform images using text and up to 10 reference images.
Meshy Remesh by Meshy rebuilds your 3D model's geometry with clean triangle or quad topology and a target polycount you control, from 100 to 300,000 polygons.
Meshy Image-to-3D by Meshy converts one photo or up to four multiview images into a textured, PBR-ready 3D mesh with full geometry coverage.
Meshy Text-to-3D by Meshy generates textured, game-ready 3D assets from a text prompt using Meshy 5 or 6.
Meshy Retexture by Meshy applies new textures to existing 3D models from a text prompt or reference image, with optional full PBR map generation.
Flux.2 [klein] 4B by Black Forest Labs. A 4B-parameter model distilled for sub-second image generation and editing, with LoRA support.
FLUX.2 [klein] 4B Base by Black Forest Labs. Compact, undistilled 4B image model built for LoRA compatibility, fine-tuning, and precise prompt control.
Flux 2 (Turbo) Edit by Black Forest Labs: fast, instruction-based image editing. Describe the change, get a transformed image in seconds.
FLUX.2 [klein] 9B Base by Black Forest Labs. Undistilled 9B foundation model for text-to-image and image-to-image, built for LoRA fine-tuning and high output diversity.
FLUX 2 (Max) by Black Forest Labs. Top-tier image editing with up to 8 reference images, up to 4 MP output, and precise consistency across colors, faces, and objects.
FLUX 2 (Flex) by Black Forest Labs. Precision image editing with up to 10 reference images, up to 4 MP output, and strong typography and detail control.
FLUX 2 (Pro) Edit by Black Forest Labs. Instruction-based image editor with up to 8 reference images, 4 MP output, and flexible aspect ratios.
LongCat Image by Meituan edits photos via plain-language instructions, no masks needed. Supports 15 edit types and accurate text rendering in Chinese and English.
Sparc3D (Portrait) by Hitem3D converts 1-4 face photos into detailed 3D head models, up to 1536³ Pro resolution, as mesh-only or fully textured assets.
Minimax Hailuo 2.3 (Fast) by MiniMax. Text-to-video and image-to-video with optimized speed. Choose 6 or 10-second clips at 768p, or 6 seconds at 1080p.
Minimax Hailuo 2.3 by MiniMax generates text and image-to-video at 768p or 1080p, with strong motion coherence, temporal consistency, and anime style support.
A Flux Kontext LoRA that ages clean building images into worn, crumbling ruins, preserving the original structure, layout, and perspective.
Turns a character image into a 9-emotion expression sheet, preserving their look and art style.
Reve Remix by Reve AI combines up to 6 reference images into a single result using a text prompt, with 7 aspect ratios to choose from.
Clarity Crystal Upscaler by Clarity AI enhances images with an adjustable scale factor and a creativity dial, optimized for portraits, faces, and product photos.
Flux Kontext LoRA that transforms 3D blockout shapes into photorealistic renders, preserving the original geometry and spatial layout.
Veo 3.1 (Fast) by Google. Fast text-to-video and image-to-video with native audio, first/last frame anchoring, and subject-consistent generation from reference images.
Generate a four-view character turnaround sheet from any character image, showing every side on a clean background.
Flux Kontext LoRA that converts photos of buildings into isometric 3D models placed on square tiles, ideal for game assets and top-down scene design.
Hunyuan Image 3 by Tencent. An 80B MoE text-to-image model with exceptional multilingual text rendering, complex scene understanding, and photorealistic output.
SeedVR2 - Image Upscale by Bytedance scales images up to 4K using one-step diffusion super-resolution, with factor or target resolution modes and a noise slider.
SeedVR2 - Video Upscale by ByteDance upscales videos up to 4K using one-step diffusion, restoring fine detail and temporal consistency with minimal hallucination.
Omni Human 1.5 by ByteDance animates a single photo into a talking avatar, syncing lip movements, expressions, and natural gestures to your audio.
Wan 2.5 I2V by Alibaba animates any image into a 720p or 1080p video clip of 5 or 10 seconds, with optional audio synchronization.
Wan 2.5 T2V by Alibaba generates 720p or 1080p videos from text, in 16:9 or 9:16, with synchronized audio support, in 5 or 10 second clips.