Minimax / minimax-h3
MiniMax H3
MiniMax H3 for text, image, frame-pair, and multimodal reference video with native stereo audio.
Products / Models
Browse video and image models on MachGen - including Text to Video, Image to Video, Reference to Video, Video Upscaling, Text to Image, and Image Editing.
Showing 37 of 39 models
Minimax / minimax-h3
MiniMax H3 for text, image, frame-pair, and multimodal reference video with native stereo audio.
ByteDance / seedance-2.5
Seedance 2.5 for 30-second native video with up to 50 multimodal references and directed editing.
Lightricks / ltx-2.3-pro
LTX 2.3 Pro for cinematic 4K video, directorial control, and synchronized native audio.
Vidu / q3-turbo
Vidu Q3 Turbo for fast, balanced video iteration across text, image, reference, and frame-pair workflows.
Wan / wan-2.2-i2v-a14b
Wan 2.2 A14B for rich visuals and detailed motion from text or stills, served on low-cost managed inference.
ByteDance / seedance-2.0
Seedance 2.0 for cinematic multimodal video, coherent multi-shot storytelling, and synced native audio.

Kuaishou / kling-o3
Compose video from prompts, reference images, and Kling elements
Kuaishou / kling-v3-video
Kling Video 3.0 for cinematic text-to-video and image-to-video with optional native audio.
Alibaba / wan-3.0
Wan 3.0 for production-scale video generation with 30-second shots, multimodal direction, and synchronized sound.
Vidu / q3-r2v
Vidu Q3 for high-quality reference-to-video with look-locked subjects, prompt-directed camera motion, and optional audio.
Vidu / q3
Vidu Q3 Pro for high-fidelity text, image, and frame-pair video with optional synchronized audio.
Vidu / q3-pro-fast
Vidu Q3 Pro Fast for low-latency, high-quality image animation with prompt-directed camera motion and optional audio.
Google / veo-3.1
Google Veo 3.1 Standard for high-quality cinematic text-to-video with synchronized native audio.
Google / veo-3.1-fast
Google Veo 3.1 Fast for cinematic native-audio video with quicker, lower-cost prompt iteration.
PixVerse / pixverse-v6
PixVerse V6 for cinematic, stylized video with references, prompt-directed camera motion, and optional audio.
PixVerse / pixverse-c1
PixVerse C1 for cinematic action, VFX-heavy scenes, and reference-guided visual continuity.
Alibaba / happyhorse-1.1
HappyHorse 1.1 for high-quality general video and stronger multi-reference subject consistency.
Alibaba / happyhorse-1.0
HappyHorse 1.0 for balanced-cost, general-purpose video from text, a first frame, or references.
xAI / grok-imagine-video
Grok Imagine Video for fast, expressive text-to-video concepts and social-ready visual exploration.
xAI / grok-imagine-video-1.5
Grok Imagine Video 1.5 for realistic still-image motion, object interaction, and synchronized generated audio.
Topaz / video-precision
Topaz Precision for detail-preserving video upscaling to 1080p, 2K, 4K, or 8K.
Topaz / video-precision-animation
Topaz Precision Animation for upscaling hand-drawn and cel animation.
Topaz / video-generative
Topaz Generative for diffusion-based video restoration up to 4K.
Topaz / video-generative-fast
Topaz Generative Fast for cheaper generative video restoration on longer clips.

Black Forest Labs / flux-2-dev
FLUX.2 Dev for high-quality open-weight image generation, multi-image editing, and fast experimentation.

Hidream / hidream-o1
HiDream O1 for high-detail text-to-image across fantasy, illustration, and photoreal visual styles.

xAI / grok-imagine-image
Grok Imagine Image for fast, low-cost text-to-image with a bold, expressive visual style.

xAI / grok-imagine-image-quality
Grok Imagine Image Quality for higher-fidelity generation, cleaner detail, and reference-based editing.

Google / nano-banana-2
Nano Banana 2 for fast, high-fidelity Gemini image generation, natural-language edits, and clear typography.

Google / nano-banana-pro
Nano Banana Pro for high-quality Gemini image generation, complex briefs, typography, and controlled edits.

ByteDance / seedream-5.0-lite
Seedream 5.0 Lite for efficient high-resolution image creation, layout-aware prompts, and reference edits.

OpenAI / gpt-image-2
GPT Image 2 for high-fidelity image generation, precise instruction following, typography, and revisions.

Topaz / image-precision
Topaz Image Precision for faithful 2x or 4x image upscaling.

Topaz / image-generative
Topaz Image Generative for prompt-steerable image upscaling.
Elevenlabs / eleven-v3
ElevenLabs v3 for expressive speech and multi-speaker dialogue generated in one pass.
Elevenlabs / eleven-sfx-v2
ElevenLabs SFX v2 for foley and designed sound effects written as a description.
Elevenlabs / eleven-music-v2
ElevenLabs Music v2 for instrumental and vocal tracks, from a prompt or a section plan.