ByteDance audio-video model with camera and native audio controls
A smart video model that balances quality, speed, and cost
MiniMax multimodal text-to-video with native stereo audio, 4–15 second clips, and up to 2K output
ByteDance multimodal video model with image and audio references for clips up to 30 seconds
ByteDance flagship video model, ultra-high quality
ByteDance fast version, speed optimized
Kling 2.6 text-to-video, smooth and natural
Kling 3.0, supports multi-shot & element references
HappyHorse text-to-video, 720p or 1080p
Latest xAI preview model for 1-15 second text-to-video generation
xAI Grok model, supports long videos
MiniMax Hailuo AI text-to-video
Examples
Configure parameters and click Generate