What MiniMax built
Text-to-video and image-to-video, with generation lengths comparable to the rest of the field.
Its noted strength is motion quality — camera movement that behaves like a camera rather than a drifting viewport, and human action that does not dissolve mid-gesture.
Worth being specific about why that matters. A viewer cannot articulate what is wrong with implausible motion, but they register it immediately: a walk cycle that floats, a gesture that reverses, a camera that moves without parallax. Fixing those is worth more to perceived quality than resolution, and it is the axis on which this model competes.
Subject reference, and the consistency problem
Subject reference features aim to keep a specific person or object recognisable across generations.
This is the hardest unsolved problem in the category, and every vendor’s attempt at it should be read as a mitigation rather than a solution. Faces drift. The further a generation moves from the reference pose, framing and lighting, the more it drifts.
What it does achieve is enough consistency for a viewer to accept two shots as the same subject when they are not scrutinised side by side — which is sufficient for social content and insufficient for anything where the character is the point.
Working from a still
Image-to-video is the more controllable path, as everywhere in this category.
Generate or supply the still, approve it, then direct the motion. That separates composition from animation instead of asking the model to solve both, and it is the difference between a usable hit rate and burning credits on compositions you would have rejected anyway.
Access and pricing
- A browser and an account. A free tier with daily credits, generous enough for real evaluation.
- Priced well below the Western tools, which alongside Kling is why both are used at volume.
- Nothing local; no hardware requirement.
- Queue waits at busy times, longer on free credits.
- An API is available through MiniMax.
The same question Kling raises
MiniMax is a Chinese company and generation runs on its infrastructure. Prompts and uploaded images are processed there.
For most creative work that is no different from any other hosted service. Where it matters is contractual: client agreements specifying processing location, data residency requirements, or work for regulated and public-sector clients. Settle it before uploading material, not afterwards.
Content policy also differs from Western tools in both directions, and the boundaries are not always predictable.
Choosing between Hailuo and Kling
These two are close enough that a general recommendation is not useful, and the comparison is worth running yourself.
Both have free daily credits. Take one source still and one prompt, run them on both, and judge the motion — which is what both compete on. The result depends on your subject matter more than on any benchmark: models differ in what they handle well, and human figures, animals, vehicles and camera moves are not equally represented.
Twenty minutes of testing gives a better answer than any comparison article, this one included.
Its strengths
- Motion quality, particularly human movement and camera work.
- Cost, well below the Western alternatives.
- A genuinely usable free tier with daily renewal.
- Strong image-to-video, the controllable path.
- Subject reference, a partial answer to the consistency problem.
Its weaknesses
- Data residency — processing on Chinese infrastructure.
- Thin English documentation and support.
- Queue waits, especially without paying.
- Short clips and unreliable continuity, as everywhere.
- Subject reference drifts away from the reference conditions.
- No editing tooling around the generation.
What stays with you
Downloaded clips only. Prompts, reference images and history stay in the platform.
Keep the source still and prompt for anything you may need to match later. Given that continuity across generations is already unreliable, losing the exact inputs means a shot cannot be approximated again — here or anywhere else.
The direct comparisons
- Kling AI — the closest equivalent, and the comparison worth running yourself.
- Luma Dream Machine — faster, lower fidelity, Western-hosted.
- Runway ML — the toolkit, at several times the cost.
- Google Veo — comparable quality with generated audio.
Compiled from MiniMax’s documentation and public sources. We have not hands-on tested this tool. Last reviewed 16 August 2026.