Built for image to video
This model always starts from a picture you provide. That makes it predictable: the character, product or scene in the first frame is the one you uploaded, and the prompt only has to describe what changes.
MiniMaxNew
MiniMax Hailuo H3: expressive image-to-video
Hailuo H3 brings a still image to life with dynamic, expressive motion at 768P or 2K. Upload a start frame (and optionally an end frame) to control the shot.
Try Hailuo H3Hot prompts from the community — open any of them and remix it in one click.
FAQ
Hailuo H3 starts at 80 credits. The exact cost for 768P or 2K and 6 or 10 seconds is shown before you generate, and failed generations are refunded automatically. New accounts receive 50 free credits; Pro includes 1,000 credits per month for $9.90.
Yes. Generated content can be used in ads, online stores and social media, subject to NeoMundo's terms of service and the model provider's policies.
No. A start frame is required on every run. If you only have an idea, create the first image with an image model on NeoMundo, then upload it here.
It fixes the last frame of the clip. The model animates from your first image to the end frame, which is useful for transformations and for landing on a specific pose.
There is no audio setting for this model. If you need a clip with dialogue or ambience generated together, use Veo 3.1; otherwise add music in your editor.
There is no separate aspect ratio control. The start frame defines the shot, so crop your image to 16:9, 9:16 or square before uploading.
The motion is the same idea at both settings. Use 768P to iterate on the prompt quickly and cheaply, then switch to 2K for the final, sharper render.
Upload one image, write one line about how it moves, and get a 6 second clip in minutes. Your free sign-up credits are enough to try it.
This model always starts from a picture you provide. That makes it predictable: the character, product or scene in the first frame is the one you uploaded, and the prompt only has to describe what changes.
Facial expressions, gestures, flowing hair and energetic camera moves are where MiniMax Hailuo stands out. It suits shots that need emotion and energy rather than a static pan across a still.
Add a second image and the clip travels from your first image to that end frame. Use it for transformations, before-and-after reveals or to make sure a character finishes in a specific pose.
768P is the fast, economical setting for testing motion. 2K gives a sharper result for final delivery, large displays and edits that will be cropped or reframed later.
Six seconds covers a single gesture or reaction. Ten seconds leaves space for a camera move plus an action, such as a slow push-in while the subject turns and smiles.
Hailuo H3 is a video model from MiniMax that brings still images to life with dynamic, expressive motion. Unlike text-first models, it is image-only in practice: a start frame is required on every generation, and a text prompt tells the model how that picture should move.
You can upload one or two images. The first image is always the start frame. The optional second image becomes the end frame, so the model animates the path between them. There are no video or audio inputs, and there is no aspect ratio selector: the first image you upload defines the shot, so crop it to the format you want before you begin.
Output is 768P or 2K at 6 or 10 seconds. Hailuo H3 starts at 80 credits, and the exact cost for your chosen resolution and duration appears before you generate. There is no sound setting on this model, so plan to add music or voice in your editor if the clip needs audio.
On NeoMundo the model lives in the same video studio as Veo 3.1 and Kling 3.0. A common workflow is to generate a still with Nano Banana 2 or GPT Image 2 first and animate it here, keeping the whole photo to video process on one credit balance.
Click Try Hailuo H3 on this page or open /create/video and choose it from the model list. New accounts get 50 free credits on sign-up.
Use a sharp JPG, PNG or WebP up to 10 MB with the subject clearly visible. Crop it to the framing you want in the video. Add an end frame only if you need a fixed final pose or composition.
Do not re-describe the image; describe what changes. For example: she laughs and tosses her hair back, petals drift across the frame, the camera slowly pushes in. One main action per clip gives the cleanest result.
Pick 6 or 10 seconds. Test at 768P, then re-run the version you like at 2K. Videos usually take 1–5 minutes, and you can download the result straight from the studio.
Turn a portrait, avatar or illustrated character into a short clip with a smile, a glance or a laugh. Expressive faces are what this model handles with the most character.
Animate a product photo with steam rising from a cup, a bottle catching light, or confetti falling around a box, while the product itself stays exactly as photographed.
Upload a before image and an after image, such as a sketch and the finished painting, or a day scene and the same place at night, and let the model animate the change.
A simple photo to video treatment gives a still scene subtle motion: moving clouds, rippling water, a person turning toward the camera. A 6 second clip is often enough for a social post.