Two video models landed in ComfyUI mid this year, MiniMax H3, the model behind Hailuo 3, launched on 31 July 2026 and gained native ComfyUI support with open weights on 3 August. Lightricks released LTX 2.5 open weights on 11 August, with ComfyUI workflows. Both generate video and sound together, both run locally, and both accept a still image as the starting frame.
We ran LTX 2.5 vs MiniMax H3 through ComfyUI image to video tests in the studio to see how each one handles the jobs we actually get asked to do: a person holding a product, a stylised 3D character and an anime character. Below are the clips, the settings, and what we would use each model for.
How We Set Up the ComfyUI Image to Video Test
Every clip starts from a single still image loaded as the first frame, with a written prompt describing the motion. The recordings below show the full ComfyUI graph, so you can see the source image on the left, the prompt in the image to video node, and the finished clip playing in the Save Video node.
Both models ran on our own localhosted ComfyUI , on our 9950X3D, 5090 workstation in our studio. ComfyUI is a privately hosted app: the model weights, the source images and the finished clips all stay on local hardware. Nothing in these tests was uploaded to a cloud video generator.
MiniMax H3 ran on the official image to video template: the pruned int8 FL2VA diffusion model, the Qwen3-VL 32B text encoder, separate video and audio VAEs, the res_multistep sampler at 20 steps, 24 fps, and a five second duration. We kept it at 0.4 megapixels (640 x 640) as a quick draft pass. LTX 2.5 ran through its image to video node at 0.9 megapixels (960 x 960). That means LTX had more than twice the pixels to work with in the head-to-head, which is worth keeping in mind when you compare sharpness.
LTX 2.5 vs MiniMax H3: Same Image, Same Prompt
The direct comparison uses a skincare shot: a model in a cream shirt holding a blue cleanser tube to camera. Both models got the same prompt: the woman looks at the facial cleanser in her hand, smiles, and looks at the camera, with a locked-off tripod shot and no camera movement.
LTX 2.5 keeps the frame almost identical to the source. The camera stays locked, her face stays on-model, and the tube stays upright and facing the lens, so the label remains readable from start to finish. The performance is gentle: a head tilt, a warm smile, the tube lifted slightly toward camera. What it does not do clearly is the first beat of the prompt. She never really looks down at the product.
MiniMax H3 acts the prompt out in order. She turns her head to look at the tube, brings her second hand up to hold it, smiles at it, then returns her eyes to the camera. It reads like a direction given to a real model on set. The trade-off is the packshot: as she handles the tube it turns side-on, and the label is no longer square to camera for most of the clip.
For a product film this is the real decision. If the brand needs the pack front and centre, LTX 2.5 gave us the safer result. If the brief is about a person reacting to the product, MiniMax H3 gave the more believable performance.
MiniMax H3 on Character Acting
Because MiniMax H3 followed direction so closely, we pushed it further with three character tests and one very short prompt, all at the same 0.4 megapixel draft setting.
Stylised 3D character: a smile that builds
A 3D-style character with a blue and pink hair in a bob, with a short prompt. The expression builds from a closed-mouth smile to a full laugh with the eyes squeezing shut, and the hair, choker and pendant stay consistent with the source render throughout.
Timed beats: a sneeze in four stages
Here the prompt was written like an animator’s timing sheet, with timestamps: the sneeze builds from 0 to 2 seconds, holds at the peak, releases at 2.5 seconds with the head snapping forward, then settles into an embarrassed look with messy hair. MiniMax H3 hit each stage in sequence, including the hair follow-through after the snap. Writing prompts as timed beats is the single most useful habit we found with this model.
1990s cel anime: comedy timing
A male anime character in glasses, and a prompt asking for classic anime comedy acting: a smug chin stroke, a snap upright as an idea lands, the lenses flashing white, then the realisation that the idea is wrong, with a sweat drop and the glasses sliding down his nose. The model held the flat colours and limited animation look rather than drifting toward 3D, and it delivered the held poses that make this style read.
A one-line prompt: spin and hair motion
The last test was a simple prompt, a manga woman spins around, her hair flowing in the wind. She turns fully away from camera, the hair sweeps across the frame, and when her face comes back round it is still the same character. Holding identity through a full turn is where many image to video models slip, and H3 handled it cleanly on a 0.4 megapixel draft.
Where LTX 2.5 Pulls Ahead
Our tests favour consistency & acting, which is MiniMax H3’s strength. LTX 2.5 is built around different priorities. It is a 22 billion parameter model that outputs up to 4K with HDR, at up to 50 fps, and its native multi-shot mode produces several connected shots in one generation while holding the character, lighting and voice across the cuts. MiniMax H3 tops out at 2K and 15 seconds a clip.
LTX 2.5 also comes in a dev model and a smaller distilled model, with image to video, text to video and first and last frame workflows in ComfyUI. For work that ends up on a large event screen, or where the brief is a sequence of shots rather than a single performance, that resolution and multi-shot control matter more than acting nuance.
Running LTX 2.5 and MiniMax H3 in ComfyUI
Both models load through ComfyUI’s native nodes, and both ship official templates you can open from the workflow browser. MiniMax H3 offers image to video, text to video and reference to video. The reference mode accepts images, video clips and audio at once, which is useful when a character, a camera move and a voice all need to match an existing brand asset.
MiniMax H3 is heavy at full precision. The released weights prune around 40 percent of the parameters and use int8 quantisation, which brings the footprint down from 123.6 GB to 42.5 GB. With ComfyUI’s dynamic VRAM offloading, it will run on a consumer 5090 graphics card, but expect long waits on smaller cards. A low megapixel draft pass, like the ones above, is the practical way to test motion before committing to a full resolution render. We cover how AI video fits into a wider production schedule in our post on how AI video is reshaping video production.
Why we keep client work off the cloud
Most online AI video tools work by uploading your image to someone else’s server. For a test portrait that does not matter. For client work it often does: an unreleased product pack, a campaign key visual under embargo, a storyboard for a launch that has not been announced, or talent likeness covered by an agreement. Once those frames sit on a third-party service, how long they are kept and what they may be used for depends on that service’s terms, not yours.
Running open weights in a privately hosted ComfyUI setup removes that question. The brand assets never leave the machine, the workflow is saved as a file we control, and the same graph can be re-run months later for a revision.
Bespoke workflows built around the brief
The other reason we build in ComfyUI is control. A hosted video tool gives you a prompt box and a fixed set of options. ComfyUI exposes every step as a node, so a workflow can be customised for one project and shaped around the brief rather than the other way round. The graphs in the recordings above are a good example: the resolution, duration, sampler and frame rate are all set by hand, per shot.
On client projects that means we can chain an image model such as Flux or Qwen Image to build the key visual (we compared four of them in our AI image generation test), pass it into LTX 2.5 or MiniMax H3 for motion, then upscale and interpolate in the same graph. We can lock first and last frames to match an approved storyboard, add a custom LoRA trained on a brand character or product where the model supports it, and output 16:9, 9:16 and square versions from one run. For the still-image side, see how we handle ComfyUI image editing for pose, light, backgrounds and style. Once a workflow is right, it is saved and reused, so a campaign keeps the same look from the first clip to the last.
Which Model We Would Use in our hybrid pipeline
MiniMax H3 when the shot is about a performance: a character reacting, a mascot with comic timing, a presenter-style moment, or any brief where the prompt reads like direction to an actor. Write the prompt as timed beats and it follows them.
LTX 2.5 when the shot is about the frame: a product that must stay on-pack and legible, high resolution delivery, or a short sequence of connected shots from one generation. It stays closer to the source image, which is what a brand manager usually wants from a packshot.
In practice we use both, alongside conventional 3D animation and 2D animation. AI image to video is fast at exploring motion and building animatics, and it can deliver finished shots for social. Hero brand work still gets the same art direction, compositing and colour pass as everything else that leaves the studio.
Because the whole pipeline is privately hosted, it also fits the confidentiality terms that finance, healthcare and pre-launch product clients usually ask for.
CRITICA is a creative production studio in Singapore. We take brands from research and concept through storyboard to final delivery, producing motion graphics, video, and experiential work for Finance, Healthcare, Technology, Hospitality, Tourism, Oil & Gas, and Renewables. Contact us to discuss a project you have in mind.

FAQ
Which is better for image to video, LTX 2.5 or MiniMax H3?
It depends on the shot. In our ComfyUI tests MiniMax H3 followed acting direction more closely, while LTX 2.5 stayed closer to the source image and kept a product label square to camera. LTX 2.5 also outputs higher resolution, up to 4K, against 2K for MiniMax H3.
Can you run MiniMax H3 in ComfyUI?
Yes. MiniMax H3 has native ComfyUI support with open weights since 3 August 2026, with official templates for image to video, text to video and reference to video.
Can you run LTX 2.5 in ComfyUI?
Yes. Lightricks released LTX 2.5 as open weights on 11 August 2026 with day-one ComfyUI support, including image to video, text to video and first and last frame workflows.
Is MiniMax H3 the same as Hailuo 3?
Yes. MiniMax H3 is the model name and Hailuo 3 is the name of MiniMax's video product built on it.
Do LTX 2.5 and MiniMax H3 generate sound?
Both models generate audio together with the video. MiniMax H3 produces native stereo sound, and LTX 2.5 produces synchronised audio and video in one generation.
How do you write a good prompt for MiniMax H3 image to video?
Describe the motion as timed beats, for example what happens from 0 to 2 seconds, then 2 to 3.5 seconds, and lock the camera if you do not want it to move. In our tests MiniMax H3 followed each beat in sequence.
Can I use ComfyUI without uploading client work to the cloud?
Yes. ComfyUI is a privately hosted app that runs on your own computer or server. With open weights such as LTX 2.5 and MiniMax H3 installed locally, source images and finished clips never leave your hardware, which suits confidential and pre-launch client work.
Can ComfyUI workflows be customised for a specific project?
Yes. ComfyUI is node-based, so every step can be changed per project: which image and video models are used, resolution, duration, first and last frames, custom LoRAs, upscaling and output formats. A finished workflow is saved as a file and reused, which keeps a campaign consistent.