preloader
post-thumb

Last Update: August 4, 2026


BYauthor-thumberic

|Loading...

Keywords

I ended the last post with a promise: the same RTX 3060 that learned to make short video would next try to make video with sound — using MiniMax H3, an open-weights model that generates its own audio. The weights landed publicly, I put them on the same card, and this is what actually happened.

The headline first: it works. A single product photo goes in, and a few-second clip comes out — the model turning, the fabric settling, a soft spoken line and room tone — audio the model generated itself, not dubbed on afterwards. On a 12GB consumer card. That part genuinely surprised me.

The machine (still the same one)

Same box as the image retrospective and the video post: an RTX 3060 with 12GB of VRAM, 64GB of system RAM, Debian. Nothing upgraded. The card is the constant this whole series is built around.

What "video with sound" means here

Most open video models are mute — you get moving pictures and add sound in an editor. MiniMax H3 is different: video and audio come out of the same model, generated together, so the soundtrack matches the scene. Describe waves and a spoken line, and you get waves and a voice. That single capability is why it was worth the effort.

Fitting 54GB into 12GB

Here is the honest part. H3 is big: about 54GB of weights — a 21GB video model, a 27GB text encoder, and two decoders for picture and sound. That does not fit in 12GB of VRAM. Not close.

It runs anyway because the weights live in the 64GB of system RAM and stream onto the card a slice at a time, keeping just enough on the GPU to do the next step. It is slow, but it finishes. Two smaller things also got in the way: the model is new enough that its picture and sound decoders each had a small bug that had to be patched before a clip would complete, and the card needed a healthy slice of memory kept free so the compressed weights had room to unpack. Fiddly, but one-time — once it was set up, it just ran.

The result: a product clip that turns and talks

This is the workflow a small business would actually use: start from one product photo, and let the model animate it — the person turns to camera, the knit settles, a quiet line of audio plays — while the product stays readable. Here is a real clip generated on this card. Press play with the sound on — the audio came out of the model, not an editor.

And the same clip laid out frame by frame — a natural turn from side-on to camera, the knit texture and anatomy holding through the whole move:

Frames from a MiniMax H3 clip: a model in a cream knit sweater turning to camera, generated with sound on an RTX 3060

The honest numbers

On the RTX 3060
Clip length
A few seconds (2-5s)
Resolution
512-640px
Audio
Yes - native stereo, generated with the video
Model size
~54GB of weights, streamed from system RAM
Time per clip
Roughly 20-27 minutes
Best for
A short product clip with a voice or ambient soundtrack

Twenty-odd minutes for a few seconds is the real cost of streaming that much weight through a small card. This is not something you batch out by the hundred on a 3060 — it is something you run overnight for one good clip. But the clip has sound, and it cost electricity.

What I learned about the "AI look"

The most useful lesson wasn't about speed — it was about why AI video still reads as AI, and how to fight it.

Image-to-video can never look better than the photo you start from; it inherits that frame's every quirk and then adds motion on top. My first attempts started from bright, high-contrast outdoor stills — and they had that unmistakable "AI feel": waxy skin, hard outlines, the joins between body parts subtly wrong. The motion clung to those hard edges and made them worse.

The clip above started from a soft, studio-lit photo instead — diffused light, a plain background, gentle edges. And it held together: real fabric texture, natural movement, believable anatomy through the whole turn. Same model, same card, completely different result.

So the practical rule, if you try this yourself: start from soft, evenly-lit photos, not harsh high-contrast ones. The model can't cling to hard edges that were never there. It's the single biggest lever on how "real" the output looks — bigger than any setting.

Where this leaves a small business

Three posts, one unchanged computer. The same RTX 3060 went from struggling with a single 512px still, to short silent video, to a short clip with its own soundtrack — a product turning to camera and speaking a line, generated from one photo, locally, for the cost of the power it drew.

It is not one-click, and it is not fast. But the gap between "I have a product photo" and "I have a short product video with sound" has quietly closed on hardware a small business can already afford. That's the whole point of this series: the tools that used to need a studio now need a modest computer and a bit of patience.

The card hasn't changed in two years. What it can do has.

Comments (0)

Leave a Comment
Your email won't be published. We'll only use it to notify you of replies to your comment.
Loading comments...
Previous Article
post-thumb

Oct 03, 2021

Setting up Ingress for a Web Service in a Kubernetes Cluster with NGINX Ingress Controller

A simple tutorial that helps configure ingress for a web service inside a kubernetes cluster using NGINX Ingress Controller

Next Article
post-thumb

Aug 03, 2026

Now It Moves: Video Generation on the Same RTX 3060

The follow-up to our image retrospective: the same RTX 3060 that struggled with a single 512px still two years ago now generates short video — including product promos animated from a single photo — with open models like Wan 2.2 and LTX-Video.

agico

We transform visions into reality. We specializes in crafting digital experiences that captivate, engage, and innovate. With a fusion of creativity and expertise, we bring your ideas to life, one pixel at a time. Let's build the future together.

Copyright ©  2026  TYO Lab · v0.0.17