{"id":24849,"date":"2026-08-13T16:25:15","date_gmt":"2026-08-13T16:25:15","guid":{"rendered":"https:\/\/jurn.link\/dazposer\/?p=24849"},"modified":"2026-08-13T16:26:23","modified_gmt":"2026-08-13T16:26:23","slug":"success-minimax-h3-video-at-1376-x-768px-on-a-3060-12gb-card","status":"publish","type":"post","link":"https:\/\/jurn.link\/dazposer\/index.php\/2026\/08\/13\/success-minimax-h3-video-at-1376-x-768px-on-a-3060-12gb-card\/","title":{"rendered":"Success &#8211; Minimax H3 video at 1376 x 768px, on a 3060 12Gb card"},"content":{"rendered":"<p>Hurrah, success with Minimax text-to-video on my RTX 3060 card! I can now generate video to Minimax&#8217;s 1.0 resolution. The sweet spot seems to be 12 minutes to get a seven second video at 0.9 resolution.<\/p>\n<div style=\"width: 640px;\" class=\"wp-video\"><video class=\"wp-video-shortcode\" id=\"video-24849-1\" width=\"640\" height=\"368\" preload=\"metadata\" controls=\"controls\"><source type=\"video\/mp4\" src=\"https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/MiniMax_H3_00042_.mp4?_=1\" \/><a href=\"https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/MiniMax_H3_00042_.mp4\">https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/MiniMax_H3_00042_.mp4<\/a><\/video><\/div>\n<p>&nbsp;<\/p>\n<p>Windows 11, RTX 3060 12Gb with just 24Gb DDR3 system RAM (most people run 32Gb of DDR4). It&#8217;s amazing I can do large coherent videos on this, in a few minutes, on an old server. How far we&#8217;ve come since the days when Stable Diffusion 1.5 was mostly gloopy mis-fires at 512px, and took ages. <\/p>\n<p>Minimax&#8217;s prompt adherence is now much better, now I&#8217;ve moved on from using a supercrushed Q2 GGUF for the prompt-processing. Next step is to see how well it works in combination with Poser \/ natural media emulation. But for now, a full setup tutorial. <\/p>\n<p><strong>TUTORIAL:<\/strong><\/p>\n<p>Here&#8217;s how to do it on this humble entry-level card, with the latest software, file-types and turbo boosters. The workflow is linked at the bottom of the post. All free, as is the way with generative AI.<\/p>\n<p>First, install <a href=\"https:\/\/github.com\/YanWenKun\/ComfyUI-Windows-Portable\/\">a new ComfyUI Portable<\/a>, then use the Manager to update ComfyUI to the latest 0.32.x or higher (needed!). If the ComfyUI console throws a <em>polars<\/em> error from Python when starting up, then also <em>pip install polars-lts-cpu<\/em> (it&#8217;s <em>polars<\/em> for older CPUs, which lack the latest whizz-bang multimedia processing extensions).<\/p>\n<p>Second, clear 40Gb of space on the SSD for the files you need, and also leave plenty of headroom for the Windows swap-file etc. I have my swap-file capped at 24Gb to match the system RAM.<\/p>\n<p><strong>Models:<\/strong> At <a href=\"https:\/\/huggingface.co\/Winnougan\/MiniMax-H3-INT4_Convrot_ComfyUI\/tree\/main\">HF<\/a> get <em>minimax_h3_fl2va_pruned-w4a8_convrot_pruned.safetensors<\/em> (11.6Gb) and <em>minimax_h3_ref2va_pruned-w4a8_convrot_pruned.safetensors<\/em> (11.6Gb). The first does text-to-video and first frame or first frame\/last frame. The second takes character \/ environment reference images (e.g. Poser character renders), and combines the character(s) and scenes into the video, if you have the prompting done correctly. Correct prompting is <em>everything<\/em> with Minimax H3. Put the files in <em>..\\ComfyUI\\models\\diffusion_models\\<\/em><\/p>\n<p><strong>Clip:<\/strong> At <a href=\"https:\/\/huggingface.co\/Winnougan\/MiniMax-H3-INT4_Convrot_ComfyUI\/tree\/main\">HF<\/a> get <em>qwen3vl_32b_minimax_h3-w4a8_convrot.safetensors<\/em> (14.6Gb). Put the file in ..\\models\\text_encoders\\<\/p>\n<p>The above are Winnougan&#8217;s highly efficient &#8216;pruned&#8217; blends of the FP4 and Int8 format, suitable for the 3060 12Gb card and matched with the very latest ComfyUI features.<\/p>\n<p><strong>VAE video:<\/strong> At <a href=\"https:\/\/huggingface.co\/Kijai\/MiniMax-H3-experimental\/tree\/main\">HF<\/a> get Kijai&#8217;s <em>minimax_h3_video_vae_int8_convrot.safetensors<\/em> (3.7Gb). This decodes the video frames into a video, and is far faster than the original. Place the file in <em>..\\ComfyUI\\models\\vae\\<\/em><\/p>\n<p><strong>VAE audio:<\/strong> At <a href=\"https:\/\/huggingface.co\/Comfy-Org\/MiniMax-H3\/tree\/main\/vae\">HF<\/a> get <em>minimax_h3_audio_vae_fp32.safetensors<\/em> (577Mb). This is the standard from ComfyUI, and works so quickly it&#8217;s not worth speeding up. Place the file in <em>..\\ComfyUI\\models\\vae\\<\/em><\/p>\n<p><strong>Turbo LoRA:<\/strong> At <a href=\"https:\/\/huggingface.co\/Abiray\/MiniMax-H3-Turbo-Lora-Pruned-ComfyUI\/tree\/main\">HF<\/a> get Abiray&#8217;s <em>minimax_h3_turbo_4step_ckpt600_V4.safetensors<\/em> turbo LoRA (590Mb) and place it in <em>..\\ComfyUI\\models\\loras\\ <\/em><\/p>\n<p>Having ComfyUI at version 0.32.0 or higher allows you to run the turbo video VAE and also to pair the turbo LoRA with their new &#8216;Sage Attention replacement&#8217; speed-up called Comfy Kitchen Attention. Simply add their new <strong>ModelAttentionBackend<\/strong> node between the model and the LoRA, and toggle it to use &#8216;Kitchen Attention&#8217;. That&#8217;s it, though note it&#8217;s best for lower-end cards and that 50-series cards may not see much difference. I think I&#8217;m getting a 20-30% speed-up from it.<\/p>\n<p><a href=\"https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/minimax-h3-3060-workflow.jpg\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/minimax-h3-3060-workflow.jpg\" alt=\"\" width=\"640\" height=\"264\" class=\"aligncenter size-large wp-image-24851\" srcset=\"https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/minimax-h3-3060-workflow.jpg 2000w, https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/minimax-h3-3060-workflow-300x124.jpg 300w, https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/minimax-h3-3060-workflow-1024x423.jpg 1024w, https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/minimax-h3-3060-workflow-768x318.jpg 768w, https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/minimax-h3-3060-workflow-1536x635.jpg 1536w\" sizes=\"auto, (max-width: 640px) 100vw, 640px\" \/><\/a><\/p>\n<p>With this setup I get good results at 5 steps with the turbo LoRA at 1.2, Euler \/ beta. Image quality is still slightly crispy and granulated, but better than smushed and gloopy. The crispiness could be because I&#8217;m over-forcing the turbo LoRA, of course.<\/p>\n<p>7 mins = 7 seconds, at 0.6 (1056 x 608px).<br \/>\n12 mins = 7 seconds, at 0.9 (1280 x 736px).<br \/>\n18 mins = 8 seconds, at 1.0 (1376 x 768px) (the model&#8217;s native trained resolution) and using 6 steps.<\/p>\n<p>All 24fps at 16:9 widescreen ratio. 0.9 seems the sweet spot for me re: the balance of time \/ quality \/ size. That&#8217;s the video which heads this tutorial.<\/p>\n<p><strong>Workflow as a .JSON file<\/strong> for ComfyUI: <a href=\"https:\/\/jurn.link\/dazposer\/wp-content\/uploads\/2026\/08\/MiniMax_H3_RTX3060-12Gb-workflow.zip\">MiniMax_H3_RTX3060-12Gb-workflow.ZIP<\/a> (225Kb).<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Hurrah, success with Minimax text-to-video on my RTX 3060 card! I can now generate video to Minimax&#8217;s 1.0 resolution. The sweet spot seems to be 12 minutes to get a seven second video at 0.9 resolution. &nbsp; Windows 11, RTX 3060 12Gb with just 24Gb DDR3 system RAM (most people run 32Gb of DDR4). It&#8217;s [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[13,9,7,12],"tags":[],"class_list":["post-24849","post","type-post","status-publish","format-standard","hentry","category-companion-software","category-freebies","category-the-animation-industry","category-tutorials"],"_links":{"self":[{"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/posts\/24849","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/comments?post=24849"}],"version-history":[{"count":3,"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/posts\/24849\/revisions"}],"predecessor-version":[{"id":24855,"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/posts\/24849\/revisions\/24855"}],"wp:attachment":[{"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/media?parent=24849"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/categories?post=24849"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/jurn.link\/dazposer\/index.php\/wp-json\/wp\/v2\/tags?post=24849"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}