7.4 KiB
H3 Blackwell Runtime
Direct MiniMax H3 audiovisual inference for NVIDIA Blackwell and Grace Blackwell. The runtime consumes the current Comfy-format NVFP4 checkpoints while keeping ComfyUI out of the inference critical path. ComfyUI remains the checkpoint and correctness oracle.
Current Status
Implemented and validated:
- Prompt-only T2VA with jointly generated video and stereo audio.
- First-frame I2VA, last-frame L2VA, and first/last-frame FL2VA conditioning.
- Qwen3-VL text and image conditioning, token refinement, H3 packed denoising, beta/RES sampling, video and audio VAE decode, H.264/AAC encoding, and muxing.
- Resident HTTP runtime with startup warmup, model reuse, request timing, and request-level attention selection.
- SageAttention2 correctness default plus SDPA, forced cuDNN SDPA, FlashAttention-4, Sage3, KJ Sage, Comfy Kitchen INT8, head-sliced, and Sol-Attn experimental backends.
- Official FL2VA Turbo 4-step and 8-step adapters.
- Optional H3-native 3D latent upscaling.
- Opt-in EasyCache and H3-Cache experiments.
- Ragged Ulysses sequence parallelism and true NVFP4 tensor parallelism with automatic 1/2/4/6/8-GPU launch tooling.
The main remaining feature gap is full arbitrary Ref2VA, especially reference
video/audio, identity, and voice conditioning. Multi-GPU code is CPU- and
one-GPU-validated; real 2/4/6/8-GPU NCCL scaling measurements are still pending.
See CURRENT_STATE.md for the canonical detailed status.
Checkpoints
The default runtime paths are:
/models/minimax_h3_fl2va_pruned_nvfp4.safetensors
/text-encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
/vae/ae.safetensors
/vae/mini_vae.safetensors
The denoiser and Qwen checkpoints use Comfy Kitchen NVFP4 layouts. Do not convert or dequantize them during loading.
DGX Spark
Dockerfile.spark and compose.spark.yml provide the ARM64 CUDA 13 runtime used
on GB10. Build and start the resident service from the Spark checkout:
docker compose -f compose.spark.yml build h3-hot-runtime
docker compose -f compose.spark.yml up -d h3-hot-runtime
curl http://127.0.0.1:8001/ready
The service listens on container port 8000 and Spark host port 8001. It keeps Qwen, H3, both VAEs, Turbo adapters, and the optional latent upscaler resident. SageAttention2 is the default because it matches the established Comfy quality baseline.
HTTP API
GET /health and GET /ready report readiness, loaded options, warmup results,
and attention backend status. POST /generate performs one serialized request:
curl -X POST http://127.0.0.1:8001/generate \
-H 'Content-Type: application/json' \
-d '{
"prompt": "A quiet medium shot of a woman by a rain-streaked window. She says, \"We should leave before dawn.\"",
"output": "/output/h3-blackwell-runtime/example.mp4",
"width": 864,
"height": 480,
"frames": 141,
"steps": 12,
"seed": 440420,
"attention": "sage2"
}'
Optional request fields include first_frame, last_frame, turbo, upscale,
mux_audio, save_latent, keep_intermediates, cache_mode, and cache tuning
parameters. Keyframes accept an on-disk image path, base64 payload, or data URL.
The presence of first and last frames selects I2VA, L2VA, or FL2VA behavior; a
named task field is not yet exposed.
Set turbo to "4step" or "8step"; the service enforces the corresponding
step count and Turbo schedule. Set upscale to 2.0 to run the resident learned
latent upscaler before video decode. Turbo and denoiser caching cannot be
combined.
Quoted dialogue is the project prompt default. Tagged <d>[English]... dialogue
is a repeatable startup-audio trigger and should not be used as the default.
Never apply unconditional audio trimming or fading because valid sound can begin
at the first sample. See H3_PROMPT_GUIDE.md and
AUDIO_BOUNDARY_INVESTIGATION.md.
Distributed Execution
Two batch-generation modes are available through torchrun:
ulysses: token-sharded blocks with ragged sequence-to-head all-to-all around attention.tensor: sequence-sharded residuals plus NVFP4-sharded QKV, attention output, MLP FC1, and MLP FC2 projections.
Use all visible GPUs or pass an explicit world size:
tools/run_distributed_t2va.sh ulysses
tools/run_distributed_t2va.sh tensor 8 sdpa
H3_SAVE_LATENTS=0 tools/run_distributed_matrix.sh
The matrix runs feasible 1/2/4/6/8-GPU configurations and skips counts larger
than the visible device count. Six ranks use ragged head ownership
[10, 10, 9, 9, 9, 9]; no semantic token padding is introduced.
See DISTRIBUTED.md for the collective contracts, validation
evidence, environment variables, RunPod provisioning client, and x86 packaging
status.
Validation
Run the contract suite with:
python -m unittest discover -s tests -v
The current suite has 38 passing tests. Distributed tests cover 2/4/6/8-rank transport identity, ragged collectives, SDPA parity, final projection parity, NVFP4 shard layout, and 2/6-rank TP attention/MLP math. On GB10, real-checkpoint one-rank Ulysses and tensor paths produced exactly equal video and audio latents at 864x480, 141 frames, 12 steps, seed 440420.
A matched one-GPU RunPod RTX PRO 6000 Blackwell Server SDPA test averaged
28.50s sampling over two runs versus 126.66s for the same tensor runner on
GB10, a 4.44x speedup. The two RTX runs were within 0.49% and produced
identical checksums. See
benchmarks/rtxpro6000-server-vs-gb10-sdpa-864x480-141f-base12-seed440420.json.
Two-GPU Ulysses scaling improves as the packed sequence grows: 1.21x at
864x480/141 frames, 1.55x at 1344x768/124 frames, and 1.72x at
1344x768/243 frames. These RTX PRO 6000 Server cards have no NVLink. CUDA P2P
read/write is available and NCCL selected P2P/CUMEM; topology varied between
SYS and NODE across RunPod allocations. See
benchmarks/rtxpro6000-server-1v2-sdpa-scaling-seed440420.json.
Matched GB10 backend results and parity evidence are recorded in PLAN.md
and PARITY.md. Standalone tools/compare_*, tools/trace_*,
tools/inspect_*, and tools/patch_comfy_* commands are diagnostic utilities,
not runtime startup requirements.
Documentation
CURRENT_STATE.md: canonical implemented and missing scope.DISTRIBUTED.md: Ulysses, TP, launchers, and RunPod workflow.H3_PROMPT_GUIDE.md: audiovisual prompting contract.AUDIO_BOUNDARY_INVESTIGATION.md: startup audio localization and prompt-format evidence.TURBO.md: official Turbo adapters and schedules.H3_LATENT_UPSCALER.md: learned latent upscaler.FLASH4.md: FlashAttention-4 integration and benchmark.PARITY.md: direct-versus-Comfy evidence ledger.PLAN.md: historical investigation and future kernel plan.
Known Gaps
- Full arbitrary Ref2VA reference video/audio and identity/voice conditioning.
- Explicit task schemas and stricter production request validation.
- Real multi-GPU NCCL parity and 1/2/4/6/8 scaling results.
- x86 SageAttention2 packaging for the generic RunPod image; SDPA is the initial cloud validation backend.
- Production queueing, cancellation, authentication, TLS, metrics, and durable job state.
- Full quality sweeps for approximate attention, cache, Turbo, and upscaler paths.