Build a direct MiniMax H3 Ref2VA runtime for Blackwell and Grace Blackwell that consumes the current Comfy safetensors checkpoints while removing ComfyUI and Raylight from the denoising critical path.
The runtime must support one GPU first, then correct 2/4/6/8 GPU execution. It must retain the current NVFP4 model artifacts and use SageAttention3 where quality validation permits.
## Future LTX 2.5 Track
Add a separate LTX 2.5 direct-runtime adapter after H3 single-GPU parity is stable. Target the gated Lightricks `ltx-2.5-22b-distilled-transformer-nvfp4.safetensors` artifact (18.7 GB, release commit `dd53cc2cd45bbeaa3563dfb575cba3f49cf44761`).
- Keep LTX model loading, conditioning, scheduler, VAE, and validation isolated from H3; this is a second model family, not an H3 checkpoint variant.
- Inspect the safetensors header and published architecture/configuration before sharing H3 modules or kernels.
- Establish a LTX SDPA/Sage2 correctness baseline before evaluating Sage3, FlashAttention-4, Sol-Attn, cache methods, or distributed layouts.
- Respect the LTX 2 Community License Agreement and gated-access requirements; do not automate downloads without authorized access.
## Reference Baseline
The first acceptance target is the clean one-GPU ComfyUI baseline in `../h3-lab/h3-raylight-usp2-results.json`:
The direct runner must first match the model contract and output quality. Beating this timing comes after correctness is established.
## Architecture
1. Checkpoint adapter: read Comfy safetensors metadata, preserve packed low-precision weights and scale tensors, and map them into a canonical H3 state dictionary.
2. Conditioning service: execute and cache Qwen layer-50 embeddings, modality tags, and reference VAE latents once per request.
3. H3 denoiser: implement the packed Ref2VA DiT, 3-axis RoPE, AdaLN, dual audio/video schedule, and RES multistep solver without node-graph orchestration.
4. Kernel layer: retain the known-good NVFP4 linear path initially; add explicit SageAttention3 and CUDA-graph buckets after exact single-GPU output validation.
5. Distributed layer: use ragged Ulysses all-to-all for Q/K/V head exchange, Sage3 on full packed tokens per local head shard, then inverse exchange. Add tensor parallelism only after sequence parallel correctness is proven.
## Milestones
1. Inspect the actual local `pruned_nvfp4` checkpoint header and classify every tensor/scale layout.
2. Create a direct single-GPU denoiser step matching ComfyUI for a fixed captured payload.
3. Implement full single-GPU Ref2VA and compare per-step tensors plus final AV output against ComfyUI.
4. Apply SageAttention3 and CUDA graphs; benchmark against the 49.893s reference.
- Optional backends and execution strategies to evaluate behind the same per-step quality gate: FlashAttention-4 (the Blackwell successor to Hopper-only FlashAttention-3), EasyCache/H3-Cache, Sol-Attn, and KJ exact memory-lifetime patches.
- Keep backend selection explicit per run; retain only candidates that match the validated direct correctness path and improve the measured denoising bottleneck.
- Current correctness baseline: SageAttention2 (`sage2`), which exactly matches the captured ComfyUI `--use-sage-attention` output. SDPA is a fallback; SageAttention3 remains experimental and must pass the same quality gate.
- Sol-Attn and KJ Sage have prior H3 test evidence and are supported experimental candidates. Integrate each as an isolated standalone adapter, record the exact mode/version, and gate it against the Sage2 per-step reference before combining it with caches or other approximation strategies.
5. Implement ragged Ulysses Sage3 with transport-identity and distributed-versus-single-Sage3 tests.
6. Sweep Ulysses/tensor-parallel layouts on 2/4/6/8 GPUs in an NVLink/NVSwitch domain.
Prompt-only FL2VA is now at warm Comfy parity with the direct Sage2 baseline. Feature and performance work should proceed in this order:
1. Validate and benchmark the existing `sage3` backend against the same cat prompt, seed, dimensions, and FP16 VAE runtime path used for Sage2 parity.
- First cat benchmark result: Sage3 runs successfully but is slower than Sage2 in this direct path. Sampling was `123.675s` versus Sage2 `114.414s`; warm after text conditioning was `158.111s` versus Sage2 `149.304s`. Same-seed MP4 frame diff versus Sage2 was mean `46.563`, max `255`, so keep Sage3 experimental pending human visual review and stricter tensor gates.
2. Add exact memory/lifetime optimizations next: `kj_head_sliced` and `kj_chunked_ffn`. These must preserve the validated direct outputs before being kept.
3. Evaluate prior H3-tested attention candidates as standalone adapters: `sol_attn` and `kj_sage`.
4. Evaluate approximate denoiser caches only after exact baselines are recorded: `easycache` and `h3_cache`.
5. Keep every backend explicit per run, with separate quality and timing records for sampling, VAE, audio, and end-to-end output.
Resolved on `vae-decode-optimization`: direct VAE temporal overlap constants now match upstream, audio decode/mux is implemented, and Comfy-equivalent FP16 video VAE is the default runtime path. The 960x544x124 cat benchmark now matches warm Comfy performance: Comfy `150.26s`, direct `149.304s` after text conditioning, direct VAE decode `25.085s`.
- Current direct standalone FP32/Sage VAE decode candidate: `direct-cat-house-backflip-disco-shades-960x544-5s-direct-fp32-sage-vae-decode.mp4`
What is proven:
- The latent is good. Same latent decoded through Comfy/upstream VAE is visually clean.
- The sampler/model path is already exact against Comfy free-run parity for the fixed baseline.
- Direct VAE BF16/checkpoint-dtype loading was wrong. Loading direct VAE weights as FP32 reduced isolated decoder clip drift from roughly `mean=1.5e-3, max=1e-1` to roughly `mean=1e-6, max=6e-5`.
- Individual direct tiled VAE clips match upstream closely (`mean` around `4e-7`).
- Full direct `decode_temporal()` still differs from upstream at global frames `17, 34, 51, 68, 85, 102`, i.e. temporal join boundaries.
- Temporal assembly tracer shows pre-blend current chunks and overlap tails each match upstream, but blended join output differs hugely (`mean` around `0.15`, max around `5`) when comparing direct blend result to upstream blend result using their respective near-identical inputs.
- Direct and upstream `blend()` return identical results on the exact same inputs, so the remaining issue is likely an input/aliasing/dtype/shape subtlety at the temporal join, not the blend formula itself.
Relevant debug tools committed:
-`tools/decode_video_latent.py`
-`tools/compare_frame_dirs.py`
-`tools/compare_vae_decoder_clip.py`
-`tools/compare_vae_full_decode.py`
-`tools/compare_vae_tiled_clip.py`
-`tools/compare_vae_temporal_assembly.py`
Next VAE debugging steps:
- In `tools/compare_vae_temporal_assembly.py`, compare `prev_d` vs `prev_u` and `part_d` vs `part_u` after casting both pairs to a shared dtype and before blending. The current stats say they are close, but the blend of respective inputs explodes, which suggests a subtle shape/stride/dim broadcasting mismatch.
- Log `shape`, `stride`, `dtype`, `is_contiguous`, and `storage_offset` for `prev_*`, `part_*`, blend weights, and slices at `chunk_1_blended_at_17`.
- Try forcing `prev_d`, `prev_u`, `part_d`, and `part_u` to `.contiguous()` immediately before temporal blend in both direct and tracer paths.
- If that fixes it, patch direct `decode_temporal()` only. If not, compare exact selected slices (`prev[..., -9:, :, :]`, `part[..., :9, :, :]`) elementwise before and after multiplying weights.
Backends/experiments to evaluate later:
- Attention backend sweeps after VAE quality is fixed: `sage2`, `sage3`, `sdpa`, plus future FlashAttention-4/Sol-Attn/KJ candidates.
- Spectrum MiniMax H3 repo to inspect later: https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
- Comfy decode path for nested H3 AV latent is effectively: select nested audio tensor, call `audio_vae.decode(audio_latent)`, then return waveform with sample rate 32000. VideoHelperSuite/Comfy audio save nodes use ffmpeg or audio helpers to save/mux.
Direct runtime implementation path:
- Change `sample_video_res_multistep()` to return both final video and final audio latents, or add a `sample_av_res_multistep()` wrapper that preserves the existing video-only API.
- Port `MiniMaxH3AudioVAE` from Comfy into a standalone `audio_vae_decoder.py`, starting decoder-only if we only need generated audio.
- Load `/vae/minimax_h3_audio_vae_fp32.safetensors` strictly in FP32, mirroring the video VAE precision lesson.
- Add a latent-only decode tool that accepts saved audio latent `[1,32,2,T]`, writes WAV/FLAC at 32 kHz, and optionally muxes with MP4 using ffmpeg.
- Validate first by capturing or decoding the same final audio latent through Comfy and direct, comparing waveform tensors before muxing.