101 lines
7.4 KiB
Markdown
101 lines
7.4 KiB
Markdown
# H3 Blackwell Runtime Plan
|
|
|
|
## Goal
|
|
|
|
Build a direct MiniMax H3 Ref2VA runtime for Blackwell and Grace Blackwell that consumes the current Comfy safetensors checkpoints while removing ComfyUI and Raylight from the denoising critical path.
|
|
|
|
The runtime must support one GPU first, then correct 2/4/6/8 GPU execution. It must retain the current NVFP4 model artifacts and use SageAttention3 where quality validation permits.
|
|
|
|
## Future LTX 2.5 Track
|
|
|
|
Add a separate LTX 2.5 direct-runtime adapter after H3 single-GPU parity is stable. Target the gated Lightricks `ltx-2.5-22b-distilled-transformer-nvfp4.safetensors` artifact (18.7 GB, release commit `dd53cc2cd45bbeaa3563dfb575cba3f49cf44761`).
|
|
|
|
- Keep LTX model loading, conditioning, scheduler, VAE, and validation isolated from H3; this is a second model family, not an H3 checkpoint variant.
|
|
- Inspect the safetensors header and published architecture/configuration before sharing H3 modules or kernels.
|
|
- Establish a LTX SDPA/Sage2 correctness baseline before evaluating Sage3, FlashAttention-4, Sol-Attn, cache methods, or distributed layouts.
|
|
- Respect the LTX 2 Community License Agreement and gated-access requirements; do not automate downloads without authorized access.
|
|
|
|
## Reference Baseline
|
|
|
|
The first acceptance target is the clean one-GPU ComfyUI baseline in `../h3-lab/h3-raylight-usp2-results.json`:
|
|
|
|
- RTX PRO 6000 Blackwell, 96 GB
|
|
- Ref2VA, 960x544, 124 frames, 24 fps
|
|
- 12 steps, `beta`, `res_multistep`, seed `440202`
|
|
- SageAttention3, 1 GPU
|
|
- ComfyUI execution time: `49.893s`
|
|
|
|
The direct runner must first match the model contract and output quality. Beating this timing comes after correctness is established.
|
|
|
|
## Architecture
|
|
|
|
1. Checkpoint adapter: read Comfy safetensors metadata, preserve packed low-precision weights and scale tensors, and map them into a canonical H3 state dictionary.
|
|
2. Conditioning service: execute and cache Qwen layer-50 embeddings, modality tags, and reference VAE latents once per request.
|
|
3. H3 denoiser: implement the packed Ref2VA DiT, 3-axis RoPE, AdaLN, dual audio/video schedule, and RES multistep solver without node-graph orchestration.
|
|
4. Kernel layer: retain the known-good NVFP4 linear path initially; add explicit SageAttention3 and CUDA-graph buckets after exact single-GPU output validation.
|
|
5. Distributed layer: use ragged Ulysses all-to-all for Q/K/V head exchange, Sage3 on full packed tokens per local head shard, then inverse exchange. Add tensor parallelism only after sequence parallel correctness is proven.
|
|
|
|
## Milestones
|
|
|
|
1. Inspect the actual local `pruned_nvfp4` checkpoint header and classify every tensor/scale layout.
|
|
2. Create a direct single-GPU denoiser step matching ComfyUI for a fixed captured payload.
|
|
3. Implement full single-GPU Ref2VA and compare per-step tensors plus final AV output against ComfyUI.
|
|
4. Apply SageAttention3 and CUDA graphs; benchmark against the 49.893s reference.
|
|
- Optional backends and execution strategies to evaluate behind the same per-step quality gate: FlashAttention-4 (the Blackwell successor to Hopper-only FlashAttention-3), EasyCache/H3-Cache, Sol-Attn, and KJ exact memory-lifetime patches.
|
|
- Keep backend selection explicit per run; retain only candidates that match the validated direct correctness path and improve the measured denoising bottleneck.
|
|
- Current correctness baseline: SageAttention2 (`sage2`), which exactly matches the captured ComfyUI `--use-sage-attention` output. SDPA is a fallback; SageAttention3 remains experimental and must pass the same quality gate.
|
|
- Sol-Attn and KJ Sage have prior H3 test evidence and are supported experimental candidates. Integrate each as an isolated standalone adapter, record the exact mode/version, and gate it against the Sage2 per-step reference before combining it with caches or other approximation strategies.
|
|
5. Implement ragged Ulysses Sage3 with transport-identity and distributed-versus-single-Sage3 tests.
|
|
6. Sweep Ulysses/tensor-parallel layouts on 2/4/6/8 GPUs in an NVLink/NVSwitch domain.
|
|
|
|
## Non-Negotiable Validation
|
|
|
|
- Never silently pad semantic H3 tokens for unmasked attention.
|
|
- Compare distributed output against the identical single-GPU Sage3 path before comparing to SDPA.
|
|
- Validate denoiser outputs at each scheduler step, not only encoded video.
|
|
- Record attention, GEMM, communication, VAE, and end-to-end timings separately.
|
|
- Treat SageAttention3 as an experimental quality-gated kernel for H3.
|
|
|
|
## 2026-08-13 VAE Debug Handoff
|
|
|
|
Current saved latent and comparison assets live under:
|
|
|
|
- Spark: `/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime`
|
|
- Share: `\\192.168.1.162\StoryStudioAssets\H3-output\h3-blackwell-runtime`
|
|
|
|
Key assets:
|
|
|
|
- Good sampled latent: `direct-cat-house-backflip-disco-shades-960x544-5s-latent.pt`
|
|
- Known-good same-latent Comfy/upstream VAE decode: `direct-cat-house-backflip-disco-shades-960x544-5s-upstream-comfy-tiled-decode.mp4`
|
|
- Current direct standalone FP32/Sage VAE decode candidate: `direct-cat-house-backflip-disco-shades-960x544-5s-direct-fp32-sage-vae-decode.mp4`
|
|
|
|
What is proven:
|
|
|
|
- The latent is good. Same latent decoded through Comfy/upstream VAE is visually clean.
|
|
- The sampler/model path is already exact against Comfy free-run parity for the fixed baseline.
|
|
- Direct VAE BF16/checkpoint-dtype loading was wrong. Loading direct VAE weights as FP32 reduced isolated decoder clip drift from roughly `mean=1.5e-3, max=1e-1` to roughly `mean=1e-6, max=6e-5`.
|
|
- Individual direct tiled VAE clips match upstream closely (`mean` around `4e-7`).
|
|
- Full direct `decode_temporal()` still differs from upstream at global frames `17, 34, 51, 68, 85, 102`, i.e. temporal join boundaries.
|
|
- Temporal assembly tracer shows pre-blend current chunks and overlap tails each match upstream, but blended join output differs hugely (`mean` around `0.15`, max around `5`) when comparing direct blend result to upstream blend result using their respective near-identical inputs.
|
|
- Direct and upstream `blend()` return identical results on the exact same inputs, so the remaining issue is likely an input/aliasing/dtype/shape subtlety at the temporal join, not the blend formula itself.
|
|
|
|
Relevant debug tools committed:
|
|
|
|
- `tools/decode_video_latent.py`
|
|
- `tools/compare_frame_dirs.py`
|
|
- `tools/compare_vae_decoder_clip.py`
|
|
- `tools/compare_vae_full_decode.py`
|
|
- `tools/compare_vae_tiled_clip.py`
|
|
- `tools/compare_vae_temporal_assembly.py`
|
|
|
|
Next VAE debugging steps:
|
|
|
|
- In `tools/compare_vae_temporal_assembly.py`, compare `prev_d` vs `prev_u` and `part_d` vs `part_u` after casting both pairs to a shared dtype and before blending. The current stats say they are close, but the blend of respective inputs explodes, which suggests a subtle shape/stride/dim broadcasting mismatch.
|
|
- Log `shape`, `stride`, `dtype`, `is_contiguous`, and `storage_offset` for `prev_*`, `part_*`, blend weights, and slices at `chunk_1_blended_at_17`.
|
|
- Try forcing `prev_d`, `prev_u`, `part_d`, and `part_u` to `.contiguous()` immediately before temporal blend in both direct and tracer paths.
|
|
- If that fixes it, patch direct `decode_temporal()` only. If not, compare exact selected slices (`prev[..., -9:, :, :]`, `part[..., :9, :, :]`) elementwise before and after multiplying weights.
|
|
|
|
Backends/experiments to evaluate later:
|
|
|
|
- Attention backend sweeps after VAE quality is fixed: `sage2`, `sage3`, `sdpa`, plus future FlashAttention-4/Sol-Attn/KJ candidates.
|
|
- Spectrum MiniMax H3 repo to inspect later: https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
|