149 lines
12 KiB
Markdown
149 lines
12 KiB
Markdown
# H3 Blackwell Runtime Plan
|
|
|
|
## Goal
|
|
|
|
Build a direct MiniMax H3 Ref2VA runtime for Blackwell and Grace Blackwell that consumes the current Comfy safetensors checkpoints while removing ComfyUI and Raylight from the denoising critical path.
|
|
|
|
The runtime must support one GPU first, then correct 2/4/6/8 GPU execution. It must retain the current NVFP4 model artifacts and use SageAttention3 where quality validation permits.
|
|
|
|
## Future LTX 2.5 Track
|
|
|
|
Add a separate LTX 2.5 direct-runtime adapter after H3 single-GPU parity is stable. Target the gated Lightricks `ltx-2.5-22b-distilled-transformer-nvfp4.safetensors` artifact (18.7 GB, release commit `dd53cc2cd45bbeaa3563dfb575cba3f49cf44761`).
|
|
|
|
- Keep LTX model loading, conditioning, scheduler, VAE, and validation isolated from H3; this is a second model family, not an H3 checkpoint variant.
|
|
- Inspect the safetensors header and published architecture/configuration before sharing H3 modules or kernels.
|
|
- Establish a LTX SDPA/Sage2 correctness baseline before evaluating Sage3, FlashAttention-4, Sol-Attn, cache methods, or distributed layouts.
|
|
- Respect the LTX 2 Community License Agreement and gated-access requirements; do not automate downloads without authorized access.
|
|
|
|
## Reference Baseline
|
|
|
|
The first acceptance target is the clean one-GPU ComfyUI baseline in `../h3-lab/h3-raylight-usp2-results.json`:
|
|
|
|
- RTX PRO 6000 Blackwell, 96 GB
|
|
- Ref2VA, 960x544, 124 frames, 24 fps
|
|
- 12 steps, `beta`, `res_multistep`, seed `440202`
|
|
- SageAttention3, 1 GPU
|
|
- ComfyUI execution time: `49.893s`
|
|
|
|
The direct runner must first match the model contract and output quality. Beating this timing comes after correctness is established.
|
|
|
|
## Architecture
|
|
|
|
1. Checkpoint adapter: read Comfy safetensors metadata, preserve packed low-precision weights and scale tensors, and map them into a canonical H3 state dictionary.
|
|
2. Conditioning service: execute and cache Qwen layer-50 embeddings, modality tags, and reference VAE latents once per request.
|
|
3. H3 denoiser: implement the packed Ref2VA DiT, 3-axis RoPE, AdaLN, dual audio/video schedule, and RES multistep solver without node-graph orchestration.
|
|
4. Kernel layer: retain the known-good NVFP4 linear path initially; add explicit SageAttention3 and CUDA-graph buckets after exact single-GPU output validation.
|
|
5. Distributed layer: use ragged Ulysses all-to-all for Q/K/V head exchange, Sage3 on full packed tokens per local head shard, then inverse exchange. Add tensor parallelism only after sequence parallel correctness is proven.
|
|
|
|
## Milestones
|
|
|
|
1. Inspect the actual local `pruned_nvfp4` checkpoint header and classify every tensor/scale layout.
|
|
2. Create a direct single-GPU denoiser step matching ComfyUI for a fixed captured payload.
|
|
3. Implement full single-GPU Ref2VA and compare per-step tensors plus final AV output against ComfyUI.
|
|
4. Apply SageAttention3 and CUDA graphs; benchmark against the 49.893s reference.
|
|
- Optional backends and execution strategies to evaluate behind the same per-step quality gate: FlashAttention-4 (the Blackwell successor to Hopper-only FlashAttention-3), EasyCache/H3-Cache, Sol-Attn, and KJ exact memory-lifetime patches.
|
|
- Keep backend selection explicit per run; retain only candidates that match the validated direct correctness path and improve the measured denoising bottleneck.
|
|
- Current correctness baseline: SageAttention2 (`sage2`), which exactly matches the captured ComfyUI `--use-sage-attention` output. SDPA is a fallback; SageAttention3 remains experimental and must pass the same quality gate.
|
|
- Sol-Attn and KJ Sage have prior H3 test evidence and are supported experimental candidates. Integrate each as an isolated standalone adapter, record the exact mode/version, and gate it against the Sage2 per-step reference before combining it with caches or other approximation strategies.
|
|
5. Implement ragged Ulysses Sage3 with transport-identity and distributed-versus-single-Sage3 tests.
|
|
6. Sweep Ulysses/tensor-parallel layouts on 2/4/6/8 GPUs in an NVLink/NVSwitch domain.
|
|
|
|
## Performance Backend Plan
|
|
|
|
Prompt-only FL2VA is now at warm Comfy parity with the direct Sage2 baseline. Feature and performance work should proceed in this order:
|
|
|
|
1. Validate and benchmark the existing `sage3` backend against the same cat prompt, seed, dimensions, and FP16 VAE runtime path used for Sage2 parity.
|
|
- First cat benchmark result: Sage3 runs successfully but is slower than Sage2 in this direct path. Sampling was `123.675s` versus Sage2 `114.414s`; warm after text conditioning was `158.111s` versus Sage2 `149.304s`. Same-seed MP4 frame diff versus Sage2 was mean `46.563`, max `255`, so keep Sage3 experimental pending human visual review and stricter tensor gates.
|
|
2. Build a persistent hot runtime service instead of measuring only process-per-run CLIs. A warm container must preload and retain Qwen, H3, video VAE, and audio VAE in GPU memory, then accept video jobs without model-load latency. Add explicit startup warmup, readiness reporting, request-level attention selection (`sage2`, `sdpa`, `sage3` initially), and timing fields that separate resident-model request latency from cold startup.
|
|
3. Add exact memory/lifetime optimizations next: `kj_head_sliced` and `kj_chunked_ffn`. These must preserve the validated direct outputs before being kept.
|
|
4. Evaluate prior H3-tested attention candidates as standalone adapters: `sol_attn` and `kj_sage`.
|
|
- `kj_sage` is implemented as explicit SageAttention mode backends: `kj_sage_cuda`, `kj_sage_triton`, `kj_sage_fp8`, and `kj_sage_fp8pp`; all passed hot-runtime smoke tests.
|
|
- `sol_attn` is still blocked on locating/adding the standalone Sol-Attn source or package. It is not installed in the Spark image and is not present in this repository.
|
|
5. Evaluate approximate denoiser caches only after exact baselines are recorded: `easycache` and `h3_cache`.
|
|
- Initial direct cache modes are implemented as opt-in approximate sampler modes. They reuse cached denoised deltas and report skipped-step stats; full-size quality/threshold sweeps are still required before using them for production output.
|
|
6. Keep every backend explicit per run, with separate quality and timing records for sampling, VAE, audio, and end-to-end output.
|
|
|
|
## Non-Negotiable Validation
|
|
|
|
- Never silently pad semantic H3 tokens for unmasked attention.
|
|
- Compare distributed output against the identical single-GPU Sage3 path before comparing to SDPA.
|
|
- Validate denoiser outputs at each scheduler step, not only encoded video.
|
|
- Record attention, GEMM, communication, VAE, and end-to-end timings separately.
|
|
- Treat SageAttention3 as an experimental quality-gated kernel for H3.
|
|
|
|
## 2026-08-13 VAE Debug Handoff
|
|
|
|
Resolved on `vae-decode-optimization`: direct VAE temporal overlap constants now match upstream, audio decode/mux is implemented, and Comfy-equivalent FP16 video VAE is the default runtime path. The 960x544x124 cat benchmark now matches warm Comfy performance: Comfy `150.26s`, direct `149.304s` after text conditioning, direct VAE decode `25.085s`.
|
|
|
|
Current saved latent and comparison assets live under:
|
|
|
|
- Spark: `/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime`
|
|
- Share: `\\192.168.1.162\StoryStudioAssets\H3-output\h3-blackwell-runtime`
|
|
|
|
Key assets:
|
|
|
|
- Good sampled latent: `direct-cat-house-backflip-disco-shades-960x544-5s-latent.pt`
|
|
- Known-good same-latent Comfy/upstream VAE decode: `direct-cat-house-backflip-disco-shades-960x544-5s-upstream-comfy-tiled-decode.mp4`
|
|
- Current direct standalone FP32/Sage VAE decode candidate: `direct-cat-house-backflip-disco-shades-960x544-5s-direct-fp32-sage-vae-decode.mp4`
|
|
|
|
What is proven:
|
|
|
|
- The latent is good. Same latent decoded through Comfy/upstream VAE is visually clean.
|
|
- The sampler/model path is already exact against Comfy free-run parity for the fixed baseline.
|
|
- Direct VAE BF16/checkpoint-dtype loading was wrong. Loading direct VAE weights as FP32 reduced isolated decoder clip drift from roughly `mean=1.5e-3, max=1e-1` to roughly `mean=1e-6, max=6e-5`.
|
|
- Individual direct tiled VAE clips match upstream closely (`mean` around `4e-7`).
|
|
- Full direct `decode_temporal()` still differs from upstream at global frames `17, 34, 51, 68, 85, 102`, i.e. temporal join boundaries.
|
|
- Temporal assembly tracer shows pre-blend current chunks and overlap tails each match upstream, but blended join output differs hugely (`mean` around `0.15`, max around `5`) when comparing direct blend result to upstream blend result using their respective near-identical inputs.
|
|
- Direct and upstream `blend()` return identical results on the exact same inputs, so the remaining issue is likely an input/aliasing/dtype/shape subtlety at the temporal join, not the blend formula itself.
|
|
|
|
Relevant debug tools committed:
|
|
|
|
- `tools/decode_video_latent.py`
|
|
- `tools/compare_frame_dirs.py`
|
|
- `tools/compare_vae_decoder_clip.py`
|
|
- `tools/compare_vae_full_decode.py`
|
|
- `tools/compare_vae_tiled_clip.py`
|
|
- `tools/compare_vae_temporal_assembly.py`
|
|
|
|
Next VAE debugging steps:
|
|
|
|
- In `tools/compare_vae_temporal_assembly.py`, compare `prev_d` vs `prev_u` and `part_d` vs `part_u` after casting both pairs to a shared dtype and before blending. The current stats say they are close, but the blend of respective inputs explodes, which suggests a subtle shape/stride/dim broadcasting mismatch.
|
|
- Log `shape`, `stride`, `dtype`, `is_contiguous`, and `storage_offset` for `prev_*`, `part_*`, blend weights, and slices at `chunk_1_blended_at_17`.
|
|
- Try forcing `prev_d`, `prev_u`, `part_d`, and `part_u` to `.contiguous()` immediately before temporal blend in both direct and tracer paths.
|
|
- If that fixes it, patch direct `decode_temporal()` only. If not, compare exact selected slices (`prev[..., -9:, :, :]`, `part[..., :9, :, :]`) elementwise before and after multiplying weights.
|
|
|
|
Backends/experiments to evaluate later:
|
|
|
|
- Attention backend sweeps after VAE quality is fixed: `sage2`, `sage3`, `sdpa`, plus future FlashAttention-4/Sol-Attn/KJ candidates.
|
|
- Spectrum MiniMax H3 repo to inspect later: https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
|
|
|
|
## H3 Audio VAE Notes
|
|
|
|
Comfy source references:
|
|
|
|
- `comfy/ldm/minimax/audio_vae.py`: `MiniMaxH3AudioVAE`
|
|
- `comfy_extras/nodes_minimax_h3.py`: AV latent creation and audio guide encoding
|
|
- `comfy_extras/nodes_lt_audio.py`: `LTXVAudioVAEDecode`, which also handles nested AV latents and calls `audio_vae.decode(audio_latent)`
|
|
- `comfy_extras/nodes_audio.py`: generic audio save/preview and normalization helpers
|
|
|
|
Local checkpoint:
|
|
|
|
- Spark: `/home/daniel/aeon-spark-test/h3/comfy-models/vae/minimax_h3_audio_vae_fp32.safetensors`
|
|
- Runtime mount: `/vae/minimax_h3_audio_vae_fp32.safetensors`
|
|
- Shape inspection: 917 keys; top-level names include `dec_in_proj.*`, `decoder.*`, `encoder.*`, `latents_mean`, `latents_std`, `mean_proj.*`, `logs_proj.*`, and `pre_block.*`.
|
|
|
|
Architecture summary:
|
|
|
|
- This is not similar to the video ViT3D VAE. It is a DAC-lineage waveform encoder plus BigVGAN decoder.
|
|
- Latents are normalized `[B, 32, 2, T]`: 32 channels, 2 stereo channels, 40 latent frames/sec.
|
|
- Decoder output is stereo waveform `[B, 2, L]` at 32 kHz.
|
|
- One audio latent frame equals 800 samples (`32000 / 40`).
|
|
- Comfy decode path for nested H3 AV latent is effectively: select nested audio tensor, call `audio_vae.decode(audio_latent)`, then return waveform with sample rate 32000. VideoHelperSuite/Comfy audio save nodes use ffmpeg or audio helpers to save/mux.
|
|
|
|
Direct runtime implementation path:
|
|
|
|
- Change `sample_video_res_multistep()` to return both final video and final audio latents, or add a `sample_av_res_multistep()` wrapper that preserves the existing video-only API.
|
|
- Port `MiniMaxH3AudioVAE` from Comfy into a standalone `audio_vae_decoder.py`, starting decoder-only if we only need generated audio.
|
|
- Load `/vae/minimax_h3_audio_vae_fp32.safetensors` strictly in FP32, mirroring the video VAE precision lesson.
|
|
- Add a latent-only decode tool that accepts saved audio latent `[1,32,2,T]`, writes WAV/FLAC at 32 kHz, and optionally muxes with MP4 using ffmpeg.
|
|
- Validate first by capturing or decoding the same final audio latent through Comfy and direct, comparing waveform tensors before muxing.
|