# H3 Runtime Current State Status date: 2026-08-26 This document is the canonical snapshot of implemented scope and remaining work. Historical handoffs in `PLAN.md` and `PARITY.md` may describe older states. ## Implemented And Validated - Single-GPU prompt-only T2VA with joint video/audio generation. - First-frame I2VA, last-frame L2VA, and first/last FL2VA through the shared keyframe-conditioning path. - Qwen text and vision conditioning, token refinement, video VAE encoding, H3 packed denoising, beta/RES sampling, video/audio decoding, and final MP4 mux. - Resident HTTP runtime with warmup, readiness reporting, request-level backend selection, timing stages, canonical-benchmark-only per-step CUDA timings, peak sampling memory and latent checksums, FC2 dispatch deltas, optional latent saving, and diagnostics. - SageAttention2 as the default quality backend. - SDPA, forced cuDNN SDPA, FlashAttention-4, Sage3, Comfy Kitchen INT8, KJ Sage, head-sliced, and Sol-Attn experimental backends. - Official FL2VA Turbo 4-step and 8-step adapters. - Optional resident H3-native latent upscaling. - Experimental EasyCache and H3-Cache delta-reuse modes. - Quoted dialogue as the project prompt default. In a matched 10-seed test, quoted dialogue eliminated immediate first-100ms activity in all ten cases and all ten quoted WAVs passed subjective review. - Ragged Ulysses sequence parallelism with 2/4/6/8-rank transport tests. - Sequence-sharded 50-block execution and distributed final projection. - True H3 NVFP4 tensor parallelism for attention QKV/output and MLP FC1/FC2. - Bit-exact fused H3 modulation and residual gates on GB10: `3.0%` faster over the canonical 12-step 1344x768/124-frame sampling run with identical video and audio checksums, `3.6-3.9%` faster individual blocks, approximately 390 MiB lower peak allocated memory, and 32 passing deployed tests. - Automatic visible-GPU launchers and 1/2/4/6/8 benchmark matrix tooling. - Real-checkpoint one-rank Ulysses-versus-TP identity at 864x480, 141 frames, and 12 steps, including exact video and audio latent equality. - Matched one-GPU RTX PRO 6000 Blackwell Server SDPA sampling averaged `28.50s` over two runs versus `126.66s` for the same tensor runner on GB10 (`4.44x`). RTX repeat variance was `0.49%` and checksums were identical between repeats. - Two-GPU Ulysses SDPA speedup grows with sequence size: `1.21x` at 864x480/141 frames, `1.55x` at 1344x768/124 frames, and `1.72x` at 1344x768/243 frames. The tested cards have no NVLink; P2P read/write is available and NCCL uses `P2P/CUMEM`. ## Primary Missing Scope ### Full Ref2VA - Arbitrary reference image, video, and audio inputs. - Reference-audio encoder and reference soundtrack conditioning. - Reference identity/voice blocks in the standalone packer. - Ref2VA position, modality, and scheduling contracts. - Direct-versus-Comfy full Ref2VA per-step and final-output parity benchmark. ### Explicit Task API - Named `task` selection for T2VA, I2VA, L2VA, FL2VA, and Ref2VA. - Mode-specific request schemas and incompatible-input validation. - Intermediate keyframe anchors beyond the current first/last restriction. ### Distributed Execution - Output parity above one rank; real two-rank NCCL transport is validated. - 2/4/6/8-GPU topology and performance sweeps on one Blackwell machine. - Distributed resident-service orchestration; the current launcher is batch generation through `torchrun`. - x86 SageAttention2 packaging; RunPod validation initially uses SDPA. ### Owned Performance Kernels - H3-specific attention backend optimized for real GB10 tensor shapes. - Blackwell-native CUTLASS/CuTe or cuBLASLt NVFP4 GEMMs. - CUDA graph capture and shape buckets. The four-GEMM roofline study and canonical FC2 cuBLASLt integration are complete. Remaining owned-GEMM work concerns QKV, attention output, and FC1 only and must be justified against the post-FC2 profile. The active NVFP4 fusion profile is now measured on one canonical GB10 block: 32 of 53 launches belong to the four scale/pack/GEMM paths. Native packed data, block scales, and linear outputs are exact at every H3 projection width after generalizing the block-scale swizzle. The next implementation must remove intermediate traffic across these exact boundaries rather than deploy the standalone packer, which is not consistently faster than Comfy Kitchen. The active subcomponent is direct QKV projection into Sage2's required layout, followed by direct attention output into the NVFP4 output projection's token-major layout. The existing `scaled_mm_nvfp4` wrapper exposes only a contiguous BF16 output, so true copy elimination requires either a supported strided epilogue from its underlying CUTLASS kernel or an owned projection collective. A post-GEMM copy kernel is useful only as a diagnostic and does not satisfy this boundary-removal target. For single-GPU Sage2, the supported strided NHD interface made that post-GEMM kernel unnecessary: Q/K/V remain strided views of the interleaved projection output and Sage emits contiguous token-major-compatible NHD output. The bit-exact path reduces canonical sampling from `301.05 s` to `290.23 s` and is enabled for Spark deployments with `H3_SAGE_QKV_LAYOUT=strided_nhd`. NVFP4 streaming feasibility is confirmed but not yet deployable. Comfy Kitchen's cuBLAS interface requires complete activation and scale pointers, while CUTLASS DSL 4.6.2 runs block-scaled FP4 on SM121 and accepts the same logical H3 data. The experimental alpha-before-BF16 epilogue is bit-exact for QKV, attention output, and FC1 on 128-row real tiles. FC2 still differs because its reference uses a different reduction policy, so no streamed producer is enabled. The explicit implementation policy is to retain FC2 on cuBLAS and develop the streamed CuTe path only for QKV, attention output, and FC1. This is a numerical fallback, not a silent compatibility path: FC2's reference reduction order is part of the exactness contract. The fixed 128-row P1 producer-consumer checkpoint is complete. The existing 32-thread DMA warp now produces four BF16 activation rows per lane directly into the owned GEMM's staged E2M1 A and E4M3 SFA shared-memory layouts. B/SFB remain on TMA, and their completion publishes the stage to the unchanged MMA consumer. No complete global activation QDATA or SFA tensor is passed to the streamed kernel. Every real 128-K tile and the complete BF16 GEMM output are bit-exact for QKV, attention output, and FC1. Matching Comfy requires its `--use_fast_math` `rcp.approx.ftz.f32` encode-scale operation. FC2 streaming is explicitly rejected and retains the cuBLAS fallback. The prototype is validator-only and still needs canonical M/padding support and runtime packaging. See `benchmarks/gb10-cute-p1-stream-a-summary.json`. The first timing gate rejects direct per-CTA streaming. For 128 rows, the exact streamed kernel is `9.14-12.98x` slower than the complete Vortex-scale plus Comfy-pack/GEMM reference because every output-N CTA rereads and repacks A. Measured producer cost is approximately `0.56-0.60 us` per `(N,K)` tile, and break-even would require reusing A across 28-102 N tiles. Duplicating that many accumulators is not viable. The active design is now a bounded global packed-tile ring or persistent work queue that produces each `(M,K)` tile once, shares it across N consumers, and recycles the slot without materializing the complete activation. See `benchmarks/gb10-cute-p1-stream-a-timing-summary.json`. The bounded-ring implementation uses caller-owned native QDATA/SFA buffers and an allocation-free `_into` producer. A capacity sweep with one full-activation scale selected 2048 rows: 6.19 MB for QKV/FC1 and 8.26 MB for attention output. At that capacity, complete 37,810-row projection parity is exact for QKV, attention output, and FC1 in blocks 0, 24, and 49, including the final 946-row chunk. The subsequent single-model alternating gate rejects runtime QKV dispatch: all three block outputs are bit-exact, but median block time regresses by `0.52%` to `0.81%`. QKV capacity checks at 3072, 4096, 8192, and 37888 rows also fail to produce a block-level gain; 4096 is closest at `0.24%` slower. Attention output and FC1 remain experimental, and FC2 remains on Comfy/cuBLAS. Do not run trajectory validation or enable the backend until launch fusion or a different persistent scheduler passes this block gate. See `benchmarks/gb10-cute-qkv-runtime-block-gate-summary.json`. The historical pre-FC2 post-optimization canonical run was `288.93 s` with unchanged video and audio checksums. Block 24 is `468.22 ms` median, of which production NHD Sage2 attention consumes `258.47 ms`. Internal attribution places `238.81 ms` in the SM89 attention mainloop, versus `7.68 ms` Q/K quantization and `10.65 ms` V quantization. The deployed SM121 path therefore still spends most of the block in an Ada-style MMA kernel. Recompiling SageAttention's Hopper WGMMA mainloop for SM121 is not possible: CUDA 13 ptxas rejects WGMMA instructions for `sm_121a`. CUTLASS SM120/121 UMMA supports F8/F6/F4, not the INT8 QK operation required for exact Sage2 parity, so a native exact attention rewrite is paused. The next practical boundary was the two approximately 10.5 ms AdaLN modulation passes. Their exact BF16 values now feed NVFP4 scale/pack for QKV and FC1 without materializing the modulated values, while retaining the current Comfy GEMMs. The producer is byte-exact for complete block 0, 24, and 49 inputs. Integrated block medians improve by `0.28-0.79%`; warmed two-step and canonical 12-step trajectories improve by `0.52%` and `0.56%`, respectively, with bit-identical video and audio tensors. Spark enables the path with `H3_NVFP4_MODULATE_FUSION=1`. See `NVFP4_MODULATE_FUSION_DESIGN.md`, `SAGE2_BLACKWELL_DESIGN.md`, and `benchmarks/gb10-post-optimization-profile-summary.json`. The FC2 cuBLASLt scheduling study and guarded canonical-shape integration are complete. The production heuristic's `_stream_k` kernel requests the same `25.664 GB` of operands as the retained public split-K-1 schedule, but its L2 hit rate is only `53.32%` versus `91.10%`; it incurs `9.853 GB` more L2 read misses and spends heavily in synchronization polling. Algorithm 70, tile 20, stages 37, split-K 1 is byte-exact with zero workspace. It improves complete blocks 0, 24, and 49 by `8.16-8.88%`, the two-step trajectory by `7.50%`, and the canonical 12-step trajectory from `278.201 s` to `255.371 s` (`8.21%`) with exact video and audio latents. The final hardened production method improves 20-round blocks by `7.86-9.36%` and the canonical 12-step trajectory from `286.431 s` to `262.979 s` (`8.19%`), with 600 successful dispatches, zero fallback, and exact latents. `H3_NVFP4_FC2_LT_SPLITK1=1` enables only the validated `M=37,810` descriptor; nearby row counts can differ by two BF16 elements and therefore retain the existing Comfy fallback. The extension and measured runtime ABI are prepared during H3 model loading rather than on the first canonical request. The resident service has now passed post-FC2 deployment and repeated canonical baseline validation and is active on Spark at port `8001`. See `research/fc2_nvfp4_scheduling/RESULTS.md`. The Spark hot runtime was rebuilt and recreated with image `sha256:5f879c43374bcedb89745971d7d95d82afc8fcf9c41d30f11b257c2863b9fe28`. Health and startup warmup pass with modulation fusion enabled. A resident real generation smoke completed in `2.14 s` (`0.227 s` sampling) and produced a valid 22-frame 320x192 H.264 file. The first startup warmup includes one-time CUDA extension compilation. See `benchmarks/gb10-nvfp4-modulate-fusion-deployment-smoke.json`. Exact SwiGLU-to-FC2 NVFP4 producer fusion is also complete. It preserves the two reference BF16 boundaries, emits byte-identical tensor scale/QDATA/SFA, and retains the exact Comfy FC2 GEMM. Blocks 0, 24, and 49 improve by `2.14-2.26%`. The warmed canonical 12-step run improves from `289.14 s` to `277.36 s` (`4.07%`) with bit-identical video and audio tensors. Spark enables it with `H3_NVFP4_SWIGLU_FUSION=1`. See `NVFP4_SWIGLU_FUSION_DESIGN.md` and `benchmarks/gb10-nvfp4-swiglu-fusion-summary.json`. Optional BF16 materialization inside both fused producers was tested for active Turbo LoRA requests. The isolated canonical Turbo-4 trajectory was bit-exact, but regressed from `131.11 s` to `135.14 s` (`3.08%`), so the prototype was rejected. Active LoRA retains the exact materialized fallback instead of using either producer fusion. See `benchmarks/gb10-nvfp4-lora-producer-fusion-turbo4-isolated.json`. The authoritative profile image is `sha256:a15d0c09dd8cc82aaf2b564d3da760ea5ab8dec974f73075f7d30ac3a504815c`. The active production overlay is `sha256:6d880d628334c981c3d155bf5244e65e26e22cc9273c80145f646eee3c3698c2`; it inherits the profiled binary/ABI layers and makes telemetry canonical-benchmark-only. Production flags enable both accepted producers and `H3_NVFP4_FC2_LT_SPLITK1=1`; ordinary requests do not create step events or copy latents for checksums. `/ready`, startup warmup, repeated exact canonical generation, profiler captures, and 58 tests pass. A previous pre-FC2 fully fused Nsight recapture superseded the old approximately 515 ms block profile. Block 24 is `458.78 ms` median uninstrumented and `465.78 ms` across the Nsight GPU span, with 41 kernels and only `0.084 ms` of inter-kernel idle time. Sage2 is `57.60%` of kernel time, the four NVFP4 GEMMs are `26.87%`, packing is `8.92%`, norm/RoPE is `4.27%`, and the two remaining gate/add kernels are `2.35%`. A complete warmed step takes `23.708 s` and shows the same distribution. Hardware counters attribute `56.25%` of the warm-cache off-chip request proxy to the NVFP4 GEMMs even though Sage2 remains the time bottleneck. See `benchmarks/gb10-fully-fused-fresh-nsight-summary.json`. The authoritative post-FC2 resident baseline is now the median of three warmed, unprofiled canonical runs: `256.464`, `255.447`, and `255.135 s`, giving `255.447 s`. Every run produced the established video SHA-256 `c62d23a42972eab907ba42f93c50247ff17a9c454b4a53fe93d2e34f9fefe578` and audio SHA-256 `852005383770480a6503504e1ffec86dd1fb63a69c6400f92da18e39e0986de2`, with 600/600 FC2 dispatches and zero fallback. Peak sampling allocation and reservation were `44,445,830,144` and `48,708,452,352` bytes. The same-script generic block-24 decomposition control is `427.410 ms`; it bypasses guarded FC2 and is not production-FC2 timing. A complete warmed step has a `20.907 s` Nsight GPU span, `20.895 s` kernel time, 2,744 kernels, `11.519 ms` total launch gaps, and `99.98%` launch-API/GPU overlap. The new time ranking is Sage2 `62.36%`, NVFP4 GEMMs `19.95%`, packing `9.70%`, norm/RoPE `5.19%`, and gate/add `2.63%`. Sage2 is therefore the next-ranked investigation, but no new optimization has begun. See `benchmarks/gb10-post-fc2-production-profile-summary-20260826.json`. The real block-24 Sage2 scheduler study and exact SM89 P0 retune are complete. Manual preparation plus the unchanged prequantized SM89 mainloop is byte-exact against public SageAttention 2.2.0. Uninstrumented medians are `2.37 ms` for K mean/smoothing, `3.80 ms` for Q quantization, `3.83 ms` for K subtract-mean quantization, `5.08 ms` for V transpose/pad/permute, `5.53 ms` for V scale/FP8 quantization, and `237.09 ms` for the fused mainloop. Nsight Compute reports 255 registers/thread, 32 KiB dynamic shared memory/CTA, `16.83%` achieved occupancy, and no eligible warp in `63.53%` of scheduler cycles. INT8 QK and FP8 PV each use `37.77%` of elapsed tensor-pipe capacity; combined tensor activity is `75.54%`. The kernel is scheduler/compute limited rather than off-chip-bandwidth limited: L2 hit rate is `98.84%`, while fixed-latency dependency and math-pipe stalls dominate. Tail CTAs add less than 1 ms. The P0 mapped the exact register cliff: caps from 255 through 170 registers remain at two CTAs and `16.67%` theoretical occupancy; only 168 registers reaches three CTAs and `25%`, while generating `4.95 billion` local spill requests and worsening no-eligible cycles to `78.79%`. Narrowed scopes reduced static spills from 44/44 to 12/12 bytes and dynamic spill requests from `1.46 million` to `0.40 million`, but changed interleaved latency by only `+0.06%` and worsened no-eligible cycles. In-place score reuse, early K prefetch, and independent softmax-chain interleaving were also byte-exact and neutral or slower. The `7.68%` shared excess maps entirely to repeated V-staging `LDGSTS.128` instructions. Padding V to a 128-byte shared stride increased shared memory to 40 KiB but left all `626,970,624` excessive wavefronts unchanged and changed latency by `-0.04%`. No variant crossed the 3% complete-block gate, so none was integrated or deployed. See `benchmarks/gb10-sage2-p0-register-scheduler-analysis.json` and the associated P0 latency JSON and NCU reports. The exact Sage2 entry-fusion P1 is also complete and rejected. A single CUDA kernel fused strided-NHD Q/K RMSNorm, split-half RoPE, and Sage2 Q INT8 quantization while leaving K mean/quantization, V preparation, and the SM89 mainloop unchanged. Randomized edge lengths and canonical block 0/24/49 tensors were bit-exact through prepared Q/K, Q/K quantization, scales, K mean, Sage2 output, and complete block output. Entry-only median latency improved by `20.9-23.4%`, but canonical complete-block median improvement was only `0.73%`, `0.86%`, and `0.53%` for blocks 0, 24, and 49. The candidate removed one launch (`87` to `86`), did not change peak memory, and reduced complete-block L2 traffic by only `0.136-0.155%`. It therefore failed the required 1% gate; the opt-in runtime branch was removed and two-step/12-step validation was skipped. See `benchmarks/gb10-sage2-p1-entry-fusion-analysis.json` and its referenced parity, timing, and Nsight reports. The exact Sage2 V-preparation P2 is complete and rejected at its isolated gate. An owned three-stage CUDA path consumes projection-strided NHD BF16 V and emits Sage2's padded/permuted E4M3 V plus FP32 per-channel scales without materializing the approximately 517 MiB BF16 transpose tensor. FP8 bytes and scales are exact for 13 boundary lengths from 1 through 37,810 tokens with 56 heads. On the canonical shape, median V preparation improves from `10.56 ms` to `6.39 ms` (`39.48%`), but the `4.17 ms` absolute saving projects to only `0.91%` of the `458.78 ms` complete block and misses the required `6.0 ms` isolated go gate. Complete-block and trajectory validation were therefore skipped, and production dispatch remains unchanged. See `benchmarks/gb10-sage2-p2-vprep-analysis.json`. The exact Sage2 mainloop P3 temporal-pair experiment is also complete and rejected. Two warp pairs alternated QK/online-softmax and prior-tile PV while retaining private per-warp scores, softmax state, and output accumulators. The isolated extension is sanitizer-clean and byte-exact over 13 adversarial short shapes plus the real 37,810-token block-24 SHA. In a 50-sample alternating run, mainloop median changed from `245.44 ms` to `245.20 ms`, only `0.10%`, and missed the absolute `<220 ms` gate. Ptxas reports 254 registers/thread and 32/24-byte static store/load spills versus baseline 255 registers and 24/24-byte spills. NCU, block integration, and trajectory validation were skipped. Production remains unchanged. See `benchmarks/gb10-sage2-p3-temporal-pair-analysis.json` and `research/sage2_temporal_pair/`. Vortex Exact Attention is initialized as an isolated clean-sheet research project under `research/vortex_exact_attention/`. Phase 0 imports and verifies the retained SageAttention 2.2.0 exactness contract at commit `d1a57a546c3d395b1ffcbeecc66d81db76f3b4b5`; it does not repeat P0-P3 or alter dispatch. Ten retained oracle artifacts match their recorded SHA-256 values. Phase 1 selects the VEA-B Q128 paired-owner architecture: four QK/softmax warps own `RS/RS_f8/m/d`, four separate PV warps own `RO`, and two producer warps own K/V staging. Its `180-207 ms` mainloop range is a heuristic screen, not achieved. Phase 2A capability probes now pass with CUDA block-scope mbarriers: 96 registers/thread for handoff, 139 for synthetic combined ownership, zero spills, one ten-warp CTA/SM, `0.076800 ms` barrier-only p50, and concurrent INT8/FP8 progress in 48/48 blocks. Canonical Q/K/V/output fixtures are captured and reload-verified against the locked Sage2 output hash. The inline named-barrier primitive is rejected on racecheck. One isolated aligned-shape prototype is authorized; no attention kernel or attention speedup exists. Production remains Sage2. The four-GEMM roofline selected FC2, and the guarded public split-K-1 schedule closed that target. Fresh NCU values at block 24 are `26.38 ms` QKV, `8.48 ms` attention output, `36.17 ms` FC1, and `16.20 ms` FC2. Sage2 remains much larger: its NCU-replayed mainloop is `259.00 ms`, with 255 registers/thread, `16.65%` achieved occupancy, `98.85%` L2 hit rate, and `161.50 GB` L2 requests. No new kernel work starts until this ranking is accepted. ## Quality Work Remaining - Generate full quoted-dialogue videos and validate wording, voice consistency, speech timing, and lip-sync before closing the startup-audio work. - Complete strict per-step LightX2V parity for Turbo adapters. - Add real-adapter Turbo end-to-end fixtures. - Add the optional target-resolution refinement stage after latent upscaling. - Resolve or formally bound upscaler ringing, texture, chromatic-edge, and identity changes. - Run full-size cache threshold and quality sweeps before enabling caches for production output. - Keep Sage3, Sol-Attn, INT8, and other approximate backends quality-gated. - Fix the inactive fused Sol QKV-layout path, which currently references an undefined `qkv` value. The deployed native Sol layout does not use this path. ## Production Work Remaining - Asynchronous jobs, queueing, progress, cancellation, and timeouts. - Strict request validation, including Boolean fields and mode combinations. - Input/output path sandboxing, request-size limits, authentication, and TLS. - Configurable FPS, video codec, audio codec, sample rate, and media policy. - Container healthcheck, restart policy, resource limits, durable structured request logs, and runtime metrics. - Batch generation and an intentional worker/concurrency model. ## Validation And Packaging Gaps - GPU end-to-end fixtures for T2VA, I2VA, L2VA, and FL2VA. - Full Ref2VA, AudioVAE waveform, cache, HTTP API, real Turbo, real upscaler, attention-quality, CUDA-graph, and distributed tests. - Reproducible local fixtures for parity evidence currently stored on Spark/SMB. - Explicit package declarations/checks for NumPy, SciPy, Pillow, and FFmpeg. - A standalone base image if removing the Comfy-derived image becomes a product requirement; the current denoising path still intentionally uses Comfy Kitchen kernels. - Align Docker `H3_MODEL_PATH` and `RuntimeConfig`; the environment variable is currently not consumed by the runtime default. ## Recommended Execution Order 1. Treat the post-FC2 profile as the GB10 baseline; investigate Sage2 only under a separately approved experiment with exactness and absolute latency gates. 2. Capture matched SM120 and SM100 component profiles and package Sage2 on SM120. 3. Resume NVFP4 GEMM/epilogue work only with a design that preserves the accepted producer fusions and exact BF16 boundaries. 4. Validate full quoted-dialogue video lip-sync and close the audio prompt change. 5. Add explicit task schemas and automated single-GPU mode tests. 6. Implement full Ref2VA, including reference-audio encoding. 7. Add CUDA graph buckets after the kernel and shape policies stabilize. 8. Harden the service API and operational deployment. 9. Complete RunPod NCCL validation and distributed scaling benchmarks. See `PERFORMANCE_ROADMAP.md` for measured component costs, architecture-specific targets, quality gates, and the rationale for this ordering. The current single-GPU T2VA/FL2VA runtime is mature. Distributed execution is implemented and CPU/one-GPU validated, with real multi-GPU NCCL results still blocked on an eight-GPU host. The other largest gap is standalone Ref2VA.