24 KiB
H3 Runtime Current State
Status date: 2026-08-26
This document is the canonical snapshot of implemented scope and remaining work.
Historical handoffs in PLAN.md and PARITY.md may describe older states.
Implemented And Validated
- Single-GPU prompt-only T2VA with joint video/audio generation.
- First-frame I2VA, last-frame L2VA, and first/last FL2VA through the shared keyframe-conditioning path.
- Qwen text and vision conditioning, token refinement, video VAE encoding, H3 packed denoising, beta/RES sampling, video/audio decoding, and final MP4 mux.
- Resident HTTP runtime with warmup, readiness reporting, request-level backend selection, timing stages, canonical-benchmark-only per-step CUDA timings, peak sampling memory and latent checksums, FC2 dispatch deltas, optional latent saving, and diagnostics.
- SageAttention2 as the default quality backend.
- SDPA, forced cuDNN SDPA, FlashAttention-4, Sage3, Comfy Kitchen INT8, KJ Sage, head-sliced, and Sol-Attn experimental backends.
- Official FL2VA Turbo 4-step and 8-step adapters.
- Optional resident H3-native latent upscaling.
- Experimental EasyCache and H3-Cache delta-reuse modes.
- Quoted dialogue as the project prompt default. In a matched 10-seed test, quoted dialogue eliminated immediate first-100ms activity in all ten cases and all ten quoted WAVs passed subjective review.
- Ragged Ulysses sequence parallelism with 2/4/6/8-rank transport tests.
- Sequence-sharded 50-block execution and distributed final projection.
- True H3 NVFP4 tensor parallelism for attention QKV/output and MLP FC1/FC2.
- Bit-exact fused H3 modulation and residual gates on GB10:
3.0%faster over the canonical 12-step 1344x768/124-frame sampling run with identical video and audio checksums,3.6-3.9%faster individual blocks, approximately 390 MiB lower peak allocated memory, and 32 passing deployed tests. - Automatic visible-GPU launchers and 1/2/4/6/8 benchmark matrix tooling.
- Real-checkpoint one-rank Ulysses-versus-TP identity at 864x480, 141 frames, and 12 steps, including exact video and audio latent equality.
- Matched one-GPU RTX PRO 6000 Blackwell Server SDPA sampling averaged
28.50sover two runs versus126.66sfor the same tensor runner on GB10 (4.44x). RTX repeat variance was0.49%and checksums were identical between repeats. - Two-GPU Ulysses SDPA speedup grows with sequence size:
1.21xat 864x480/141 frames,1.55xat 1344x768/124 frames, and1.72xat 1344x768/243 frames. The tested cards have no NVLink; P2P read/write is available and NCCL usesP2P/CUMEM.
Primary Missing Scope
Full Ref2VA
- Arbitrary reference image, video, and audio inputs.
- Reference-audio encoder and reference soundtrack conditioning.
- Reference identity/voice blocks in the standalone packer.
- Ref2VA position, modality, and scheduling contracts.
- Direct-versus-Comfy full Ref2VA per-step and final-output parity benchmark.
Explicit Task API
- Named
taskselection for T2VA, I2VA, L2VA, FL2VA, and Ref2VA. - Mode-specific request schemas and incompatible-input validation.
- Intermediate keyframe anchors beyond the current first/last restriction.
Distributed Execution
- Output parity above one rank; real two-rank NCCL transport is validated.
- 2/4/6/8-GPU topology and performance sweeps on one Blackwell machine.
- Distributed resident-service orchestration; the current launcher is batch
generation through
torchrun. - x86 SageAttention2 packaging; RunPod validation initially uses SDPA.
Owned Performance Kernels
- H3-specific attention backend optimized for real GB10 tensor shapes.
- Blackwell-native CUTLASS/CuTe or cuBLASLt NVFP4 GEMMs.
- CUDA graph capture and shape buckets.
The four-GEMM roofline study and canonical FC2 cuBLASLt integration are complete. Remaining owned-GEMM work concerns QKV, attention output, and FC1 only and must be justified against the post-FC2 profile.
The active NVFP4 fusion profile is now measured on one canonical GB10 block: 32 of 53 launches belong to the four scale/pack/GEMM paths. Native packed data, block scales, and linear outputs are exact at every H3 projection width after generalizing the block-scale swizzle. The next implementation must remove intermediate traffic across these exact boundaries rather than deploy the standalone packer, which is not consistently faster than Comfy Kitchen.
The active subcomponent is direct QKV projection into Sage2's required layout,
followed by direct attention output into the NVFP4 output projection's
token-major layout. The existing scaled_mm_nvfp4 wrapper exposes only a
contiguous BF16 output, so true copy elimination requires either a supported
strided epilogue from its underlying CUTLASS kernel or an owned projection
collective. A post-GEMM copy kernel is useful only as a diagnostic and does not
satisfy this boundary-removal target. For single-GPU Sage2, the supported
strided NHD interface made that post-GEMM kernel unnecessary: Q/K/V remain
strided views of the interleaved projection output and Sage emits contiguous
token-major-compatible NHD output. The bit-exact path reduces canonical sampling
from 301.05 s to 290.23 s and is enabled for Spark deployments with
H3_SAGE_QKV_LAYOUT=strided_nhd.
NVFP4 streaming feasibility is confirmed but not yet deployable. Comfy Kitchen's cuBLAS interface requires complete activation and scale pointers, while CUTLASS DSL 4.6.2 runs block-scaled FP4 on SM121 and accepts the same logical H3 data. The experimental alpha-before-BF16 epilogue is bit-exact for QKV, attention output, and FC1 on 128-row real tiles. FC2 still differs because its reference uses a different reduction policy, so no streamed producer is enabled.
The explicit implementation policy is to retain FC2 on cuBLAS and develop the streamed CuTe path only for QKV, attention output, and FC1. This is a numerical fallback, not a silent compatibility path: FC2's reference reduction order is part of the exactness contract.
The fixed 128-row P1 producer-consumer checkpoint is complete. The existing
32-thread DMA warp now produces four BF16 activation rows per lane directly into
the owned GEMM's staged E2M1 A and E4M3 SFA shared-memory layouts. B/SFB remain
on TMA, and their completion publishes the stage to the unchanged MMA consumer.
No complete global activation QDATA or SFA tensor is passed to the streamed
kernel. Every real 128-K tile and the complete BF16 GEMM output are bit-exact for
QKV, attention output, and FC1. Matching Comfy requires its --use_fast_math
rcp.approx.ftz.f32 encode-scale operation. FC2 streaming is explicitly
rejected and retains the cuBLAS fallback. The prototype is validator-only and
still needs canonical M/padding support and runtime packaging.
See benchmarks/gb10-cute-p1-stream-a-summary.json.
The first timing gate rejects direct per-CTA streaming. For 128 rows, the exact
streamed kernel is 9.14-12.98x slower than the complete Vortex-scale plus
Comfy-pack/GEMM reference because every output-N CTA rereads and repacks A.
Measured producer cost is approximately 0.56-0.60 us per (N,K) tile, and
break-even would require reusing A across 28-102 N tiles. Duplicating that many
accumulators is not viable. The active design is now a bounded global packed-tile
ring or persistent work queue that produces each (M,K) tile once, shares it
across N consumers, and recycles the slot without materializing the complete
activation. See benchmarks/gb10-cute-p1-stream-a-timing-summary.json.
The bounded-ring implementation uses caller-owned native QDATA/SFA buffers and
an allocation-free _into producer. A capacity sweep with one full-activation
scale selected 2048 rows: 6.19 MB for QKV/FC1 and 8.26 MB for attention output.
At that capacity, complete 37,810-row projection parity is exact for QKV,
attention output, and FC1 in blocks 0, 24, and 49, including the final 946-row
chunk. The subsequent single-model alternating gate rejects runtime QKV dispatch:
all three block outputs are bit-exact, but median block time regresses by 0.52%
to 0.81%. QKV capacity checks at 3072, 4096, 8192, and 37888 rows also fail to
produce a block-level gain; 4096 is closest at 0.24% slower. Attention output
and FC1 remain experimental, and FC2 remains on Comfy/cuBLAS. Do not run
trajectory validation or enable the backend until launch fusion or a different
persistent scheduler passes this block gate. See
benchmarks/gb10-cute-qkv-runtime-block-gate-summary.json.
The historical pre-FC2 post-optimization canonical run was 288.93 s with unchanged video
and audio checksums. Block 24 is 468.22 ms median, of which production NHD
Sage2 attention consumes 258.47 ms. Internal attribution places 238.81 ms
in the SM89 attention mainloop, versus 7.68 ms Q/K quantization and 10.65 ms
V quantization. The deployed SM121 path therefore still spends most of the
block in an Ada-style MMA kernel. Recompiling SageAttention's Hopper WGMMA
mainloop for SM121 is not possible: CUDA 13 ptxas rejects WGMMA instructions for
sm_121a. CUTLASS SM120/121 UMMA supports F8/F6/F4, not the INT8 QK operation
required for exact Sage2 parity, so a native exact attention rewrite is paused.
The next practical boundary was the two approximately 10.5 ms AdaLN modulation
passes. Their exact BF16 values now feed NVFP4 scale/pack for QKV and FC1 without
materializing the modulated values, while retaining the current Comfy GEMMs.
The producer is byte-exact for complete block 0, 24, and 49 inputs. Integrated
block medians improve by 0.28-0.79%; warmed two-step and canonical 12-step
trajectories improve by 0.52% and 0.56%, respectively, with bit-identical
video and audio tensors. Spark enables the path with
H3_NVFP4_MODULATE_FUSION=1. See NVFP4_MODULATE_FUSION_DESIGN.md,
SAGE2_BLACKWELL_DESIGN.md, and
benchmarks/gb10-post-optimization-profile-summary.json.
The FC2 cuBLASLt scheduling study and guarded canonical-shape integration are
complete. The production
heuristic's _stream_k kernel requests the same 25.664 GB of operands as the
retained public split-K-1 schedule, but its L2 hit rate is only 53.32% versus
91.10%; it incurs 9.853 GB more L2 read misses and spends heavily in
synchronization polling. Algorithm 70, tile 20, stages 37, split-K 1 is
byte-exact with zero workspace. It improves complete blocks 0, 24, and 49 by
8.16-8.88%, the two-step trajectory by 7.50%, and the canonical 12-step
trajectory from 278.201 s to 255.371 s (8.21%) with exact video and audio
latents. The final hardened production method improves 20-round blocks by
7.86-9.36% and the canonical 12-step trajectory from 286.431 s to
262.979 s (8.19%), with 600 successful dispatches, zero fallback, and exact
latents. H3_NVFP4_FC2_LT_SPLITK1=1 enables only the validated M=37,810
descriptor; nearby row counts can differ by two BF16 elements and therefore
retain the existing Comfy fallback. The extension and measured runtime ABI are
prepared during H3 model loading rather than on the first canonical request.
The resident service has now passed post-FC2 deployment and repeated canonical
baseline validation and is active on Spark at port 8001. See
research/fc2_nvfp4_scheduling/RESULTS.md.
The Spark hot runtime was rebuilt and recreated with image
sha256:5f879c43374bcedb89745971d7d95d82afc8fcf9c41d30f11b257c2863b9fe28.
Health and startup warmup pass with modulation fusion enabled. A resident real
generation smoke completed in 2.14 s (0.227 s sampling) and produced a valid
22-frame 320x192 H.264 file. The first startup warmup includes one-time CUDA
extension compilation. See
benchmarks/gb10-nvfp4-modulate-fusion-deployment-smoke.json.
Exact SwiGLU-to-FC2 NVFP4 producer fusion is also complete. It preserves the
two reference BF16 boundaries, emits byte-identical tensor scale/QDATA/SFA, and
retains the exact Comfy FC2 GEMM. Blocks 0, 24, and 49 improve by 2.14-2.26%.
The warmed canonical 12-step run improves from 289.14 s to 277.36 s
(4.07%) with bit-identical video and audio tensors. Spark enables it with
H3_NVFP4_SWIGLU_FUSION=1. See NVFP4_SWIGLU_FUSION_DESIGN.md and
benchmarks/gb10-nvfp4-swiglu-fusion-summary.json.
Optional BF16 materialization inside both fused producers was tested for active
Turbo LoRA requests. The isolated canonical Turbo-4 trajectory was bit-exact,
but regressed from 131.11 s to 135.14 s (3.08%), so the prototype was
rejected. Active LoRA retains the exact materialized fallback instead of using
either producer fusion. See
benchmarks/gb10-nvfp4-lora-producer-fusion-turbo4-isolated.json.
The authoritative profile image is
sha256:a15d0c09dd8cc82aaf2b564d3da760ea5ab8dec974f73075f7d30ac3a504815c.
The active production overlay is
sha256:6d880d628334c981c3d155bf5244e65e26e22cc9273c80145f646eee3c3698c2;
it inherits the profiled binary/ABI layers and makes telemetry
canonical-benchmark-only. Production flags enable both accepted producers and
H3_NVFP4_FC2_LT_SPLITK1=1; ordinary requests do not create step events or
copy latents for checksums. /ready, startup warmup, repeated exact canonical
generation, profiler captures, and 58 tests pass.
A previous pre-FC2 fully fused Nsight recapture superseded the old approximately
515 ms block profile. Block 24 is 458.78 ms median uninstrumented and
465.78 ms across the Nsight GPU span, with 41 kernels and only 0.084 ms of
inter-kernel idle time. Sage2 is 57.60% of kernel time, the four NVFP4 GEMMs
are 26.87%, packing is 8.92%, norm/RoPE is 4.27%, and the two remaining
gate/add kernels are 2.35%. A complete warmed step takes 23.708 s and shows
the same distribution. Hardware counters attribute 56.25% of the warm-cache
off-chip request proxy to the NVFP4 GEMMs even though Sage2 remains the time
bottleneck. See benchmarks/gb10-fully-fused-fresh-nsight-summary.json.
The authoritative post-FC2 resident baseline is now the median of three warmed,
unprofiled canonical runs: 256.464, 255.447, and 255.135 s, giving
255.447 s. Every run produced the established video SHA-256
c62d23a42972eab907ba42f93c50247ff17a9c454b4a53fe93d2e34f9fefe578 and audio
SHA-256 852005383770480a6503504e1ffec86dd1fb63a69c6400f92da18e39e0986de2,
with 600/600 FC2 dispatches and zero fallback. Peak sampling allocation and
reservation were 44,445,830,144 and 48,708,452,352 bytes.
The same-script generic block-24 decomposition control is 427.410 ms; it
bypasses guarded FC2 and is not production-FC2 timing. A complete warmed step
has a 20.907 s Nsight GPU span, 20.895 s kernel time, 2,744 kernels,
11.519 ms total launch gaps, and 99.98% launch-API/GPU overlap. The new time
ranking is Sage2 62.36%, NVFP4 GEMMs 19.95%, packing 9.70%, norm/RoPE
5.19%, and gate/add 2.63%. Sage2 is therefore the next-ranked investigation,
but no new optimization has begun. See
benchmarks/gb10-post-fc2-production-profile-summary-20260826.json.
The real block-24 Sage2 scheduler study and exact SM89 P0 retune are complete.
Manual preparation plus the
unchanged prequantized SM89 mainloop is byte-exact against public SageAttention
2.2.0. Uninstrumented medians are 2.37 ms for K mean/smoothing, 3.80 ms for
Q quantization, 3.83 ms for K subtract-mean quantization, 5.08 ms for V
transpose/pad/permute, 5.53 ms for V scale/FP8 quantization, and 237.09 ms
for the fused mainloop. Nsight Compute reports 255 registers/thread, 32 KiB
dynamic shared memory/CTA, 16.83% achieved occupancy, and no eligible warp in
63.53% of scheduler cycles. INT8 QK and FP8 PV each use 37.77% of elapsed
tensor-pipe capacity; combined tensor activity is 75.54%. The kernel is
scheduler/compute limited rather than off-chip-bandwidth limited: L2 hit rate is
98.84%, while fixed-latency dependency and math-pipe stalls dominate. Tail
CTAs add less than 1 ms.
The P0 mapped the exact register cliff: caps from 255 through 170 registers
remain at two CTAs and 16.67% theoretical occupancy; only 168 registers reaches
three CTAs and 25%, while generating 4.95 billion local spill requests and
worsening no-eligible cycles to 78.79%. Narrowed scopes reduced static spills
from 44/44 to 12/12 bytes and dynamic spill requests from 1.46 million to
0.40 million, but changed interleaved latency by only +0.06% and worsened
no-eligible cycles. In-place score reuse, early K prefetch, and independent
softmax-chain interleaving were also byte-exact and neutral or slower.
The 7.68% shared excess maps entirely to repeated V-staging LDGSTS.128
instructions. Padding V to a 128-byte shared stride increased shared memory to
40 KiB but left all 626,970,624 excessive wavefronts unchanged and changed
latency by -0.04%. No variant crossed the 3% complete-block gate, so none was
integrated or deployed. See
benchmarks/gb10-sage2-p0-register-scheduler-analysis.json and the associated
P0 latency JSON and NCU reports.
The exact Sage2 entry-fusion P1 is also complete and rejected. A single CUDA
kernel fused strided-NHD Q/K RMSNorm, split-half RoPE, and Sage2 Q INT8
quantization while leaving K mean/quantization, V preparation, and the SM89
mainloop unchanged. Randomized edge lengths and canonical block 0/24/49 tensors
were bit-exact through prepared Q/K, Q/K quantization, scales, K mean, Sage2
output, and complete block output. Entry-only median latency improved by
20.9-23.4%, but canonical complete-block median improvement was only 0.73%,
0.86%, and 0.53% for blocks 0, 24, and 49. The candidate removed one launch
(87 to 86), did not change peak memory, and reduced complete-block L2 traffic
by only 0.136-0.155%. It therefore failed the required 1% gate; the opt-in
runtime branch was removed and two-step/12-step validation was skipped. See
benchmarks/gb10-sage2-p1-entry-fusion-analysis.json and its referenced parity,
timing, and Nsight reports.
The exact Sage2 V-preparation P2 is complete and rejected at its isolated gate.
An owned three-stage CUDA path consumes projection-strided NHD BF16 V and emits
Sage2's padded/permuted E4M3 V plus FP32 per-channel scales without materializing
the approximately 517 MiB BF16 transpose tensor. FP8 bytes and scales are exact
for 13 boundary lengths from 1 through 37,810 tokens with 56 heads. On the
canonical shape, median V preparation improves from 10.56 ms to 6.39 ms
(39.48%), but the 4.17 ms absolute saving projects to only 0.91% of the
458.78 ms complete block and misses the required 6.0 ms isolated go gate.
Complete-block and trajectory validation were therefore skipped, and production
dispatch remains unchanged. See
benchmarks/gb10-sage2-p2-vprep-analysis.json.
The exact Sage2 mainloop P3 temporal-pair experiment is also complete and
rejected. Two warp pairs alternated QK/online-softmax and prior-tile PV while
retaining private per-warp scores, softmax state, and output accumulators. The
isolated extension is sanitizer-clean and byte-exact over 13 adversarial short
shapes plus the real 37,810-token block-24 SHA. In a 50-sample alternating run,
mainloop median changed from 245.44 ms to 245.20 ms, only 0.10%, and
missed the absolute <220 ms gate. Ptxas reports 254 registers/thread and
32/24-byte static store/load spills versus baseline 255 registers and 24/24-byte
spills. NCU, block integration, and trajectory validation were skipped.
Production remains unchanged. See
benchmarks/gb10-sage2-p3-temporal-pair-analysis.json and
research/sage2_temporal_pair/.
Vortex Exact Attention is initialized as an isolated clean-sheet research
project under research/vortex_exact_attention/. Phase 0 imports and verifies
the retained SageAttention 2.2.0 exactness contract at commit
d1a57a546c3d395b1ffcbeecc66d81db76f3b4b5; it does not repeat P0-P3 or alter
dispatch. Ten retained oracle artifacts match their recorded SHA-256 values.
Phase 1 selects the VEA-B Q128 paired-owner architecture: four QK/softmax warps
own RS/RS_f8/m/d, four separate PV warps own RO, and two producer warps own
K/V staging. Its 180-207 ms mainloop range is a heuristic screen, not achieved.
Phase 2A capability probes now pass with CUDA block-scope mbarriers: 96
registers/thread for handoff, 139 for synthetic combined ownership, zero spills,
one ten-warp CTA/SM, 0.076800 ms barrier-only p50, and concurrent INT8/FP8
progress in 48/48 blocks. Canonical Q/K/V/output fixtures are captured and
reload-verified against the locked Sage2 output hash. The inline named-barrier
primitive is rejected on racecheck. One isolated aligned-shape prototype is
authorized; no attention kernel or attention speedup exists. Production remains
Sage2.
The four-GEMM roofline selected FC2, and the guarded public split-K-1 schedule
closed that target. Fresh NCU values at block 24 are 26.38 ms QKV, 8.48 ms
attention output, 36.17 ms FC1, and 16.20 ms FC2. Sage2 remains much larger:
its NCU-replayed mainloop is 259.00 ms, with 255 registers/thread, 16.65%
achieved occupancy, 98.85% L2 hit rate, and 161.50 GB L2 requests. No new
kernel work starts until this ranking is accepted.
Quality Work Remaining
- Generate full quoted-dialogue videos and validate wording, voice consistency, speech timing, and lip-sync before closing the startup-audio work.
- Complete strict per-step LightX2V parity for Turbo adapters.
- Add real-adapter Turbo end-to-end fixtures.
- Add the optional target-resolution refinement stage after latent upscaling.
- Resolve or formally bound upscaler ringing, texture, chromatic-edge, and identity changes.
- Run full-size cache threshold and quality sweeps before enabling caches for production output.
- Keep Sage3, Sol-Attn, INT8, and other approximate backends quality-gated.
- Fix the inactive fused Sol QKV-layout path, which currently references an
undefined
qkvvalue. The deployed native Sol layout does not use this path.
Production Work Remaining
- Asynchronous jobs, queueing, progress, cancellation, and timeouts.
- Strict request validation, including Boolean fields and mode combinations.
- Input/output path sandboxing, request-size limits, authentication, and TLS.
- Configurable FPS, video codec, audio codec, sample rate, and media policy.
- Container healthcheck, restart policy, resource limits, durable structured request logs, and runtime metrics.
- Batch generation and an intentional worker/concurrency model.
Validation And Packaging Gaps
- GPU end-to-end fixtures for T2VA, I2VA, L2VA, and FL2VA.
- Full Ref2VA, AudioVAE waveform, cache, HTTP API, real Turbo, real upscaler, attention-quality, CUDA-graph, and distributed tests.
- Reproducible local fixtures for parity evidence currently stored on Spark/SMB.
- Explicit package declarations/checks for NumPy, SciPy, Pillow, and FFmpeg.
- A standalone base image if removing the Comfy-derived image becomes a product requirement; the current denoising path still intentionally uses Comfy Kitchen kernels.
- Align Docker
H3_MODEL_PATHandRuntimeConfig; the environment variable is currently not consumed by the runtime default.
Recommended Execution Order
- Treat the post-FC2 profile as the GB10 baseline; investigate Sage2 only under a separately approved experiment with exactness and absolute latency gates.
- Capture matched SM120 and SM100 component profiles and package Sage2 on SM120.
- Resume NVFP4 GEMM/epilogue work only with a design that preserves the accepted producer fusions and exact BF16 boundaries.
- Validate full quoted-dialogue video lip-sync and close the audio prompt change.
- Add explicit task schemas and automated single-GPU mode tests.
- Implement full Ref2VA, including reference-audio encoding.
- Add CUDA graph buckets after the kernel and shape policies stabilize.
- Harden the service API and operational deployment.
- Complete RunPod NCCL validation and distributed scaling benchmarks.
See PERFORMANCE_ROADMAP.md for measured component costs, architecture-specific
targets, quality gates, and the rationale for this ordering.
The current single-GPU T2VA/FL2VA runtime is mature. Distributed execution is implemented and CPU/one-GPU validated, with real multi-GPU NCCL results still blocked on an eight-GPU host. The other largest gap is standalone Ref2VA.