h3-blackwell-runtime/research/fc2_nvfp4_scheduling
2026-08-25 22:32:48 +07:00
..
__init__.py Record exact FC2 cuBLASLt schedule 2026-08-25 22:32:48 +07:00
candidate_splitk1.json Record exact FC2 cuBLASLt schedule 2026-08-25 22:32:48 +07:00
fc2_nvfp4_lt.cpp Record exact FC2 cuBLASLt schedule 2026-08-25 22:32:48 +07:00
README.md Record exact FC2 cuBLASLt schedule 2026-08-25 22:32:48 +07:00
RESULTS.md Record exact FC2 cuBLASLt schedule 2026-08-25 22:32:48 +07:00

FC2 NVFP4 Library Scheduling

This directory contains the first isolated library-scheduling study for H3 FC2. It does not register a PyTorch operator, alter Nvfp4Linear, or participate in production dispatch. The extension is loaded only by tools/benchmark_fc2_nvfp4_algorithms.py.

The completed measurements and decision are in RESULTS.md. Algorithm 70 with public split-K 1 is byte-exact and passes the block, two-step, and canonical 12-step gates, but remains research-only until production integration and deployment validation are performed.

Exact Operation

The extension reproduces Comfy Kitchen 0.2.31's cublas_gemm_nvfp4.cu descriptors for row-major packed activation [M,K] times packed weight [N,K] to BF16 [M,N]:

  • cuBLASLt sees column-major weight.T @ activation, so Lt m=N, n=M, and k=K; the output storage is the row-major [M,N] view.
  • A and B are CUDA_R_4F_E2M1, with CUBLASLT_MATMUL_MATRIX_SCALE_VEC16_UE4M3 block scales.
  • Compute, scalar scale type, alpha, and beta are FP32. Alpha and zero beta are device pointers.
  • Output is BF16, epilogue is default, and there is no bias.
  • The activation producer is the deployed, accepted vortex_native_quantize_swiglu_nvfp4. Alpha is its FP32 tensor scale times FC2's FP32 weight tensor scale.

The canonical 37,810-row input pads to 37,824 rows in the accepted producer. The report distinguishes logical [M,N,K] from descriptor/padded dimensions; comparison slices back to the logical output exactly as Comfy Kitchen does.

Current Baseline

The accepted production path fuses exact BF16 SwiGLU into NVFP4 production and retains Comfy Kitchen's cuBLASLt FC2 GEMM. Existing validated metadata reports:

  • Producer tensor scale, QDATA, and SFA are byte-identical at blocks 0, 24, and 49.
  • Canonical block FC2 input is logically 37,810 x 14,336 BF16 before packing.
  • Fully fused block-24 profiling attributes 26.87% of kernel time to all four NVFP4 GEMMs; this is not claimed as an FC2-only percentage.
  • Spark's accepted image is sha256:1d340e14cb6fc45ccfdbe63dde8db2a2b3aea94b493702c8a08e1f8d5b4f7b83 and enables H3_NVFP4_SCALE_BACKEND=vortex, H3_FUSED_ELEMENTWISE=1, H3_NVFP4_MODULATE_FUSION=1, and H3_NVFP4_SWIGLU_FUSION=1.

Every run also records a fresh profiler-derived baseline kernel list. Kernel names and times in that list are run metadata, not hard-coded claims.

Commands

Run inside image 1d340e14cb6f... from the Spark checkout, without starting the hot service. Preserve its production environment switches, for example:

export H3_NVFP4_SCALE_BACKEND=vortex H3_NVFP4_SCALE_VERSION=1
export H3_FUSED_ELEMENTWISE=1 H3_NVFP4_MODULATE_FUSION=1
export H3_NVFP4_SWIGLU_FUSION=1 H3_SAGE_QKV_LAYOUT=strided_nhd

Compile only:

python tools/benchmark_fc2_nvfp4_algorithms.py --mode compile --verbose-build \
  --build-directory /tmp/fc2-nvfp4-build

Default characterization enumerates workspace budgets 0, 4, 8, 16, 32, and 64 MiB, checks a bounded set of public split-K configurations, and times one valid probe with alternating AB/BA order:

python tools/benchmark_fc2_nvfp4_algorithms.py --mode characterize \
  --build-directory /tmp/fc2-nvfp4-build \
  --output /output/h3-blackwell-runtime/benchmarks/fc2-nvfp4-characterize.json

The same workload can come from an existing capture:

python tools/benchmark_fc2_nvfp4_algorithms.py --mode characterize \
  --capture /artifacts/capture --build-directory /tmp/fc2-nvfp4-build \
  --output /output/h3-blackwell-runtime/benchmarks/fc2-nvfp4-capture.json

A bounded timing sweep defaults to eight candidates, not a full combinatorial gate:

python tools/benchmark_fc2_nvfp4_algorithms.py --mode sweep \
  --candidate-limit 8 --rounds 6 --workspace-bytes 33554432 \
  --build-directory /tmp/fc2-nvfp4-build \
  --output /output/h3-blackwell-runtime/benchmarks/fc2-nvfp4-sweep.json

Re-run one selected candidate. --candidate accepts an inline JSON object, a JSON file containing one object, or enumerate:<budget-index>:<result-index>:

python tools/benchmark_fc2_nvfp4_algorithms.py --mode selected \
  --candidate '{"algorithm_id":23,"tile_id":42,"stages_id":35,"split_k":1,"reduction_scheme":0}' \
  --workspace-bytes 33554432 --build-directory /tmp/fc2-nvfp4-build \
  --output /output/h3-blackwell-runtime/benchmarks/fc2-nvfp4-selected.json

NCU capture uses cudaProfilerStart/Stop; only the selected FC2 cuBLASLt call is inside the range:

ncu --target-processes all --profile-from-start off --set full \
  --export /output/h3-blackwell-runtime/benchmarks/fc2-nvfp4-selected \
  python tools/benchmark_fc2_nvfp4_algorithms.py --mode profile \
  --candidate /tmp/fc2-candidate.json --workspace-bytes 33554432 \
  --build-directory /tmp/fc2-nvfp4-build \
  --output /output/h3-blackwell-runtime/benchmarks/fc2-nvfp4-ncu.json

Complete block gate for blocks 0, 24, and 49:

python tools/benchmark_fc2_nvfp4_algorithms.py --mode block-gate \
  --candidate /tmp/fc2-candidate.json --rounds 6 \
  --workspace-bytes 33554432 --build-directory /tmp/fc2-nvfp4-build \
  --output /output/h3-blackwell-runtime/benchmarks/fc2-nvfp4-block-gate.json

Block-gate monkeypatches only block.mlp.fc2.forward_swiglu. It still invokes the accepted SwiGLU producer on the actual FC1 result. Norm, FC1, modulation, gate, residual, attention, and all other code remain the accepted path.

Requested NCU Traffic Counters

Use ncu --query-metrics on the installed NCU version before replacing --set full with an explicit list. Request the available equivalents of:

gpu__time_duration.sum
dram__bytes_read.sum
dram__bytes_write.sum
lts__t_bytes.sum
lts__t_sector_hit_rate.pct
sm__throughput.avg.pct_of_peak_sustained_elapsed
smsp__inst_executed.sum
smsp__pipe_tensor_cycles_active.avg.pct_of_peak_sustained_active
launch__registers_per_thread
launch__shared_mem_per_block_allocated

Metric spelling and availability vary with NCU and GB10. Missing counters must be reported as unavailable rather than silently substituted. Also retain kernel duration, grid/cluster dimensions, achieved occupancy, waves, required workspace, and the extension's tile/stage/inner/cluster configuration.

Validation Ladder

  1. Compile against the image's CUDA 13 headers and print build_info().
  2. Characterize heuristics and verify every executed candidate passes cublasLtMatmulAlgoCheck within caller-supplied workspace.
  3. Compare direct-probe candidate FC2 against a cloned fc2.forward_swiglu(gate_up) output byte-for-byte.
  4. Alternate baseline/candidate AB and BA each round; record p50, p95, dense TFLOP/s, producer-plus-FC2 time, workspace, and output SHA-256.
  5. NCU one FC2 call and inspect traffic/resource counters.
  6. Gate complete blocks 0, 24, and 49 with only FC2 GEMM replaced. Require BF16 exactness against both the paired baseline and retained traversal.
  7. Only after those gates should a separate, explicitly authorized trajectory experiment be considered.

That trajectory authorization was granted for the retained candidate. Both the two-step and canonical 12-step gates passed with byte-exact video and audio latents; see RESULTS.md and the linked benchmark artifacts.

Prohibited Experiments

  • Do not alter production source, configuration, dispatch, or existing research files.
  • Do not expose this extension through Nvfp4Linear or any production operator.
  • Do not start or perturb the hot service.
  • Do not add bias, accumulation, host scalars, another output dtype, a different producer, or approximate validation to this study.
  • Do not claim a checkpoint hash unless --checkpoint-sha256 actually computes it.
  • Do not label undocumented behavior Stream-K. CUDA 13 exposes no documented public cublasLtMatmulAlgoConfig attribute that directly selects Stream-K; the report states that limitation explicitly.
  • Do not promote a direct-GEMM result without the complete block gate and later separately authorized trajectory validation.

Limitations

  • Heuristics are library, driver, GPU, shape, and workspace specific.
  • Explicit split-K is attempted only through documented AlgoInit, ConfigSet, and AlgoCheck. Unsupported factors/reduction schemes are recorded; they are not emulated.
  • Negative SPLITK_NUM values returned by the heuristic are retained as raw signed library sentinels. They are not interpreted or labeled as Stream-K.
  • Characterization deliberately bounds explicit checks and timing candidates. Increase limits consciously because each full canonical FC2 call is costly.
  • Caller-owned output and workspace are reused during timing. The accepted producer still owns its quantized activation allocations.
  • A capture supplies packed-denoiser inputs, not pre-captured FC2 operands; the harness traverses the loaded H3 model once to construct exact current boundaries.