Direct MiniMax H3 CUDA runtime
Find a file
2026-08-20 17:39:44 +07:00
benchmarks Record Sage3 cat benchmark 2026-08-14 20:02:14 +07:00
src/h3_blackwell_runtime Add selectable hot attention backends 2026-08-20 17:39:44 +07:00
tests Add selectable hot attention backends 2026-08-20 17:39:44 +07:00
tools Add selectable hot attention backends 2026-08-20 17:39:44 +07:00
wheels Cache GB10 SageAttention3 wheel 2026-08-12 15:21:03 +07:00
.dockerignore Add selectable hot attention backends 2026-08-20 17:39:44 +07:00
.gitignore Initial direct H3 runtime 2026-08-12 14:12:42 +07:00
compose.spark-comfy-lab.yml Initial direct H3 runtime 2026-08-12 14:12:42 +07:00
compose.spark.yml Add selectable hot attention backends 2026-08-20 17:39:44 +07:00
Dockerfile.spark Add selectable hot attention backends 2026-08-20 17:39:44 +07:00
PARITY.md Default video VAE decode to FP16 2026-08-14 19:50:36 +07:00
PLAN.md Document NVFP4 native pack status 2026-08-15 03:00:06 +07:00
pyproject.toml Add selectable hot attention backends 2026-08-20 17:39:44 +07:00
README.md Add selectable hot attention backends 2026-08-20 17:39:44 +07:00

H3 Blackwell Runtime

Direct MiniMax H3 Ref2VA runtime research project. ComfyUI is the checkpoint and correctness oracle, not the target runtime.

First Gate

Inspect the mounted H3 NVFP4 safetensors headers before designing an importer:

python .\tools\inspect_safetensors.py /runpod-volume/ComfyUI/models/diffusion_models/minimax_h3_ref2va_pruned_nvfp4.safetensors

Write the output to artifacts/checkpoints/ on the mounted volume. The result must identify packed weights, scales, and tensor naming before any kernel conversion work begins.

Benchmark Contract

benchmarks/ref2va-960x544-124f.json is the single-GPU performance contract. Record direct-runner results as JSON and compare them with:

python .\tools\compare_benchmark.py --result direct-result.json

DGX Spark

Dockerfile.spark and compose.spark.yml prepare an ARM64 GB10 development image using the existing AEON CUDA 13/SageAttention3 base. The compose target opens a shell only; it does not start inference.

Forgejo Pulls From Spark

The Spark checkout uses Forgejo through the host's published local SSH port and a dedicated key:

cd /home/daniel/aeon-spark-test/h3/h3-blackwell-runtime
git config core.sshCommand 'ssh -i ~/.ssh/id_ed25519_forgejo_h3 -o IdentitiesOnly=yes'
git remote set-url origin ssh://git@127.0.0.1:2222/daniel/h3-blackwell-runtime.git
git pull --ff-only origin master

The private key remains on Spark at ~/.ssh/id_ed25519_forgejo_h3; only its public key is registered in Forgejo.

Runtime Output

Generation and latent-decode tools are quiet by default: they suppress ffmpeg banners and only print compact JSON summaries. Use these flags when debugging:

  • --progress: print per-step sampler timing in tools/direct_t2v_preview.py.
  • --profile-memory: print memory checkpoints in tools/direct_t2v_preview.py.
  • --ffmpeg-loglevel info: show ffmpeg details instead of the default error level.
  • --quiet: suppress JSON summary lines.
  • --vae-dtype float16: use Comfy-style FP16 video VAE decode in tools/direct_t2v_preview.py or tools/decode_video_latent.py; this is the default runtime path. Use --vae-dtype float32 only for exact direct-path diagnostics. tools/direct_t2v_preview.py also accepts H3_VAE_DTYPE.
  • --vae-tile-size 256: set the direct video VAE spatial tile size. tools/direct_t2v_preview.py also accepts H3_VAE_TILE_SIZE.

Standalone tools/compare_*, tools/trace_*, tools/inspect_*, and tools/patch_comfy_* scripts are debugging utilities and remain opt-in by being separate commands.

Hot Runtime Service

tools/serve_hot_runtime.py keeps Qwen, H3, video VAE, and audio VAE resident in one process. Start the optional Spark service with:

docker compose -f compose.spark.yml up -d h3-hot-runtime

Use GET /ready to confirm resident model readiness. Use POST /generate with JSON fields like prompt, output, width, height, frames, steps, seed, and optional attention. Supported request-level attention values are reported by /ready; switching attention does not reload model weights. The hot image includes Sage2, forced cuDNN SDPA, and Comfy Kitchen INT8 attention. Sage2 is the default based on the 960x544x124 GB10 benchmark and the existing parity baseline.

Exact memory/lifetime options:

  • attention: "kj_head_sliced" slices attention heads and runs the slice backend from H3_HEAD_SLICE_BACKEND (sage2 by default) with H3_HEAD_SLICE_SIZE heads per slice (8 by default).
  • attention: "cudnn_sdpa" forces cuDNN SDPA with no fallback to another PyTorch kernel.
  • attention: "ck_int8" uses Comfy Kitchen's approximate INT8 Q/K/V attention kernel.
  • attention: "sol_attn" routes eligible H3 attention calls through the pinned ComfyUI Sol-Attn Triton kernel vendored into the Spark image. Configure with H3_SOL_TAU (1.3), H3_SOL_MIN_TOKENS (4096), H3_SOL_THRESH_TYPE (diag), H3_SOL_INT8_QK, H3_SOL_INT8_PV, H3_SOL_FALLBACK (sage2), and H3_SOL_STRICT.
  • --mlp-chunks N on tools/serve_hot_runtime.py or tools/direct_t2v_preview.py chunks H3 SwiGLU rows exactly to reduce peak activation memory. Default is 1 (disabled).

Approximate cache options are opt-in and must be quality-gated per prompt:

  • cache_mode: "easycache" reuses cached denoised deltas while cumulative latent input change stays below cache_threshold.
  • cache_mode: "h3_cache" reuses cached denoised deltas when the current per-step latent input change is below cache_threshold.
  • Both modes accept cache_start_percent, cache_end_percent, and cache_subsample_factor in POST /generate; the CLI exposes equivalent --cache-* flags.