72 lines
3.9 KiB
Markdown
72 lines
3.9 KiB
Markdown
# H3 Blackwell Runtime
|
|
|
|
Direct MiniMax H3 Ref2VA runtime research project. ComfyUI is the checkpoint and correctness oracle, not the target runtime.
|
|
|
|
## First Gate
|
|
|
|
Inspect the mounted H3 NVFP4 safetensors headers before designing an importer:
|
|
|
|
```powershell
|
|
python .\tools\inspect_safetensors.py /runpod-volume/ComfyUI/models/diffusion_models/minimax_h3_ref2va_pruned_nvfp4.safetensors
|
|
```
|
|
|
|
Write the output to `artifacts/checkpoints/` on the mounted volume. The result must identify packed weights, scales, and tensor naming before any kernel conversion work begins.
|
|
|
|
## Benchmark Contract
|
|
|
|
`benchmarks/ref2va-960x544-124f.json` is the single-GPU performance contract. Record direct-runner results as JSON and compare them with:
|
|
|
|
```powershell
|
|
python .\tools\compare_benchmark.py --result direct-result.json
|
|
```
|
|
|
|
## DGX Spark
|
|
|
|
`Dockerfile.spark` and `compose.spark.yml` prepare an ARM64 GB10 development image using the existing AEON CUDA 13/SageAttention3 base. The compose target opens a shell only; it does not start inference.
|
|
|
|
### Forgejo Pulls From Spark
|
|
|
|
The Spark checkout uses Forgejo through the host's published local SSH port and a dedicated key:
|
|
|
|
```bash
|
|
cd /home/daniel/aeon-spark-test/h3/h3-blackwell-runtime
|
|
git config core.sshCommand 'ssh -i ~/.ssh/id_ed25519_forgejo_h3 -o IdentitiesOnly=yes'
|
|
git remote set-url origin ssh://git@127.0.0.1:2222/daniel/h3-blackwell-runtime.git
|
|
git pull --ff-only origin master
|
|
```
|
|
|
|
The private key remains on Spark at `~/.ssh/id_ed25519_forgejo_h3`; only its public key is registered in Forgejo.
|
|
|
|
## Runtime Output
|
|
|
|
Generation and latent-decode tools are quiet by default: they suppress ffmpeg banners and only print compact JSON summaries. Use these flags when debugging:
|
|
|
|
- `--progress`: print per-step sampler timing in `tools/direct_t2v_preview.py`.
|
|
- `--profile-memory`: print memory checkpoints in `tools/direct_t2v_preview.py`.
|
|
- `--ffmpeg-loglevel info`: show ffmpeg details instead of the default `error` level.
|
|
- `--quiet`: suppress JSON summary lines.
|
|
- `--vae-dtype float16`: use Comfy-style FP16 video VAE decode in `tools/direct_t2v_preview.py` or `tools/decode_video_latent.py`; this is the default runtime path. Use `--vae-dtype float32` only for exact direct-path diagnostics. `tools/direct_t2v_preview.py` also accepts `H3_VAE_DTYPE`.
|
|
- `--vae-tile-size 256`: set the direct video VAE spatial tile size. `tools/direct_t2v_preview.py` also accepts `H3_VAE_TILE_SIZE`.
|
|
|
|
Standalone `tools/compare_*`, `tools/trace_*`, `tools/inspect_*`, and `tools/patch_comfy_*` scripts are debugging utilities and remain opt-in by being separate commands.
|
|
|
|
## Hot Runtime Service
|
|
|
|
`tools/serve_hot_runtime.py` keeps Qwen, H3, video VAE, and audio VAE resident in one process. Start the optional Spark service with:
|
|
|
|
```bash
|
|
docker compose -f compose.spark.yml up -d h3-hot-runtime
|
|
```
|
|
|
|
Use `GET /ready` to confirm resident model readiness. Use `POST /generate` with JSON fields like `prompt`, `output`, `width`, `height`, `frames`, `steps`, `seed`, and optional `attention`. Supported request-level attention values are reported by `/ready`; switching attention does not reload model weights.
|
|
|
|
Exact memory/lifetime options:
|
|
|
|
- `attention: "kj_head_sliced"` slices attention heads and runs the slice backend from `H3_HEAD_SLICE_BACKEND` (`sage2` by default) with `H3_HEAD_SLICE_SIZE` heads per slice (`8` by default).
|
|
- `--mlp-chunks N` on `tools/serve_hot_runtime.py` or `tools/direct_t2v_preview.py` chunks H3 SwiGLU rows exactly to reduce peak activation memory. Default is `1` (disabled).
|
|
|
|
Approximate cache options are opt-in and must be quality-gated per prompt:
|
|
|
|
- `cache_mode: "easycache"` reuses cached denoised deltas while cumulative latent input change stays below `cache_threshold`.
|
|
- `cache_mode: "h3_cache"` reuses cached denoised deltas when the current per-step latent input change is below `cache_threshold`.
|
|
- Both modes accept `cache_start_percent`, `cache_end_percent`, and `cache_subsample_factor` in `POST /generate`; the CLI exposes equivalent `--cache-*` flags.
|