Wire full fl2va into the direct H3 runtime so first/last keyframes flow through VAE encode -> Qwen vision tokens -> DiT cond segments: - vae_encoder.py: direct encoder-only H3 video VAE (causal 3D convs, reflect spatial padding, causal temporal padding, single-frame tap truncation, tiling, FP32 moments + mean/std normalization). - qwen3vl_vision.py: Qwen3-VL-32B visual tower (27 blocks, 2D rotary, deepstack mergers) ported to match the Comfy reference exactly (head_dim=72, no-bias proj, LayerNorm blocks, split-half apply_rope), plus Qwen image preprocess, mrope ids/freqs, DiT token tags, keyframe resize (first=stretch / last=center cover-crop matching Comfy common_upscale), and build_fl2va_presentation. - qwen3vl_text.py: split-half apply_rope, _embed_rows/_run_layers, optional mrope position_ids + DeepStack injection at the first three decoder layers at visual positions. - packing.py: H3PromptPacker builds [text | cond | audio | video] with tag-run text spans, cond rows (first/last cond_t anchors, VISUAL_COND_TIMESTEP=0.999 noise augmentation via CPU-seeded RNG), three-timestep row table (t_row*3 + modality_tag), and rope positions. - runtime.py: load VAE encoder + vision tower; generate() accepts first_frame/last_frame, builds the fl2va presentation, encodes keyframes, and passes text_token_tags/cond_latents/frame_count/seed to the sampler. - sampler.py: thread pack kwargs + seed. - serve_hot_runtime.py / direct_t2v_preview.py: /generate and --first-frame/--last-frame accept image paths or base64. |
||
|---|---|---|
| benchmarks | ||
| src/h3_blackwell_runtime | ||
| tools | ||
| wheels | ||
| .gitignore | ||
| compose.spark-comfy-lab.yml | ||
| compose.spark.yml | ||
| Dockerfile.spark | ||
| PARITY.md | ||
| PLAN.md | ||
| pyproject.toml | ||
| README.md | ||
H3 Blackwell Runtime
Direct MiniMax H3 Ref2VA runtime research project. ComfyUI is the checkpoint and correctness oracle, not the target runtime.
First Gate
Inspect the mounted H3 NVFP4 safetensors headers before designing an importer:
python .\tools\inspect_safetensors.py /runpod-volume/ComfyUI/models/diffusion_models/minimax_h3_ref2va_pruned_nvfp4.safetensors
Write the output to artifacts/checkpoints/ on the mounted volume. The result must identify packed weights, scales, and tensor naming before any kernel conversion work begins.
Benchmark Contract
benchmarks/ref2va-960x544-124f.json is the single-GPU performance contract. Record direct-runner results as JSON and compare them with:
python .\tools\compare_benchmark.py --result direct-result.json
DGX Spark
Dockerfile.spark and compose.spark.yml prepare an ARM64 GB10 development image using the existing AEON CUDA 13/SageAttention3 base. The compose target opens a shell only; it does not start inference.
Forgejo Pulls From Spark
The Spark checkout uses Forgejo through the host's published local SSH port and a dedicated key:
cd /home/daniel/aeon-spark-test/h3/h3-blackwell-runtime
git config core.sshCommand 'ssh -i ~/.ssh/id_ed25519_forgejo_h3 -o IdentitiesOnly=yes'
git remote set-url origin ssh://git@127.0.0.1:2222/daniel/h3-blackwell-runtime.git
git pull --ff-only origin master
The private key remains on Spark at ~/.ssh/id_ed25519_forgejo_h3; only its public key is registered in Forgejo.
Runtime Output
Generation and latent-decode tools are quiet by default: they suppress ffmpeg banners and only print compact JSON summaries. Use these flags when debugging:
--progress: print per-step sampler timing intools/direct_t2v_preview.py.--profile-memory: print memory checkpoints intools/direct_t2v_preview.py.--ffmpeg-loglevel info: show ffmpeg details instead of the defaulterrorlevel.--quiet: suppress JSON summary lines.--vae-dtype float16: use Comfy-style FP16 video VAE decode intools/direct_t2v_preview.pyortools/decode_video_latent.py; this is the default runtime path. Use--vae-dtype float32only for exact direct-path diagnostics.tools/direct_t2v_preview.pyalso acceptsH3_VAE_DTYPE.--vae-tile-size 256: set the direct video VAE spatial tile size.tools/direct_t2v_preview.pyalso acceptsH3_VAE_TILE_SIZE.
Standalone tools/compare_*, tools/trace_*, tools/inspect_*, and tools/patch_comfy_* scripts are debugging utilities and remain opt-in by being separate commands.
Hot Runtime Service
tools/serve_hot_runtime.py keeps Qwen, H3, video VAE, and audio VAE resident in one process. Start the optional Spark service with:
docker compose -f compose.spark.yml up -d h3-hot-runtime
Use GET /ready to confirm resident model readiness. Use POST /generate with JSON fields like prompt, output, width, height, frames, steps, seed, and optional attention. Supported request-level attention values are reported by /ready; switching attention does not reload model weights.
Exact memory/lifetime options:
attention: "kj_head_sliced"slices attention heads and runs the slice backend fromH3_HEAD_SLICE_BACKEND(sage2by default) withH3_HEAD_SLICE_SIZEheads per slice (8by default).attention: "sol_attn"routes eligible H3 attention calls through the pinned ComfyUI Sol-Attn Triton kernel vendored into the Spark image. Configure withH3_SOL_TAU(1.3),H3_SOL_MIN_TOKENS(4096),H3_SOL_THRESH_TYPE(diag),H3_SOL_INT8_QK,H3_SOL_INT8_PV,H3_SOL_FALLBACK(sage2), andH3_SOL_STRICT.--mlp-chunks Nontools/serve_hot_runtime.pyortools/direct_t2v_preview.pychunks H3 SwiGLU rows exactly to reduce peak activation memory. Default is1(disabled).
Approximate cache options are opt-in and must be quality-gated per prompt:
cache_mode: "easycache"reuses cached denoised deltas while cumulative latent input change stays belowcache_threshold.cache_mode: "h3_cache"reuses cached denoised deltas when the current per-step latent input change is belowcache_threshold.- Both modes accept
cache_start_percent,cache_end_percent, andcache_subsample_factorinPOST /generate; the CLI exposes equivalent--cache-*flags.