149 lines
5.5 KiB
Markdown
149 lines
5.5 KiB
Markdown
# H3 Native Latent Upscaler
|
|
|
|
The direct runtime can apply LBH-123-AI's learned 3D upscaler to a saved,
|
|
normalized MiniMax H3 video latent without a ComfyUI dependency. This is an
|
|
experimental spatial upscale. It preserves latent time and does not interpolate
|
|
frames.
|
|
|
|
Upstream implementation and checkpoint:
|
|
|
|
- <https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler>
|
|
- <https://huggingface.co/LBH-123-AI/Minimax_h3_latent_Upscaler>
|
|
- Revision: `13ccf95d85d120bdbc92c05b1247a6e147bf54bf`
|
|
- FP16 SHA-256: `043e5a48e161610ef6c3ea974645220354d06fa618abca15f76d084812eb55c2`
|
|
- License: Apache-2.0
|
|
|
|
## Usage
|
|
|
|
The hot service supports the complete generation, latent upscale, video/audio
|
|
decode, and mux process in one request:
|
|
|
|
```json
|
|
{
|
|
"prompt": "A simple cinematic scene.",
|
|
"output": "/output/h3-native-upscale/example.mp4",
|
|
"width": 960,
|
|
"height": 544,
|
|
"frames": 124,
|
|
"turbo": "4step",
|
|
"upscale": 2.0,
|
|
"mux_audio": true,
|
|
"keep_intermediates": false
|
|
}
|
|
```
|
|
|
|
The response reports `source_width`, `source_height`, final `width` and
|
|
`height`, and a `latent_upscaled` timing stage. H3 audio bypasses the spatial
|
|
upscaler and is decoded and muxed normally.
|
|
|
|
Only the final muxed MP4 is retained by default. Set `keep_intermediates` to
|
|
`true` only when a separate WAV and video-only MP4 are needed for diagnostics;
|
|
`save_latent` remains a separate opt-in field.
|
|
|
|
For offline diagnostics, save the normalized latent during generation by
|
|
setting `save_latent`, then run:
|
|
|
|
```bash
|
|
python tools/upscale_video_latent.py \
|
|
--latent /output/source.pt \
|
|
--model /upscaler/minimax_h3_latent_upscaler_3d_fp16.safetensors \
|
|
--output /output/source-h3latent2x.pt \
|
|
--scale 2 \
|
|
--precision float16
|
|
|
|
python tools/decode_video_latent.py \
|
|
--latent /output/source-h3latent2x.pt \
|
|
--output /output/source-h3latent2x.mp4 \
|
|
--vae-dtype float16 \
|
|
--vae-tile-size 256
|
|
```
|
|
|
|
The upscaled state retains other source metadata and any audio latent. Its
|
|
`width` and `height` fields are updated from the resulting H3 latent dimensions.
|
|
|
|
## GB10 Validation
|
|
|
|
Validated on 2026-08-20 with the Posly hero image, Turbo 4-step v1.1, Sage2,
|
|
seed `440408`, and 124 frames:
|
|
|
|
| Stage | Result |
|
|
| --- | --- |
|
|
| Source | `960x544`, latent `[1,24,37,34,60]` |
|
|
| Learned 3D upscale | `10.94s`, latent `[1,24,37,68,120]` |
|
|
| Checkpoint load | `1.36s` |
|
|
| Tiled FP16 decode command | `137.67s`, `1920x1088` H.264 |
|
|
| Source vs downscaled result | SSIM `0.924986` |
|
|
|
|
The learned stage adds visible edge, face, and food detail. It also produces a
|
|
harsher illustrated texture, ringing/chromatic edges in some high-contrast
|
|
areas, and small facial changes. It is not fidelity-safe for protected UI,
|
|
text, QR codes, or identity-critical material without visual review.
|
|
|
|
The upstream example follows latent upscaling with a second target-resolution
|
|
H3 sampling pass using sigmas `0.9035, 0.6316, 0.3158, 0.0`. That refinement
|
|
was deliberately excluded from this first test so the upscaler itself could be
|
|
evaluated. It adds generation cost and another opportunity to alter content.
|
|
|
|
## Proofs
|
|
|
|
Spark output directory:
|
|
|
|
```text
|
|
/home/daniel/StoryStudioAssets/H3-output/posly/upscale-h3-native/
|
|
```
|
|
|
|
Key files:
|
|
|
|
- `01-posly-hero-turbo4-source-960x544.mp4`
|
|
- `01-posly-hero-turbo4-h3latent2x-1920x1088.mp4`
|
|
- `source-frame60-right-crop-lanczos.png`
|
|
- `h3latent2x-frame60-right-crop.png`
|
|
|
|
This result is promising for cinematic footage but does not replace the proven
|
|
LTX 2.3 chain yet. Run a matched LTX 2.3 comparison and a carefully bounded H3
|
|
refinement test before selecting a default production path.
|
|
|
|
## Integrated T2V Validation
|
|
|
|
The single-request path was validated with a matched 5.17-second talking-head
|
|
T2V pair at seed `440410`. Both requests sampled at `960x544`, applied the
|
|
resident 2x upscaler, decoded at `1920x1088`, and muxed native H3 audio into the
|
|
only retained artifact.
|
|
|
|
| Variant | Sampling | Latent upscale | VAE decode | Wall time |
|
|
| --- | ---: | ---: | ---: | ---: |
|
|
| Turbo 4-step v1.1 | `56.54s` | `7.52s` | `82.74s` | `163.02s` |
|
|
| Turbo 8-step v1.0 | `113.41s` | `7.28s` | `83.69s` | `220.06s` |
|
|
| Base 12-step beta/RES | `114.67s` | `7.26s` | `83.34s` | `221.66s` |
|
|
|
|
Both outputs contain H.264 video at 24 fps and stereo AAC audio at 32 kHz. The
|
|
4-step audio measured `-31.4 dB` mean / `-10.9 dB` peak; the 8-step audio
|
|
measured `-23.0 dB` mean / `-5.0 dB` peak; and the base 12-step audio measured
|
|
`-22.4 dB` mean / `-5.2 dB` peak.
|
|
|
|
```text
|
|
/home/daniel/StoryStudioAssets/H3-output/h3-native-upscale-t2v/talking-woman-h3native2x-1920x1088.mp4
|
|
/home/daniel/StoryStudioAssets/H3-output/h3-native-upscale-t2v/talking-woman-turbo8-h3native2x-1920x1088.mp4
|
|
/home/daniel/StoryStudioAssets/H3-output/h3-native-upscale-t2v/talking-woman-base12-h3native2x-1920x1088.mp4
|
|
```
|
|
|
|
### Attention Audio Isolation
|
|
|
|
A matched base 12-step run used direct PyTorch SDPA instead of Sage2. Upscaling
|
|
was disabled, but that does not affect the audio comparison because the learned
|
|
upscaler runs after joint AV sampling and only receives the separated video
|
|
latent.
|
|
|
|
| Attention | Sampling | Final size | Audio level | Wall time |
|
|
| --- | ---: | ---: | ---: | ---: |
|
|
| Sage2 | `114.67s` | `1920x1088` after post-sample 2x | `-22.4 dB` mean / `-5.2 dB` peak | `221.66s` |
|
|
| PyTorch SDPA | `144.40s` | `960x544`, no upscale | `-22.5 dB` mean / `-2.3 dB` peak | `171.21s` |
|
|
|
|
The SDPA comparison file is:
|
|
|
|
```text
|
|
/home/daniel/StoryStudioAssets/H3-output/h3-native-upscale-t2v/talking-woman-base12-sdpa-960x544.mp4
|
|
```
|
|
|
|
Signal levels alone cannot establish dialogue quality or lip sync; the matched
|
|
clips require listening review to determine whether Sage2 is the cause.
|