Skip to content

vllm_omni.diffusion.models.minimax_h3.latent_upscaler

Learned latent super-resolution for MiniMax-H3 video latents.

The released upscaler is a pure-3D convolutional resizer trained on H3's 24-channel VAE latents: it lifts a latent to a target spatial size without a decode/encode round trip through the ~5B-parameter H3 VAE. Weights come from LBH-123-AI/Minimax_h3_latent_Upscaler and are a plain state_dict, so the module layout below reproduces the checkpoint's parameter names exactly.

The network works in a space one normalization below the pipeline latent. A vLLM-Omni H3 latent is already normalized -- :class:MiniMaxH3VideoVAE denormalizes at decode time -- and the reference ComfyUI node applies the VAE's latents_mean / latents_std on top of a latent in that same convention, so the upscaler was trained on the doubly-normalized tensor. Feeding it the pipeline latent directly inflates the output by roughly 5x (measured: output std 5.4 against an input std of 1.0), which decodes as a magenta grid at one tile per latent cell. :meth:MiniMaxH3LatentUpscaler.upscale therefore normalizes in and denormalizes out, exactly as the reference node does.

MINIMAX_H3_LATENT_UPSCALE_MAX_SCALE module-attribute

MINIMAX_H3_LATENT_UPSCALE_MAX_SCALE = 4.0

MINIMAX_H3_VAE_SPATIAL_DOWNSAMPLE module-attribute

MINIMAX_H3_VAE_SPATIAL_DOWNSAMPLE = 16

logger module-attribute

logger = init_logger(__name__)

MiniMaxH3LatentRefineSpec dataclass

How much of the sigma schedule a second denoise pass re-runs.

strength instance-attribute

strength: float

start_index

start_index(num_sigma_points: int) -> int

The schedule position the refine pass starts from.

strength is the fraction of the request's steps to re-run, the img2img convention, so the cost of the pass is predictable from the request. At least one step always runs.

MiniMaxH3LatentResizer3D

Bases: Module

The released pure-3D latent resizer.

in_blocks run at the source resolution, a trilinear interpolation moves the feature volume to the target size, and out_blocks refine it there. The requested scale enters every residual block through a two-layer embedding of scale - 1.

arch instance-attribute

arch = arch

conv_in instance-attribute

conv_in = nn.Conv3d(
    arch.in_channels, arch.channels, 3, padding=1
)

conv_out instance-attribute

conv_out = nn.Conv3d(
    arch.channels, arch.in_channels, 3, padding=1
)

embed instance-attribute

embed = nn.Sequential(
    nn.Linear(1, _EMBED_DIM),
    nn.SiLU(),
    nn.Linear(_EMBED_DIM, _EMBED_DIM),
)

in_blocks instance-attribute

in_blocks = self._build_stage(arch, arch.in_blocks)

norm_out instance-attribute

norm_out = _normalization(arch.channels)

out_blocks instance-attribute

out_blocks = self._build_stage(arch, arch.out_blocks)

temporal_kernel property

temporal_kernel: int

forward

forward(
    latent: Tensor,
    *,
    scale: float,
    target_size: tuple[int, int, int],
) -> Tensor

Resize [B, C, T, H, W] to target_size under a scale hint.

MiniMaxH3LatentUpscaleTarget dataclass

A resolved upscale target in latent cells, plus the scale hint to feed the network.

height property

height: int

latent_height instance-attribute

latent_height: int

latent_width instance-attribute

latent_width: int

scale instance-attribute

scale: float

width property

width: int

MiniMaxH3LatentUpscaler

Bases: Module

Serving wrapper around :class:MiniMaxH3LatentResizer3D.

Owns the temporal chunking that keeps long clips inside VRAM and the residency policy that keeps the extra ~700MB off the device between requests.

chunk_frames instance-attribute

chunk_frames = chunk_frames

chunk_overlap instance-attribute

chunk_overlap = (
    resizer.temporal_kernel
    if chunk_overlap is None
    else chunk_overlap
)

device instance-attribute

device = device

dtype instance-attribute

dtype = dtype

latents_mean instance-attribute

latents_mean = tuple(float(value) for value in latents_mean)

latents_std instance-attribute

latents_std = tuple(float(value) for value in latents_std)

resident instance-attribute

resident = resident

resizer instance-attribute

resizer = (
    resizer.to(dtype=dtype).eval().requires_grad_(False)
)

load_to_device

load_to_device() -> None

offload_to_cpu

offload_to_cpu() -> None

upscale

upscale(
    latent: Tensor, target: MiniMaxH3LatentUpscaleTarget
) -> Tensor

Upscale a normalized [B, 24, T, H, W] H3 latent to target.

MiniMaxH3LatentUpscalerArch dataclass

Shape of a released latent upscaler, recovered from its state_dict.

channels class-attribute instance-attribute

channels: int = 512

in_blocks class-attribute instance-attribute

in_blocks: int = 12

in_channels class-attribute instance-attribute

in_channels: int = 24

out_blocks class-attribute instance-attribute

out_blocks: int = 12

temporal_every class-attribute instance-attribute

temporal_every: int = 2

temporal_kernel class-attribute instance-attribute

temporal_kernel: int = 5

MiniMaxH3LatentUpscalerError

Bases: ValueError

A latent upscaler artifact that cannot be served.

detect_minimax_h3_upscaler_arch

detect_minimax_h3_upscaler_arch(
    state_dict: Mapping[str, Tensor],
) -> MiniMaxH3LatentUpscalerArch

Recover the block layout of a released upscaler from its tensors.

load_minimax_h3_latent_upscaler

load_minimax_h3_latent_upscaler(
    path: str | Path,
    *,
    device: device,
    dtype: dtype,
    latents_mean: Sequence[float],
    latents_std: Sequence[float],
    chunk_frames: int = 0,
    chunk_overlap: int | None = None,
    resident: bool = False,
) -> MiniMaxH3LatentUpscaler

Build a serving-ready upscaler from a released checkpoint.

parse_minimax_h3_latent_refine_request

parse_minimax_h3_latent_refine_request(
    value,
) -> MiniMaxH3LatentRefineSpec | None

Normalize extra_args['latent_refine'] into a refine spec.

Accepts a bare strength (0.4), false/null to opt out of a server-side default, or {"strength": 0.4}.

parse_minimax_h3_latent_upscale_request

parse_minimax_h3_latent_upscale_request(
    value,
) -> dict[str, float | int] | None

Normalize extra_args['latent_upscale'] into resolver keywords.

Accepts a bare multiplier (2.0), false/null to opt out of a server-side default, or an object naming one sizing mode: {"scale": 2}, {"width": 2688, "height": 1536} or {"megapixels": 4}, each optionally with align.

resolve_minimax_h3_latent_upscale_target

resolve_minimax_h3_latent_upscale_target(
    *,
    latent_height: int,
    latent_width: int,
    scale: float | None = None,
    height: int | None = None,
    width: int | None = None,
    megapixels: float | None = None,
    align: int = 32,
) -> MiniMaxH3LatentUpscaleTarget

Resolve one of the three sizing modes into latent dimensions.

scale multiplies the source size; height/width name the target in pixels; megapixels names a total pixel budget and keeps the aspect ratio. For the two size-based modes the network still needs a single scale hint, which is the mean of the per-axis ratios -- what the reference node feeds the same weights. Targets snap to align pixels and then to the VAE's 16px cell.

resolve_minimax_h3_latent_upscaler

resolve_minimax_h3_latent_upscaler(
    od_config,
    *,
    device: device,
    latent_stats: Callable[
        [], tuple[Sequence[float], Sequence[float]]
    ],
) -> MiniMaxH3LatentUpscaler | None

Build the upscaler declared by --additional-config, if any.

latent_stats yields the VAE's per-channel latents_mean / latents_std and is called only once a checkpoint is configured, so a deployment without this stage never reaches for them.

Recognized keys: latent_upscaler_path (required to enable the stage), latent_upscaler_dtype, latent_upscaler_chunk_frames (0 disables temporal chunking) and latent_upscaler_resident (keep the weights on the device between requests instead of parking them in host memory).