Skip to content

vllm_omni.diffusion.models.minimax_h3.adaln_cache

Default exact AdaLN memoization and optional weight-bound startup sidecars.

This is not a cross-step approximation. The main blocks' and final layer's AdaLN outputs are reused only for identical inputs and unchanged weights. All weights remain available for new schedules and adapters. Legacy experimental sidecars without weight identities are rejected.

FORMAT_VERSION module-attribute

FORMAT_VERSION = '4'

MODES module-attribute

MODES = {
    "t2va": ("video", "audio"),
    "fl2va": ("video", "audio", "image"),
    "ref2va-image": ("video", "audio", "image"),
    "ref2va-audio": ("video", "audio", "audio_ref"),
    "ref2va-mixed": (
        "video",
        "audio",
        "image",
        "audio_ref",
    ),
}

PAYLOAD_NAMES module-attribute

PAYLOAD_NAMES = frozenset(
    {
        "plan_timesteps",
        "plan_lengths",
        "time_embeddings",
        "block_params",
        "final_params",
    }
)

MiniMaxH3AdalnCache

Bases: Module

Validate an optional sidecar, then verify its weights at load.

Hashes are checked against the post-fusion stream, before TP conversion. They cover all inputs of the cached computation, not unrelated QKV/MLP weights. A fixed adapter has an additional complete-file identity.

arch instance-attribute

arch = arch

block_params instance-attribute

block_params: Tensor

final_params instance-attribute

final_params: Tensor

manifest instance-attribute

manifest = manifest

path instance-attribute

path = str(Path(path).expanduser().resolve())

plan_lengths instance-attribute

plan_lengths: Tensor

plan_timesteps instance-attribute

plan_timesteps: Tensor

projection_names instance-attribute

projection_names = projection_names(arch.num_layers)

time_embeddings instance-attribute

time_embeddings: Tensor

bind_adapter

bind_adapter(adapter_sha256: str | None) -> None

check_request

check_request(
    *,
    mode: str,
    num_steps: int,
    base_schedule: Sequence[float] | None,
    flow_shift: float,
    audio_flow_shift: float,
) -> None

finish_loading

finish_loading(device: device) -> None

lookup

lookup(
    timesteps: Tensor,
) -> tuple[
    tuple[tuple[Tensor, ...], ...], tuple[Tensor, ...]
]

verify_weight

verify_weight(name: str, weight: Tensor) -> None

MiniMaxH3RuntimeAdalnCache

Bases: ExactProjectionCache

Exact projection caching with optional H3 schedule-bound sidecar results.

sidecar instance-attribute

sidecar: MiniMaxH3AdalnCache | None = None

clear

clear() -> None

prepare

prepare(embedding: Tensor) -> None

seed

seed(
    sidecar: MiniMaxH3AdalnCache,
    modules: Mapping[str, Module],
) -> None

architecture

architecture(
    arch: MiniMaxH3DiTArchConfig,
) -> dict[str, int]

build_cache

build_cache(
    arch: MiniMaxH3DiTArchConfig,
    weights: Iterable[tuple[str, Tensor]],
    *,
    contract: dict[str, Any],
    model_variant: str,
    adapter_sha256: str | None,
    device: device,
) -> tuple[dict[str, Tensor], dict[str, Any]]

Stream one projection at a time; do not instantiate the full DiT.

The four time-embedder tensors must precede the AdaLN projections. Callers fuse any fixed adapter before supplying this stream, just like serving.

file_digest

file_digest(path: str | Path) -> str

input_names

input_names(num_layers: int) -> frozenset[str]

math_identity

math_identity(device: device) -> dict[str, Any]

Conservatively bind the builder's numerical environment (TP1 only).

projection_names

projection_names(num_layers: int) -> frozenset[str]

schedule_contract

schedule_contract(
    *,
    mode: str,
    num_steps: int,
    base_schedule: Sequence[float] | None,
    flow_shift: float,
    audio_flow_shift: float,
) -> dict[str, Any]

timestep_plans

timestep_plans(contract: Mapping[str, Any]) -> list[Tensor]

Use the serving scheduler and its conditioning timesteps, without interpolation.