vllm_omni.diffusion.models.minimax_h3.adaln_cache ¶
Default exact AdaLN memoization and optional weight-bound startup sidecars.
This is not a cross-step approximation. The main blocks' and final layer's AdaLN outputs are reused only for identical inputs and unchanged weights. All weights remain available for new schedules and adapters. Legacy experimental sidecars without weight identities are rejected.
MODES module-attribute ¶
MODES = {
"t2va": ("video", "audio"),
"fl2va": ("video", "audio", "image"),
"ref2va-image": ("video", "audio", "image"),
"ref2va-audio": ("video", "audio", "audio_ref"),
"ref2va-mixed": (
"video",
"audio",
"image",
"audio_ref",
),
}
PAYLOAD_NAMES module-attribute ¶
PAYLOAD_NAMES = frozenset(
{
"plan_timesteps",
"plan_lengths",
"time_embeddings",
"block_params",
"final_params",
}
)
MiniMaxH3AdalnCache ¶
Bases: Module
Validate an optional sidecar, then verify its weights at load.
Hashes are checked against the post-fusion stream, before TP conversion. They cover all inputs of the cached computation, not unrelated QKV/MLP weights. A fixed adapter has an additional complete-file identity.
MiniMaxH3RuntimeAdalnCache ¶
Bases: ExactProjectionCache
Exact projection caching with optional H3 schedule-bound sidecar results.
build_cache ¶
build_cache(
arch: MiniMaxH3DiTArchConfig,
weights: Iterable[tuple[str, Tensor]],
*,
contract: dict[str, Any],
model_variant: str,
adapter_sha256: str | None,
device: device,
) -> tuple[dict[str, Tensor], dict[str, Any]]
Stream one projection at a time; do not instantiate the full DiT.
The four time-embedder tensors must precede the AdaLN projections. Callers fuse any fixed adapter before supplying this stream, just like serving.
math_identity ¶
Conservatively bind the builder's numerical environment (TP1 only).