Skip to content

vllm_omni.diffusion.cache.exact_projection_cache

Model-independent exact caching of small conditioning projections.

ExactProjectionCache

Bounded, exact memoization for deterministic conditioning projections.

Models call :meth:prepare once for an immutable conditioning tensor, then :meth:project for each leaf projection using a stable, model-defined name. The supplied computation must depend only on that tensor, the projection's parameters/buffers and fixed preprocessing. Include all request conditioning in the tensor; timestep alone is insufficient for prompt/guidance-dependent projections. Outputs returned on cache hits must be treated as read-only.

The original projection computes misses, preserving its quantization, LoRA and TP behavior. Parameter/buffer versions guard reuse, and TP ranks vote before skipping collectives. Call :meth:clear for adapter/configuration changes or model moves. Gradient-enabled and compiled execution bypass reuse.

Each cache belongs to one model instance executing forwards serially. It retains a bounded subset across requests, so it is independent of the request-scoped approximate cache backends. It does not offload weights or suppress block prefetch. Models may supply validated precomputed results by overriding :meth:_lookup_precomputed; artifact loading stays model-owned.

hits instance-attribute

hits = 0

max_bytes instance-attribute

max_bytes = max_bytes

misses instance-attribute

misses = 0

clear

clear() -> None

prepare

prepare(embedding: Tensor) -> None

project

project(
    name: str,
    module: Module,
    embedding: Tensor,
    compute: Callable[[], Tensor],
) -> Tensor

canonical_json

canonical_json(value: Any) -> str

tensor_digest

tensor_digest(tensor: Tensor) -> str

Hash dtype, shape and actual bytes, including BF16, without widening.