vllm_omni.diffusion.cache.exact_projection_cache ¶
Model-independent exact caching of small conditioning projections.
ExactProjectionCache ¶
Bounded, exact memoization for deterministic conditioning projections.
Models call :meth:prepare once for an immutable conditioning tensor, then :meth:project for each leaf projection using a stable, model-defined name. The supplied computation must depend only on that tensor, the projection's parameters/buffers and fixed preprocessing. Include all request conditioning in the tensor; timestep alone is insufficient for prompt/guidance-dependent projections. Outputs returned on cache hits must be treated as read-only.
The original projection computes misses, preserving its quantization, LoRA and TP behavior. Parameter/buffer versions guard reuse, and TP ranks vote before skipping collectives. Call :meth:clear for adapter/configuration changes or model moves. Gradient-enabled and compiled execution bypass reuse.
Each cache belongs to one model instance executing forwards serially. It retains a bounded subset across requests, so it is independent of the request-scoped approximate cache backends. It does not offload weights or suppress block prefetch. Models may supply validated precomputed results by overriding :meth:_lookup_precomputed; artifact loading stays model-owned.