Skip to content

vllm_omni.worker_v2.output_snapshot

Compatibility imports for the shared model/runner output snapshot contract.

PackedOutputSnapshot

Bases: dict

A normal payload mapping with an internal batched-copy plan.

The runner's snapshot-slot event protects the device slabs until D2H is complete. Host slabs are allocated per output, so downstream consumers can retain their views without depending on reuse of the device ring slot.

producer_event instance-attribute

producer_event: Event | None = None

copy_to_cpu

copy_to_cpu(
    copy_tensor: Callable[[Tensor], Tensor],
) -> dict[str, Any]

record_producer_event

record_producer_event(stream: Stream) -> None

Publish readiness when a model produces on a separate CUDA stream.

RequestOutputSnapshot dataclass

Already partitioned, CPU-owned payloads in the current request order.

Model finalizers may return this after D2H completes to avoid recursively cloning and partitioning data whose request ownership is already known.

client class-attribute instance-attribute

client: list[dict[str, Any] | None] | None = None

inter_stage instance-attribute

inter_stage: list[dict[str, Any] | None]

pack_output_snapshot

pack_output_snapshot(
    payload: dict[str, Any],
    slot: dict[tuple[Any, ...], Tensor],
    *,
    max_buckets: int,
    reuse_existing_storage: bool = False,
) -> PackedOutputSnapshot | None

Copy tensor leaves once per dtype/device, retaining their exact values.

No CUDA synchronization is introduced here. The caller must select the producer stream and wait for the previous consumer before reusing a slot.