Skip to content

vllm_omni.model_executor.common.request_outputs

Reading a vLLM RequestOutput the way model-side projection code needs to.

Stage outputs arrive as duck-typed objects (a RequestOutput, an OmniRequestOutput wrapping one, or a test double): these helpers unwrap them, read the multimodal_output mapping, coerce tensor scalars, and turn cumulative audio/text into the delta since the last read. No duplex vocabulary lives here.

audio_sample_count

audio_sample_count(audio: object | None) -> int | None

audio_value

audio_value(
    multimodal: Mapping[str, object],
) -> object | None

The audio carried by a multimodal_output (audio / model_outputs / latent).

coerce_int

coerce_int(value: object) -> int | None

int(value) for scalars, one-element tensors/arrays and numeric strings; None otherwise.

coerce_int_list

coerce_int_list(value: object) -> list[int]

first_completion

first_completion(output: object) -> object | None

multimodal_output

multimodal_output(
    output: object, completion: object | None = None
) -> dict[str, object]

The first non-empty multimodal_output mapping of the output or its completion, copied.

sample_rate_hz

sample_rate_hz(
    multimodal: Mapping[str, object], *, default: int
) -> int

slice_audio_delta

slice_audio_delta(
    audio: object | None, offset: int
) -> object | None

The samples of a cumulative audio past offset; the whole audio when it restarted.

text_delta

text_delta(text: str, previous: str) -> str

What text adds over previous when outputs are cumulative; the whole text on a restart.

text_value

text_value(
    multimodal: Mapping[str, object],
    completion: object | None,
) -> str

unwrap_request_output

unwrap_request_output(
    output: object,
) -> tuple[object, object | None]

Return (request_output, first_completion) for a stage output or a wrapper around one.