vllm_omni.model_executor.common.request_outputs ¶
Reading a vLLM RequestOutput the way model-side projection code needs to.
Stage outputs arrive as duck-typed objects (a RequestOutput, an OmniRequestOutput wrapping one, or a test double): these helpers unwrap them, read the multimodal_output mapping, coerce tensor scalars, and turn cumulative audio/text into the delta since the last read. No duplex vocabulary lives here.
audio_value ¶
The audio carried by a multimodal_output (audio / model_outputs / latent).
coerce_int ¶
int(value) for scalars, one-element tensors/arrays and numeric strings; None otherwise.
multimodal_output ¶
The first non-empty multimodal_output mapping of the output or its completion, copied.
slice_audio_delta ¶
The samples of a cumulative audio past offset; the whole audio when it restarted.
text_delta ¶
What text adds over previous when outputs are cumulative; the whole text on a restart.