vllm_omni.worker_v2.first_audio_sender ¶
Deliver in-stage first audio as soon as its device copy completes.
The regular output of a step is published only after the whole step's outputs are copied, materialized for every request, popped by the engine loop and serialized together. A stream's first audio chunk does not need any of that, so a dedicated thread waits for the chunk's own copy event and hands a one-request output straight to the engine core's output queue.
FirstAudioSender ¶
submit ¶
submit(
request_ids: list[str],
pcm: Tensor,
sample_rate: Tensor,
valid: Tensor | None = None,
) -> list[str]
Queue a D2H copy of pcm [n, samples] on the current stream and deliver it once done.
valid [n] (on device) drops rows whose frame turned out not to be audio (e.g. a codec EOS sample); it is read only after the copy.
Return the request IDs whose delivery was accepted. Rows without an engine output route must retain regular codec delivery instead of promising the orchestrator a first frame that can never arrive.