vllm_omni.worker_v2.model_states.eager_mtp ¶
Opt-in completion of a Talker frame immediately after sampling codebook zero.
EagerMTPState ¶
Used only by models explicitly declaring mtp_eager_frames.
has_pending_replay ¶
has_pending_replay() -> bool
Whether any preempted request still has saved inputs to replay.
prepare_audio_output ¶
prepare_audio_output(
input_batch: InputBatch,
req_states: RequestState,
outputs: dict[str, Any],
) -> StreamingAudioOutput | None
replay_inputs ¶
Rebuild freed Talker KV from the exact conditioned inputs, without rerunning MTP.
run_eager_mtp ¶
run_eager_mtp(
input_batch: InputBatch,
text_hidden: Tensor,
sampled_token_ids: Tensor,
multimodal_outputs: dict[str, Any],
mtp_batch_descriptor_dispatcher: Callable[[int], Any]
| None = None,
) -> None
Complete each sampled row's frame in this step and publish it with this step's output.
Consumes the rows recorded by run_preprocess. Writes the frame codes and validity into the last token row of each request span of the retained multimodal output, and keeps the frame's codec embedding sum for the next step's input. Issued on the main stream before the async output copy, so no extra synchronization is needed.