Skip to content

vllm_omni.worker_v2.model_states.eager_mtp

Opt-in completion of a Talker frame immediately after sampling codebook zero.

logger module-attribute

logger = init_logger(__name__)

EagerMTPState

Used only by models explicitly declaring mtp_eager_frames.

owner instance-attribute

owner = owner

finish_audio

finish_audio(req_ids: set[str]) -> None

has_pending_replay

has_pending_replay() -> bool

Whether any preempted request still has saved inputs to replay.

prepare_audio_output

prepare_audio_output(
    input_batch: InputBatch,
    req_states: RequestState,
    outputs: dict[str, Any],
) -> StreamingAudioOutput | None

record_inputs

record_inputs(
    input_batch: InputBatch, embeds: Tensor
) -> None

replay_inputs

replay_inputs(
    req_id: str,
    req_idx: int,
    offset: int,
    ids: Tensor,
    embeds: Tensor,
) -> bool

Rebuild freed Talker KV from the exact conditioned inputs, without rerunning MTP.

resume_audio

resume_audio(req_id: str, req_idx: int) -> None

run_eager_mtp

run_eager_mtp(
    input_batch: InputBatch,
    text_hidden: Tensor,
    sampled_token_ids: Tensor,
    multimodal_outputs: dict[str, Any],
    mtp_batch_descriptor_dispatcher: Callable[[int], Any]
    | None = None,
) -> None

Complete each sampled row's frame in this step and publish it with this step's output.

Consumes the rows recorded by run_preprocess. Writes the frame codes and validity into the last token row of each request span of the retained multimodal output, and keeps the frame's codec embedding sum for the next step's input. Issued on the main stream before the async output copy, so no extra synchronization is needed.

set_first_audio_sink

set_first_audio_sink(sink: Any) -> None

suspend_audio

suspend_audio(req_id: str, req_idx: int) -> None

SuspendedAudio dataclass

decoder_state instance-attribute

decoder_state: list[Tensor]

frame_embed instance-attribute

frame_embed: Tensor

generator instance-attribute

generator: Generator | None

position instance-attribute

position: int

runtime instance-attribute

runtime: dict[str, Any]

TalkerInputs dataclass

embeds instance-attribute

embeds: Tensor

length class-attribute instance-attribute

length: int = 0