Skip to content

vllm_omni.model_executor.models.personaplex.duplex

PersonaPlex full-duplex integration.

PersonaPlex (nvidia/personaplex-7b-v1) is a Moshi finetune: a pure-lockstep speech-to-speech model. This package plugs it into the unified full-duplex framework through the one seam the framework has, PipelineConfig.duplex_plugin:

  • :class:PersonaPlexDuplexPlugin the DuplexModelPlugin (engine + session policy)
  • :class:PersonaPlexStage0DuplexRuntime worker-side lockstep state and first-append prefill
  • :class:PersonaPlexPcmAppendBuffer 80 ms PCM input framing

Modules:

Name Description
capabilities
config

Constants of the PersonaPlex full-duplex integration.

data_plane

PersonaPlex output projection: cumulative Code2Wav audio and inner-monologue text into deltas.

input

PCM input framing for PersonaPlex: 24 kHz pcm_f32le into 1920-sample (80 ms) units.

plugin

PersonaPlex full-duplex model plugin: engine policy and session policy in one class.

policy

PersonaPlex frame/token contract (the model policy, no engine state).

stage0

PersonaPlexDuplexPlugin

Bases: DuplexModelPlugin

PersonaPlex-owned sampling policy, append planning, session state and output projection.

data_plane instance-attribute

data_plane = PersonaPlexDataPlaneSession(encode_audio)

plugin_id class-attribute instance-attribute

plugin_id = 'personaplex'

private_runtime_config_keys class-attribute instance-attribute

private_runtime_config_keys = PRIVATE_RUNTIME_CONFIG_KEYS

silence_continuation_sample_rate_hz class-attribute instance-attribute

silence_continuation_sample_rate_hz = SAMPLE_RATE

silence_continuation_samples class-attribute instance-attribute

silence_continuation_samples = FRAME_SIZE

capabilities

capabilities(*, max_sessions: int) -> DuplexCapabilities

configure_sampling_params

configure_sampling_params(
    *,
    runtime_config: dict[str, object],
    defaults: tuple[object, ...],
) -> tuple[object, ...]

create_session_state

create_session_state() -> PersonaPlexSessionState

decide_output

decide_output(
    *,
    stage_id: int,
    final_stage_id: int,
    segment_finished: bool,
    segment_token_ids: tuple[int, ...],
    segment_output_metadata: dict[str, object],
    output: object,
) -> DuplexOutputDecision | None

plan_append

plan_append(
    *,
    request_id: str,
    fence: DuplexFence,
    session_config: dict[str, object],
    runtime_config: dict[str, object],
    seq: int,
    turn_seq: int,
    payload: object,
    final: bool,
    sampling_params: object,
) -> DuplexAppendPlan

prepare_runtime_config async

prepare_runtime_config(
    config: DuplexSessionConfig,
    *,
    model_config: ModelConfig | None,
) -> dict[str, object]

runtime_config_for_update

runtime_config_for_update(
    config: DuplexSessionConfig,
    current: Mapping[str, object],
) -> dict[str, object]

PersonaPlexPcmAppendBuffer

Bases: FixedFramePcmAppendBuffer

Transactionally frame 24 kHz float PCM into PersonaPlex 80 ms units.

PersonaPlexStage0DuplexRuntime

Own the shared streaming Mimi encoder and each session's first-append prefill.

One encoder holds max_sessions streaming rows; a live (session, epoch) leases one row for its lifetime. encode_appends encodes every new append of a scheduler step in one batched call (rows without a new append are inactive and keep their state), and prepare_append consumes the result.

device instance-attribute

device = device

max_sessions instance-attribute

max_sessions = max_sessions

model_path instance-attribute

model_path = model_path

request_sessions instance-attribute

request_sessions: dict[str, tuple[str, int]] = {}

sessions instance-attribute

stage_model instance-attribute

stage_model = stage_model

close_request

close_request(request_id: str) -> None

close_session

close_session(session_id: str, epoch: int) -> None

encode_appends

encode_appends(appends: list[dict[str, Any]]) -> None

Encode the new user frame of each append in one batched encoder call.

Called once per scheduler step before the per-request prepare_append calls. An append whose (epoch, seq) is already encoded or prepared (a chunked first prefill spans several steps) is skipped, so a row's streaming state advances exactly once per frame. Appends that cannot be admitted are left for prepare_append to reject.

prepare_append

prepare_append(
    duplex: dict[str, Any],
    *,
    prompt_len: int,
    request_id: str | None = None,
) -> PersonaPlexStage0PreparedAppend

record_sample

record_sample(
    *, request_id: str, text_token: Any, agent_codes: Any
) -> None

Commit one sampled temporal frame for the next live append.

PrefillStep dataclass

One tick of a recycled slot's system-prompt replay (see batched serving).

embedding class-attribute instance-attribute

embedding: Any = None

kind instance-attribute

kind: str

moshi_tokens class-attribute instance-attribute

moshi_tokens: Any = None

text_token class-attribute instance-attribute

text_token: int | None = None

user_sine class-attribute instance-attribute

user_sine: bool = False