vllm_omni.model_executor.models.personaplex.duplex ¶
PersonaPlex full-duplex integration.
PersonaPlex (nvidia/personaplex-7b-v1) is a Moshi finetune: a pure-lockstep speech-to-speech model. This package plugs it into the unified full-duplex framework through the one seam the framework has, PipelineConfig.duplex_plugin:
- :class:
PersonaPlexDuplexPlugintheDuplexModelPlugin(engine + session policy) - :class:
PersonaPlexStage0DuplexRuntimeworker-side lockstep state and first-append prefill - :class:
PersonaPlexPcmAppendBuffer80 ms PCM input framing
Modules:
| Name | Description |
|---|---|
capabilities | |
config | Constants of the PersonaPlex full-duplex integration. |
data_plane | PersonaPlex output projection: cumulative Code2Wav audio and inner-monologue text into deltas. |
input | PCM input framing for PersonaPlex: 24 kHz |
plugin | PersonaPlex full-duplex model plugin: engine policy and session policy in one class. |
policy | PersonaPlex frame/token contract (the model policy, no engine state). |
stage0 | |
PersonaPlexDuplexPlugin ¶
Bases: DuplexModelPlugin
PersonaPlex-owned sampling policy, append planning, session state and output projection.
private_runtime_config_keys class-attribute instance-attribute ¶
private_runtime_config_keys = PRIVATE_RUNTIME_CONFIG_KEYS
silence_continuation_sample_rate_hz class-attribute instance-attribute ¶
silence_continuation_sample_rate_hz = SAMPLE_RATE
silence_continuation_samples class-attribute instance-attribute ¶
silence_continuation_samples = FRAME_SIZE
configure_sampling_params ¶
configure_sampling_params(
*,
runtime_config: dict[str, object],
defaults: tuple[object, ...],
) -> tuple[object, ...]
decide_output ¶
decide_output(
*,
stage_id: int,
final_stage_id: int,
segment_finished: bool,
segment_token_ids: tuple[int, ...],
segment_output_metadata: dict[str, object],
output: object,
) -> DuplexOutputDecision | None
plan_append ¶
plan_append(
*,
request_id: str,
fence: DuplexFence,
session_config: dict[str, object],
runtime_config: dict[str, object],
seq: int,
turn_seq: int,
payload: object,
final: bool,
sampling_params: object,
) -> DuplexAppendPlan
prepare_runtime_config async ¶
prepare_runtime_config(
config: DuplexSessionConfig,
*,
model_config: ModelConfig | None,
) -> dict[str, object]
PersonaPlexPcmAppendBuffer ¶
Bases: FixedFramePcmAppendBuffer
Transactionally frame 24 kHz float PCM into PersonaPlex 80 ms units.
PersonaPlexStage0DuplexRuntime ¶
Own the shared streaming Mimi encoder and each session's first-append prefill.
One encoder holds max_sessions streaming rows; a live (session, epoch) leases one row for its lifetime. encode_appends encodes every new append of a scheduler step in one batched call (rows without a new append are inactive and keep their state), and prepare_append consumes the result.
encode_appends ¶
Encode the new user frame of each append in one batched encoder call.
Called once per scheduler step before the per-request prepare_append calls. An append whose (epoch, seq) is already encoded or prepared (a chunked first prefill spans several steps) is skipped, so a row's streaming state advances exactly once per frame. Appends that cannot be admitted are left for prepare_append to reject.
PrefillStep dataclass ¶
One tick of a recycled slot's system-prompt replay (see batched serving).