Skip to content

vllm_omni.diffusion.models.boogu_image.sp_layout

Shard-layout arithmetic for BOOGU Image sequence parallelism.

BOOGU splits three independent sequences (noise image, reference images, and instruction context) at one SP boundary. The framework records each sequence's pre-padding global length under a shard_group key; everything else -- how that global length maps onto a given rank's shard -- is arithmetic, so every rank can derive the layout of all ranks without communicating.

That property is what lets the model build attention masks for the global, rank-concatenated sequence that Ulysses produces after its all-to-all, instead of all-gathering rank-local masks inside the attention layer.

RotaryEmbedding module-attribute

RotaryEmbedding = tuple[torch.Tensor, ...]

ShardLayout dataclass

How one globally padded sequence maps onto the SP ranks.

local_seq_len property

local_seq_len: int

original_seq_len instance-attribute

original_seq_len: int

padded_seq_len property

padded_seq_len: int

padding_size property

padding_size: int

rank instance-attribute

rank: int

world_size instance-attribute

world_size: int

bounds

bounds(rank: int) -> tuple[int, int]

Half-open [start, end) span of the global sequence owned by rank.

resolve classmethod

resolve(
    shard_group: str, *, local_seq_len: int
) -> ShardLayout

Read a boundary's layout from the ForwardContext.

Falls back to an unsharded layout (world_size=1) when SP is off or the tensor was never split, in which case local_seq_len is the whole sequence.

segment_lengths

segment_lengths(
    global_segments: list[list[int]], *, rank: int
) -> list[list[int]]

Per-sample contiguous segments intersected with rank's shard.

valid_lengths

valid_lengths(
    global_lengths: list[int], *, rank: int
) -> list[int]

Per-sample valid prefix lengths clipped to rank's shard.

pack_local_rotary

pack_local_rotary(
    context_rotary_emb: RotaryEmbedding,
    ref_img_rotary_emb: RotaryEmbedding,
    noise_rotary_emb: RotaryEmbedding,
    encoder_seq_lengths: list[int],
    ref_img_seq_lengths: list[list[int]],
    img_seq_lengths: list[int],
) -> tuple[
    RotaryEmbedding, RotaryEmbedding, list[int], list[int]
]

Pack local RoPE as [context, references, noise].

Capacities stay equal across ranks; returned lengths exclude padding. Rotary embeddings arrive as (cos, sin) pairs -- the same packing is applied component-wise so both halves stay index-aligned.

rank_concat_mask_or_none

rank_concat_mask_or_none(
    per_rank_lengths: list[list[int]],
    capacity: int,
    *,
    like: Tensor,
    required: bool = False,
) -> Tensor | None

Mask over the rank-concatenated sequence Ulysses builds post all-to-all.

per_rank_lengths[r][i] is sample i's valid length inside rank r's shard, and every shard has the same capacity (guaranteed by auto_pad), so the global sequence is world_size * capacity long. Returns None for an all-valid mask unless a padding contract requires one.