vllm_omni.diffusion.models.boogu_image.sp_layout ¶
Shard-layout arithmetic for BOOGU Image sequence parallelism.
BOOGU splits three independent sequences (noise image, reference images, and instruction context) at one SP boundary. The framework records each sequence's pre-padding global length under a shard_group key; everything else -- how that global length maps onto a given rank's shard -- is arithmetic, so every rank can derive the layout of all ranks without communicating.
That property is what lets the model build attention masks for the global, rank-concatenated sequence that Ulysses produces after its all-to-all, instead of all-gathering rank-local masks inside the attention layer.
ShardLayout dataclass ¶
How one globally padded sequence maps onto the SP ranks.
bounds ¶
Half-open [start, end) span of the global sequence owned by rank.
resolve classmethod ¶
resolve(
shard_group: str, *, local_seq_len: int
) -> ShardLayout
Read a boundary's layout from the ForwardContext.
Falls back to an unsharded layout (world_size=1) when SP is off or the tensor was never split, in which case local_seq_len is the whole sequence.
segment_lengths ¶
Per-sample contiguous segments intersected with rank's shard.
pack_local_rotary ¶
pack_local_rotary(
context_rotary_emb: RotaryEmbedding,
ref_img_rotary_emb: RotaryEmbedding,
noise_rotary_emb: RotaryEmbedding,
encoder_seq_lengths: list[int],
ref_img_seq_lengths: list[list[int]],
img_seq_lengths: list[int],
) -> tuple[
RotaryEmbedding, RotaryEmbedding, list[int], list[int]
]
Pack local RoPE as [context, references, noise].
Capacities stay equal across ranks; returned lengths exclude padding. Rotary embeddings arrive as (cos, sin) pairs -- the same packing is applied component-wise so both halves stay index-aligned.
rank_concat_mask_or_none ¶
rank_concat_mask_or_none(
per_rank_lengths: list[list[int]],
capacity: int,
*,
like: Tensor,
required: bool = False,
) -> Tensor | None
Mask over the rank-concatenated sequence Ulysses builds post all-to-all.
per_rank_lengths[r][i] is sample i's valid length inside rank r's shard, and every shard has the same capacity (guaranteed by auto_pad), so the global sequence is world_size * capacity long. Returns None for an all-valid mask unless a padding contract requires one.