vllm.models.deepseek_v41.decoder_replay_layers
¶
Decoder-side SWA bounded replay: running the replay layers on their batch.
The layers past the last KV-source layer own nothing but sliding-window KV, so
in eager prefill steps they run on each request's last window rows only.
DeepseekV41ModelState prepares those rows as a sub-batch with attention
metadata and a forward context of its own, like a microbatch;
DecoderReplayLayers gathers the layer inputs by its rows, runs the layers
under that context and scatters the outputs back to batch rows. FULL graphs and
small PIECEWISE ones keep the layers on the whole batch; larger PIECEWISE graphs
break out to the replay batch, which may run in a graph of its own
(decoder_replay_cudagraph.py).
Classes:
-
DecoderReplayLayers–Runs the replay layers on the step's replay batch.
-
ReplayBatch–One forward's replay batch;
run_graphis set when it runs in a graph.
DecoderReplayLayers
¶
Runs the replay layers on the step's replay batch.
run_layers takes a batch's layer inputs and returns its per-row
outputs. row_buffers hold per-row results the source layer's indexer
publishes for the layers after it; they are compacted to the replay rows
in place. metadata_prefixes are the attention metadata keys the layers
read, so the replay batch builds only those.
Source code in vllm/models/deepseek_v41/decoder_replay_layers.py
_run_in_graph_break(hidden_states, *states)
¶
A graph break: the outputs land in fixed buffers the next segment reads.
Source code in vllm/models/deepseek_v41/decoder_replay_layers.py
ReplayBatch
dataclass
¶
One forward's replay batch; run_graph is set when it runs in a graph.