Skip to content

vllm.compilation.cudagraph_pool

Graph pool routing for CUDA graph capture under sleep mode.

Functions:

  • capture_outside_cumem_pool –

    CUDA graphs captured here stay out of cuMem, so that memory profiling can

  • capture_pool –

    Yield the pool to capture into, the cuMem graph pool when sleep offloads

capture_outside_cumem_pool()

CUDA graphs captured here stay out of cuMem, so that memory profiling can free its throwaway graphs normally.

Source code in vllm/compilation/cudagraph_pool.py
@contextmanager
def capture_outside_cumem_pool() -> Iterator[None]:
    """CUDA graphs captured here stay out of cuMem, so that memory profiling can
    free its throwaway graphs normally."""
    token = _outside_cumem.set(True)
    try:
        yield
    finally:
        _outside_cumem.reset(token)

capture_pool(pool, vllm_config)

Yield the pool to capture into, the cuMem graph pool when sleep offloads graph memory, and point NCCL's graph allocator at it.

Source code in vllm/compilation/cudagraph_pool.py
@contextmanager
def capture_pool(
    pool: tuple[int, int] | None, vllm_config: VllmConfig
) -> Iterator[tuple[int, int] | None]:
    """Yield the pool to capture into, the cuMem graph pool when sleep offloads
    graph memory, and point NCCL's graph allocator at it."""
    if vllm_config.use_cumem_cudagraph_pool and not _outside_cumem.get():
        allocator = cast("CuMemAllocator", get_mem_allocator_instance())
        with allocator.cudagraph_pool() as cumem_pool:
            set_graph_pool_id(cumem_pool)
            yield cumem_pool
    else:
        set_graph_pool_id(pool or current_platform.graph_pool_handle())
        yield pool