Skip to content

7. Offloaded Tensors

Goals

At the end of this tutorial you will be able to:

  1. Use offloaded tensors when wished

Prerequisite

  • PyTorch and Python basics
  • 5 min to read this page

Offload tensors have been developed to allow to manipulate and export more easily large neural network models.

Recall that if you only want to export a LLM model offloaded you can look at our related LLM tutorial and do not need to look at what happen behind.

This class is defined as such:

  • torch_to_nnef.tensor.offload.OffloadedTensor

    OffloadedTensor(elem, device, offload_dir: Path, name: str, offloaded_tensor_type: Type[Tensor], force_gc_collect: bool = False, storage_id: Optional[str] = None)
    

    Bases: OpaqueTensor

    Tensor subclass that maintains data on disk.

    It hold an virtual internal memory storage (permanent) and a temporary instantiation at each operation accessing it on targeted device.

    Warning

    we recommend to version of PyTorch > 1.12 for best compatibility.

    Methods:

    Name Description
    from_original_tensor

    Take a torch.Tensor or OpaqueTensor and offload it to disk.

    reload

    Reload the stored value on device.

    set_

    Implement tensor-style storage replacement for offloaded payloads.

    to

    Change the target device when reloaded in memory.

    update_values

    Replace offloaded tensor by new 'values' tensor.

    Attributes:

    Name Type Description
    is_meta bool

    Whether the tensor is on the meta device.

    is_meta property

    is_meta: bool
    

    Whether the tensor is on the meta device.

    Always False as the tensor is (off|re)loaded from disk.

    from_original_tensor classmethod

    from_original_tensor(tensor: Tensor, name: str, offload_dir: Optional[Path] = None, suffix_log_msg: str = '')
    

    Take a torch.Tensor or OpaqueTensor and offload it to disk.

    Parameters:

    Name Type Description Default
    tensor
    Tensor

    the torch.Tensor or torch_to_nnef.tensor.OpaqueTensor to dump on disk

    required
    name
    str

    the name of the tensor that will be used to create the filename store on disk

    required
    offload_dir
    Optional[Path]

    The directory where this file will be stored (temporarly)

    None
    suffix_log_msg
    str

    Added message log suffix for context

    ''

    reload

    reload(device: Optional[TDEVICE] = None)
    

    Reload the stored value on device.

    The optional override does not change the tensor's configured target device. This lets a residency manager stage the same payload on a worker-selected device without mutating shared tensor state.

    set_

    set_(source: Tensor, *args, **kwargs)
    

    Implement tensor-style storage replacement for offloaded payloads.

    OffloadedTensor uses a meta tensor as its in-memory shell, so PyTorch's native Tensor.set_ cannot replace its storage with a CPU tensor. Route the common param.set_(new_tensor) form through the offload store instead. This is important for quantizers that update a weight in-place before replacing it with a QTensor.

    to

    to(*args, **kwargs)
    

    Change the target device when reloaded in memory.

    update_values

    update_values(values: Tensor, strict_shape: bool = True, strict_dtype: bool = True)
    

    Replace offloaded tensor by new 'values' tensor.

    Parameters:

    Name Type Description Default
    values
    Tensor

    The tensor that will replace it on disk assertion are made to ensure same shape, dtype as prior

    required
    strict_shape
    bool

    if True (default) the shape of the new tensor must be the same as the prior one

    True
    strict_dtype
    bool

    if True (default) the dtype of the new tensor must be the same as the prior one

    True

You can directly load any .safetensor or .pt into this object that will mimic classical torch.Tensor except that each access will load the Tensor from disk and remove it from RAM as soon as those are not needed, allowing to manipulate very large model bit by bit. It is composable with other torch_to_nnef.tensor.opaque.OpaqueTensor such as QTensor.

To load from disk without overhead, you can call the t2n_load_checkpoint_and_dispatch with appropriate options like in the following example:

example of offload usage from disk (extracted from LLM exporter)
import tempfile
from pathlib import Path
from torch_to_nnef.tensor.offload import (
    ON_DISK_DEVICE_MAP_KEY,
    t2n_load_checkpoint_and_dispatch,
)
from torch_to_nnef.utils import init_empty_weights

from transformers import AutoModelForCausalLM
import huggingface_hub

slug = "meta-llama/Llama-3.2-1B-Instruct"
with init_empty_weights():
    # model instantiation with empty tensors
    # this can be come from any library (here transformers)
    model = AutoModelForCausalLM.from_pretrained(slug, **kwargs)
hf_repo_files = huggingface_hub.list_repo_files(slug)
weights_location = Path(
    huggingface_hub.hf_hub_download(
        slug, hf_repo_files[-1]
    )  # assume at least 1 file is in targeted repo
).parent

# here model tensors are properly loaded into
t2n_load_checkpoint_and_dispatch(
    model,
    weights_location,
    device_map=ON_DISK_DEVICE_MAP_KEY,
    offload_dir=Path(tempfile.mkdtemp(suffix="offload_t2n")),
)

These OffloadedTensor are also very useful to implement into quantization techniques to support very large model quantization with a calibration based on observed values like Hessian from activation. Indeed if we think of the Hessian example: these square matrices can be pretty large especially when multiplied by the number of activations on a big neural network.

If you only wish to maintain QTensor into OffloadedTensor if original float tensor was offloaded you can just use the helper:

  • torch_to_nnef.compress.offloaded_tensor_qtensor

    offloaded_tensor_qtensor(q_fn, tensor: Tensor, suffix_name: str) -> torch.Tensor
    

    Maintains a QTensor offloaded if original tensor is offloaded.

If this is a new tensor just use the OffloadedTensor.from_original_tensor defined upper.

Scoped residency and prefetch

Repeated operations on offloaded values can retain them in memory with TensorResidencyPool. Activating the pool for an execution scope makes normal tensor operations use resident values transparently, while prefetch starts loading a later value on a background worker.

from torch_to_nnef.tensor import ResidencyStrategy, TensorResidencyPool

with TensorResidencyPool(max_cached_bytes=2 * 1024**3) as pool:
    with pool.scope(device="cuda"):
        pool.prefetch(next_weight, device="cuda")
        output = input @ current_weight
        output = output @ next_weight

This allows a model or module hook to schedule the next tensor without changing the operations that consume the current one.

The first access to an OffloadedTensor through a pool that does not yet contain it is a cache miss: the pool materializes it from disk on the selected device. With the default ResidencyStrategy.LRU, a value that fits within max_cached_bytes becomes cached, and later accesses through the same pool reuse it without another disk load. It remains cached until explicit eviction, least-recently-used eviction to make room for another value, or pool closure. prefetch starts this first load early but otherwise follows the same retention rules. This caching is specific to access routed through the pool; other access to an OffloadedTensor keeps its normal transient materialization behavior.

LRU performs poorly for a cyclic scan whose working set is larger than the cache: values can be evicted shortly before the next forward pass needs them. ResidencyStrategy.FIXED instead admits the first completed values that fit in the budget and does not replace them during later accesses. Values that do not fit are still delivered to their callers but are not retained. This provides a stable resident subset while the rest of each forward pass streams through the pool.

with TensorResidencyPool(
    max_cached_bytes=2 * 1024**3,
    strategy=ResidencyStrategy.FIXED,
) as pool:
    with pool.scope(device="cuda"):
        for batch in calibration_batches:
            calibrate(model, batch)

With the default single worker, fixed admission follows request order. With multiple workers, concurrently prefetched values are admitted in completion order. Explicit evict can remove an admitted value and make capacity available for a later materialization. Pool closure releases the full set.

Manual pin and unpin apply only to ResidencyStrategy.LRU; the FIXED strategy rejects them with T2NErrorMisuse because it already controls fixed admission. Under LRU, pinning starts or reuses a load and is idempotent. A pinned value is excluded from LRU eviction until unpin is called or the pool closes, counts toward resident_bytes and the cache budget, and can make residency exceed the soft budget.

A lease is a scoped claim that a materialized tensor is currently in use. It lasts for the duration of the with pool.acquire(...) block. While any lease is active, the pool keeps that value resident and does not evict it. Multiple callers may lease the same resident value. Use a read_write lease when an operation changes the value; the pool writes dirty values back before eviction, during flush, or when the pool closes.

with pool.acquire(statistic, mode="read_write") as value:
    value.add_(update)

max_cached_bytes is a soft limit on payloads retained for reuse, not a hard limit on process or device memory. LRU enforces it by evicting unleased, unpinned values in least-recently-used order. FIXED preserves admitted values and removes non-admitted values after delivery. If an active caller leases a value larger than the cache budget, the pool still materializes it so the operation can proceed, keeps it resident until the lease ends, and then evicts it instead of caching it. A prefetched oversized value is similarly delivered to its waiting caller without being retained. Consequently, active values and temporary framework allocations can exceed max_cached_bytes; callers that require a hard allocation limit must validate their working-set sizes separately. The torch_to_nnef.tensor.residency logger emits these decisions at DEBUG level, including whether the value is not admitted by FIXED, delivered without cache retention, kept until its final lease ends, or retained by an LRU pin.

Scheduling decisions, such as which model block to prefetch next, remain with the caller. The pool only manages OffloadedTensor values. Passing a regular torch.Tensor to acquire, prefetch, resolve, pin, unpin, flush, or evict raises T2NErrorMisuse immediately; ordinary tensors are already materialized and do not need the residency layer.