7. Offloaded Tensors
Goals
At the end of this tutorial you will be able to:
- Use offloaded tensors when wished
Prerequisite
- PyTorch and Python basics
- 5 min to read this page
Offload tensors have been developed to allow to manipulate and export more easily large neural network models.
Recall that if you only want to export a LLM model offloaded you can look at our related LLM tutorial and do not need to look at what happen behind.
This class is defined as such:
-
torch_to_nnef.tensor.offload.OffloadedTensorOffloadedTensor(elem, device, offload_dir: Path, name: str, offloaded_tensor_type: Type[Tensor], force_gc_collect: bool = False, storage_id: Optional[str] = None)Bases:
OpaqueTensorTensor subclass that maintains data on disk.
It hold an virtual internal memory storage (permanent) and a temporary instantiation at each operation accessing it on targeted device.
Warning
we recommend to version of PyTorch > 1.12 for best compatibility.
Methods:
Name Description from_original_tensorTake a torch.Tensor or OpaqueTensor and offload it to disk.
reloadReload the stored value on
device.set_Implement tensor-style storage replacement for offloaded payloads.
toChange the target device when reloaded in memory.
update_valuesReplace offloaded tensor by new 'values' tensor.
Attributes:
Name Type Description is_metaboolWhether the tensor is on the meta device.
is_metapropertyWhether the tensor is on the meta device.
Always False as the tensor is (off|re)loaded from disk.
from_original_tensorclassmethodfrom_original_tensor(tensor: Tensor, name: str, offload_dir: Optional[Path] = None, suffix_log_msg: str = '')Take a torch.Tensor or OpaqueTensor and offload it to disk.
Parameters:
Name Type Description Default tensorTensorthe torch.Tensor or torch_to_nnef.tensor.OpaqueTensor to dump on disk
required namestrthe name of the tensor that will be used to create the filename store on disk
required offload_dirOptional[Path]The directory where this file will be stored (temporarly)
Nonesuffix_log_msgstrAdded message log suffix for context
''reloadReload the stored value on
device.The optional override does not change the tensor's configured target device. This lets a residency manager stage the same payload on a worker-selected device without mutating shared tensor state.
set_Implement tensor-style storage replacement for offloaded payloads.
OffloadedTensoruses a meta tensor as its in-memory shell, so PyTorch's nativeTensor.set_cannot replace its storage with a CPU tensor. Route the commonparam.set_(new_tensor)form through the offload store instead. This is important for quantizers that update a weight in-place before replacing it with a QTensor.update_valuesReplace offloaded tensor by new 'values' tensor.
Parameters:
Name Type Description Default valuesTensorThe tensor that will replace it on disk assertion are made to ensure same shape, dtype as prior
required strict_shapeboolif True (default) the shape of the new tensor must be the same as the prior one
Truestrict_dtypeboolif True (default) the dtype of the new tensor must be the same as the prior one
True
You can directly load any .safetensor or .pt into this object that will mimic classical
torch.Tensor except that each access will load the Tensor from disk and remove it from RAM as
soon as those are not needed, allowing to manipulate very large model bit by bit.
It is composable with other torch_to_nnef.tensor.opaque.OpaqueTensor such as QTensor.
To load from disk without overhead,
you can call the t2n_load_checkpoint_and_dispatch with appropriate options like in the following example:
import tempfile
from pathlib import Path
from torch_to_nnef.tensor.offload import (
ON_DISK_DEVICE_MAP_KEY,
t2n_load_checkpoint_and_dispatch,
)
from torch_to_nnef.utils import init_empty_weights
from transformers import AutoModelForCausalLM
import huggingface_hub
slug = "meta-llama/Llama-3.2-1B-Instruct"
with init_empty_weights():
# model instantiation with empty tensors
# this can be come from any library (here transformers)
model = AutoModelForCausalLM.from_pretrained(slug, **kwargs)
hf_repo_files = huggingface_hub.list_repo_files(slug)
weights_location = Path(
huggingface_hub.hf_hub_download(
slug, hf_repo_files[-1]
) # assume at least 1 file is in targeted repo
).parent
# here model tensors are properly loaded into
t2n_load_checkpoint_and_dispatch(
model,
weights_location,
device_map=ON_DISK_DEVICE_MAP_KEY,
offload_dir=Path(tempfile.mkdtemp(suffix="offload_t2n")),
)
These OffloadedTensor are also very useful to implement into quantization techniques to
support very large model quantization with a calibration based on observed values like Hessian from activation.
Indeed if we think of the Hessian example: these square matrices can be pretty large especially
when multiplied by the number of activations on a big neural network.
If you only wish to maintain QTensor into OffloadedTensor if original float tensor was offloaded you can just use the helper:
If this is a new tensor just use the OffloadedTensor.from_original_tensor defined upper.
Scoped residency and prefetch
Repeated operations on offloaded values can retain them in memory with
TensorResidencyPool. Activating the pool for an execution scope makes normal
tensor operations use resident values transparently, while prefetch starts
loading a later value on a background worker.
from torch_to_nnef.tensor import ResidencyStrategy, TensorResidencyPool
with TensorResidencyPool(max_cached_bytes=2 * 1024**3) as pool:
with pool.scope(device="cuda"):
pool.prefetch(next_weight, device="cuda")
output = input @ current_weight
output = output @ next_weight
This allows a model or module hook to schedule the next tensor without changing the operations that consume the current one.
The first access to an OffloadedTensor through a pool that does not yet
contain it is a cache miss: the pool materializes it from disk on the selected
device. With the default ResidencyStrategy.LRU, a value that fits within
max_cached_bytes becomes cached, and later accesses through the same pool
reuse it without another disk load. It remains cached until explicit eviction,
least-recently-used eviction to make room for another value, or pool closure.
prefetch starts this first load early but otherwise follows the same retention
rules. This caching is specific to access routed through the pool; other access
to an OffloadedTensor keeps its normal transient materialization behavior.
LRU performs poorly for a cyclic scan whose working set is larger than the
cache: values can be evicted shortly before the next forward pass needs them.
ResidencyStrategy.FIXED instead admits the first completed values that fit in
the budget and does not replace them during later accesses. Values that do not
fit are still delivered to their callers but are not retained. This provides a
stable resident subset while the rest of each forward pass streams through the
pool.
with TensorResidencyPool(
max_cached_bytes=2 * 1024**3,
strategy=ResidencyStrategy.FIXED,
) as pool:
with pool.scope(device="cuda"):
for batch in calibration_batches:
calibrate(model, batch)
With the default single worker, fixed admission follows request order. With
multiple workers, concurrently prefetched values are admitted in completion
order. Explicit evict can remove an admitted value and make capacity
available for a later materialization. Pool closure releases the full set.
Manual pin and unpin apply only to ResidencyStrategy.LRU; the FIXED
strategy rejects them with T2NErrorMisuse because it already controls fixed
admission. Under LRU, pinning starts or reuses a load and is idempotent. A
pinned value is excluded from LRU eviction until unpin is called or the pool
closes, counts toward resident_bytes and the cache budget, and can make
residency exceed the soft budget.
A lease is a scoped claim that a materialized tensor is currently in use. It
lasts for the duration of the with pool.acquire(...) block. While any lease
is active, the pool keeps that value resident and does not evict it. Multiple
callers may lease the same resident value. Use a read_write lease when an
operation changes the value; the pool writes dirty values back before
eviction, during flush, or when the pool closes.
max_cached_bytes is a soft limit on payloads retained for reuse, not a hard
limit on process or device memory. LRU enforces it by evicting unleased,
unpinned values in least-recently-used order. FIXED preserves admitted values
and removes non-admitted values after delivery. If an active caller leases a
value larger than the cache budget, the pool still materializes it so the
operation can proceed, keeps it resident until the lease ends, and then evicts
it instead of caching it. A prefetched oversized value is similarly delivered
to its waiting caller without being retained. Consequently, active values and
temporary framework allocations can exceed max_cached_bytes; callers that
require a hard allocation limit must validate their working-set sizes
separately.
The torch_to_nnef.tensor.residency logger emits these decisions at DEBUG
level, including whether the value is not admitted by FIXED, delivered without
cache retention, kept until its final lease ends, or retained by an LRU pin.
Scheduling decisions, such as which model block to prefetch next, remain with
the caller. The pool only manages OffloadedTensor values. Passing a regular
torch.Tensor to acquire, prefetch, resolve, pin, unpin, flush, or
evict raises T2NErrorMisuse immediately; ordinary tensors are already
materialized and do not need the residency layer.