torch_to_nnef.tensor
Modules:
| Name | Description |
|---|---|
named |
|
offload |
OffLoad Tensor. |
opaque |
|
quant |
Advanced QTensor (<= 8bits) with complex quant scheme non torch native. |
residency |
Scoped residency and asynchronous prefetch for offloaded tensors. |
updater |
|
utils |
|
Classes:
| Name | Description |
|---|---|
NamedTensor |
Tensor enriched with name attribute. |
OffloadedTensor |
Tensor subclass that maintains data on disk. |
OpaqueTensorRef |
Allow to pass through 'tracing'. |
QScalePerGroupF16 |
f16 scale only per group. |
QTensor |
Common interface for all Compressed storage. |
QTensorTractScaleOnly |
Tract data format it serializes to: Q4_0. |
ResidencyStrategy |
Retention strategy for unleased values in a residency pool. |
TensorPrefetch |
Handle for an asynchronous tensor prefetch. |
TensorResidencyLease |
Scoped access to a value managed by a tensor residency pool. |
TensorResidencyPool |
Keep offloaded tensors resident across a bounded operation scope. |
Functions:
| Name | Description |
|---|---|
apply_name_to_tensor_in_module |
Transform torch.Tensor or Parameters into NamedTensor. |
set_opaque_tensor_in_params_as_ref |
Transform OpaqueTensor Parameters into OpaqueTensorRef. |
NamedTensor
Bases: Tensor
Tensor enriched with name attribute.
Attributes:
| Name | Type | Description |
|---|---|---|
data |
Very important to keep access to all special attr of NamedTensor. |
OffloadedTensor
OffloadedTensor(elem, device, offload_dir: Path, name: str, offloaded_tensor_type: Type[Tensor], force_gc_collect: bool = False, storage_id: Optional[str] = None)
Bases: OpaqueTensor
Tensor subclass that maintains data on disk.
It hold an virtual internal memory storage (permanent) and a temporary instantiation at each operation accessing it on targeted device.
Warning
we recommend to version of PyTorch > 1.12 for best compatibility.
Methods:
| Name | Description |
|---|---|
from_original_tensor |
Take a torch.Tensor or OpaqueTensor and offload it to disk. |
reload |
Reload the stored value on |
set_ |
Implement tensor-style storage replacement for offloaded payloads. |
to |
Change the target device when reloaded in memory. |
update_values |
Replace offloaded tensor by new 'values' tensor. |
Attributes:
| Name | Type | Description |
|---|---|---|
is_meta |
bool
|
Whether the tensor is on the meta device. |
is_meta
property
Whether the tensor is on the meta device.
Always False as the tensor is (off|re)loaded from disk.
from_original_tensor
classmethod
from_original_tensor(tensor: Tensor, name: str, offload_dir: Optional[Path] = None, suffix_log_msg: str = '')
Take a torch.Tensor or OpaqueTensor and offload it to disk.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
Tensor
|
the torch.Tensor or torch_to_nnef.tensor.OpaqueTensor to dump on disk |
required |
|
str
|
the name of the tensor that will be used to create the filename store on disk |
required |
|
Optional[Path]
|
The directory where this file will be stored (temporarly) |
None
|
|
str
|
Added message log suffix for context |
''
|
reload
Reload the stored value on device.
The optional override does not change the tensor's configured target device. This lets a residency manager stage the same payload on a worker-selected device without mutating shared tensor state.
set_
Implement tensor-style storage replacement for offloaded payloads.
OffloadedTensor uses a meta tensor as its in-memory shell, so
PyTorch's native Tensor.set_ cannot replace its storage with a CPU
tensor. Route the common param.set_(new_tensor) form through the
offload store instead. This is important for quantizers that update a
weight in-place before replacing it with a QTensor.
update_values
Replace offloaded tensor by new 'values' tensor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
Tensor
|
The tensor that will replace it on disk assertion are made to ensure same shape, dtype as prior |
required |
|
bool
|
if True (default) the shape of the new tensor must be the same as the prior one |
True
|
|
bool
|
if True (default) the dtype of the new tensor must be the same as the prior one |
True
|
OpaqueTensorRef
Bases: Tensor
Allow to pass through 'tracing'.
QScalePerGroupF16
QTensor
QTensor(fp_tensor: Tensor, qscheme: QScheme, dequant_to_dtype=torch.float32, u8_compressors: Optional[List[U8Compressor]] = None)
Bases: OpaqueTensor
Common interface for all Compressed storage.
Methods:
| Name | Description |
|---|---|
to_device |
Specific device handling. |
write_in_file |
Called at NNEF write time. |
QTensorTractScaleOnly
Bases: QTensorTract, SupportsOffloadState
Tract data format it serializes to: Q4_0.
Methods:
| Name | Description |
|---|---|
decompress |
Tract dequantization depends on hardware. |
ResidencyStrategy
Bases: Enum
Retention strategy for unleased values in a residency pool.
LRU replaces the least recently used eligible value and permits manual
pins. FIXED preserves its first-fit admitted set for the pool lifetime
and rejects manual pin operations.
TensorPrefetch
TensorPrefetch(pool: TensorResidencyPool, source: OffloadedTensor, device: device, future: Optional[Future], value: Optional[Tensor] = None)
TensorResidencyLease
Scoped access to a value managed by a tensor residency pool.
TensorResidencyPool
TensorResidencyPool(max_cached_bytes: Optional[int] = None, max_workers: int = 1, strategy: ResidencyStrategy = ResidencyStrategy.LRU)
Keep offloaded tensors resident across a bounded operation scope.
Values can be prefetched by background workers and acquired through a
lease. A lease is a scoped claim that a materialized value is in active
use, so its value remains resident until the lease ends. The default LRU
strategy evicts other values in least-recently-used order when the optional
cache budget is exceeded and supports manual pins. The FIXED strategy
retains the first completed values that fit and streams later values
without displacing that resident set. A read_write lease writes the
value back before eviction.
max_cached_bytes is a soft cache-retention limit, not a hard bound on
process memory. Leased values, and values manually pinned under LRU, remain
available even when they exceed it. An oversized value can therefore be
materialized for an active caller but is evicted as soon as its final lease
ends. Prefetched oversized values are returned to their waiting caller
without being retained by the pool unless they are pinned.
Methods:
| Name | Description |
|---|---|
acquire |
Return a scoped lease for |
close |
Flush all values, stop workers, and release resident memory. |
evict |
Write back and remove an unleased, unpinned value from the pool. |
flush |
Wait for loads and write dirty resident values back to storage. |
is_pinned |
Return whether |
is_resident |
Return whether |
pin |
Load |
prefetch |
Begin loading |
resolve |
Return a resident value, loading it when necessary. |
scope |
Make this pool transparently serve offloaded tensor reloads. |
unpin |
Return a manually pinned value to normal LRU management. |
Attributes:
| Name | Type | Description |
|---|---|---|
resident_bytes |
int
|
Return the approximate bytes held by resident values. |
acquire
acquire(source: OffloadedTensor, device: Optional[TDEVICE] = None, mode: str = 'read') -> TensorResidencyLease
Return a scoped lease for source.
Use mode="read_write" when the returned value may be modified.
The pool writes dirty values back before eviction or flushing.
evict
Write back and remove an unleased, unpinned value from the pool.
flush
Wait for loads and write dirty resident values back to storage.
is_pinned
Return whether source is manually pinned under LRU.
is_resident
Return whether source has completed loading into the pool.
pin
Load source and exclude it from LRU eviction until unpinned.
Pinning is idempotent and the returned handle can be used to wait for
materialization. A pin is cache policy rather than an active-use lease,
so closing the pool releases pinned values automatically. Manual pins
are available only with :attr:ResidencyStrategy.LRU.
prefetch
Begin loading source and return a completion handle.
resolve
Return a resident value, loading it when necessary.
This is the read-only path used automatically by OffloadedTensor
inside :meth:scope. Use :meth:acquire for active-use protection or
mutation tracking. Under LRU, use :meth:pin for manual retention.
scope
Make this pool transparently serve offloaded tensor reloads.
apply_name_to_tensor_in_module
Transform torch.Tensor or Parameters into NamedTensor.
This is applied at export time of torch_to_nnef
Just before doing any tracing and allow to keep
variable naming identical to PyTorch one
This consistent naming unlock subsequent manipulations such as LORA applications @ inference or such.