Skip to content

torch_to_nnef.tensor

Modules:

Name Description
named
offload

OffLoad Tensor.

opaque
quant

Advanced QTensor (<= 8bits) with complex quant scheme non torch native.

residency

Scoped residency and asynchronous prefetch for offloaded tensors.

updater
utils

Classes:

Name Description
NamedTensor

Tensor enriched with name attribute.

OffloadedTensor

Tensor subclass that maintains data on disk.

OpaqueTensorRef

Allow to pass through 'tracing'.

QScalePerGroupF16

f16 scale only per group.

QTensor

Common interface for all Compressed storage.

QTensorTractScaleOnly

Tract data format it serializes to: Q4_0.

ResidencyStrategy

Retention strategy for unleased values in a residency pool.

TensorPrefetch

Handle for an asynchronous tensor prefetch.

TensorResidencyLease

Scoped access to a value managed by a tensor residency pool.

TensorResidencyPool

Keep offloaded tensors resident across a bounded operation scope.

Functions:

Name Description
apply_name_to_tensor_in_module

Transform torch.Tensor or Parameters into NamedTensor.

set_opaque_tensor_in_params_as_ref

Transform OpaqueTensor Parameters into OpaqueTensorRef.

NamedTensor

NamedTensor(fp_tensor: Tensor, nnef_name: str)

Bases: Tensor

Tensor enriched with name attribute.

Attributes:

Name Type Description
data

Very important to keep access to all special attr of NamedTensor.

data property writable

data

Very important to keep access to all special attr of NamedTensor.

OffloadedTensor

OffloadedTensor(elem, device, offload_dir: Path, name: str, offloaded_tensor_type: Type[Tensor], force_gc_collect: bool = False, storage_id: Optional[str] = None)

Bases: OpaqueTensor

Tensor subclass that maintains data on disk.

It hold an virtual internal memory storage (permanent) and a temporary instantiation at each operation accessing it on targeted device.

Warning

we recommend to version of PyTorch > 1.12 for best compatibility.

Methods:

Name Description
from_original_tensor

Take a torch.Tensor or OpaqueTensor and offload it to disk.

reload

Reload the stored value on device.

set_

Implement tensor-style storage replacement for offloaded payloads.

to

Change the target device when reloaded in memory.

update_values

Replace offloaded tensor by new 'values' tensor.

Attributes:

Name Type Description
is_meta bool

Whether the tensor is on the meta device.

is_meta property

is_meta: bool

Whether the tensor is on the meta device.

Always False as the tensor is (off|re)loaded from disk.

from_original_tensor classmethod

from_original_tensor(tensor: Tensor, name: str, offload_dir: Optional[Path] = None, suffix_log_msg: str = '')

Take a torch.Tensor or OpaqueTensor and offload it to disk.

Parameters:

Name Type Description Default

tensor

Tensor

the torch.Tensor or torch_to_nnef.tensor.OpaqueTensor to dump on disk

required

name

str

the name of the tensor that will be used to create the filename store on disk

required

offload_dir

Optional[Path]

The directory where this file will be stored (temporarly)

None

suffix_log_msg

str

Added message log suffix for context

''

reload

reload(device: Optional[TDEVICE] = None)

Reload the stored value on device.

The optional override does not change the tensor's configured target device. This lets a residency manager stage the same payload on a worker-selected device without mutating shared tensor state.

set_

set_(source: Tensor, *args, **kwargs)

Implement tensor-style storage replacement for offloaded payloads.

OffloadedTensor uses a meta tensor as its in-memory shell, so PyTorch's native Tensor.set_ cannot replace its storage with a CPU tensor. Route the common param.set_(new_tensor) form through the offload store instead. This is important for quantizers that update a weight in-place before replacing it with a QTensor.

to

to(*args, **kwargs)

Change the target device when reloaded in memory.

update_values

update_values(values: Tensor, strict_shape: bool = True, strict_dtype: bool = True)

Replace offloaded tensor by new 'values' tensor.

Parameters:

Name Type Description Default

values

Tensor

The tensor that will replace it on disk assertion are made to ensure same shape, dtype as prior

required

strict_shape

bool

if True (default) the shape of the new tensor must be the same as the prior one

True

strict_dtype

bool

if True (default) the dtype of the new tensor must be the same as the prior one

True

OpaqueTensorRef

OpaqueTensorRef(meta_tensor: Tensor, opaque_tensor: OpaqueTensor)

Bases: Tensor

Allow to pass through 'tracing'.

QScalePerGroupF16

QScalePerGroupF16(group_size: int, scale: Tensor, n_bits: int)

Bases: QScheme

f16 scale only per group.

Tract aligned using negative scales.

QTensor

QTensor(fp_tensor: Tensor, qscheme: QScheme, dequant_to_dtype=torch.float32, u8_compressors: Optional[List[U8Compressor]] = None)

Bases: OpaqueTensor

Common interface for all Compressed storage.

Methods:

Name Description
to_device

Specific device handling.

write_in_file

Called at NNEF write time.

to_device

to_device(new_device)

Specific device handling.

write_in_file

write_in_file(dirpath: Union[str, Path], label: str)

Called at NNEF write time.

Each specific inference engine format should implement the file dump prefered.

QTensorTractScaleOnly

QTensorTractScaleOnly(*args, specific_machine: Optional[str] = None, **kwargs)

Bases: QTensorTract, SupportsOffloadState

Tract data format it serializes to: Q4_0.

Methods:

Name Description
decompress

Tract dequantization depends on hardware.

decompress

decompress()

Tract dequantization depends on hardware.

Typically dequantization happen with ops in f16 on ARM and f32 (scale directly casted) on others so we overwrite the function to be consistant with tract.

ResidencyStrategy

Bases: Enum

Retention strategy for unleased values in a residency pool.

LRU replaces the least recently used eligible value and permits manual pins. FIXED preserves its first-fit admitted set for the pool lifetime and rejects manual pin operations.

TensorPrefetch

TensorPrefetch(pool: TensorResidencyPool, source: OffloadedTensor, device: device, future: Optional[Future], value: Optional[Tensor] = None)

Handle for an asynchronous tensor prefetch.

Methods:

Name Description
done

Return whether the prefetch has finished.

wait

Wait for prefetch completion and return the resident value.

done

done() -> bool

Return whether the prefetch has finished.

wait

wait() -> torch.Tensor

Wait for prefetch completion and return the resident value.

TensorResidencyLease

TensorResidencyLease(pool: TensorResidencyPool, source: OffloadedTensor, device: device, mode: str)

Scoped access to a value managed by a tensor residency pool.

TensorResidencyPool

TensorResidencyPool(max_cached_bytes: Optional[int] = None, max_workers: int = 1, strategy: ResidencyStrategy = ResidencyStrategy.LRU)

Keep offloaded tensors resident across a bounded operation scope.

Values can be prefetched by background workers and acquired through a lease. A lease is a scoped claim that a materialized value is in active use, so its value remains resident until the lease ends. The default LRU strategy evicts other values in least-recently-used order when the optional cache budget is exceeded and supports manual pins. The FIXED strategy retains the first completed values that fit and streams later values without displacing that resident set. A read_write lease writes the value back before eviction.

max_cached_bytes is a soft cache-retention limit, not a hard bound on process memory. Leased values, and values manually pinned under LRU, remain available even when they exceed it. An oversized value can therefore be materialized for an active caller but is evicted as soon as its final lease ends. Prefetched oversized values are returned to their waiting caller without being retained by the pool unless they are pinned.

Methods:

Name Description
acquire

Return a scoped lease for source.

close

Flush all values, stop workers, and release resident memory.

evict

Write back and remove an unleased, unpinned value from the pool.

flush

Wait for loads and write dirty resident values back to storage.

is_pinned

Return whether source is manually pinned under LRU.

is_resident

Return whether source has completed loading into the pool.

pin

Load source and exclude it from LRU eviction until unpinned.

prefetch

Begin loading source and return a completion handle.

resolve

Return a resident value, loading it when necessary.

scope

Make this pool transparently serve offloaded tensor reloads.

unpin

Return a manually pinned value to normal LRU management.

Attributes:

Name Type Description
resident_bytes int

Return the approximate bytes held by resident values.

resident_bytes property

resident_bytes: int

Return the approximate bytes held by resident values.

acquire

acquire(source: OffloadedTensor, device: Optional[TDEVICE] = None, mode: str = 'read') -> TensorResidencyLease

Return a scoped lease for source.

Use mode="read_write" when the returned value may be modified. The pool writes dirty values back before eviction or flushing.

close

close() -> None

Flush all values, stop workers, and release resident memory.

evict

evict(source: OffloadedTensor) -> None

Write back and remove an unleased, unpinned value from the pool.

flush

flush(source: Optional[OffloadedTensor] = None) -> None

Wait for loads and write dirty resident values back to storage.

is_pinned

is_pinned(source: OffloadedTensor) -> bool

Return whether source is manually pinned under LRU.

is_resident

is_resident(source: OffloadedTensor) -> bool

Return whether source has completed loading into the pool.

pin

pin(source: OffloadedTensor, device: Optional[TDEVICE] = None) -> TensorPrefetch

Load source and exclude it from LRU eviction until unpinned.

Pinning is idempotent and the returned handle can be used to wait for materialization. A pin is cache policy rather than an active-use lease, so closing the pool releases pinned values automatically. Manual pins are available only with :attr:ResidencyStrategy.LRU.

prefetch

prefetch(source: OffloadedTensor, device: Optional[TDEVICE] = None) -> TensorPrefetch

Begin loading source and return a completion handle.

resolve

resolve(source: OffloadedTensor, device: Optional[TDEVICE] = None) -> torch.Tensor

Return a resident value, loading it when necessary.

This is the read-only path used automatically by OffloadedTensor inside :meth:scope. Use :meth:acquire for active-use protection or mutation tracking. Under LRU, use :meth:pin for manual retention.

scope

scope(device: Optional[TDEVICE] = None)

Make this pool transparently serve offloaded tensor reloads.

unpin

unpin(source: OffloadedTensor) -> None

Return a manually pinned value to normal LRU management.

apply_name_to_tensor_in_module

apply_name_to_tensor_in_module(model: Module)

Transform torch.Tensor or Parameters into NamedTensor.

This is applied at export time of torch_to_nnef Just before doing any tracing and allow to keep variable naming identical to PyTorch one

This consistent naming unlock subsequent manipulations such as LORA applications @ inference or such.

set_opaque_tensor_in_params_as_ref

set_opaque_tensor_in_params_as_ref(model: Module)

Transform OpaqueTensor Parameters into OpaqueTensorRef.

This is applied at export time of torch_to_nnef Just before doing any tracing