Skip to content

torch_to_nnef.tensor.residency

Scoped residency and asynchronous prefetch for offloaded tensors.

Classes:

Name Description
ResidencyStrategy

Retention strategy for unleased values in a residency pool.

TensorPrefetch

Handle for an asynchronous tensor prefetch.

TensorResidencyLease

Scoped access to a value managed by a tensor residency pool.

TensorResidencyPool

Keep offloaded tensors resident across a bounded operation scope.

Functions:

Name Description
active_tensor_residency

Return the pool and default device active in this context.

ResidencyStrategy

Bases: Enum

Retention strategy for unleased values in a residency pool.

LRU replaces the least recently used eligible value and permits manual pins. FIXED preserves its first-fit admitted set for the pool lifetime and rejects manual pin operations.

TensorPrefetch

TensorPrefetch(pool: TensorResidencyPool, source: OffloadedTensor, device: device, future: Optional[Future], value: Optional[Tensor] = None)

Handle for an asynchronous tensor prefetch.

Methods:

Name Description
done

Return whether the prefetch has finished.

wait

Wait for prefetch completion and return the resident value.

done

done() -> bool

Return whether the prefetch has finished.

wait

wait() -> torch.Tensor

Wait for prefetch completion and return the resident value.

TensorResidencyLease

TensorResidencyLease(pool: TensorResidencyPool, source: OffloadedTensor, device: device, mode: str)

Scoped access to a value managed by a tensor residency pool.

TensorResidencyPool

TensorResidencyPool(max_cached_bytes: Optional[int] = None, max_workers: int = 1, strategy: ResidencyStrategy = ResidencyStrategy.LRU)

Keep offloaded tensors resident across a bounded operation scope.

Values can be prefetched by background workers and acquired through a lease. A lease is a scoped claim that a materialized value is in active use, so its value remains resident until the lease ends. The default LRU strategy evicts other values in least-recently-used order when the optional cache budget is exceeded and supports manual pins. The FIXED strategy retains the first completed values that fit and streams later values without displacing that resident set. A read_write lease writes the value back before eviction.

max_cached_bytes is a soft cache-retention limit, not a hard bound on process memory. Leased values, and values manually pinned under LRU, remain available even when they exceed it. An oversized value can therefore be materialized for an active caller but is evicted as soon as its final lease ends. Prefetched oversized values are returned to their waiting caller without being retained by the pool unless they are pinned.

Methods:

Name Description
acquire

Return a scoped lease for source.

close

Flush all values, stop workers, and release resident memory.

evict

Write back and remove an unleased, unpinned value from the pool.

flush

Wait for loads and write dirty resident values back to storage.

is_pinned

Return whether source is manually pinned under LRU.

is_resident

Return whether source has completed loading into the pool.

pin

Load source and exclude it from LRU eviction until unpinned.

prefetch

Begin loading source and return a completion handle.

resolve

Return a resident value, loading it when necessary.

scope

Make this pool transparently serve offloaded tensor reloads.

unpin

Return a manually pinned value to normal LRU management.

Attributes:

Name Type Description
resident_bytes int

Return the approximate bytes held by resident values.

resident_bytes property

resident_bytes: int

Return the approximate bytes held by resident values.

acquire

acquire(source: OffloadedTensor, device: Optional[TDEVICE] = None, mode: str = 'read') -> TensorResidencyLease

Return a scoped lease for source.

Use mode="read_write" when the returned value may be modified. The pool writes dirty values back before eviction or flushing.

close

close() -> None

Flush all values, stop workers, and release resident memory.

evict

evict(source: OffloadedTensor) -> None

Write back and remove an unleased, unpinned value from the pool.

flush

flush(source: Optional[OffloadedTensor] = None) -> None

Wait for loads and write dirty resident values back to storage.

is_pinned

is_pinned(source: OffloadedTensor) -> bool

Return whether source is manually pinned under LRU.

is_resident

is_resident(source: OffloadedTensor) -> bool

Return whether source has completed loading into the pool.

pin

pin(source: OffloadedTensor, device: Optional[TDEVICE] = None) -> TensorPrefetch

Load source and exclude it from LRU eviction until unpinned.

Pinning is idempotent and the returned handle can be used to wait for materialization. A pin is cache policy rather than an active-use lease, so closing the pool releases pinned values automatically. Manual pins are available only with :attr:ResidencyStrategy.LRU.

prefetch

prefetch(source: OffloadedTensor, device: Optional[TDEVICE] = None) -> TensorPrefetch

Begin loading source and return a completion handle.

resolve

resolve(source: OffloadedTensor, device: Optional[TDEVICE] = None) -> torch.Tensor

Return a resident value, loading it when necessary.

This is the read-only path used automatically by OffloadedTensor inside :meth:scope. Use :meth:acquire for active-use protection or mutation tracking. Under LRU, use :meth:pin for manual retention.

scope

scope(device: Optional[TDEVICE] = None)

Make this pool transparently serve offloaded tensor reloads.

unpin

unpin(source: OffloadedTensor) -> None

Return a manually pinned value to normal LRU management.

active_tensor_residency

active_tensor_residency() -> T.Optional[T.Tuple[TensorResidencyPool, T.Optional[torch.device]]]

Return the pool and default device active in this context.