NeFut Logo NeFut
中 Admin Login

[CS.AI] DanLing NestedTensor: Composable Multi‑Ragged Tensors for Deep Learning

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #Compiler

Variable‑size inputs are ubiquitous in deep learning, yet dense batching forces a shared envelope for all samples, incurring heavy padding overhead. The cost grows across each varying axis: an explicit pair state allocates $BN_{\max}^2$ positions while the true requirement is $\sum_i N_i^2$.

Packing eliminates this waste, but composing packed operations remains difficult because a flat buffer no longer exposes logical axes and sample boundaries. We introduce DanLing NestedTensor, a PyTorch tensor abstraction that treats multi‑ragged structure as an intrinsic property of the tensor.

Packed values carry tensor‑backed partition metadata and logical dimension order, so broadcasting creates ragged axes, feature transformations preserve them, and reductions consume them. The same representation flows seamlessly through autograd, eager execution, and compiled execution.

On an A100 GPU, experiments across four BERT scales show a geometric‑mean speedup of 2.74× in eager mode and 3.39× in compiled mode; across four FCN backbones the eager speedup is 1.97×. A four‑block Pairformer‑style workload runs 2.40‑4.32× faster than a padded reference using native PyTorch kernels in eager execution, with peak memory dropping from 38.08 GiB to 5.41 GiB.

The unified tensor interface lets model code built from supported operators compose efficient variable‑size computation without managing offsets at any call site. The code will be released publicly upon publication.

Review

Original Source: https://arxiv.org/abs/2609.30379

[h] Back to Home