Nunchux AI 推出 VC-Attention:免训练低比特注意力内核,加速视频 Diffusion Transformer

MarkTechPost(RSS)·2026-09-17 08:45·26分钟前·Asif Razzaq
AI 导读

Nunchux AI 发布免训练低比特注意力内核 VC-Attention,针对视频 Diffusion Transformer 的值量化误差与 softmax 瓶颈,由 V-Smooth 和 ExpCast-FP8 两部分组成。

MarkTechPost(RSS)
44AI 编辑部评分,满分 100

Nunchux AI 推出 VC-Attention:免训练低比特注意力内核,加速视频 Diffusion Transformer

2026-09-17 08:45· 26分钟前· Asif Razzaq
AI 导读

Nunchux AI 发布免训练低比特注意力内核 VC-Attention,针对视频 Diffusion Transformer 的值量化误差与 softmax 瓶颈,由 V-Smooth 和 ExpCast-FP8 两部分组成。

Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (DiTs). It targets 2 problems at once: value quantization error and a slow softmax stage.

Why Attention is the Video Bottleneck

Video DiTs flatten a clip into 1 sequence of spatiotemporal tokens and run full self-attention at every layer. A 5-second 720p Wan2.2-14B clip spans about 70K tokens. On the RTX 5090, attention takes more than 64% of generation time. The research team states that attention is about two thirds of every MiniMax-H3 denoising step on a single B200.

Low-bit Tensor Cores speed up the 2 matrix products, QK and PV. 2 obstacles remain. First, prior methods like SageAttention2 smooth queries and keys. After QK smoothing and rotation, the value term accounts for 82% of output error on Wan2.2. Second, the softmax between the products still runs in FP32. On B200 and H200, that exponential and its FP8 cast become the longest pipeline stage.

V-Smooth: Fixing Value Outliers

Value outliers sit in a few tokens, and their channels shift across heads, layers, and steps. A Hadamard rotation preserves token norms, so it does not remove them. Rotating V changes value error by just 0.2%.

V-Smooth takes a different route:

  • Group: An online k-means clusters value tokens per batch and head. Keys and values are permuted together, so non-causal attention output is unchanged.
  • Demean: Each 128-token hardware block subtracts its mean. Only the residual is quantized, using per-channel E4M3 at 8 bits or NVFP4 at 4 bits.
  • Restore: The mean is added back using the row sum online softmax already keeps. No second pass or extra buffer is needed.

Averaged over 100 Wan2.2 heads, the block mean removes 8% of block energy in sequence order. It removes 12% under DeltaQuant’s static cube and 36% after sorting. Each mean costs 0.125 bit per value element.

Grouping runs only on the first 25% of denoising steps. The permutation is reused across 4 adjacent steps. Averaged over the full schedule, grouping costs 3 to 4% of attention time.

ExpCast-FP8: Removing the Softmax Bottleneck

An E4M3 byte is already close to a logarithm of the value it stores. Read as an integer, it equals roughly 8 log2(v) + 56. So ExpCast-FP8 writes the byte directly from the log-domain score with 1 fused multiply-add. The constant β = -0.35 centers the leftover error, and no constant is fitted per model.

The direct path writes the same byte as the FP32 exponent-then-cast path on 79.6% of each doubling. Elsewhere it lands 1 code away. The paper proves a per-row total variation bound under 3.64%, plus any underflow tail. Across 204.8K Wan2.2 attention rows, the measured average is 1.6%. ExpCast-FP8 applies only to the 8-bit kernel, since NVFP4 has no single affine log-to-code map.

Hand-written CuTe/CUDA fusion of the preprocessing chain cuts 1 V-Smooth call from 42.2 ms to 4.8 ms on B200.

Explainer: How VC-Attention Works

Benchmarks

Tests cover 4 open-weight video DiTs: Wan2.2-T2V-A14B, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3. Fidelity is scored against BF16 FlashAttention-4 outputs over 100 prompts.

GPU (Wan2.2)PrecisionAttention speedupEnd-to-end speedup
B2008-bit1.59×1.19×
H2008-bit1.46×1.13×
RTX PRO 60004-bit2.27×1.36×
RTX 50904-bit3.58×1.70×

On B200, VC-Attention is 6.02× faster than SageAttention2, which ships no Blackwell kernel. On H200, the gap is 1.16×. On workstation cards, 4-bit V-Smooth matches SageAttention3 on the RTX PRO 6000. It stays within 5% on the RTX 5090, so fidelity separates them.

Fidelity results:

  • At 8 bits, V-Smooth adds 2.3 dB PSNR over SageAttention2 on Wan2.2 and 2.8 dB on HunyuanVideo-1.5.
  • Adding ExpCast-FP8 gives back 0.7 to 2.1 dB but still beats SageAttention2 on all 4 models.
  • At 4 bits, V-Smooth beats SageAttention3 by 2.9 dB on Wan2.2 and 3.6 dB on LongCat-Video.
  • Run training-free, Attn-QAT falls 3.4 to 6.7 dB below SageAttention2.

On MiniMax-H3 at 1344×768, attention runs 1.60× faster than BF16 FlashAttention-4 on B200. PSNR is 20.2 dB versus 19.9 dB for SageAttention2. On B300, the paper reports 1.47× versus 1.31× for a naive FP8 kernel. The blog chart lists 1.51× for B300.

Nunchux Attention, the company's proprietary extension, reaches 1.91× on B200 and 1.83× on B300 for MiniMax-H3 attention.

The method changes only per-interaction cost. So it can compose with sparse attention like Sparse VideoGen and Radial Attention, and distillation and multi-GPU execution. Nunchux says free MiniMax-H3 access is coming through its Modelverse waitlist.

Key Takeaways

  • VC-Attention is training-free low-bit attention for video DiTs from Nunchux AI.
  • V-Smooth clusters value tokens, then quantizes only residuals after block-mean subtraction.
  • ExpCast-FP8 replaces the FP32 exponential and cast with 1 multiply-add.
  • Wan2.2 attention runs 1.59× faster on B200 and 3.58× on RTX 5090.
  • No public kernel release yet; Nunchux runs a proprietary extension in its stack.

来源:MarkTechPost(RSS)· marktechpost.com