跳到正文
Apple Machine Learning Research·· 23 小时前AI 评分47

Apple 研究:扩散模型置信度排序的局限性

Limits of Confidence in Diffusion

AI 导读

Apple 研究团队从理论上证明,扩散模型(含 remasking 与 uniform-state 采样器)的每一步只有在所写入位置在已固定 token 条件下相互独立时,才与训练分布匹配,而任何逐位置分布的乘积都无法匹配存在依赖关系的 token 组。

正文

AuthorsRuss Webb, Amitis Shidani, Alice Bizeul, Dan Busbridge

Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest (pixels, phonemes, or words) there are inherent dependencies between tokens. We show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, that no product of per-position distributions can match a dependent group, and that per-position distributions do not determine whether a group is dependent: two joint distributions can have identical per-position marginals while differing in which combinations of values occur. On ScanAndAdd, a synthetic task whose joint distribution is available in closed form, we verify that every group of two or more undetermined positions a confidence ranking writes is dependent, and measure the generated distribution to be 29× the sampling-noise floor total variation while per-sample metrics are 1.0.

Related readings and updates.

On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher model should be used, and in the case of self-distillation, which specific context should serve as the supervisory signal? Does the optimal choice vary from one token to the next? At present, addressing these questions typically…

Read more

Transformers have gained increasing popularity in a wide range of applications, including Natural Language Processing (NLP), Computer Vision and Speech Recognition, because of their powerful representational capacity. However, harnessing this representational capacity effectively requires a large amount of data, strong regularization, or both, to mitigate overfitting. Recently, the power of the Transformer has been unlocked by self-supervised…

Read more

来源:Apple Machine Learning Research · machinelearning.apple.com