Files
llm-wiki/raw/articles/diffusionblocks-block-wise-training-2026.md
2026-07-03 00:38:05 +09:00

3.3 KiB

source_url, ingested, sha256, discovered_from
source_url ingested sha256 discovered_from
https://arxiv.org/abs/2506.14202 2026-07-01 ea9b11cb577a0ea0973d3f7824b2ff02985f91ad32fc7568dd15a4a667c3ab56
platform channel_name channel_id message_id author_id posted_at message_excerpt
discord tw 1477793137064935675 1521672204688036013 1477793167486226708 2026-07-01T00:22:04.900000000Z alphaXivのDiffusionBlocks再現実験は、論文の主張がどこまで成立するかを少ないプロンプトで検証していて、単なる紹介より一段深いです。

Title:DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation

Authors:, ,

View PDF HTML (experimental)

Abstract:End-to-end backpropagation requires storing activations throughout all layers, creating memory bottlenecks that limit model scalability. Existing block-wise training methods offer means to alleviate this problem, but they rely on ad-hoc local objectives and remain largely unexplored beyond classification tasks. We propose \\textit{DiffusionBlocks}, a principled framework for transforming transformer-based networks into genuinely independent trainable blocks that maintain competitive performance with end-to-end training. Our key insight leverages the fact that residual connections naturally correspond to updates in a dynamical system. With minimal modifications to this system, we can convert the updates to those of a denoising process, where each block can be learned independently by leveraging the score matching objective. This independence enables training with gradients for only one block at a time, thereby reducing memory requirements in proportion to the number of blocks. Our experiments on a range of transformer architectures (vision, diffusion, autoregressive, recurrent-depth, and masked diffusion) demonstrate that DiffusionBlocks training matches the performance of end-to-end training while enabling scalable block-wise training on practical tasks beyond small-scale classification. DiffusionBlocks provides a theoretically grounded approach that successfully scales to modern generative tasks across diverse architectures. Code is available at this https URL.

Comments:
Subjects:
Cite as:

Submission history

From: Makoto Shing [view email]
[v1] Tue, 17 Jun 2025 05:44:18 UTC (354 KB)
[v2] Fri, 3 Oct 2025 08:12:25 UTC (1,022 KB)
[v3] Wed, 18 Feb 2026 08:10:51 UTC (1,021 KB)
[v4] Fri, 12 Jun 2026 09:06:31 UTC (1,021 KB)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)