39 lines
3.3 KiB
Markdown
39 lines
3.3 KiB
Markdown
---
|
|
source_url: "https://arxiv.org/abs/2506.14202"
|
|
ingested: 2026-07-01
|
|
sha256: ea9b11cb577a0ea0973d3f7824b2ff02985f91ad32fc7568dd15a4a667c3ab56
|
|
discovered_from:
|
|
platform: discord
|
|
channel_name: tw
|
|
channel_id: "1477793137064935675"
|
|
message_id: "1521672204688036013"
|
|
author_id: "1477793167486226708"
|
|
posted_at: "2026-07-01T00:22:04.900000000Z"
|
|
message_excerpt: "alphaXivのDiffusionBlocks再現実験は、論文の主張がどこまで成立するかを少ないプロンプトで検証していて、単なる紹介より一段深いです。"
|
|
---
|
|
|
|
## Title:DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
|
|
|
|
Authors:, ,
|
|
|
|
[View PDF](https://arxiv.org/pdf/2506.14202) [HTML (experimental)](https://arxiv.org/html/2506.14202v4)
|
|
|
|
> Abstract:End-to-end backpropagation requires storing activations throughout all layers, creating memory bottlenecks that limit model scalability. Existing block-wise training methods offer means to alleviate this problem, but they rely on ad-hoc local objectives and remain largely unexplored beyond classification tasks. We propose $\\textit{DiffusionBlocks}$, a principled framework for transforming transformer-based networks into genuinely independent trainable blocks that maintain competitive performance with end-to-end training. Our key insight leverages the fact that residual connections naturally correspond to updates in a dynamical system. With minimal modifications to this system, we can convert the updates to those of a denoising process, where each block can be learned independently by leveraging the score matching objective. This independence enables training with gradients for only one block at a time, thereby reducing memory requirements in proportion to the number of blocks. Our experiments on a range of transformer architectures (vision, diffusion, autoregressive, recurrent-depth, and masked diffusion) demonstrate that DiffusionBlocks training matches the performance of end-to-end training while enabling scalable block-wise training on practical tasks beyond small-scale classification. DiffusionBlocks provides a theoretically grounded approach that successfully scales to modern generative tasks across diverse architectures. Code is available at [this https URL](https://github.com/SakanaAI/DiffusionBlocks).
|
|
|
|
| Comments: |
|
|
| --- |
|
|
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML) |
|
|
| Cite as: | [arXiv:2506.14202](https://arxiv.org/abs/2506.14202) \[cs.LG\] |
|
|
| | (or [arXiv:2506.14202v4](https://arxiv.org/abs/2506.14202v4) \[cs.LG\] for this version) |
|
|
| | [https://doi.org/10.48550/arXiv.2506.14202](https://doi.org/10.48550/arXiv.2506.14202) |
|
|
|
|
## Submission history
|
|
|
|
From: Makoto Shing \[[view email](https://arxiv.org/show-email/6d714b0e/2506.14202)\]
|
|
**[\[v1\]](https://arxiv.org/abs/2506.14202v1)** Tue, 17 Jun 2025 05:44:18 UTC (354 KB)
|
|
**[\[v2\]](https://arxiv.org/abs/2506.14202v2)** Fri, 3 Oct 2025 08:12:25 UTC (1,022 KB)
|
|
**[\[v3\]](https://arxiv.org/abs/2506.14202v3)** Wed, 18 Feb 2026 08:10:51 UTC (1,021 KB)
|
|
**\[v4\]** Fri, 12 Jun 2026 09:06:31 UTC (1,021 KB)
|
|
|
|
[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2506.14202) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
|