Generating Long Sequences with Sparse Transformers

TM
Trevor McFedries
@trevvyboi

Transformers are powerful sequence models, but require time and memory that grows quadratically with the sequence length. In this paper we introduce sparse factorizations of the attention matrix which reduce this to $O(n \sqrt{n})$. We also introduce a) a v...

Uploaded
Uploaded Jul 10, 2026
Queried
Queried 0 times

No preview text is available for this document yet.

Want to learn more?

Ask a question