Sliding Window Attention(SWA)の理論と実装 — 固定窓による局所注意とMistralでの採用 2026年7月2日 Transformer Transformerベースの大規模言語モデル(LLM)を使っていると、ある壁に... GQAKVキャッシュLongformerMistralSliding Window Attention効率的Transformer機械学習
Sparse Attention(Longformer・BigBird)の理論と実装 — 長系列を効率的に処理する 2026年7月2日 Transformer 論文1本(数万トークン)を丸ごとTransformerに入力して要約したい。数十... BigBirdLongformerSparse Attention効率的Transformer機械学習長系列