Sliding Window Attention(SWA)の理論と実装 — 固定窓による局所注意とMistralでの採用 2026年7月2日 Transformer Transformerベースの大規模言語モデル(LLM)を使っていると、ある壁に... GQAKVキャッシュLongformerMistralSliding Window Attention効率的Transformer機械学習
Sparse Attention(Longformer・BigBird)の理論と実装 — 長系列を効率的に処理する 2026年7月2日 Transformer 論文1本(数万トークン)を丸ごとTransformerに入力して要約したい。数十... BigBirdLongformerSparse Attention効率的Transformer機械学習長系列
Linear Attention / Performerの理論と実装 — カーネル近似で計算量をO(n)にする 2026年6月27日 Transformer Transformerは自然言語処理や画像認識で驚異的な性能を発揮していますが、... FAVOR+Linear AttentionPerformerTransformerカーネル法効率的Transformer機械学習