Sliding Window Attention(SWA)の理論と実装 — 固定窓による局所注意とMistralでの採用 2026年7月2日 Transformer Transformerベースの大規模言語モデル(LLM)を使っていると、ある壁に... GQAKVキャッシュLongformerMistralSliding Window Attention効率的Transformer機械学習
Mistral/Mixtralのアーキテクチャ — Sliding Window AttentionとMoEの融合 2026年6月7日 Transformer 7Bのパラメータで13Bクラスのモデルを上回る性能を出せるとしたら、どうでしょう... LLMMistralMixtralMoENLPSliding Window AttentionTransformer