Transformer Decoderの構造とMasked Self-Attentionの仕組み 2026年7月4日 Transformer ChatGPTに文章を書かせるとき、モデルは「文章全体」をいきなり出力しているわ... Cross-AttentionDecoderMasked AttentionTransformer深層学習