Sliding Window Attention
See the related terms and live BenchGecko data for current examples.
Text reviewed October 5, 2026
See the related terms and live BenchGecko data for current examples.
Basic
In standard transformer attention, every token attends to every previous token · O(N²) memory and compute. Sliding window attention limits each token to the previous K tokens (window size). Enables long-context handling without quadratic blowup. See the related terms and live BenchGecko data for current examples.
Deep
Sliding window works because most useful context is local. Information that needs to travel far uses the "receptive field" expansion through stacked layers · a 32-layer model with 4K window has effective context of 128K through depth. Compute cost: sliding window attention is ~K/N cheaper than full attention when N >> K. See the related terms and live BenchGecko data for current examples.
Expert
Sliding window attention integrates with FlashAttention for memory-efficient implementation. The window size K is a hyperparameter: too small loses context, too large loses efficiency. Hybrid architectures (e.g., Gemma 2: sliding + global layers interleaved) balance quality vs speed. Mistral's effective context via sliding was questionable in practice · retrieval from far positions degraded. Modern approaches pair sliding with explicit long-context mechanisms (KV-cache compression, retrieval) to preserve quality.
Depending on why you're here
- ·A way to make AI only look at recent words for speed
- ·Lets AI handle long documents faster
- ·Used in some of Google's and Mistral's models
- ·Efficient long-context models often use sliding window
- ·Quality degrades at far positions · pair with retrieval
- ·Check model card for window size + hybrid layer config
- ·Sliding window enables cheaper long-context serving
- ·Compute savings of 3-10× on long contexts
- ·Quality trade-off mostly mitigated in hybrid designs
- ·Limits attention to K nearest tokens · O(N·K) vs O(N²)
- ·Hybrid local+global architectures restore quality
Sliding window attention is the quiet architecture that makes long-context serving economically possible.