Qwen3Apr 2025
20192026
Qwen3 normalized queries and keys before RoPE, so attention logits cannot grow without bound and the loss spikes that plagued large-batch training stop happening. It bought training stability at almost no inference cost, and it is now standard enough that a new family without QK-norm is the surprising one.
L0–3

