Byte-level state-space models. That sounded pretty scary for a scientist decades ago. Now we have:
1. Knowledge that deeper layers train smoothly. 2. Knowledge that Transformers work but is quadratic on sequence length. 3. Knowledge that SSMs work even better. Numerically unstable sometimes. 4. Speculative-decoding. 5. Open high-quality data. 6. Knowledge that KD works.