CoolingResearch
Paper: byte-level language models beat subword models at scale in compute-optimal fits
An Oct. 5 arXiv preprint by Jie Wang, Shiwei Luo, Qi Zhang and Yuanbin Wu trains tokenizer-free byte models. Under compute-optimal fitting they reach lower loss than subword models at equal parameter count, though subword models win at equal training compute. On CUTE, byte models score 99.19 to 99.93 versus 69.89 to 80.83 for subword.
- HEAT
- OUTLETS
- 0
- FIRST SEEN
- LAST SIGNAL
Primary source
arXivarxiv.org/html/2610.05978v1Coverage · 1 article, oldest first
Saturday, October 10, 2026