Skip to content
LiveNext 12:04:12
CoolingResearch

Paper: byte-level language models beat subword models at scale in compute-optimal fits

An Oct. 5 arXiv preprint by Jie Wang, Shiwei Luo, Qi Zhang and Yuanbin Wu trains tokenizer-free byte models. Under compute-optimal fitting they reach lower loss than subword models at equal parameter count, though subword models win at equal training compute. On CUTE, byte models score 99.19 to 99.93 versus 69.89 to 80.83 for subword.

HEAT
OUTLETS
0
FIRST SEEN
LAST SIGNAL