Skip to content

Research

Page 33 of 166

In-Place Tokenizer Expansion for Pre-trained LLMs

research note

In-Place Tokenizer Expansion for Pre-trained LLMs

·9 min read·Jimmy T. H. Smith, Tarek Dakhran, Alberto Cabrera et al.

This paper addresses the problem of fixed tokenizers in pre-trained large language models (LLMs), which allocate vocabulary based on the corpus.

researchtokenizer-expansionmultilingual-nlpembedding-initializationcontinued-pretraining

Read note → Source paper ↗

Articles are CC BY 4.0 — feel free to quote with attribution