Blog

Most people who use an LLM have a rough sense that "the model reads your text in blocks," but few can say what actually happens between your prompt and the answer. Here is the plain-language version, in numbers you can follow.

Say we have ten words. The tokenizer turns them into twenty tokens:

[...