Blog

Yesterday — 11 September 2026 — one of my long generations came back at roughly 250 tokens per second. Not a spec sheet and not a vendor benchmark: one run, on my own account, with the rate visible in the client. It was fast enough that I stopped and checked the number twice.

Here is what a rate...

A great prompt isn't about magic words — it's about making the outcome unambiguous. Most prompts that fail don't fail because the model is stupid; they fail because the request was vague about what done means, what's off-limits, and what the result is actually for.

This is the template I use whe...

Most people who use an LLM have a rough sense that "the model reads your text in blocks," but few can say what actually happens between your prompt and the answer. Here is the plain-language version, in numbers you can follow.

Say we have ten words. The tokenizer turns them into twenty tokens:

[...