You're mid-way through something useful and the AI tells you you've hit your limit. You've asked maybe fifteen questions. Someone else asks a hundred a day and never sees this message. It feels arbitrary.
It is measured in a unit that is rarely explained. You're not billed by the question. You're billed by the amount of text involved. And the text involved is usually far more than what you typed.
A token is a chunk of text
Models don't read letters and they don't quite read words. They read tokens.
"Shopify" might be one token. "Unsubscribing" might be three. A page of A4 is around 500 tokens. A long email thread is a couple of thousand. A forty-page PDF is twenty thousand or so.
That's the whole concept: a unit of text, slightly smaller than a word, and everything is priced and limited in them.
The model re-reads everything
This is the mechanic that explains almost every "why did I run out so fast" moment.
A model has no memory between messages. To answer your fifth question in a conversation, it is given the entire conversation so far, including your first four questions, its four answers and the document you pasted at the start. It reads the whole thing again before responding.
So a chat doesn't cost a steady amount per message. Every message costs more than the one before it, because the conversation keeps growing and gets re-read each time. A twenty-message conversation about a long document can easily use more than a hundred short, separate questions.
Which is why attaching a big document is expensive repeatedly, why the last few exchanges of a marathon session cost several times what the first ones did, and why starting a fresh chat is a real optimisation rather than tidiness.
What actually burns tokens
Roughly in order of how often it catches people out:
| What burns them | Why |
|---|---|
| Long conversations | The re-reading above. This is the big one |
| Large attachments | A 40-page PDF is ~20,000 tokens every turn it stays attached |
| Connected tools returning a lot | A connector that returns 300 rows just spent 300 rows of budget |
| Reasoning | Models that "think" first produce text you never see, and pay for |
| Long outputs | Real, but the smallest of the five |
Four of the five have nothing to do with how much you typed.
The habits that make the limit stop mattering
Expensive by default
- One long conversation carried across topics
- The whole contract attached
- "Show me the orders"
- The biggest model for every task
Cheap by habit
- A new chat when the topic changes
- The one clause you need, pasted in
- "The last 20 orders"
- The model the job actually calls for
Narrow requests are cheaper and produce better answers, because the model isn't hunting through noise. And reformatting a list does not need your most capable model; defaulting to the biggest one means paying for deliberation the task never required.
One more to watch: if an AI is working through a long task on its own, it spends tokens on every step. Long autonomous jobs are where limits vanish quickly and quietly.
The reframe worth keeping
Once you know the unit, "I ran out" becomes a straightforward question: what was I making it read?
Nine times out of ten the answer is a conversation that got long, or a document that stayed attached after it was needed. Your limit is usually not too small; your context is too big.