The largest AI token saving is a shorter conversation, not a shorter prompt. Every turn resends the whole thread and every attached file, so the habits that matter are about context: one task per chat, select the method instead of attaching the file, name the symbol you mean, and ask for a diff instead of a rewritten file. On one real fix in a Flutter app, that took a single question from about 212,000 tokens to about 20,000.
A tester files a short bug against the MSDevBuild Eats food delivery app: “Service fee jumps by RM 2 after placing the order.” Checkout said RM 1.50. The order tracking screen says RM 3.50. The total is the same on both.
In the previous part the month was divided into a daily budget and every task was routed to a tier before typing. Routing is the big lever. This article is the small ones, and there are more of them than expected. They are 15 habits, and almost none of them are features you can switch on. They follow from how these tools work underneath.
None of them are about using AI less. That is the wrong trade. They are about paying for context once instead of twelve times.
The thread that cost 212,000 tokens
The developer who picks up the bug does what most of us do on a Monday. They open Copilot Chat and attach what looks related: checkout_screen.dart, order_tracking_screen.dart and the Order entity. Then they ask, “Why is the service fee wrong after ordering?”
- Turn 1. The answer blames rounding in the currency formatter.
- Turns 2 to 5. The developer explains that the difference is exactly RM 2, every time. The answer moves on to the promo discount.
- Turns 6 to 11. Back and forth about the cart BLoC, which is not attached either.
- Turn 12. The developer finds the cause by hand and asks Copilot to “rewrite the four files with the fix”. It returns four complete files.
The thread worked in the end. It also cost about 212,000 tokens, more than half of a 400,000-token month, for a 15-line change.
Why did it cost that much?
The answer was never in the context. The merge happens in a fourth file nobody attached. One plain search finds it:
grep -rn "smallOrderFee" lib
lib/features/checkout/presentation/checkout_screen.dart:447: if (cart.smallOrderFee > 0)
lib/features/checkout/presentation/checkout_screen.dart:450: value: Formatters.currency(cart.smallOrderFee),
lib/features/cart/presentation/cart_screen.dart:164: if (cart.smallOrderFee > 0)
lib/features/cart/presentation/cart_screen.dart:167: value: Formatters.currency(cart.smallOrderFee),
lib/data/repositories/order_repository_impl.dart:43: serviceFee: cart.serviceFee + cart.smallOrderFee,
lib/domain/entities/cart.dart:46: static const double smallOrderFee = 2;
lib/domain/entities/cart.dart:95: double get smallOrderFee =>
lib/domain/entities/cart.dart:96: isBelowMinimum && isNotEmpty ? PricingPolicy.smallOrderFee : 0;
lib/domain/entities/cart.dart:102: subtotal + deliveryFeeMyr + serviceFee + smallOrderFee - promoDiscount;
Checkout shows the service fee and the small-order fee as two rows. createOrder in order_repository_impl.dart folds them into one serviceFee on line 43, and the tracking screen shows that single field. Nothing is overcharged. The label just changes meaning after the order is placed.
The cost came from three things stacked together. Measured with the GPT-4o tokenizer (o200k_base) on the real files:
- The three attached files are 7,591 tokens. With about 4,000 tokens of system prompt and tools and the 1,096-token
AGENTS.md, each turn started at about 12,700 tokens before the question. - Twelve turns resent that, plus a history that grew by roughly 700 tokens a turn.
- The last turn returned four whole files: 7,279 output tokens for 15 changed lines.
def thread_cost(fixed, turns, history_per_turn=700, out_per_turn=600):
return sum(fixed + history_per_turn * (k - 1) for k in range(1, turns + 1)) + out_per_turn * turns
thread_cost(12_687, 12) # ≈ 205,000, plus 6,700 extra output for the full files ≈ 212,000
Why does a long AI chat cost so much?
Seven of the fifteen habits follow from this one fact, so it is worth seeing in numbers.
A chat with an AI model is stateless underneath. The model does not remember turn 4 when you send turn 5. The client resends the whole conversation every time: your first question, its first answer, everything since, and every file you attached along the way.
So an N-turn conversation does not cost N times the first turn. It grows with the square of N. Here it is with two 300-line Dart files attached (about 3,900 tokens at the measured 6.5 tokens per line), a 100-token question and a 900-token answer per turn.
| Turn | What gets sent | Tokens that turn | Running total |
|---|---|---|---|
| 1 | Files + question, plus the answer | 4,900 | 4,900 |
| 3 | Files + 3 questions + 3 answers | 6,900 | 17,700 |
| 6 | Files + 6 questions + 6 answers | 9,900 | 44,400 |
| 10 | Files + 10 questions + 10 answers | 13,900 | 94,000 |
Ten turns cost about 94,000 tokens. The same ten questions in three short threads cost about 61,000. Batched into two turns of five questions each, about 23,000. The questions did not change. Only the container did.

1. One task, one chat
Close the thread when the task is done. Not when you log off, not when the topic drifts: when the task is done. A new thread is free. Continuing an old one is not.
2. When an answer is wrong, restart instead of arguing
This habit saves the most and feels the least natural.
When turn 3 comes back wrong, the reflex is to explain the mistake in turn 4. That turn buys the entire thread again, including the wrong answer, which now sits in context and steers everything after it. The fee thread spent nine turns this way.
Instead: read why it went wrong, close the chat, and rewrite the first prompt with the missing fact in it. “The difference is always exactly RM 2” would have pointed straight at a fixed fee. It feels like conceding. It costs a fraction as much, and the answer is cleaner because the model is not reasoning around its own earlier mistake.
3. Batch related questions into one turn
Five small questions about the same file, asked as five turns, pay for that file five times. Asked as one numbered list in one turn, once. Keep a scratch note during a task and ask in groups.
Context discipline: pay for what you need
4. Select, do not attach
The biggest per-turn saving. checkout_screen.dart is 499 lines and 3,014 tokens. The createOrder method that held the answer is 38 lines and 269 tokens, and the two fee getters in Cart add 112. Selecting them costs 381 tokens instead of 7,591, and it contains the line the first thread never saw.
There is a quality argument too. A model given 500 lines when 40 are relevant has 460 lines of distraction, and it will sometimes answer about the wrong method.
5. Name the symbol, not the area
“Why is the service fee wrong?” makes the model guess where to look, and in agent mode it makes the tool search and read half the repository. This question is 63 tokens and names everything it needs:
Checkout shows 'Service fee' and 'Small order fee' as two rows. After placing
the order, OrderTrackingScreen shows one 'Service fee' that is RM 2 higher.
Where in createOrder do they merge, and what is the smallest change that keeps
them separate? Reply with a unified diff only.
If you do not know the symbol, find it with your editor’s search or grep in two seconds, then ask. Repository-wide context is the most expensive thing you can request, and the question rarely needed it.
6. Paste the failure, not the log
Stack traces, build output and flutter test runs are long and mostly repetition. A failing test run can be hundreds of lines of which fifteen matter.
Paste the exception type, the message, and the frames that are your code. Delete the framework frames, the payloads and the timestamps. It takes about ten seconds by hand and routinely removes most of a debugging turn.
7. Keep instruction files small, and push detail into references
Your copilot-instructions.md or AGENTS.md is prepended to every request. The reference app’s AGENTS.md is 108 lines and 1,096 tokens, so Markdown runs near 10 tokens a line. A 400-line instruction file is a 4,000-token surcharge on every turn, all month.
This is why progressive disclosure matters in Copilot Skills. It reads like an organising nicety and is really a cost design: the SKILL.md body stays short, the detail sits in reference files, and the agent loads a reference only when the task needs it.
The rule: the always-on file stays near 100 lines and holds only rules that apply to nearly every change. Everything else moves into a Skill or a reference file. The complete AGENTS.md playbook covers the split in detail.
8. Exclude generated and vendored code from context
Lock files, *.g.dart and *.freezed.dart, build output, minified bundles and vendored libraries are token weight with almost no signal. They are also the files most likely to get swept into a broad context request.
Copilot Business and Enterprise support content exclusion at the repository and organisation level, and most AI gateways support similar path filters. Set them once. Nobody has ever needed pubspec.lock explained to them.
Output discipline
9. Ask for a diff, not a file
The fee fix touches four files: Order, OrderModel, createOrder and the tracking screen’s summary. Fifteen changed lines. Returned as four complete files it is 7,279 output tokens. As a unified diff it is 916.
The diff is also easier to review, and it cannot hide a quiet edit somewhere else in a 500-line file.
10. Cap the explanation
By default these tools explain themselves at length, and every word is metered. “Code only, no explanation” or “two sentences of rationale, then the diff” removes several hundred tokens a turn, several times a day. Keep the explanation on when you are learning something and off when you are not.
11. Do the thinking in the prompt
A vague prompt is cheap to type and expensive to fix. “Add caching to this method” invites the wrong cache, the wrong lifetime and the wrong invalidation, then three correction turns at full thread price.
Sixty words of constraints up front, naming the cache type, key shape, expiry and what happens on a miss, cost about 80 tokens. One avoided correction turn on a file-heavy thread is worth more than a hundred times that.
12. Match the model to the mechanical work
Most tools let you pick a model per request. Renaming, writing a DTO, converting JSON to a class and writing a straightforward test do not need the most capable model, and on metered plans the cheaper models usually carry lower multipliers. Save the strong model for architecture, concurrency and debugging something genuinely strange.
Agent mode: where the real money goes
13. Stop an agent the moment it goes off course
An agent that has misunderstood the task does not self-correct. It compounds. Every step reads more files, produces more tool output, and carries all of it into the next step.
Watch the first two or three actions. If it opened the wrong folder or started editing a file you never mentioned, stop it. Fix the prompt and start again. Letting a wrong run finish because it might recover is the most expensive habit in this list.
14. Prepare before an agent run, the way you would for a deploy
An agent that has to discover your conventions spends its first steps exploring, reading files and resending them, and may still get the shape wrong and need a full rerun. An agent told the conventions up front starts writing on step two.
Commit your work first, so a bad run is one git checkout away. Name the exact files in scope. State the acceptance test. These are quality habits too, but they belong on a token budget because a rerun is the largest single cost event in a normal month.
15. Measure yourself once, then trust the numbers
Twenty minutes and you never guess again. Paste a typical file into the OpenAI tokenizer page, or count locally with tiktoken, and divide by the line count:
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
text = open("lib/domain/entities/order.dart").read()
print(len(enc.encode(text)), len(text.splitlines())) # 1682 tokens, 247 lines
Measured this way, the reference app’s Dart came in near 6.5 tokens per line and its Markdown near 10. Dense C# or TypeScript runs higher, so budget up to 10 per line until you have your own figure.
Then check your usage dashboard weekly rather than monthly. Weekly is early enough to change behaviour. Monthly is a post-mortem.
The same question, asked again
The second time, the developer starts a new chat, selects createOrder and the two fee getters, and asks the 63-token question from habit 5.
The first answer points at line 43. The second turn attaches the four files that need to change and asks for a diff only. It comes back as 916 tokens: a smallOrderFee field on Order, its JSON in OrderModel, line 43 split into two fields, and a “Small order fee” row on the tracking screen.

What I did not bother with
Two things that were tried and dropped, because an honest list includes the misses.
Abbreviating prompts. Writing “impl repo pattern for Order w/ EF” instead of a proper sentence saves maybe 15 tokens and costs accuracy. The prompt is never the expensive part. The context is.
Trimming history mid-thread. Some tools let you remove earlier turns. By the time a thread is long enough to need trimming, restarting is faster and cheaper. Habit 1 makes this one unnecessary.
One genuine unknown is worth flagging. Prompt caching, where a repeated context prefix is billed at a discount, changes some of this arithmetic, but whether your plan or your company’s gateway passes the discount through varies. If it does, keeping a stable prefix (same instruction file, same file order) becomes worth doing on purpose. Check before building habits around it; the GitHub Copilot documentation is the place to confirm what your plan counts.
How much do the habits save?
Same bug, same fix, same developer. Sizes measured on the reference app; the 4,000-token system prompt and the per-turn history and answer sizes are stated assumptions.
| First thread | Second thread | |
|---|---|---|
| Context per turn | 7,591 tokens of attached files | 381 tokens of selection |
| Question | “Why is the service fee wrong?” | 63 tokens naming both fees and the screen |
| Turns to find the cause | 12 | 1 |
| Output for the fix | 7,279 tokens, four whole files | 916 tokens, one diff |
| Thread total | about 212,000 tokens | about 20,000 tokens |
| Share of a 400,000-token month | about 53% | about 5% |
Questions a tech lead will ask about this
- “Which habit do we teach first?” One task per chat, and restart instead of arguing. Those two alone remove most of the cost of a long thread.
- “Is selecting code instead of attaching files risky?” Only if the selection misses the cause. That is why habit 5 comes with a search first: find the symbol, then select it.
- “Do we need content exclusion if everyone selects carefully?” Yes. Agent mode reads files on its own, and exclusion is the only control that applies to every request without anyone remembering it.
- “How do we know the numbers hold for our code?” Run the tokenizer on three typical files from your repository. It takes twenty minutes and replaces every rule of thumb in this article.
- “Do these habits matter on a request-based plan?” Fewer of them. Turn count still matters because each turn is a request, but attachment size matters less. Ask how you are metered before optimising.
What to do on day one
- Close every chat whose task is finished.
- Measure tokens per line on three files of your own code.
- Add generated files and lock files to content exclusion.
- Trim your always-on instruction file towards 100 lines and move the rest into references.
- The next time an answer is wrong, start a new thread with the missing fact instead of a correction.
Key takeaways
- Conversation length, not prompt length, is what drains a token budget. Every turn resends the whole thread and every attached file.
- One task per chat, and restart rather than argue when an answer comes back wrong.
- Select the method instead of attaching the file. On the reference app that was 381 tokens instead of 7,591, and the selection held the answer.
- Name the symbol you mean, and search for it first if you do not know it. Repository-wide context is the most expensive request there is.
- Keep always-on instruction files near 100 lines and push detail into reference files that load on demand.
- Ask for diffs, cap explanations, and put your constraints in the first prompt rather than in three corrections.
- In agent mode, stop a wrong run immediately and prepare properly before starting one.
- Measure tokens per line in your own codebase once, then check usage weekly instead of monthly.
The next ticket
A week later the same tester files another one: the delivery fee on a reopened order shows “Free” when it should not. The developer searches for deliveryFee, selects the two places it is read, and asks one question that names both. The answer is a two-line diff.
One thread, two turns, about 12,000 tokens. The 212,000-token Monday has not happened again.
