400K AI Tokens a Month: My Daily GitHub Copilot Workflow That Makes the Budget Last

A 400K monthly AI token budget sounds generous until agent mode eats it in a week. Here is the daily Copilot workflow I use to make it last the month.

By Suthahar Jegatheesan Updated October 1, 202619 min read —views
Article banner. On the left, the eyebrow "GitHub Copilot, AGENTS.md" above the title "400K tokens a month" and the line "About 19,000 a day. Choose the mode before you type.", with the MSDEVBUILD wordmark and the author name below. On the right, three stacked boxes joined by arrows: a grey box reading "One thread, four files pinned, arguing with an agent for 14 turns", an arrow labelled "agent by reflex" to an amber box reading "Every turn resends everything, about 225,000 tokens", and an arrow labelled "one bug" to a red box reading "61% gone on day nine, and three weeks still to go".

A 400,000-token monthly AI allowance is about 19,000 tokens per working day: roughly two chat turns with a couple of files attached, or a fraction of one agent run. It lasts the month only if you choose the mode before you type: no model at all, an unmetered tool for generic questions, Copilot on a selection for work in one file, and agent mode for the few multi-file jobs that deserve it. This is the daily workflow that makes it last.

Nine working days into a month, the usage dashboard said 61% was gone. Not on anything clever. On one chat, arguing with an agent about an HttpClient registration, with four files pinned to the context for fourteen turns.

Four hundred thousand of anything sounds generous. It took two months to see it is a small daily budget, and one more to see that the fix is not typing less. The fix is deciding, before typing, which mode of Copilot the task belongs to, and whether it needs Copilot at all.

That decision happens twenty or thirty times a day. Inline completion, chat on a selection and agent mode are not three levels of one feature; they differ in cost by orders of magnitude.

If the word “token” is still fuzzy, start with what a token is and how AI cost is calculated: everything below is arithmetic built on it, including the measured file sizes.

Nine days, one thread

DayWhat happened
1 to 7Normal work: completions, a few chat questions. About 5% of the month
8, 10:00An HttpClient registration bug. Agent mode, four service files pinned
8, all dayThe agent gets it wrong on turn 3. The fix is explained on turn 4, again on turn 6
9, 11:00Turn 14. The bug is fixed
9, 14:00Dashboard: 61% used. Three weeks of the month left

The estimate that explained it

The cost of a thread is its fixed part, sent on every turn, plus the history, which grows every turn. With four files of about 1,500 tokens each, a 902-token instructions file, around 4,000 tokens of system prompt and tool definitions, and replies of about 600 tokens:

def thread_cost(fixed, turns, history_per_turn=700, reply=600):
    """Input for every turn, including all earlier turns, plus the replies."""
    return sum(fixed + history_per_turn * (k - 1) for k in range(1, turns + 1)) + reply * turns

fixed = 4000 + 902 + 4 * 1500 + 50      # system + tools, instructions, 4 files, question
print(thread_cost(fixed, 14))           # ≈ 225,000 tokens

About 225,000 tokens: 56% of the month, from one thread. The other 5% was a normal week. Nothing in the thread was wasteful turn by turn. The thread itself was the waste. All of it.

Flowchart for routing a task before you type. It starts with a task arriving. First decision: does a deterministic tool give the exact answer? Yes leads to Tier 0, the IDE refactor, the CLI or the docs, with no tokens. No continues to the second decision: can it be asked without any company code? Yes leads to Tier 1, an unmetered chat rewritten against generic code. No continues to the third decision: does it fit inside one open file or selection? Yes leads to Tier 2, Copilot on the selection in one short thread. No continues to the fourth decision: is it planned multi-file work worth a budget? Yes leads to Tier 3, agent mode, with instructions and Skills in place. No leads to: split it until a piece fits Tier 2.

Figure 1 — route the task before you type. Each “yes” leaves to the cheapest tier that can do the job.

If you would rather see the idea first, it is a short video:

Is 400,000 tokens a month actually enough?

The answer is arithmetic, not opinion.

Twenty-one working days in a typical month gives me about 19,000 tokens per day. To turn that into something I can feel, I need two conversion rules I now use constantly.

Rule one: measure tokens per line once. The Dart in the food delivery app measured about 6.5 tokens per line with the GPT-4o tokenizer; denser C# or TypeScript runs higher, so budget up to 10. At 6.5, a 300-line class is about 2,000 tokens. Every time it enters the context.

Rule two: every turn pays for the whole conversation again. This is the one that quietly destroyed my first two months. A chat is stateless underneath. Turn 5 does not send your fifth question, it sends questions 1 through 5, all five answers, and every attached file, again. Cost grows with the square of the turn count, not in a straight line.

Put those together and the day looks like this.

What I doRough token costHow many fit in one day (19K)
Inline completion, single line100 – 400Hundreds
Short question, no files attached500 – 1,500~15
One question with two 300-line files attached8,000 – 10,0002
An 8-turn chat over those same two files85,000 – 95,000Four to five days of budget
One agent task: read repo, edit 4 files, run tests twice60,000 – 150,0003 to 8 days of budget

So the honest verdict on 400,000 a month, for a working developer:

  • Enough, comfortably, if AI is a coding assistant. Completions all day, tight chat turns, one file at a time.
  • Not enough, not close, if agent mode is your default reflex. Four hundred thousand tokens is three to six real agent runs. That is one per week, and you still have the other four days to get through.
  • The deciding variable is not how much you use AI. It is how much repository you hand it, how many times.

That reframing is what changed my month. I stopped thinking about a monthly quota and started thinking about a daily one, and the daily one is small enough that I have to be deliberate.

One caveat before the routing model, and it matters more than anything else here. Find out how your company meters you. A seat-based GitHub Copilot plan counts premium requests, not tokens, and inline completions are usually outside that count. A company gateway sitting in front of Azure OpenAI or the Anthropic API counts every token in both directions, completions included. Those two worlds need opposite habits. Ask your admin four questions: do completions count, do input and output count the same, does cached input get discounted, and does the pool reset monthly or roll over. I optimised the wrong thing for six weeks because I never asked.

Step flow dividing a token allowance. Four hundred thousand tokens a month becomes about nineteen thousand tokens per working day across twenty-one days with no rollover, which is about four file-heavy chat turns or roughly one third of a single agent-mode run.

Figure 2 — the allowance divided down to something you can actually plan against.

Which AI tool for which task?

Here is the model I landed on. Four tiers, and the work moves down a tier only when the tier above genuinely cannot do it.

Tier 0: the task that needs no model at all

The cheapest token is the one you never spend. A surprising share of what I used to ask an AI is faster without one.

Renaming a symbol across a solution is a keyboard shortcut in Visual Studio, not a prompt. Extracting an interface is a refactor menu item. Scaffolding a project is dotnet new. Finding every caller of a method is a right-click. Checking whether IAsyncEnumerable supports cancellation is a documentation page that will be correct, where a model may guess.

I keep a rule for this: if a deterministic tool gives the exact answer, a probabilistic one is the wrong instrument. Not because of cost. Because it is also slower and occasionally wrong.

Tier 1: small, generic questions with no company code in them

This is the tier the user of a tight budget lives in, and it is where the unmetered consumer tools earn their place.

What goes here: syntax I have forgotten, a regular expression, the difference between two framework APIs, a sample JSON payload, naming ideas, a Bicep snippet from a blank page, “what does this compiler error mean”, boilerplate I could have typed but would rather not. Short in, short out, no repository context.

For these I use whatever is not metered against my work pool. ChatGPT or Gemini in a browser tab does this perfectly well, and it costs my allowance nothing.

The hard boundary, and it is not a cost boundary: nothing proprietary crosses into Tier 1. No source file from the repository. No customer data, no connection string, no internal service name, no architecture diagram, no unreleased product detail. Route by sensitivity first and cost second. If a question can only be asked by pasting company code, it is not a Tier 1 question, and the tight budget is not a reason to make it one. I rewrite the question against a generic Order and Customer instead, which usually takes twenty seconds and gets a cleaner answer anyway.

Tier 2: work inside code I already have open

This is Copilot’s home ground, and where most of my metered spend should sit.

Completions as I type. A selected method plus a short chat instruction: “rewrite this to use AsNoTracking and keep the projection”. A test for the class in front of me. A specific exception in a specific file. The context is small and I chose it deliberately, so each turn is cheap.

The habit that matters at this tier: select, do not attach. Highlight the 40 lines the question is about instead of handing over the 600-line file. Same answer, a tenth of the tokens.

Tier 3: multi-file work that needs the architecture in its head

Agent mode. New feature across data, domain, and presentation layers. A migration that touches twelve call sites. Backfilling tests for a class nobody covered. Legacy modernisation.

This tier is expensive and it is also where AI is worth the most. I do not avoid it. I ration it. I plan for two or three agent runs a month, and I prepare for them like a deployment, because an agent run that fails halfway costs the same as one that succeeds.

The preparation is exactly what the earlier articles in this series were about, and it turns out those files are a cost control as much as a quality control. An agent with a good custom instructions file does not spend 20,000 tokens exploring your folder structure to guess at your conventions, because the conventions are stated. An agent with a proper Skill gets the shape right on the first attempt instead of the third. Every correction round trip you avoid is a full context resend you did not pay for. I wrote those files to stop architectural drift. They ended up saving more budget than any prompt trick I know.

The routing table

This is the version pinned above my desk.

TaskTierWhere it goes
Rename, extract, move, scaffold0IDE refactor tools, CLI
API behaviour, framework semantics0Official docs
Regex, syntax recall, generic snippet1Unmetered chat, no company code
Naming, wording, commit message1Unmetered chat
Explain this compiler error1Unmetered chat, paste the error only
Write or fix code in the open file2Copilot inline and chat, on a selection
Unit test for one class2Copilot chat, class selected
Review a diff before I push2Copilot on the diff, not the repo
New feature across layers3Agent mode, with instructions and Skill in place
Repo-wide migration or test backfill3Agent mode, planned and budgeted
Anything containing customer data—Approved tooling only, or not at all

What does a normal Copilot day actually look like?

The table above is the rule. This is the rule applied to a Tuesday, because a routing model you cannot run under pressure is just a diagram.

Morning, picking up a bug. I read the failing test myself first. Copilot is not faster than me at reading one assertion. When I have the suspect method on screen, I select it, and ask chat one specific question about it. One turn, about 700 tokens. If the answer is wrong I close the thread and re-ask with the missing constraint instead of arguing, because arguing buys the whole thread again.

Mid-morning, writing the fix. Inline completion, nothing else. This is where Copilot earns its seat and it is the cheapest thing it does. I let it finish lines, and I stop accepting when it starts inventing a method that does not exist, which is usually the signal that I have not decided what I want yet.

Before lunch, tests. Select the class, ask for tests for the two branches I care about, specify the framework and the naming convention. One turn. I never ask for “tests for this file” because the answer is fifteen tests, twelve of which I delete, all of them paid for.

Afternoon, someone else’s pull request. Copilot on the diff. Not the repository, not the branch, the diff. A 200-line diff is 2,000 tokens and gives a genuinely useful second opinion. Pointing it at the whole project to “review this PR properly” costs thirty times more and reads worse.

The question that arrives at 3pm. Something generic, no company code in it: how a SemaphoreSlim behaves on cancellation, or the exact syntax for a Bicep loop. That goes to an unmetered browser tab, rewritten against a generic example. Nothing from the repository leaves the approved tooling.

Agent mode: not today. On a normal day I do not open it at all. It runs on the days I have planned for it, on the work that deserves it, with the repository committed first so a bad run is a git checkout rather than an afternoon.

Add that day up and it lands between 8,000 and 15,000 tokens, comfortably inside the 19,000 I have. The days that break the budget are the days I skipped the decision and reached for the most powerful mode by reflex.

What changed once I chose the mode instead of just typing less

Three things, and only one of them is about tokens.

The budget stopped being the constraint. Most of my week is Tier 0, 1, and 2 work now. Tier 3 is planned, not reflexive. The monthly number stopped being something I watched nervously in the third week.

My prompts got better. Deciding which tier a task belongs to forces you to state what the task actually is. Half the time the act of classifying it tells me I do not need a model, and the other half I write a sharper prompt because I already know what kind of answer I want.

I stopped arguing with agents. That 61% month happened because when the agent got something wrong on turn 3, I explained why on turn 4, and again on turn 6. Every one of those turns bought the entire thread again. The correct move is to close the thread, fix the original prompt, and start clean. It feels like giving up. It costs a fifth as much and lands faster.

How I plan the month

I split 400,000 into weekly buckets rather than one monthly pool, because a monthly pool always gets spent in the first half.

  • Roughly 40,000 per week for normal Tier 2 work: completions, selections, tests, small fixes. Four weeks is 160,000.
  • One Tier 3 agent run per fortnight, budgeted at 100,000, chosen deliberately. The work with the worst ratio of tedium to thinking wins.
  • A 40,000 reserve, untouched until the last week. The three lines add up to the full 400,000. Production incidents do not respect budgets, and reading unfamiliar code fast is exactly when AI earns its keep.

The reserve is the part I would recommend hardest. Running out on a day when an integration breaks, with no allowance left to help read someone else’s service, is a bad way to learn this lesson.

If your company gives you a different number, the shape still holds. Divide by 21, convert to lines of code at your measured tokens per line, and you will know within a minute whether you can afford to be casual about agent mode. Most developers cannot.

How efficient is routing by tier?

The same nine days, the same bug, handled by tier instead of by reflex:

Agent by reflexRouted by tier
How the bug was handledagent mode, four files pinned, 14 turnsTier 2: the registration method selected, three short threads
Tokens for the bugabout 225,000about 30,000
Share of the month after day nine61%about 13%
Agent runs left for planned worknone, realisticallytwo or three

The rule underneath: the cost of a task is set by the tier you pick, not by how carefully you type. The same bug costs about seven times more in the wrong tier. Seven.

Questions a tech lead will ask about this

  1. “Should we ban agent mode?” No. Plan it. Two or three well-prepared agent runs a month are where AI saves the most time.
  2. “Is it safe to use a free chat tool for Tier 1?” Only for questions with no company code, customer data or internal names in them. Sensitivity first, cost second.
  3. “How do we know how we are metered?” Ask four questions: do completions count, do input and output cost the same, is cached input discounted, and does the pool reset monthly.
  4. “Why do instruction files save budget?” Because an agent that knows your conventions does not spend tokens exploring or correcting. Every avoided correction is a whole-thread resend you did not pay for.
  5. “What is the single habit to teach first?” Close the thread and restate the question instead of arguing with a wrong answer.

What to do on day one

  • Divide your allowance by the working days in the month and write the daily number down.
  • Measure tokens per line on your own code once.
  • Pin the routing table where you can see it.
  • Keep a reserve for the last week; incidents do not respect budgets.
  • When an answer is wrong on turn 3, start a new thread with a better first prompt.

Key takeaways

  • 400,000 tokens a month is about 19,000 a day, which is a small daily allowance once you convert it into lines of code at your measured rate, about 6.5 per line of Dart.
  • It is enough for assistant-style coding and clearly not enough for agent-first coding. The variable is how much repository you attach and how many times you resend it.
  • Conversation history is resent on every turn, so a long thread costs far more than the same questions asked in separate threads.
  • Route work in tiers: deterministic tools first, unmetered chat for generic questions, Copilot for work inside an open file, agent mode reserved for two or three planned jobs a month.
  • Sensitivity outranks cost. A tight budget is never a reason to paste company code into an unapproved tool.
  • Instruction and Skill files are a budget control, not just a quality control, because they remove the exploration and the correction rounds you would otherwise pay for.
  • Ask your admin how you are metered before optimising. Request-based and token-based plans reward opposite habits.

The next bug

Three weeks later, another registration bug, this time in the notification service. The developer reads the failing test, selects the one method, and asks chat a single specific question. The first answer misses a constraint, so they close the thread and ask again with it included. The second answer is right. Done.

Two threads, three turns in total, about 18,000 tokens. The dashboard is at 58% with a week to go, and the reserve is untouched.

Test yourself: answer in the comments

Was this useful?

Share

Found a mistake or an outdated step? Edit this page on GitHub

Frequently asked questions

Is 400,000 AI tokens per month enough for a developer?
It is enough for chat-assisted coding and not enough for agent-first coding. Spread over 21 working days it is about 19,000 tokens a day, which covers four or five chat turns with files attached, or roughly a third of one multi-file agent task. If agent mode is your default habit you will run out in the second week.
How many tokens does one line of code cost?
Measured with the GPT-4o tokenizer on a real Flutter app, Dart came in at about 6.5 tokens per line, so a 300-line file is roughly 2,000 tokens every time it enters the context. Dense C# or TypeScript runs higher; budget up to 10 per line. Measure your own codebase once with a tokenizer.
Why does a long AI chat cost so much more than a short one?
Every turn resends the whole conversation and every attached file, so a thread's total grows with the square of its length. Eight turns over two attached files cost around ten times the first turn. Starting a fresh chat per task is the single largest saving available to you.
Should I send company code to a free ChatGPT account to save tokens?
No. Route by data sensitivity first and cost second. Unmetered consumer tools are for questions that contain no proprietary code, no customer data, and no internal architecture. Anything touching your repository stays inside the tool your company approved, even when the metered budget is tight.
Do inline Copilot completions count against a token budget?
It depends entirely on how your company meters access. Seat-based GitHub Copilot plans count premium requests rather than raw tokens, while a company gateway in front of Azure OpenAI or the Anthropic API usually counts every token including completions. Ask your admin which model you are on before you optimise anything.
When should I use Copilot agent mode instead of chat?
For planned, multi-file work that is worth a budget: a feature across layers, a migration touching many call sites, or a test backfill. Agent runs read and resend large parts of the repository, so ration them to two or three a month on a metered plan, and use chat on a selection for everything that fits inside one file.
Article banner. On the left, the eyebrow "GitHub Copilot, AGENTS.md" above the title "Save AI tokens, 15 habits" and the line "The cost is the conversation, not the prompt.", with the MSDEVBUILD wordmark and the author name below. On the right, three stacked boxes joined by arrows: a grey box reading "Three files attached, 12 turns, and four whole files returned", an arrow labelled "the same question" to a red box reading "About 212,000 tokens for a 15-line fix", and an arrow labelled "asked again" to a green box reading "Select, name it, ask for a diff, about 20,000 tokens".

Next in this series · Part 13 of 13

17 min

How to Save AI Tokens: 15 Habits That Cut My Copilot Context Waste

Fifteen tested habits that cut monthly AI token usage without cutting how much you use AI, from context discipline to knowing when to restart a chat.

Continue the series
Part 11 of 13What Is an AI Token? Why Every AI Cost Is Calculated in Them

Get new posts by email

New technical articles, Azure AI and GitHub Copilot updates, and upcoming events. No spam, unsubscribe anytime.

Comments

Your turn

How did Suthahar's articles help you?

If something here saved you time or unblocked a real project, I'd love to hear about it. Submissions are reviewed before they appear on the site.

0/1500 · minimum 10 characters

Never published — used only to verify your feedback.

Your name, company, and role appear publicly if published. Nothing else is collected.

↑↓ navigate ↵ open