GitHub Copilot cannot automate code review 100%. It can own the mechanical layer of review, standards, naming, missing tests, obvious security smells and layer-boundary breaks, and it cannot own the judgment layer: whether this is the right code for the business, and who is accountable when it ships. This is where that line falls, shown on a pull request that passed everything and was still wrong.
Picture a pull request on the MSDevBuild Eats food delivery app: “feat: cancel order and refund”. It adds a CancelOrder use case, a button on the order screen and tests. flutter analyze is clean. The tests pass. The layering rule in CI is happy. The review Skill posts two naming suggestions and no blocking issues. A reviewer skims, agrees and merges.
Three days later support notices that customers who used a promo code are being refunded more than they paid. The use case refunds order.subtotal, the price of the dishes, instead of order.total, what the customer was actually charged after fees and the promo discount. Every line passed review. The whole thing was wrong.
That gives this article its sentence. AI can review whether your code is written well. It cannot review whether your code is the right code. You can automate most of the review work. You cannot automate the judgment. The gap between those two is where engineering happens.
If Skills are new to you, start with why Copilot Skills exist. Here we point a Skill at a harder target, pull request review itself, with the limits made explicit.
One pull request, all green
| Step | What happened |
|---|---|
| Tue 10:00 | The refund pull request goes up: a use case, a button, three tests |
| Tue 10:04 | analyze, test and the dependency rule pass in CI |
| Tue 10:06 | The review Skill posts: 0 blocking issues, 2 naming suggestions |
| Tue 11:30 | A reviewer reads the Skill’s summary, skims the diff, approves |
| Fri 15:00 | Support: promo customers are getting back more than they paid |
The commands that all said yes
The food delivery app’s CI runs three deterministic checks on every pull request. Run them locally on the refund branch:
flutter analyze # the lint and type check
flutter test # the whole suite, including the new use case
grep -rn "package:flutter/" lib/domain/ # the layering rule: must print nothing
| Check | Result on the refund branch |
|---|---|
flutter analyze | No issues found |
flutter test | All tests passed |
| Dependency rule | No Flutter imports in lib/domain |
pr-review-mechanical Skill | 0 blocking, 2 suggestions |
Four gates, four passes. Not one of them could have asked the only question that mattered: what did the customer pay? The tests passed because they were written with no promo in the order, so subtotal and total were the same number.

If you would rather see the idea first, it is a short video:
Code review is really two jobs, not one
Every review is two different jobs wearing one hat. Separating them is the whole trick.
Layer 1: mechanical. Does this follow our standards? Is the naming consistent? Are there tests? Any obvious bug, any secret in the diff, any layer boundary crossed? This is rule-based. The answer does not depend on who wrote it or why. It is a checklist, and a checklist can be automated.
Layer 2: judgment. Is this the right design for where the product is going? Does it fit the business rule that lives in someone’s head and a Jira ticket from March? Is this abstraction worth its complexity? And the one that never automates: who is accountable when this ships?
The mechanical layer asks is this code written correctly? The judgment layer asks is this the correct code? Copilot is excellent at the first and cannot do the second.
Most disappointment with AI code review comes from expecting Layer 2 out of a tool that only reaches Layer 1. And the mechanical layer is exactly where your reviewers bleed time. Watch a senior review a PR and count the comments: most are mechanical: “rename this”, “this belongs in the handler”, “where is the test”. Real value, but boring, repetitive value that a rule could do. Worse, they get tired on it, so the mechanical noise crowds out the one design comment that mattered. That was my refund bug precisely.
So the goal is not “replace reviewers.” It is “stop spending your most expensive people on a checklist.” Let us build something that owns Layer 1 and hands Layer 2 back to a human.

Building a Copilot review Skill for the mechanical layer
We build it the same way as the feature Skills earlier in the series. If you followed the .NET Clean Architecture Skill, this is its partner: that Skill teaches Copilot to write a feature our way, this one teaches it to check a diff against the same rules. It is a SKILL.md in its own folder under .github/skills.
Four parts matter. I will show the parts, not every line.
1. A bounded description. The description is the activation trigger. Copilot reads it on every task and loads the Skill only when the words match. Pack it with the phrases people actually type, and state what it must not own.
---
name: pr-review-mechanical
description: >
Use when reviewing a pull request, diff, or set of changes in the .NET
backend. Performs the mechanical, rule-based review layer only: standards,
naming, Clean Architecture layer boundaries, test coverage, structured
logging, secrets, error handling, ApiResponse<T> compliance. Trigger on
"review this PR", "review this diff", "check these changes".
Does NOT decide whether the design or business approach is correct.
---
That last line is a safety feature, not a formality. Without it, Copilot produces confident sentences about architecture and people read them as approval, which is how you merge my refund bug with a green tick.
2. The rules it enforces. Encode your team’s real standards, concretely. Vague rules produce vague reviews. These mirror the feature Skill so writing and reviewing agree.
Flag as issues:
1. Layer boundaries — no logic, DbContext, or repository calls in a
controller. Controllers send a command/query and translate ApiResponse<T>.
2. Validation — every command taking user input needs a FluentValidation
validator. Missing = blocking.
3. Tests — new handler logic needs a main-path and a failure-path xUnit
test. Zero tests on new behaviour = blocking.
4. API contract — all endpoints return ApiResponse<T>. Flag raw entities,
anonymous objects, bare status codes.
5. Logging — structured, named properties. Never log secrets or full bodies.
6. Secrets — no keys, connection strings, or tokens in the diff. Blocking.
7. Error handling — no empty catch, no swallowed exception.
3. A strict output format. This is the part almost everyone gets wrong. Free prose gives you a wall of mixed-priority comments where a hardcoded secret sits below a variable-name nitpick. Force priority first, and cap the noise.
Respond in exactly this order:
### 1. Blocking issues (must fix before merge; "None" if clean)
### 2. Suggestions (non-blocking; report AT MOST FIVE, highest impact)
### 3. Questions for the author (needs human context you cannot infer)
Never write "Approved", "LGTM", or "ready to merge". Approval is a human
decision.
The cap of five comes from pain. An uncapped AI review floods a PR with thirty low-value comments and buries the one that mattered. The cap forces the model to prioritise, which is the actual job of a review.
4. A self-check so output stays consistent: blocking issues before suggestions, suggestions capped, no approval language, anything needing business context routed to “Questions for the author” instead of guessed at, every blocking issue tied to a line.
The Skill on a real diff
The mechanical layer is easiest to see on a backend diff. Here is the same “cancel order” feature as an ASP.NET Core endpoint, written the way Copilot often writes it without a feature Skill.
// OrdersController.cs — the change under review
[HttpPost("{id}/cancel")]
public async Task<IActionResult> Cancel(Guid id)
{
var order = await _db.Orders.FindAsync(id); // data access in controller
if (order == null) return NotFound();
order.Status = "Cancelled"; // business logic in controller
order.CancelledAt = DateTime.Now; // local time, not UTC
await _db.SaveChangesAsync();
_logger.LogInformation("Order " + id + " cancelled"); // string concat log
return Ok(new { order.Id, order.Status }); // anonymous object, not ApiResponse<T>
}
It compiles. It works in a demo. It breaks four team rules at once. Here is what the Skill produces.
### 1. Blocking issues
- Logic and data access in controller (lines 4–9). FindAsync, status
change, and SaveChanges belong in a CancelOrderCommandHandler. Violates
the Clean Architecture layer boundary.
- No validator for the cancel operation. Needs a FluentValidation validator.
- No test. New cancel + not-found paths have zero xUnit coverage. Blocking.
- Response breaks the API contract (line 13). Anonymous object instead of
ApiResponse<T>.
### 2. Suggestions
- Use DateTime.UtcNow, not DateTime.Now (line 6).
- Structured logging: LogInformation("Order {OrderId} cancelled", id) (line 11).
### 3. Questions for the author
- Should cancelling an already-shipped or already-cancelled order be
rejected? The status-transition rules are business logic I cannot infer —
please confirm with the domain owner.
Read that last section. The Skill did not pretend to know the cancellation rules. It flagged the gap and handed it to a human. That is exactly right: the AI clears the mechanical rubble and shines a light on the one thing that needs judgment. No model in the repo knows your refund policy.
Does Copilot PR review actually read your Skill?
Short answer: yes, but read the fine print, and check current docs, because these features are moving fast. When Copilot reviews a PR it pulls your configuration from the head branch (the branch with the changes, not base): repository custom instructions (.github/copilot-instructions.md), path-specific *.instructions.md files, agent instructions, and Agent Skills in .github/skills/ when they are relevant to the code being reviewed.
Two honest caveats. Agent Skills and MCP servers for code review are currently in public preview and can change; repository custom instructions are the stable, generally available baseline today. And to make a Skill target reviews, give its directory a review-focused name like code-review so Copilot associates it with PR review. There is also an excludeAgent: "code-review" property to hide an instructions file from the reviewer. So the Skill we built is legitimately usable. Treat custom instructions as the reliable layer and the Skill as the richer preview layer on top.
The workflow shift that makes review faster
The payoff is not “AI reviews your code.” It is a change in who does what, and when.
BEFORE — everything hits the human first
Author opens PR → reviewer reads cold → writes 12 mechanical comments
→ reviewer now tired, gives design a tired glance → 2 days of back-and-forth
AFTER — AI clears the mechanical layer first
Author opens PR → Skill posts structured review → author fixes blocking
issues first → reviewer opens a clean PR, skips the checklist → spends full
attention on design + the author's questions → human approves and merges
The wins are concrete: the mechanical pass is done before a human opens the PR, standards land consistently no matter which reviewer is on duty, and juniors get instant feedback at 11pm instead of waiting for a reviewer to wake up. Your seniors stop spending scarce judgment on formatting. In short: the Skill does not review better than your best engineer. It removes the work that was stopping your best engineer from reviewing well.
The honest core: what AI code review misses
Now the part the hype skips. A misplaced trust here ships bugs. These are the AI code review limitations I have watched cost real hours.
- It approves code that solves the wrong problem. The refund pull request. Clean, tested, well named, wrong number. AI reviews how, not whether.
- It misses design flaws that are correct line by line. A circular dependency, a leaky abstraction, a pattern that dies past ten thousand orders: every individual line passes. Architecture is emergent across many files and the roadmap. A diff shows neither.
- Its “best practice” suggestions can be outdated or invented. I have seen it recommend a 2019 anti-pattern and an API that does not exist in our version. Treat every suggestion as a proposal from a well-read junior, not a ruling.
- On security it catches smells, not threats. It flags a hardcoded key. It misses that your new endpoint lets user A cancel user B’s order because there is no ownership check.
- It cannot weigh “is this complexity worth it.” Add a cache? A new abstraction? Split this service? Those are trade-offs against team size, timeline, and load. The model has no stake and no cost model.
- It has no accountability. When code fails at 2am, a person is paged, explains it, owns the fix. Review is, at heart, an act of accountability. You cannot delegate that to something that cannot be answerable.
One real cost worth naming: Copilot code review is metered on most plans, so running it on every pull request is not free. Check your plan’s current terms. That is another reason to keep it a tight mechanical pass on small, focused diffs, which is also where any reviewer, human or AI, is most accurate.
So, 100% automated? No. Here is the split
Chasing 100% is the wrong target; it makes you trust the tool exactly where it is weakest. The realistic goal is to shift the mechanical load off humans so their review time flows to design and risk. A good Skill can own the large majority of comments on a routine PR, because most comments are mechanical. But the few it cannot make are usually the ones that decide whether the feature is actually right.
The split, worth putting on the wall:
| Layer | Owner | Example checks |
|---|---|---|
| Mechanical | AI review Skill | Naming and formatting, layer boundaries, missing tests, missing validators, ApiResponse<T> contract, structured logging, hardcoded secrets, empty catch blocks |
| Judgment | Human reviewer | Is this the right design? Does it match the business rule? Is the abstraction worth it? Is the trade-off acceptable? Is the threat model sound? Who is accountable in production? |
Read it as a division of labour, not a competition. The AI is a different reviewer: tireless and consistent on exactly the layer where humans are slow and bored. Pair them and you get the best of both. Ask either to do the other’s job and you get the refund pull request.
Advisory vs enforceable: Copilot review reasons, it doesn’t run
One limit people miss: the review bot does not execute anything. No git hooks, no scanners, no tests. It reads the diff and reasons about it. So any check that must be deterministic (a license policy, a security scan, the test suite) belongs in CI as a required check, with Copilot review as the advisory layer on top.
Deterministic gate (CI + hooks) Advisory layer (Copilot review)
license / security / tests reads diff, reasons, nudges
→ required status check, BLOCKS merge → comments, can miss, never blocks
Take a license check on a .NET repo. Two lanes, four steps:
- Add a GitHub Actions workflow on
pull_requestthat runs a real scanner and fails on a disallowed license.
# .github/workflows/license-check.yml
on: pull_request
jobs:
licenses:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-dotnet@v4
- run: dotnet tool install --global dotnet-project-licenses
- run: dotnet-project-licenses -i . --allowed-license-types MIT Apache-2.0 --failed
- Make that workflow a required status check in a branch-protection ruleset, so a failing scan blocks merge. This is the actual gate.
- (Advisory) Add a Copilot review instruction in
.github/copilot-instructions.md, for example “flag any newly added dependency and confirm its license is on the approved list”, or point it at a checklist file. GitHub’s own docs show instructions like “apply the checks in/security/security-checklist.md”. The human-readable nudge then shows up in the PR too. - (Local) Mirror the scanner in a pre-commit or pre-push hook for fast, free feedback before push.
The rule to remember: Copilot review reasons, it doesn’t run. Deterministic checks are the gate; Copilot is the reasoning layer on top of it.
Do’s and don’ts of AI-assisted PR review
| Do | Don’t |
|---|---|
| Use it as the first pass, before a human opens the PR | Let it auto-approve or auto-merge |
| Demand a structured output: blocking / suggestions / questions | Accept a free-prose blob of mixed-priority comments |
| Cap the suggestions so signal beats noise | Let it flood the author with thirty nitpicks |
| Tie its rules to your real, written standards | Trust generic “best practice” suggestions blindly |
| Keep a human as the accountability gate on merge | Remove the human because the AI passed it |
| Scope it to small, focused diffs | Point it at a 2,000-line PR and expect accuracy |
Where this fits
This review Skill clicks into the Skills that write code: the .NET feature Skill writes backend code to a standard, the Flutter Skill does the same for mobile, the Azure Skill guards infrastructure, and this one checks all of it against the same rules. Writing and reviewing speak one language.
One metric is worth tracking from the first week: comments actioned over comments made. If the AI reviewer posts a hundred comments and developers resolve five, you have noise, not review. That number is how you tune the Skill, and how you show it earns its cost.
How efficient is AI-assisted review?
The Skill takes the mechanical comments off the human, so human review time goes to the questions only a human can answer.
| Human review only | Skill first, then a human | |
|---|---|---|
| Mechanical comments made by the human | most of them | almost none |
| Time before the first feedback | hours, until a reviewer is free | minutes, before anyone opens the PR |
| Questions about business rules asked | whenever there is time left | the reviewer’s whole job |
| Who approves and owns the merge | a human | a human |
The refund pull request is the reason for the last row. The rule underneath: automation should buy the reviewer time, never replace the reviewer’s judgment. A review that ends at “the Skill said 0 blocking” has automated the one part that should not be.
Questions a tech lead will ask about this
- “If the Skill found nothing, why do we still need a reviewer?” Because the Skill checks how the code is written. The refund bug was in what the code does.
- “Can we make the Skill check business rules?” Only the ones you can write as rules: “refunds use
order.total” can go in. Rules you have not thought of yet cannot. - “Should the Skill block merges?” No. Deterministic checks block, in CI. The Skill advises.
- “How do we stop the Skill drowning us in nitpicks?” Cap suggestions at a handful and measure comments actioned over comments made.
- “What should reviewers read first now?” The Skill’s “questions for the author” section, then the parts of the diff that change money, permissions or data.
What to do on day one
- Split review into mechanical and judgment in your team’s own words, and write it down.
- Put the mechanical rules in a review Skill with a capped, three-section output.
- Put the must-never-break rules in CI, where they block.
- Require a named human approval on every merge.
- Add a test with a promo, a discount or an edge case to any change that touches money.
Key takeaways
- GitHub Copilot cannot automate code review 100%. It owns the mechanical layer; judgment stays human.
- Mechanical (standards, tests, secrets, layer boundaries, contract) is rule-based and automatable. Judgment (design fit, business context, trade-offs, accountability) is not.
- A good review Skill needs a bounded description, real rules, a strict output format, and a suggestion cap to kill noise.
- The win is a workflow shift: AI clears mechanical comments before a human opens the PR.
- Known limits: it approves wrong-but-clean code, misses architecture and threat-model issues, gives outdated suggestions, and has no accountability.
- Never let AI approve or merge. Keep a human gate. Measure comments actioned, not comments made.
The next refund change
Two weeks later a developer adds partial refunds: cancel one dish from an order. The review Skill runs first and posts one suggestion about a method name.
The reviewer skips the naming and goes straight to the “questions for the author” section, which now carries a line the team added to the Skill after the refund bug: “Does this change how much money moves? If so, which field is it based on?” The author answers in the thread: the refund is the dish’s share of order.total, after the promo. There is a test with a promo order to prove it.
The reviewer approves with their own name on it. That was always the part no tool could do.
