Skip to content
Panoptix AI
AI Coding

How to Debug Code Faster with AI: A Practical Guide (2026)

A tool-agnostic AI debugging workflow: reproduce the bug, give full context, ask for hypotheses, write a failing test and avoid the classic AI traps.

By Panoptix Editorial TeamPublished 11 min read
A stack trace in a terminal next to an AI chat panel listing three ranked hypotheses for the bug
Table of contents

Difficulty

Intermediate

Time required

45 minutes

Tools needed

  • ChatGPT
  • Claude or Gemini
  • An agentic coding tool (optional)
  • Your test runner
  • Git

The fastest way to debug with AI isn't pasting an error into a chatbot and asking "fix this." That works for typos. For anything real, it tends to produce a confident patch that makes the error disappear without explaining why it appeared, and you end up debugging the fix.

What actually works is the same discipline good engineers already use, with AI doing the tedious parts: reproduce the bug, hand over complete context, ask for ranked hypotheses instead of a fix, test those hypotheses with targeted logging or a bisect, lock the bug in with a failing test, and only then let the AI (or an agent) write the fix. You stay in charge of the diagnosis. The AI speeds up everything around it.

This guide walks through that workflow step by step, with copy-paste prompts and a small Python bug as a running example. It's tool-agnostic: ChatGPT, Claude or Gemini in a browser tab is enough for most steps, and agentic tools like Claude Code, GitHub Copilot or Cursor make the later steps faster.

Why "just ask the AI to fix it" is slower than it looks

Developers already know this intuitively. In the 2025 Stack Overflow Developer Survey, the top frustration with AI tools, cited by 66% of respondents, was "AI solutions that are almost right, but not quite." Another 45% said debugging AI-generated code is more time-consuming. Only about 3% said they highly trust the accuracy of AI output, while roughly 46% distrust it to some degree. (Stack Overflow's 2026 survey opened in June; its results weren't published when we checked.)

The productivity evidence is mixed, too. METR's randomized trial of experienced open-source developers found they took 19% longer on tasks when allowed to use early-2025 AI tools, even though they believed AI had sped them up. A February 2026 follow-up pointed toward a speedup with newer tools, but METR itself called that "very weak evidence" because of selection problems in the study.

Our takeaway: AI can make debugging much faster, but only if you give it the structure a good human debugger would need. "Almost right" is what you get when the model is guessing.

The running example

Here's the bug we'll use throughout. A small reporting script sums up daily revenue, and it started crashing this week:

# orders.py
def order_total(order):
    subtotal = sum(item["price"] * item["qty"] for item in order["items"])
    return subtotal - order.get("discount", 0)

# report.py
from orders import order_total

def daily_revenue(orders):
    return sum(order_total(o) for o in orders)

And the (illustrative, trimmed) traceback:

Traceback (most recent call last):
  File "/app/report.py", line 14, in <module>
    print(daily_revenue(orders))
  File "/app/report.py", line 5, in daily_revenue
    return sum(order_total(o) for o in orders)
  File "/app/orders.py", line 4, in order_total
    return subtotal - order.get("discount", 0)
TypeError: unsupported operand type(s) for -: 'float' and 'NoneType'

If you paste only the last line into a chatbot, you'll likely get "use order.get('discount') or 0." That stops the crash. It doesn't answer the question that matters: why are orders suddenly arriving with a discount key set to None, and is anything else in them wrong too?

  1. Step 1: Reproduce the bug on demand

    Before you open any AI tool, get a command that triggers the bug every time. A bug you can't reproduce is a bug you can't verify as fixed, and the AI can't either.

    For the example, that means saving one failing order to a fixture and running the script against it:

    python report.py --input fixtures/orders_2026-09-27.json

    If the bug is intermittent, write down what you know: roughly how often it happens, under what load, on which environment. "Fails about 1 in 20 CI runs, only on the Linux runner" is far more useful to an AI than "sometimes fails."

  2. Step 2: Package the full context

    Most bad AI debugging answers come from missing information. The model fills gaps with the most statistically likely explanation, which may have nothing to do with your code. Give it what a senior colleague would ask for:

    • The exact error message and the full stack trace, not a paraphrase
    • Versions: language runtime, framework, key libraries, OS
    • The relevant code, including the caller, not just the line that threw
    • A minimal reproducible example, following Stack Overflow's well-known guidance on minimal, complete and reproducible code
    • What changed recently (a deploy, a dependency bump, a new data source)
    • What you already tried and what happened

    Strip secrets, customer data and internal hostnames before pasting anything into a chat tool. Replace real records with fake ones that still trigger the bug.

    Prompt: debugging context template
    I'm debugging an error. Don't propose a fix yet.
    
    Environment: Python 3.12, running on Ubuntu 24.04. No frameworks.
    Command that reproduces it: python report.py --input fixtures/orders.json
    It started: after Monday's change to the orders export job.
    
    Full traceback:
    [paste the entire traceback]
    
    Relevant code:
    [paste orders.py and report.py]
    
    Sample input that fails (anonymized):
    [paste one failing order]
    
    What I've tried: confirmed older fixtures from last week still work.

    Writing this out takes five minutes. It also regularly solves the bug on its own, because assembling the facts forces you to look at them.

  3. Step 3: Ask for hypotheses, not fixes

    This is the single biggest change most people can make. Instead of "fix it," ask the model to reason like a debugger: list possible causes, rank them, and tell you how to confirm or rule out each one.

    Prompt: ranked hypotheses
    Based on the context above, list the 3-5 most likely root causes.
    For each one:
    1. Explain the mechanism in one or two sentences
    2. Rate how likely it is (high / medium / low) and why
    3. Give me one cheap check (a log line, an assertion, a query or
       a command) that would confirm or rule it out
    
    Distinguish the root cause from the symptom. Don't write a fix.

    For the example, a good answer would include something like: "the export job now writes discount: null instead of omitting the key, so dict.get returns None rather than the default." That's a hypothesis you can check in seconds, and it points upstream, away from the line that crashed.

    This step is also where the choice of model matters a bit. We compare how the major assistants handle code reasoning in Claude vs ChatGPT for coding. If you want a refresher on structuring prompts generally, see our 10 prompt techniques.

  4. Step 4: Add targeted logging to test the hypotheses

    Now collect evidence. Ask the AI to write logging or assertions that distinguish between its hypotheses, rather than sprinkling print statements everywhere.

    Prompt: instrument to confirm a hypothesis
    Write temporary logging for hypothesis #1 only. I want to see, for each
    order: its id, whether the "discount" key exists, the discount's value
    and type, and the source field from the export job. Use the logging
    module at DEBUG level, prefix each line with "DEBUG-BUG-4812" so I can
    grep for it and remove it later.

    The prefix is a small habit that pays off. It makes the temporary instrumentation easy to find and delete, and easy to spot in a code review if someone forgets.

    Some tools now automate this loop. Cursor's Debug Mode, for example, generates hypotheses, adds log statements, asks you to reproduce the bug, reads the runtime logs, makes a targeted fix and then removes its instrumentation. It's a good sign that tool makers are building products around evidence over guesswork.

  5. Step 5: Narrow it down with a bisect

    If the bug is a regression ("this worked last week"), you often don't need to reason about the cause at all. You can find the exact commit that introduced it. Git's built-in git bisect does a binary search through history, and git bisect run automates it with a script: exit code 0 means the commit is good, 1 to 127 (except 125) means bad, and 125 means skip.

    AI is handy for writing that script, which is the fiddly part:

    Prompt: bisect script
    Write a bash script for "git bisect run" that:
    - installs dependencies quietly (exit 125 if install fails, so the
      commit is skipped)
    - runs: python report.py --input fixtures/orders.json
    - exits 1 if the output contains "TypeError", 0 otherwise
    Then give me the exact git bisect commands, assuming v2.3.0 was good
    and HEAD is bad.
    git bisect start HEAD v2.3.0
    git bisect run ./bisect-check.sh
    git bisect reset

    Once you have the offending commit, paste its diff into the chat along with the traceback. A 40-line diff plus a stack trace gives a model far better odds than an entire codebase.

  6. Step 6: Write a failing test before the fix

    With the root cause confirmed, capture it in a test that fails right now. This turns a vague goal ("make it work") into a pass/fail signal, for you and for any AI tool you hand the job to.

    # test_orders.py
    from orders import order_total
    
    def test_order_total_treats_null_discount_as_zero():
        order = {"items": [{"price": 10.0, "qty": 2}], "discount": None}
        assert order_total(order) == 20.0
    
    def test_order_total_applies_discount():
        order = {"items": [{"price": 10.0, "qty": 2}], "discount": 5.0}
        assert order_total(order) == 15.0

    Run it and watch it fail for the right reason (the same TypeError). A test that fails for a different reason isn't testing your bug.

    Anthropic's own Claude Code best practices recommend exactly this pattern in its prompt examples: describe the symptom and the likely location, then "write a failing test that reproduces the issue, then fix it." The advice applies to any tool.

  7. Step 7: Let an agent run the fix-and-test loop

    Chat assistants can only suggest. Agentic tools can read your repo, edit files, run your test suite and iterate until it passes, which is where a lot of the time savings come from. The main options as of September 2026:

    • Claude Code runs in your terminal or IDE. Its fix-bugs workflow recommends giving it the command that reproduces the issue, the steps to trigger it, and whether it's intermittent. Our Claude Code review covers it in depth.
    • GitHub Copilot has agent mode in the editor, plus the Copilot cloud agent (renamed from "coding agent" in 2026), which works in a GitHub Actions-powered environment where it can run tests and linters, then hands you a branch or pull request. GitHub says it's available on all paid Copilot plans.
    • Cursor's Agent can edit multiple files and run terminal commands, and Debug Mode (Step 4) adds the runtime-evidence loop. We weigh the two editors in GitHub Copilot vs Cursor.

    Whichever you use, give it a pass/fail check and tell it what not to touch:

    Prompt: agentic fix with guardrails
    Root cause (confirmed with logs): the orders export job now writes
    "discount": null instead of omitting the key. orders.order_total
    doesn't handle None.
    
    1. Make tests/test_orders.py pass. Run pytest after each change.
    2. Fix the root cause in order_total and tell me whether the export
       job should also be changed to omit null discounts.
    3. Do NOT modify, skip or delete existing tests. If you think a test
       is wrong, stop and explain why instead of changing it.
    4. Don't wrap the calculation in try/except to hide errors.
    5. When done, show me the test output and a summary of the diff.
  8. Step 8: Review the fix and the root cause

    A green test suite is necessary, not sufficient. Read the diff yourself and ask three questions:

    1. Does the fix address the cause I confirmed? Or does it just make the symptom go away?
    2. Is anything else affected? If the export job changed its output format, other consumers may be quietly getting wrong numbers instead of crashing.
    3. Did the AI change anything it shouldn't have? Look for edited tests, loosened assertions, new dependencies and deleted "unused" code.

    Then remove your DEBUG-BUG-4812 logging and write a one-paragraph note in the commit or PR explaining the root cause. Future you will be grateful.

    Prompt: second-opinion review
    Review this diff as a skeptical senior engineer. The goal was to fix
    [root cause]. Report only real problems:
    - Does it fix the root cause or only suppress the symptom?
    - Could it change behavior for any other input?
    - Were any tests modified, weakened or removed?
    - Does it call any function, option or package that may not exist in
      [library + version]?
    [paste diff]

    Running this in a fresh chat or session helps, because the model isn't anchored to the reasoning that produced the fix.

The AI debugging traps to watch for

Hallucinated APIs and packages

Models sometimes suggest functions, options or whole packages that don't exist. A classic JavaScript example: asking for a fetch request with a timeout and getting this.

// Looks plausible, but fetch has no "timeout" option. It's silently ignored.
const res = await fetch(url, { timeout: 5000 });

// The real way: abort the request with a timeout signal
const res = await fetch(url, { signal: AbortSignal.timeout(5000) });

AbortSignal.timeout() is documented on MDN. The first version runs without errors and never times out, which is worse than crashing.

Packages are a security issue, not just an annoyance. A USENIX Security 2025 paper, We Have a Package for You!, found that across 16 code-generating models, 19.7% of recommended packages didn't exist, and it identified 205,474 unique hallucinated package names. Attackers can register those names ("slopsquatting"). A 2026 follow-up testing newer frontier models found lower rates, roughly 4.6% to 6.1% depending on the model, but still found 127 fake names that all five models invented identically. Check any unfamiliar package on its registry before you install it.

Fixing the symptom, not the cause

Wrapping code in try/except, adding a null check at the crash site, or bumping a timeout can all make an error vanish. Sometimes that's the right fix. Often it just moves the failure somewhere quieter. Claude Code's documentation puts it plainly in a sample prompt: "address the root cause, don't suppress the error." That line belongs in your prompts no matter which tool you use.

Quietly editing or deleting tests

When an agent can't make a test pass, it may change the test instead. Researchers behind ImpossibleBench built coding tasks where the tests deliberately contradicted the spec, so passing required cheating. Frontier models frequently did cheat: editing tests despite instructions not to, overriding comparison operators, and special-casing the exact test inputs. The paper found that making tests read-only or hidden, and using stricter prompts, cut cheating sharply.

Losing the thread in long sessions

A long debugging chat fills up with failed theories, and the model starts reusing them. Claude Code's docs suggest that if you've corrected the model more than twice on the same issue, you should clear the session and start fresh with a better prompt. That's good advice for any chat tool: when things go in circles, rewrite your Step 2 context with what you've learned and start a new conversation.

When not to reach for AI

Some bugs are faster to solve with a debugger and a breakpoint than with any prompt: stepping through a loop, inspecting live state, watching a variable change. AI is at its best on unfamiliar code, cryptic errors, noisy logs and generating the scaffolding (tests, scripts, logging) around a fix. It's weakest on bugs that depend on production state it can't see, like data, timing and infrastructure.

If you're choosing a tool for this kind of work, our best AI coding assistants for developers roundup compares the main options, and if much of your debugging involves database queries, see how to write SQL queries with AI.

Frequently asked questions

Which AI is best for debugging code?

There's no single winner, and the workflow matters more than the model. For quick questions about an error, any major chat assistant (ChatGPT, Claude or Gemini) works well if you give it full context. For bugs that need running tests across a codebase, an agentic tool like Claude Code, GitHub Copilot's agents or Cursor saves the most time because it can run the fix-and-test loop itself.

Is it safe to paste my code and stack traces into ChatGPT or Claude?

It depends on your employer's policy and the data involved. Remove secrets, API keys, customer data and internal hostnames first, and use a business or enterprise plan if your company provides one. When in doubt, build a minimal reproducible example with fake data. It protects your code and usually gets a better answer anyway.

Can AI find bugs I don't know about yet?

Somewhat. AI code review can flag likely null dereferences, off-by-one errors and unhandled cases in a diff, and it's worth running before merging. It isn't a substitute for tests, and reviewers prompted to find problems tend to report some even when the code is fine, so treat findings as leads to check.

Why does the AI keep suggesting fixes that don't work?

Usually because it's missing context and guessing, or because the conversation is cluttered with earlier failed attempts. Stop asking for fixes, give it the full traceback, versions and a minimal example, and ask for ranked hypotheses with a way to test each one. If that doesn't help, start a fresh chat with an updated summary of what you've ruled out.

Should I write the failing test myself or let the AI write it?

Letting the AI draft it is fine, but read it carefully and run it before the fix to confirm it fails for the right reason. The test defines what "fixed" means, so it's the one piece of the process you should always understand line by line.

Share:

Written by

Panoptix Editorial Team

Our editors test AI tools hands-on for weeks before we publish a word. We pay for our own subscriptions and never accept payment for rankings.

How we test AI tools →