Read a running AI agent without stopping it.

Debuggers answer questions by freezing the program. That works on a checkout screen. Do it to an agent and the model request times out, the session drops, and the bug you were chasing is gone. aidebugger traps capture and continue — real values out of a live process that never pauses.

Bring your own Python bug

Upload a .py file or paste your code. Run it in your browser, inspect the actual error and captured values, then use a connected AI model to propose a repair. Every repair runs again; add your own checks to verify the behavior and download the corrected file.

Open the debugging workspace

Single-file Python scripts with the standard library. Running code needs no AI connection; automatic repairs require the site's repair service.

Run it, right here, on a real Python process

Pyodide is CPython 3.13 compiled to WebAssembly, and sys.monitoring came with it — so the engine that runs on your laptop runs unmodified in a browser tab. Arm a trap on the checkout pricing code, apply the coupon, read the values that come back. Nothing is recorded and nothing is served: the frames are captured microseconds earlier, in your own tab.

Open the live console

Tick apply the fix and press apply again: the same armed trap re-reads the repaired run, base_amount goes from 565.64 to 628.49, and the total falls from $638.72 to $511.44. That is the whole loop, with no install and no server.

What a breakpoint does to an agent

Same moment, same bug, two tools. The difference isn't convenience — one of these destroys the evidence.

Breakpoint hit

a stepping debugger
  • 0.0sexecution stops, locals readable
  • 4syou start reading the frame
  • 10smodel request times out
  • 18ssession teardown, call dropped
  • —state you wanted is gone

Trap fires

aidebugger
  • 0.0slocals captured, ~18µs
  • 0.0sfunction returns normally
  • —agent keeps talking
  • lateryou poll the buffer, unhurried
  • —bug reproduced 400 calls in

An agent finds the bug, fixes it, and proves it — five times

A checkout screen rejects a coupon it should accept: the agent attaches to the running backend, arms traps, triggers the failing request, reads the values, writes the two-line fix and re-runs the same trap — $638.72 down to $511.44, with the app serving throughout. Then the same loop on a LiveKit call filed against the wrong person, a LangGraph agent that answers with filler, an OpenAI Agents handoff that loses the customer's price cap, and a Pydantic AI refund that pays out a hundred times over. Every screen and every frame is a real capture.

Tap a phone, read the backend

A checkout app on a real phone-sized screen, talking to a Python backend with a pricing bug. Change the coupon or the shipping toggle: the frames beside it are what the trap captured from that exact request, while the backend carried on serving.

Open the phone demo

Two functions in one request disagree about what the customer spent — compute_discount sees 628.49, evaluate_promo is handed 565.64. The source line reads fine; only the live values give it away.

Same loop, on a voice agent

A LiveKit agent takes a call about somebody else's account. Nothing errors, the caller is thanked, and the wrong account is attached — because one matcher only accepted a bare word and quietly defaulted to self. Toggle before and after the fix.

Open the voice demo

The caller said "for my mother". normalise_relationship returned "self", and every step after it was confidently wrong. The transcript reads perfectly.

Ten agents, one console

Recorded runs this time, not live ones — the real console replaying real captures — each captured by running an example from the repo under aidebugger and arming that framework's trapset by name. Every adapter is here through both doors, sync and async, because they do not fire the same symbols — and the three framework agents from the video are here mid-bug, so you can open the frame that gives each one away. Pick a run from the menu at the top right, click any row for its locals, then click a non-primitive to expand the object behind it.

Nothing here is hand-written, and the pairs are the point. LangChain fires five of nine armed symbols synchronously and the other four under ainvoke. Runner.run never fires when the OpenAI Agents SDK is entered through run_sync — 5 of 6 — while Pydantic AI's run_sync funnels back through AbstractAgent.run and fires all four either way. scripts/record_timeline.py regenerates all six framework recordings, no API keys.

What it costs to leave armed

Traps are armed per function, not globally, so code you didn't trap runs at full speed. Measured on Python 3.13.

0.94×untrapped code — indistinguishable from no traps
~18nsadded per call to a trapped function
0times the process pauses
0dependencies outside the standard library

Install it

Python 3.12 or newer — the design rests on sys.monitoring, which does not exist before that. No dependencies outside the standard library.

From the repository

Works today. uv pip install takes the same URL.

pip install git+https://github.com/riyadadlani02/aidebugger

With the MCP server, so your coding agent can drive it

Adds the aidebugger-mcp entry point alongside the CLI.

pip install "aidebugger[mcp] @ git+https://github.com/riyadadlani02/aidebugger"

Or install nothing at all

The live console runs the real engine in your browser, on WebAssembly.

riyadadlani02.github.io/aidebugger/live/

Not on PyPI yet — the release workflow publishes on a v* tag once the name is claimed there.

Four commands

No code changes to the thing you're debugging. The launcher puts aidebugger on the path; the agent under test never imports it.

Start the process under aidebugger

Or set PYTHONPATH for a container you don't launch yourself.

aidebugger run -- python my_agent.py

Arm a trap while it runs

A dotted symbol, or a framework hook name so you don't hunt for line numbers.

aidebugger trap --trapset livekit --hook on-tool-call

Only capture the case you care about

A predicate over the function's own locals — for the bug that shows up on call 400.

aidebugger trap myapp.pricing.apply_discount --when "base_amount < coupon.min_spend"

Read what happened

Returns immediately with a cursor. Nothing blocks, so an AI agent can poll between other work — or open the console in a browser.

aidebugger poll 0

Framework hooks

Adapters are plain JSON mapping a hook name to real symbols — each one links to a recording of that adapter running. Every symbol below was watched firing against a running agent — resolving isn't enough, since a symbol can resolve and never be called, which is how an adapter rots silently when a framework moves a function. scripts/verify_trapset.py reproduces it with no API keys.

adapterhooksverified against
livekit on-tool-call, on-llm-request, on-handoff, on-user-turn livekit-agents 1.6.6, on a live production voice agent — 7/7
langchain on-agent-run, on-llm-request, on-tool-call, on-state-write langchain-core 1.6.2 + langgraph — 9/9 across sync + async
openai_agents on-agent-run, on-turn, on-llm-request, on-tool-call, on-handoff openai-agents 0.22.0 — 6/6 async, 5/6 via run_sync
pydantic_ai on-agent-run, on-llm-request, on-tool-call pydantic-ai 2.40.0 — 4/4 async and via run_sync

Where it stops short

Worth knowing before you point it at anything sensitive.