Anthropic API · ReAct · Agent architecture
I built a self-debugging AI agent from scratch - here's exactly how the reasoning cycle works
No LangChain. No agent framework. Just the raw Anthropic API and a Python loop - and what that teaches you about every framework you will ever use.
I built a self-debugging agent that takes a repository with failing tests and autonomously finds and fixes the bug. Not using LangChain, AutoGPT, or any agent framework - just the raw Anthropic API and about 200 lines of Python.
The reason: I wanted to actually understand how AI agents work, not just call a library and hope for the best. If you have ever looked at an "AI agent" and thought but what is it actually doing? - this article is for you.
What is a ReAct agent?
ReAct (Reasoning + Acting) is a pattern introduced in a 2022 paper by Yao et al. The core idea is elegantly simple: give a language model tools it can call, and let it alternate between three steps until it is done:
- THOUGHT - the model reasons about what it knows and what it needs to find out
- ACTION - it calls a tool to gather information or make a change
- OBSERVATION - it reads the result and updates its understanding
Then it loops. Each THOUGHT is informed by every OBSERVATION that came before it. The context window becomes a working memory that accumulates evidence across iterations.
This sounds simple. The implementation is where it gets interesting.
The core problem: the API has no memory
Here is the thing most people miss: Claude has no memory between API calls. Every single call is completely stateless. Send a message, get a response - the model has no idea what happened in the previous call.
So how does an agent that runs for 10 iterations remember what it tried in iteration 3?
The answer is almost anticlimactic: you keep a list and send the entire thing on every call.
self.messages: list[dict] = [] # this IS the agent's memory
After three iterations, this list looks like:
[
{ role: "user", content: "Debug fixtures/01_off_by_one, tests are failing..." },
{ role: "assistant", content: [TextBlock("I'll run the tests first"), ToolUseBlock("run_tests")] },
{ role: "user", content: [tool_result("1 failed: assert find_max([1,2,3,4,5]) == 5, got 4")] },
{ role: "assistant", content: [TextBlock("The test shows find_max returns 4 instead of 5..."), ToolUseBlock("read_file")] },
{ role: "user", content: [tool_result("4: for i in range(len(items) - 1): # BUG")] },
{ role: "assistant", content: [TextBlock("Found it. range(len(items) - 1) skips the last element"), ToolUseBlock("edit_file")] },
{ role: "user", content: [tool_result("Successfully replaced 1 occurrence in buggy_code.py")] },
]
Every iteration appends two things to this list: the assistant's thought plus tool call, and the tool's result. The next API call sends the whole list. The model reads the entire transcript from scratch and responds as if it remembers - but it does not. It just re-reads the conversation you handed it.
The context window is the agent's working memory. Nothing more.
How tool calls actually work at the API level
When you pass tools=TOOL_DECLARATIONS to the Anthropic API, you are sending Claude a list of JSON schemas describing what actions are available. Claude reads the descriptions to decide which tool to use. Vague descriptions lead to bad tool choices - this is prompt engineering, whether you think of it that way or not.
{
"name": "read_file",
"description": "Read the contents of a file. Optionally specify a line range. Use this to inspect code you suspect contains the bug.",
"input_schema": {
"type": "object",
"properties": {
"path": {"type": "string"},
"start_line": {"type": "integer"},
"end_line": {"type": "integer"},
},
"required": ["path"],
},
}
The API response has a content field that is a list of blocks, not a single string. Two types matter:
- TextBlock - the model's reasoning in plain English (the THOUGHT)
- ToolUseBlock - which tool to call and what arguments (the ACTION)
response.content = [
TextBlock(text="The test fails on the last element. Let me look at the loop bounds..."),
ToolUseBlock(id="toolu_abc123", name="read_file", input={"path": "buggy_code.py"}),
]
The TextBlock comes first - reasoning before acting. That is the "Reason" in ReAct. Then comes the action.
The contract you must not break
Here is the part that took me a debugging session to understand properly.
When Claude returns a ToolUseBlock, the API establishes a strict contract: before your next API call, you must send back a tool_result block whose tool_use_id matches the id in the ToolUseBlock. If you do not, you get a 400.
The conversation structure the API enforces:
Turn 1: user → task description
Turn 2: assistant → [thought text] + [tool_use id="abc" name="run_tests"]
Turn 3: user → [tool_result tool_use_id="abc" content="1 failed..."]
Turn 4: assistant → [thought text] + [tool_use id="def" name="read_file"]
Turn 5: user → [tool_result tool_use_id="def" content="line 4: range(len..."]
This broke for me when Claude returned two ToolUseBlocks in a single response - list_files and read_file in the same turn. My code only sent back one tool_result. The API rejected it immediately.
The fix: collect every ToolUseBlock from the response, execute every one, return all results in a single user message.
# WRONG: only handles the last tool call
for block in response.content:
if block.type == "tool_use":
tool_call = ToolCall(name=block.name, input=block.input, id=block.id)
# RIGHT: collect all of them
tool_calls: list[ToolCall] = []
for block in response.content:
if block.type == "tool_use":
tool_calls.append(ToolCall(name=block.name, input=block.input, id=block.id))
Claude may request parallel tool calls when it can gather multiple pieces of information at once. Every single one needs a matching result in the next message, or the entire request fails.
The loop structure in full
Here is the complete loop, simplified. This is the entire agentic pattern:
# 1. Seed the conversation with the task
messages.append({"role": "user", "content": "Debug this repo, tests are failing..."})
while iteration < max_iterations:
# 2. Ask Claude: given everything you've seen, what next?
response = client.messages.create(
model="claude-sonnet-4-6",
tools=TOOL_DECLARATIONS,
messages=messages, # the FULL history every time
)
# 3. Parse the response into thought + tool calls
thought = ""
tool_calls = []
for block in response.content:
if block.type == "text":
thought += block.text # THOUGHT: the reasoning
elif block.type == "tool_use":
tool_calls.append(block) # ACTION: what to do
# 4. Stop if the agent declares success or failure
if any(tc.name == "finish" for tc in tool_calls):
break
# 5. Execute every tool call
results = [(tc.id, execute(tc)) for tc in tool_calls]
# 6. Append everything to history
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": [
{"type": "tool_result", "tool_use_id": id, "content": result.output}
for id, result in results
]})
# 7. Loop - Claude now sees the new observations and reasons again
That is it. Seven steps. Every agent framework you will ever use - LangChain, the Claude Agent SDK, AutoGen - is some variation of this loop with more scaffolding around it.
What the agent actually did
I gave the agent a Python file with an off-by-one bug: range(len(items) - 1) instead of range(len(items)). The tests were failing because the loop skipped the last element.
The agent never saw the bug description. It had to find it.
In six iterations it:
- Ran tests, and saw
assert find_max([1,2,3,4,5]) == 5failing with value4 - Listed files, and identified
buggy_code.pyas the source - Read the file, and saw
range(len(items) - 1)with no surrounding context - Formed the hypothesis: the loop iterates indices 0 through n-2, skipping the last element
- Called
edit_filewith the fix - Called
finishwith an accurate root cause summary
Nobody told it that range(len(items) - 1) skips the last element. It connected two facts: the test fails on the last element, and the loop stops one short of the last element.
That connection happened inside a single THOUGHT block - plain English reasoning, visible in the terminal, formed by reading the test failure and the source code in sequence. The context window was doing what a developer's working memory does.
Why not just ask "what's wrong with this code?" in one shot?
You could. For a small, self-contained file, a one-shot prompt would likely find the same bug.
But the agent pattern scales to things a single prompt cannot handle:
- The bug is in file B, but you only know it exists because test A failed - you need to navigate there
- The fix requires reading context - imports, constants, how functions call each other
- Verification matters - apply the fix, then run the tests again to confirm
- Errors teach - if
edit_filefails because the string was not found exactly, the agent reads the error and tries differently
A single prompt gets one shot. The agent gets as many shots as you give it iterations. Each shot is informed by every observation that came before it.
The execution sandbox
One last piece: when the agent calls run_tests or edit_file, those calls need to execute somewhere safely. I used Docker.
The agent's edits need to persist across tool calls within a run - it cannot apply a fix and then have run_tests see the original file. So the setup is:
- Copy the repo to a temporary directory at the start of a run
- Start a Docker container with that directory mounted at
/repo - File reads and writes go directly to the tmpdir on the host
- Command execution (
pytest, arbitrary shell commands) goes throughdocker exec - Both see the same files because the tmpdir is
/repovia the volume mount
The container runs with --network none so the agent cannot make outbound HTTP requests, and --memory 512m to prevent runaway resource usage. At the end of the run, success or failure, the container stops and the tmpdir is deleted.
The clean separation is the point: the ReAct loop itself does not change between Phase 1 (fake stubs) and Phase 2 (real Docker). Only what execute_tool() calls at the bottom changes. The reasoning cycle is completely independent of the execution environment.
What I learned
The ReAct pattern is not magic. It is a while loop with a list. The intelligence is entirely the language model's - the loop just gives it a structure to gather information before committing to an answer.
The hardest part is not the loop. It is the details:
- Every
tool_useneeds a matchingtool_resultin the very next message - Parallel tool calls are real and your code must handle all of them
- Tool descriptions are prompt engineering - write them carefully
- The context window fills up - you need to truncate long observations before they blow the token budget
Build it from scratch once and you will understand every agent framework you ever use.
I build agentic tooling for regulated environments through Trustflux Ltd. This loop is the foundation the rest of the portfolio sits on - md-to-jira, azure-cost-investigator and spec-check are all the same cycle with a narrower tool set and a stricter contract.
Written with Claude Code. The code is the artefact; the article is the receipt.
Try it
The full agent, the Docker sandbox and the test fixtures are in the repo. mindaugasnakrosis/ai-react-loop →