Most agent tutorials teach tool use the same way: define a tool, Claude calls it, you return the result, Claude thinks, Claude calls the next one. It works, and it is also the reason agents that edit files feel slow and burn tokens.
Programmatic Tool Calling (PTC) changes one thing. Instead of calling your tools one at a time, Claude writes a short Python script that calls them, and the script runs in Anthropic’s code execution sandbox. Your tools still run on your side. But the loop, the branching, and the intermediate data all live in the script, not in the model’s context window.
This post builds the smallest agent I could think of that shows the difference: an agent that generates and edits an HTML page.
The agent
The agent gets one job: take a hand-written index.html that contains a single placeholder card, and turn it into a finished blog index. That means cloning the card for every post in posts.json, filling in titles and links, and then “repainting” the theme by swapping a few CSS variables.
It has three tools, all of them boring on purpose:
| Tool | What it does | Returns |
|---|---|---|
list_files | Lists the files in the site folder | JSON array of paths |
read_file | Reads one text file | The full file contents as a string |
write_file | Overwrites one text file | A one-line confirmation |
No search_file, no replace_line, no insert_after. That is the first thing PTC buys you. You stop writing micro-tools to compensate for the fact that the model can only take one step at a time.
Without PTC: the round-trip tax
Here is what the same task looks like with plain tool use, where every call is a separate trip through the model.
- Claude calls
read_file("index.html"). The whole file lands in context. - Claude calls
read_file("posts.json"). All the post data lands in context. - Claude now has to produce the new page. It has two bad options.
- Option A: rewrite the whole file. It emits the entire new HTML as output tokens in one
write_filecall. A modest page is 30 to 40 KB, so that is roughly 10,000 output tokens for what is conceptually a copy-paste job. Every later tweak means emitting the whole file again. - Option B: surgical edits. You give it
search_fileandreplace_rangetools, and it makes one round trip per edit. Eight posts plus a theme swap is easily 10 to 15 round trips. Each one re-reads or re-sends context, adds a second or two of latency, and gives the model another chance to lose its place.
- Option A: rewrite the whole file. It emits the entire new HTML as output tokens in one
Either way, the model is doing string manipulation with its attention instead of with code. You are paying frontier-model prices for str.replace.
With PTC: one script, a handful of calls
With PTC enabled, Claude looks at the same task and writes something like this. This is the script Claude generates and runs inside the sandbox, not code you write:
import re, json, asyncio# Both reads are independent, so run them in parallel.html, posts_raw = await asyncio.gather( read_file({"path": "index.html"}), read_file({"path": "posts.json"}),)posts = json.loads(posts_raw)# 1. Use the designer's single placeholder card as the template.card = re.search(r'<article class="card">.*?</article>', html, re.S).group(0)# 2. Clone it once per post.cards = "\n".join( card.replace("{{title}}", p["title"]) .replace("{{summary}}", p["summary"]) .replace("{{href}}", p["url"]) for p in posts)html = html.replace(card, cards)# 3. Repaint: swap the theme tokens in one pass.html = (html .replace("--accent: #3b82f6", "--accent: #d97706") .replace("--bg: #ffffff", "--bg: #0b0b0f") .replace("--fg: #111827", "--fg: #f5f5f4"))print(await write_file({"path": "index.html", "content": html}))print(f"rendered {len(posts)} cards, applied dark theme")
Three tool calls total: two reads and one write. Zero of the intermediate data touches Claude’s context. The only thing the model reads back is the two print lines:
wrote index.html (38214 bytes)rendered 8 cards, applied dark theme
That is the whole idea. Cloning, templating, and repainting are loops and string operations. Code is good at those. Let the code do them.
How it actually works on the wire
The mechanics are simple once you see the shape.
- You include the
code_executiontool in your request and mark your own tools as callable from code withallowed_callers. - Claude writes a script and the API starts running it in a sandbox.
- The moment the script calls one of your tools, the sandbox pauses and the API returns a normal
tool_useblock to you, tagged with acallerthat says it came from code. - You run the tool and send back a
tool_result. The script resumes. - When the script finishes, its stdout is what Claude reads. Then Claude writes its reply.
Inside the sandbox your tool shows up as an async Python function. It takes one dict of arguments and returns the string you sent back in the tool_result. That is why the script above calls await read_file({"path": ...}) and then json.loads the result.
The client code
This is the part you write. It is a standard tool-use loop with two additions: the allowed_callers field, and passing the container ID back so the paused script can resume.
import jsonfrom pathlib import Pathimport anthropicclient = anthropic.Anthropic()ROOT = Path("site").resolve()def _safe(path: str) -> Path: p = (ROOT / path).resolve() if ROOT not in p.parents and p != ROOT: raise ValueError(f"refusing to touch {path}: outside site folder") return pHANDLERS = { "list_files": lambda a: json.dumps( [str(p.relative_to(ROOT)) for p in ROOT.rglob("*") if p.is_file()] ), "read_file": lambda a: _safe(a["path"]).read_text(), "write_file": lambda a: ( _safe(a["path"]).write_text(a["content"]), f"wrote {a['path']} ({len(a['content'])} bytes)", )[1],}CODE_EXEC = "code_execution_20260120"tools = [ {"type": CODE_EXEC, "name": "code_execution"}, { "name": "list_files", "description": "List every file in the site folder. Returns a JSON array of relative paths.", "input_schema": {"type": "object", "properties": {}}, "allowed_callers": [CODE_EXEC], }, { "name": "read_file", "description": "Read one text file from the site folder. Returns the full file contents as a plain string.", "input_schema": { "type": "object", "properties": {"path": {"type": "string", "description": "Relative path, e.g. index.html"}}, "required": ["path"], }, "allowed_callers": [CODE_EXEC], }, { "name": "write_file", "description": "Overwrite one text file in the site folder with the given content. Returns a one-line confirmation.", "input_schema": { "type": "object", "properties": { "path": {"type": "string"}, "content": {"type": "string"}, }, "required": ["path", "content"], }, "allowed_callers": [CODE_EXEC], },]def run(task: str) -> str: messages = [{"role": "user", "content": task}] container = None while True: kwargs = {"container": container} if container else {} resp = client.messages.create( model="claude-opus-5", max_tokens=16000, tools=tools, messages=messages, **kwargs, ) messages.append({"role": "assistant", "content": resp.content}) if resp.container: container = resp.container.id # required while a script is paused pending = [b for b in resp.content if b.type == "tool_use"] if resp.stop_reason != "tool_use" or not pending: return "".join(b.text for b in resp.content if b.type == "text") results = [] for block in pending: try: out = HANDLERS[block.name](block.input) results.append({"type": "tool_result", "tool_use_id": block.id, "content": out}) except Exception as e: results.append({ "type": "tool_result", "tool_use_id": block.id, "content": f"Error: {e}", "is_error": True, }) # When answering a programmatic call this message must be tool_result blocks only. messages.append({"role": "user", "content": results})if __name__ == "__main__": print(run( "index.html has one placeholder <article class='card'> with {{title}}, " "{{summary}} and {{href}} tokens. Render one card per entry in posts.json, " "then switch the page to a dark theme with an amber accent." ))
Three lines carry the whole feature:
{"type": "code_execution_20260120", "name": "code_execution"}turns the sandbox on."allowed_callers": ["code_execution_20260120"]on each tool tells Claude to call it from code rather than directly.container=containeron follow-up requests lets the paused script pick up where it stopped.
Everything else is the loop you already have.
What you get back
The response for a programmatic call looks like a normal tool call with one extra field:
{ "type": "tool_use", "id": "toolu_def456", "name": "read_file", "input": { "path": "index.html" }, "caller": { "type": "code_execution_20260120", "tool_id": "srvtoolu_abc123" }}
The caller.tool_id points at the server_tool_use block that holds the script, so you can log which script made which call. When the script finishes you get a code_execution_tool_result block with stdout, stderr, and return_code, followed by Claude’s text.
When PTC helps and when it does not
PTC is not free. There is a fixed cost for spinning up the container and for the tokens Claude spends writing the script. It pays off when the work has one of these shapes:
- Fan-out. Same operation over many items: clone a card 8 times, check 50 links, fetch 20 records.
- Big intermediate data. The raw tool output is large but the useful part is small. A 40 KB HTML file where the answer is “wrote it, 8 cards.”
- Conditional chains. Read, decide, act, repeat, where the decision is mechanical enough for code to make.
It does not pay off when every step needs the model’s judgment before the next one. Anthropic’s own numbers make this concrete. On a 75-tool project-management benchmark, PTC cut billed input tokens by about 38% with no accuracy change. On a benchmark where each turn made one or two strictly sequential calls, PTC left scores unchanged and cost about 8% more. If your agent’s calls are “one at a time, think in between,” leave it off.
The rough test: if you can imagine the work as a for loop, PTC will help.
Gotchas worth knowing before you ship
- Tool results only. When you answer a programmatic call, the user message must contain only
tool_resultblocks. No text, not even after the results. The API rejects it otherwise. - Strings only. The
contentof each result must be a string or text blocks. No images or documents. - Pass the container ID. If a script is paused and you omit
container, the request fails. - Send the same
toolsarray on the continuation. The code execution tool has to still be there or the script cannot resume. - Four-minute clock. If your tool takes longer than about 4 minutes to answer, the call raises a
TimeoutErrorinside the script. Claude sees it in stderr and usually retries. - Describe your output format. Say “returns a JSON array of objects with fields x, y” in the tool description. Claude parses results in code, so the more it knows about the shape, the less it guesses.
- Not a security boundary.
allowed_callersguides Claude. It does not hard-block a direct call, so your handler still has to validate paths and inputs. That is what the_safehelper above is for. - Incompatibilities. No
strict: trueon PTC-enabled tools, no forcing a programmatic tool throughtool_choice, nodisable_parallel_tool_use, and MCP connector tools cannot be called from code. - Availability. Works on Claude Opus 5, Sonnet 5, Fable 5.1 and the 4.5+ family on the Claude API and Claude Platform on AWS. Not on Amazon Bedrock or Google Cloud, and not eligible for zero data retention.
The takeaway
Tool use gave Claude hands. PTC gives it a scripting language for those hands. For an agent that edits files, that is the difference between “read, think, replace, think, replace” and “here is a 20-line script, run it.” Fewer tools to build, fewer round trips to wait on, and a context window that holds the summary instead of the raw material.
If you already have a tool-use loop, enabling PTC is a two-field change. Try it on the next task that looks like a for loop and measure the token count before and after.
Further reading: the Programmatic tool calling docs and the Advanced tool use engineering post.
Leave a comment