2026-10-03
Killing Your Coding Agent
Coding agents are genuinely impressive right now. I can point one at a messy codebase, describe what I want in half-broken English, and watch it tear through files, write functions, refactor whole modules, and patch bugs faster than I can open the file. They will sit there for hours grinding away at a complicated task. It is the closest thing I have used to a tireless junior developer who never complains and somehow knows half the standard library by heart.
Which makes it almost funny how easy it can be to throw them off. Today, while I was working on a relatively large project of mine, I watched my coding agent crash—not once, not twice, but three separate times—while trying to patch a Python script. Same failure, same general area of the code. That did not look random.
So I dug in.
What I found was a ridiculously simple way to make Python source fail to parse: put a docstring-looking string directly beneath a function definition, but give it the wrong indentation. Python is strict about indentation; a line can look almost right to a person and still leave the parser unable to understand the function body. For an agent, that can mean it dies on the parse error, loops through retries, or produces a bad repair for the wrong problem.
I got tired of watching the same failures, so I wrote agent-pwned.py to demonstrate the dirty work. The name is tongue-in-cheek: the script modifies source files; it does not communicate with or exploit an agent. I am describing it as a warning and a defensive case study—not encouraging anyone to damage someone else’s codebase.
What the script actually does
In the version I published, the script accepts file and directory arguments. It processes Python files, recursively searches directories for .py files, looks for lines that appear to begin a function definition, and inserts a standalone string literal at the left edge of the following line. It writes the changed text back to the original path—which is exactly why this should only ever be tested on a disposable copy you own or are authorized to modify.
Python requires the body of a function to be indented. A string literal can be a perfectly valid docstring when it is indented as the first statement inside a function; placed at column zero immediately after a function header, it is not the function body Python expects. The resulting file can fail to parse with an indentation error, so tools that need to parse or import that module may stop before they can make the intended edit.
That explains the failure I was investigating, but it does not prove why every agent crash happens. My script does not communicate with a coding agent, exploit its runtime, or establish that every agent will crash. It changes source files; what an editor, agent, test runner, or build does next depends on what it reads and how it handles parse errors.
Why an agent can make the incident look stranger
When an agent is asked to patch a file, it may first inspect, parse, lint, import, or test the existing project. If one of those steps encounters malformed Python, the task can fail before the requested change is reached. An agent may then retry, attempt a repair, or report a confusing downstream error. Repeated failure around the same file is a useful clue: the source may already be broken, generated, partially written, or different from the version the developer believes is present.
This is not unique to coding agents. Compilers, language servers, linters, test frameworks, and human reviewers all depend on the integrity of their inputs. The agent adds another layer that can obscure the original error if its summary focuses on the attempted patch rather than the parser’s first failure.
The important distinction: malformed input versus model failure
A coding agent failing on a file does not, by itself, show that the model is defective or that an exploit against the agent has occurred. The repository demonstrates a way to modify Python source so it no longer parses. It does not establish that the technique bypasses an agent sandbox, gains access to another project, or affects a system without the files being changed.
Still, the operational risk is real when a developer or automated process runs a destructive transformation against a working tree. The script writes over the original file and, in the inspected version, does not create a backup, provide a dry-run mode, or validate the result with a parser before reporting the change. Applying it broadly can leave many modules unusable and cost time to identify and recover the edits.
What to do when an agent repeatedly fails on one file
- Stop retries long enough to inspect the input. Read the first syntax or indentation error from the interpreter, linter, or test output; later errors may only be consequences.
- Inspect the diff before editing again. Compare the failing file with version control and look for unexplained insertions, truncation, encoding changes, or whitespace changes.
- Check a clean copy. If the repository is under version control, compare the working tree with a trusted commit or a fresh checkout. Preserve legitimate uncommitted work before restoring anything.
- Validate the repaired source. Run the project’s normal syntax checks and focused tests before asking the agent to continue.
- Protect automated edits. Use a clean branch or disposable worktree, require reviewable diffs, and keep backups or recoverable version history before running bulk transformations.
Defensive lessons for agent workflows
Agents should treat parse and tool failures as evidence to investigate, not as a reason to blindly retry. A robust coding workflow can record the initial file state, inspect the first compiler diagnostic, limit retries, and show the proposed diff before accepting broad edits. Sandboxing and least privilege also matter: an agent should only be able to modify the workspace it is meant to handle, and destructive scripts should not be run on a shared or production checkout.
For developers, the practical defense is ordinary but effective: version control, small diffs, backups for bulk changes, and tests that run before and after automated edits. If a coding agent suddenly cannot parse a file it previously handled, investigate the file and the toolchain before blaming the model.
Bottom line
A badly placed string literal can make a Python module unparsable, and that can derail a coding task—including one being handled by an AI agent. The linked script demonstrates the mechanism by modifying source files in place. Its impact comes from the damaged inputs, not from any special control over the agent.
So if an agent keeps failing on the same file, pause and check what it is actually reading. The fastest fix may not be a cleverer prompt. It may be finding and safely restoring a malformed source file.
Source
DTRHnet/Shell-Scripts: agent-pwned.py
This article discusses defensive diagnosis and authorized testing. Do not use source-corruption tools on codebases you do not own or have explicit permission to modify.