06 · File editing and patches: what survives a failed edit
Choose patch or minimal replacement semantics while handling stale files, partial failure, and reviewability.
Chapter brief
Question to answer
When a file changes or a patch half-applies, how do you keep editing detectable, reversible, and reviewable?
By the end, you can
- Define invariants for stale files, ambiguous matches, and partial failure
- Compare patch DSLs, minimal replacement, and whole-file rewrite
- Design failure feedback both the agent and reviewer can explain
- Read this now if
- Engineers implementing Edit, Patch, SEARCH/REPLACE, or code-modification tools
- Prerequisites
- Understand file versions, diffs, and basic concurrency conflicts
- Deliverable
- Patch-protocol contract tests and rollback cases
- Evidence boundary
- Protocols reduce ambiguity but cannot prevent wrong-location choices or logic bugs hidden by weak tests
What must stay true when an edit fails
Section titled “What must stay true when an edit fails”Scenario: an agent reads config.ts and prepares a replacement while the user edits the file. The stale fragment still appears twice; the patch hits the wrong location and only half-applies. Retrying repeats the successful half.
Passing conditions: edits carry a file-version or content-hash precondition; ambiguous matches refuse to guess; a change set commits atomically or records a reversible partial failure; the final diff exposes actual changes to the reviewer rather than trusting a “success” return value.
Define the failure boundary first
Section titled “Define the failure boundary first”How each system covers the four steps (express / validate / persist / feedback):
| Step | Codex | Claude Code | OpenClaw | Hermes |
|---|---|---|---|---|
| Expression DSL | V4A inline patch DSL (5 marker types: Begin/End/Update/Add/Delete/Move) | str_replace: old_string + new_string + replace_all | Generic fs.read / fs.write / fs.edit tool group | V4A inline patch DSL (Python `tools/patch_parser.py` reimplementation) |
| Atomicity | Parse/context failure rejects the whole patch; disk apply is not a cross-file transaction | One edit per call. N changes = N tool calls | One file per write, multi-file = multi-call | V4A parser can reject the whole patch; disk failure still needs recovery |
| Validation point | Parser stage (`apply-patch/src/parser.rs`) + filesystem state recheck | `FILE_UNEXPECTEDLY_MODIFIED_ERROR`: pre-write mtime compare | `tool-fs-policy.workspaceOnly` middleware | V4A parser check; disk failures handled separately |
| Failure recovery | Parser error returned to model. Model rewrites the patch | Permission denied → deny tool_result. Model reads the error and retries | Hook blocks call. Standard tool_result error | Parse error + model rewrites |
| Side effects | rollout/* persistence + execpolicy + sandbox | LSP diagnostics invalidate + fileHistory + VS Code SDK notify | Tool event stream + session lane | memory commit + trajectory event |
Compare only implementations that change atomicity or review
Section titled “Compare only implementations that change atomicity or review”Codex · Designs a dedicated patch DSL V4A, lets the model inline whole patches in the assistant message
Section titled “Codex · Designs a dedicated patch DSL V4A, lets the model inline whole patches in the assistant message”Codex’s starting point on file editing is: models are actually very good at generating unified diffs (the format used in GitHub / GitLab PR diffs that they have seen millions of in training), but git’s standard unified diff has several features that don’t quite fit the agent scenario.
It needs precise line numbers + line counts (context 5 lines before / 5 after), which the model is prone to miscount; it can’t express “rename a file” in one diff (has to be split into Delete + Add); and it assumes the reviewer and patch author are at the same git version (requires fuzz matching to handle offsets).
So Codex decides to invent a patch DSL specifically optimised for agents, called V4A (V for “version”, A probably for “apply” or “agent”), keeping unified diff’s readability but removing the parts agents tend to get wrong, and adding file-level semantics (add, delete, move).
V4A doesn’t go through function-call JSON arguments; instead the model inline outputs the whole patch in the assistant message (wrapped in special markers).
Codex’s message parser sees the marker and intercepts the whole patch text, passing it to the apply-patch crate for processing.
Putting the patch in assistant output moves the capacity boundary away from function-call arguments. Both limits depend on provider and model, so fixed token ranges are misleading.
The Rust apply-patch crate uses Lark grammar (a parser-generator language similar to EBNF) to define V4A’s complete grammar:
Codex codex/codex-rs/apply-patch/src/parser.rs:1-22 Lark grammar for V4A
//! The official Lark grammar for the apply-patch format is://!//! start: begin_patch hunk+ end_patch//! begin_patch: "*** Begin Patch" LF//! end_patch: "*** End Patch" LF?//!//! hunk: add_hunk | delete_hunk | update_hunk//! add_hunk: "*** Add File: " filename LF add_line+//! delete_hunk: "*** Delete File: " filename LF//! update_hunk: "*** Update File: " filename LF change_move? change?//! filename: /(.+)///! add_line: "+" /(.+)/ LF -> line//!//! change_move: "*** Move to: " filename LF//! change: (change_context | change_line)+ eof_line?//! change_context: ("@@" | "@@ " /(.+)/) LF//! change_line: ("+" | "-" | " ") /(.+)/ LF//! eof_line: "*** End of File" LFReading this grammar reveals several key V4A designs. The whole patch is wrapped by *** Begin Patch / *** End Patch, so the parser can intercept the whole patch from anywhere in the assistant message (even if the model wrote explanatory text before or after the patch, parsing is unaffected).
Hunks split into three classes: add_hunk (create a new file, every line prefixed with +); delete_hunk (delete an entire file, only one line *** Delete File:); update_hunk (modify an existing file, with optional change_move for rename + a change block containing edits).
Most critically, the change-block format is almost identical to unified diff (+ prefix for additions, - prefix for deletions, space prefix for context), but with line numbers and counts removed (the parser locates changes via context lines, instead of asking the model to count line numbers). change_context uses @@ function_name @@ anchors to help the parser locate (in case context lines are too short and might match multiple places). eof_line is the *** End of File marker, telling the parser the change extends to file end.
V4A’s first trade-off is putting the patch in assistant output rather than function arguments. Both paths have provider- and model-specific limits, but the boundaries differ; measure the current API before claiming one can carry a larger diff.
Second, multi-file submission: one patch can include any combination of Update + Add + Delete + Move. During parsing, apply-patch rejects the whole patch if any hunk fails and no writes happen; the later disk-apply phase can still hit permissions, locks, or a full disk, so this is not a cross-file transaction.
Third, the patch text is itself a readable diff. Rollout persistence stores the raw patch text; replay reapplies verbatim; the audit reviewer can review all agent changes like a PR diff.
The cost is of course that the model has to learn this DSL. Codex teaches V4A in the system prompt (with a few examples for in-context learning), but even so gpt-4.1 occasionally writes wrong format (missing a space, missing a @@ anchor), so Codex’s parser has a ParseMode::Lenient mode (used for non-gpt-4.1 models with strict mode); common format errors (extra whitespace, missing anchors) are proactively fixed by the parser.
Claude Code · Break every edit to smallest unit str_replace, let reviewer see each smallest diff
Section titled “Claude Code · Break every edit to smallest unit str_replace, let reviewer see each smallest diff”Claude Code scopes edits to smaller tool calls, which suits stepwise IDE display and review. “Twenty files in one turn” is a synthetic scenario here, not a product limit; the real distinction is one concentrated patch versus a sequence of located edits.
FileEditTool edits one string at a time, with minimal params:
Claude Code claude-code/src/tools/FileEditTool/types.ts:1-30 FileEdit three params: old_string / new_string / replace_all
inputSchema: z.object({ file_path: z.string(), old_string: z.string().describe('The text to replace'), new_string: z .string() .describe( 'The text to replace it with (must be different from old_string)', ), replace_all: z.boolean().default(false).describe( 'Replace all occurrences of old_string (default false)', ),})This “minimal three params” design has several careful engineering considerations. old_string must match uniquely in the file.
If the same string appears multiple times and replace_all is false, the tool refuses and asks the model to supply more context to make old_string unique; this forces the model to Read the file before Edit, to see the context (multiple identical strings usually mean the model doesn’t understand the file structure well enough).
If the user really wants bulk replace (e.g. renaming a variable across the whole file), they can pass replace_all=true to replace all occurrences at once. new_string must differ from old_string (clearly stated in the zod schema describe), otherwise the operation is meaningless.
Before writing there is a hidden critical check: compare the file’s mtime (modification time).
If the mtime changed since reading, some other process (user manually editing in IDE, git pull, other agent) just modified this file, and FileEditTool refuses the write and throws FILE_UNEXPECTEDLY_MODIFIED_ERROR, forcing the model to Read again before Edit; this mechanism prevents the catastrophic race condition where “the model edits based on stale content and clobbers what someone else just wrote”.
Each FileEdit call also cascades through Claude Code’s full side-effect network.
LSP diagnostics invalidation (clearDeliveredDiagnosticsForFile() tells the LSP to re-analyze this file’s syntax / types / lint); file history tracking (fileHistoryTrackEdit() writes each change into the session’s history record, users can view all agent changes with /diff);
VS Code SDK notification (notifyVscodeFileUpdated() triggers VS Code editor to refresh the opened tab so users see the latest content); permission check (checkWritePermissionForTool() runs through the full permission mode system, with different behaviour for acceptEdits / plan / bypassPermissions / default).
This “every edit triggers full side-effect network” makes the IDE experience extremely smooth: the agent edits the file, VS Code refreshes immediately, LSP re-analyses immediately, error hints update immediately; the cost is that each edit pays this overhead.
The cost of course is token burn: one location per call, big refactors mean many Edit calls, and each Edit’s tool-call context repeats the same things (file_path / old_string / new_string).
Claude Code 2.1.88 doesn’t ship MultiEdit; the older batch-edit tool was folded in.
The team probably decided MultiEdit makes the model easy to dump too many changes for users to keep up with, hurting reviewer experience; they prefer letting Edit be called more times.
OpenClaw · Don’t invent a DSL for coding; make file editing a generic fs tool + workspaceOnly policy
Section titled “OpenClaw · Don’t invent a DSL for coding; make file editing a generic fs tool + workspaceOnly policy”OpenClaw’s judgement on file editing is: it is itself an agent control plane (not a coding tool); coding is just one workload among many (users may use OpenClaw to write Slack bots, customer support agents, data analysis agents; these scenarios don’t need to edit files at all); inventing a dedicated patch DSL for one scenario is wrong; fs operations should go through the generic tool stack with constraints from policy middleware.
The actual implementation hangs fs.read / fs.write / fs.list ordinary read/write tools under the fs category of tool-catalog.ts (interface fully consistent with Node.js fs module, familiar to the model); no dedicated editing protocol (no V4A, no str_replace).
Constraint is all in one boolean field of tool-fs-policy.ts:
OpenClaw openclaw/src/agents/tool-fs-policy.ts:1-32 tool-fs-policy: one switch, workspaceOnly
export type ToolFsPolicy = { workspaceOnly: boolean;};
export function createToolFsPolicy(params: { workspaceOnly?: boolean }): ToolFsPolicy { return { workspaceOnly: params.workspaceOnly === true, };}
export function resolveEffectiveToolFsWorkspaceOnly(params: { cfg?: OpenClawConfig; agentId?: string;}): boolean { return resolveToolFsConfig(params).workspaceOnly === true;}With workspaceOnly: true, any path outside the session workspace is rejected. The plugin pipeline enforces the rule in the before_tool_call hook; see Tool System.
The tradeoff: OpenClaw is a control plane. Editing files is one workload among many, so it does not invent a DSL for one workload. The cost is no atomic multi-file semantics.
Multi-file edits become multi-call sequences, and any failure mid-way leaves consistency to the caller.
Hermes · Directly reuses Codex’s V4A format, makes patches a cross-ecosystem common interface
Section titled “Hermes · Directly reuses Codex’s V4A format, makes patches a cross-ecosystem common interface”Hermes’ judgement on file editing is: do not invent another format when a patch shape is already usable in the sampled tools. Codex and Hermes both implement V4A, and other tools in the cited ecosystem reuse that shape; compatibility still needs to be tested per parser and model.
That’s a waste of ecosystem.
So Hermes’ tools/patch_parser.py is a Python reimplementation of V4A, and the docstring states the compatibility intent very directly:
Hermes hermes-agent/tools/patch_parser.py:1-29 patch_parser.py: V4A reused across the coding-agent ecosystem
"""V4A Patch Format Parser
Parses the V4A patch format used by codex, cline, and other coding agents.
V4A Format: *** Begin Patch *** Update File: path/to/file.py @@ optional context hint @@ context line (space prefix) -removed line (minus prefix) +added line (plus prefix) *** Add File: path/to/new.py +new file content +line 2 *** Delete File: path/to/old.py *** Move File: old/path.py -> new/path.py *** End Patch"""Why pick V4A over inventing a new format? Two practical reasons. First, models already understand it.
V4A already has public implementations and examples, so Hermes can reuse a parser and prompt examples. Public examples do not establish what a model saw during training; output reliability still needs model-specific evaluation.
Second, cross-ecosystem portability: users coming from codex / cline can reuse the same V4A shape in Hermes (the patch format is identical); a Hermes-generated patch can also paste into codex or other V4A-compatible tools (e.g. shipped to teammates for review or applied in CI). The amount of retraining depends on the model and prompt.
One key difference from Codex: Hermes wraps the patch as a function-call argument (the patch string is one of the tool input fields, e.g. apply_patch({patch: "*** Begin Patch\n..."})), instead of inlining it in the assistant message text.
This decision has trade-offs. It loses V4A’s original “dodging function-args token limits” advantage (giant patches still get capped by JSON args size); but gains simpler Python-side protocol handling (no need to scan the assistant message for *** Begin Patch, just read the function-call arguments directly), and gets clean visualisation in tool-call UIs (function calls render as structured cards in chat UIs; inline DSL would have to be specially parsed and styled).
This trade-off reflects Hermes’ positioning. It’s not a pure coding agent (also runs general-purpose conversation, browser, etc.); maintaining a uniform tool-call protocol is more important than coding-specific optimisation.
Protect atomicity and reviewability first
Section titled “Protect atomicity and reviewability first”Looking across the four file-edit implementations, four review questions emerge. They are useful checks for a coding workflow, not rules every agent implements in the same way:
Read before write: carry current-file evidence when the protocol supports it. Codex’s update_hunk requires @@ context @@ anchors; Claude Code requires a unique old_string and checks mtime; Hermes likewise needs context lines. OpenClaw’s workspace policy constrains paths but does not force a preceding Read in this snapshot, so the host must decide how to handle stale content.
Validate before write where possible. Codex’s parser checks patch legality on apply; Claude Code checks mtime for out-of-band modification races; Hermes runs V4A parser validation; OpenClaw validates workspace path and write permission in its policy pipeline. Early rejection can make retry explicit, but it does not remove disk-write failure paths.
Parse rejection is atomic; disk application may not be. V4A can reject all hunks before writing when parsing or context matching fails. Once disk writes begin, permissions, locks, or a full disk can still leave partial effects unless the caller provides a transaction or compensation. Claude Code’s single-edit scope narrows that blast radius but does not create cross-file atomicity.
Diff-like output is a useful feedback channel. The four snapshots expose a diff or change record in different places; returning it to the model and UI can make the applied result reviewable. The comparison does not establish a universal editing standard, so preserve enough evidence for the host’s audit and replay needs.
Put recovery cost beside the choice
Section titled “Put recovery cost beside the choice”The four systems’ divergences on file editing answer four different questions, and which one to follow depends on what scenario your agent is in.
If you want one turn to refactor 20 files: borrow from Codex’s V4A path. A single tool call carries all changes simultaneously, and the parser rejects the patch during parsing when a hunk fails; disk application still needs its own error path. The model can write large diffs without function-args token ceilings, subject to the context limit. The price is learning V4A’s grammar, so test it with the model and prompt you actually plan to run. The path suits large refactors, dependency upgrades, and migrations that need cross-file consistency.
If you want external reviewers to scrutinise every change line by line: borrow from Claude Code’s str_replace path. Each edit is the smallest unit (one location only), each tool call generates an independent minimal diff, reviewers can scroll through agent decisions one by one (no surprise of “20 files changed in one shot”). The price is big refactors burn tokens (many Edit calls), but for IDE-integrated scenarios this trade-off is worth it: IDE users care more about “I see every step the agent takes” than “the agent finished in one shot”. Bonus: the full side-effect network (LSP, file history, VS Code SDK) makes edits feel “alive” rather than “the agent did things behind my back”.
If your agent is a control plane and editing is incidental: borrow from OpenClaw’s generic fs path. Don’t invent a dedicated patch DSL for one scenario; use the generic fs.read / fs.write tools, with constraints handled by the policy middleware (workspaceOnly + allowlist). The price is no atomic multi-file semantics (model handles consistency itself), but the gain is platform generality. File editing tools also apply to non-coding scenarios (config file modification, log writing, data export), no need to maintain two sets of file operation APIs.
If you want compatibility with the codex / cline patch ecosystem: borrow from Hermes’s path of reusing V4A wrapped in function-call arguments. Patches can be reused across tools (Hermes-generated patches paste into codex, codex patches apply in Hermes); function-call wrapping integrates with tool-call UIs. The price is the function-args token ceiling, so giant patches still need splitting; test the limit in your provider and model setup.
Choose by whether a failed edit can roll back
Section titled “Choose by whether a failed edit can roll back”| Editing constraint | Route to borrow | Cost or boundary |
|---|---|---|
| One request must express several files and validate before writing | Codex V4A patch grammar | Disk application is not a cross-file transaction |
| Small diffs and race detection matter most | Claude Code str_replace plus mtime checks | Large refactors spend more turns |
| A generic agent only needs workspace-scoped fs | OpenClaw workspaceOnly policy | No multi-file atomic guarantee |
| Compatibility with existing V4A tooling matters | Hermes patch parser | Function-call arguments cap very large patches |
Start with one editor that can refuse
Section titled “Start with one editor that can refuse”When building a file-edit tool, make locations unambiguous, writes recoverable, and results reviewable first. Add multi-file patches and IDE side effects only when the workflow needs them.
Build recipe
Minimum viable
- Start with str_replace accepting only three args (old_string / new_string / file_path). It is a small starting point without DSL parsing; make single-file editing and failure reporting observable before adding more complex semantics
- Require old_string to match uniquely in the file (borrow from Claude Code). Multiple matches error out forcing the model to add more context; this constraint forces the model to first Read for precise context before Edit, avoiding "blind edit" misalignment
- Compare file mtime before write to catch races (borrow from Claude Code's FILE_UNEXPECTEDLY_MODIFIED_ERROR). User editing in IDE simultaneously, another agent editing, git switching branch can all trigger; mtime mismatch refuses write forcing model to re-Read
- Return a diff after editing (not just success/fail) so model and user can both verify: model can confirm correctness from diff; user can see diff to decide rollback; diff is the core of "auditable"
Once that works
- Evaluate V4A when related hunks need one reviewable request. Parse and context validation can reject the whole block; disk application still needs an isolated worktree, compensation, or an explicit partial-failure report.
- Parse V4A with Lark grammar or three-segment regex (Begin / hunks / End), reject whole patch on failure. Lark is more readable / extensible than regex; any line marker mismatch rejects the whole patch (don't attempt partial application, leaves inconsistent state)
- Add fs-policy middleware (borrow from OpenClaw's workspaceOnly). Restrict paths to within workspace (preventing model from accidentally editing ~/.bashrc or /etc/...); this is the filesystem layer's safety baseline
- On success trigger LSP re-analysis, file history persistence, editor notification (borrow from Claude Code's side-effect network). Successful editing isn't the endpoint; let IDE see changes, git history record it, other agents see notifications; the side-effect network done well makes IDE experience smooth
Don't do this
- Letting the model run sed / awk through bash: no diff feedback (user can't see what changed), errors can't be located (can't roll back to pre-edit state), and the model often writes sed syntax wrong (easy to skip individual cases with -i); use a dedicated Edit tool
- Using a line-number range as the only edit locator. File changes or context trimming can invalidate line numbers; verify with old_string, context anchors, or a content hash
- Stuffing a large patch into function-call arguments without testing capacity and truncation. Limits vary by provider and model; whether inline or structured, reject incomplete patches before any disk write
- Skipping mtime / hash checks. Two agents editing same file simultaneously, user editing in IDE then overwritten by agent, all produce "edits lost / mutual overwrite" incidents; mtime is the cheapest defense
Where the two editing paths split
Section titled “Where the two editing paths split”Side by side: V4A lets the model state every change once and uses a Lark parser as gatekeeper. str_replace asks the model to ship the smallest unit per call and uses uniqueness + mtime as the gatekeeper.
Neither lets the model edit blindly, but the paths could not be more different.
Source paths for atomicity and conflicts
Section titled “Source paths for atomicity and conflicts”What to carry forward and the next experiment
Section titled “What to carry forward and the next experiment”Safe editing depends on preconditions, unique location, atomicity, and an auditable result. A patch DSL matters because conflict, partial failure, and rollback become explicit state—not because the syntax looks elegant.
Next experiment: trigger four failures on one edit task: stale file, two matches, second-hunk failure, and process exit after write. Verify restoration, no duplicate successful hunks, a complete final diff, and feedback that causes a reread rather than blind retry.
Appendix: exercises and review
Section titled “Appendix: exercises and review”Open the exercises and ten review questions
Exercises
Section titled “Exercises”- 🟢 Build str_replace: implement
file_path / old_string / new_string. Enforce thatold_stringmatches uniquely in the file; otherwise error. Return a diff. - 🟠 V4A parser: implement a minimal V4A subset in your favorite language (only
*** Update File+ add/delete/context lines). Verify your parser handles at least one test case fromapply-patch/tests/suite/scenarios.rs. - 🟠 mtime check: extend exercise 1 with mtime validation. Simulate two processes editing one file; the second write should hit
FILE_UNEXPECTEDLY_MODIFIED_ERROR. - 🔴 Cross-system compatibility: feed your V4A parser the test patches from Codex
apply-patch/tests/and Hermespatch_parsertests. Which cases diverge?
Review questions
Section titled “Review questions”Q1 · Concept: V4A inline DSL vs normal function-call tools: what’s the real difference?
V4A is Codex’s patch description language. The model emits the entire patch in assistant text (not tool_use args), and the harness parses it via a Lark grammar. Three real differences:
1. Protocol location. Function call arguments go through tool_use.input (JSON-wrapped); V4A goes through assistant.content text alongside natural-language output. The former passes through Anthropic / OpenAI’s protocol serializer; the latter skips it.
2. Capacity boundary. Both tool_use.input and assistant output are capped, with limits varying by provider, model, and SDK. V4A moves the patch to another output path; it does not make it unbounded.
3. Error recovery path. Function-call errors are protocol-layer (bad arg, schema mismatch); V4A errors are text-format errors that the model can simply re-emit, no tool_use restart needed.
Why choose different protocols? V4A fits multiple hunks in one reviewable patch; str_replace fits stepwise, uniquely located edits. The source does not establish universal limits of 20 files or 100 lines.
Practical: choose between str_replace and V4A by edit structure, conflict risk, and measured payload behaviour. A fixed 1K token boundary does not transfer across models, providers, or SDKs.
Source: codex/codex-rs/apply-patch/src/parser.rs:1-22 (Lark grammar); hermes-agent/tools/patch_parser.py:1-29 (Python reimplementation).
Follow-up: “Is V4A a de-facto standard?” It appears in several implementations, but there is no standards body or compatibility specification. “A patch shape adopted by multiple coding agents” is more precise.
Q2 · Architecture: Claude Code’s str_replace forces unique old_string matches. Why? Can users turn it off?
Forced unique matching kills ambiguous edits. If old_string appears 3 times and the model says “change this to that,” the harness can’t tell which instance was meant.
Demanding either replace_all=true or enough context for unique match pushes the judgment onto the model where it belongs.
Why not let users toggle it off? Because “guess the model’s intent” is dangerous.
The current source confirms the unique-match constraint. It does not ship before/after error-rate samples, so this chapter does not claim a specific reduction.
Codex’s V4A solves it differently: every update_hunk carries 3 context lines (@@ context @@), and seek_sequence.rs finds the anchor via context uniqueness rather than string uniqueness.
If you implement str_replace:
- Default enforces unique. Low-level API never silently picks the first match.
- Provide
replace_allopt-in. Let the model explicitly say “all of them”. - Error includes line ranges: “old_string matches at lines 12, 47, 89; please add context to disambiguate.” The model reads it, then issues Read for more context.
Source: claude-code/src/tools/FileEditTool/FileEditTool.ts:1-130; codex/codex-rs/apply-patch/src/seek_sequence.rs.
Follow-up: “Why not have the model give line numbers?” Because line numbers drift in multi-turn: file is edited by another process, by a previous edit, by the user. Line-number protocols are fragile by design.
Q3 · Engineering: How does FILE_UNEXPECTEDLY_MODIFIED_ERROR work? Why not use fcntl file locks?
Implementation: each Read records the file mtime; each Edit stats before write, mismatch = error. Pseudocode:
const { mtime: readMtime } = await stat(file_path);trackFileRead(file_path, readMtime);
// later in Editconst { mtime: currentMtime } = await stat(file_path);if (currentMtime !== trackedMtimeFor(file_path)) { throw new Error('FILE_UNEXPECTEDLY_MODIFIED_ERROR');}// proceed to writeWhy not fcntl? Three reasons:
1. Locks miss out-of-band writers. Vim or VS Code can write without acquiring a lock; lock is meaningless. mtime is passive observation, every writer trips it.
2. Different concurrency model. Agent isn’t a database; “conflict” means “what I read is stale,” not “I want exclusive write.” mtime maps to optimistic concurrency control, the right semantic.
3. Cross-platform. Windows / macOS / Linux fcntl behavior differs; mtime is the POSIX + Windows common denominator.
Gotchas:
- The effective mtime resolution and the value exposed by the runtime vary by filesystem, platform, and API. Do not rely on mtime alone for closely spaced writes; pair it with a content hash or another version marker when the race matters.
- For multi-worker, store tracked mtimes in shared state. Claude Code is single-process, in-memory is fine.
Source: claude-code/src/tools/FileEditTool/utils.ts has findActualString + mtime details.
Follow-up: “Why not git hashes?” Could replace mtime. But hashing is slower (SHA-256 per check) and small edits may yield identical hashes for unchanged regions. Claude Code picked mtime for speed.
Q4 · Architecture: A V4A patch contains Update + Add + Delete. Where does atomicity stop?
Only the first phase has all-or-nothing rejection semantics; this is not a transactional two-phase commit:
Phase 1 · Parse + Validate. The whole patch parses; every hunk becomes a structured object in memory. Any parse failure rejects the entire patch, nothing reaches disk.
Phase 2 · Apply. After all hunks parse, disk writes can still fail because of permissions, locks, or a full disk. The parser does not make this phase a cross-file transaction; callers need isolation or explicit compensation.
Why not strict all-or-nothing? POSIX file systems don’t offer cross-file atomic writes. You can atomic-rename one file, not commit a group. Doing it properly requires:
- Write to temp directory / temp filename.
- After all succeed, rename each into place.
- On mid-failure, clean up temp dir.
Neither Codex nor Hermes implements this fully because:
- Complexity high: temp dir management, rename edge cases, cross-fs rename failures.
- Real demand low: phase 1 catches syntax and context errors; phase 2 can still fail on disk-full or permission errors, which require user intervention. This chapter has no failure sample from which to claim a percentage.
- Disposable environments can be discarded: an isolated worktree or temporary copy can be abandoned after failure. Do not use a broad reset in a worktree that may contain user changes.
Practical: start with phase 1 validation, then design recovery for phase 2. For cross-file atomicity, use an isolated worktree, temporary copy, or transactional file layer; do not rely on a broad git reset that could overwrite unrelated user changes.
Source: codex/codex-rs/apply-patch/src/lib.rs, the apply_patch_to_disk function.
Follow-up: “Is git apply better?” Its error output identifies the failed hunk; that is an observable format difference. Whether it helps a person or agent diagnose the failure more easily than V4A output must be tested in the target editing workflow, so this article does not assign a general usability ranking.
Q5 · Concept: What does “diff is the feedback format” mean? Why return a diff after every edit?
“Diff is the feedback format” means the edit tool’s result is not “ok” or a boolean. It’s the textual diff between before and after:
--- before+++ after@@ -10,3 +10,3 @@- const name = "foo"+ const name = "bar"Three reasons to return a diff:
1. Lets the model verify its own change. The model thought it was editing line 12, but the old_string match may have landed at line 47. Returning the diff lets it see “yes, that’s what I meant” or “no, roll back.”
2. Lets humans review. A diff shows the exact additions and removals, which is more reviewable than “edit success.” Claude Code exposes /diff to users; CI can also put the diff in a PR description.
3. Gives downstream tools (LSP / linter) a hook point. Diff triggers LSP diagnostics recompute, linter re-run, test runner re-run. With just “ok,” downstream doesn’t know what changed.
Implementation:
- Diff should be unified diff format (git-style), every programmer reads it.
- Make the returned-diff limit configurable. If truncated, preserve a path to the complete diff and include a per-file summary rather than replacing evidence with a summary.
- For multi-file patches (V4A), group by file with a diff block per file.
V4A bonus: rollout stores the patch text. Replay reuses the patch string directly, no diff re-derivation.
Source: claude-code/src/tools/FileEditTool/ has diff generation in utils.ts.
Follow-up: “Which diff algorithm?” Typically Myers diff (O(ND)); modern use patience diff. On Node, the diff package is standard. Codex uses Rust’s similar crate.
Q6 · Practical: Your agent must fix a bug touching 5 files. Design the edit tool schema and workflow.
Schema (V4A-compatible + str_replace fallback):
// Option A: V4A bulk patchinterface ApplyPatchInput { patch: string; // *** Begin Patch ... *** End Patch}
// Option B: str_replace single pointinterface FileEditInput { file_path: string; old_string: string; // must be unique new_string: string; replace_all?: boolean;}Route by structure: use FileEdit for one uniquely located change, and ApplyPatch for multiple files or a single reviewable patch. Line count is an observation, not a hard routing rule.
Workflow:
- Model reads every relevant file first. System prompt enforces: “Before fixing the bug, Read every involved file.”
- Model writes a think block describing root cause and plan. This lives in trajectory for later review.
- Model emits ApplyPatch or multiple FileEdits. Coordinated 5-file fix → one ApplyPatch; isolated 1-2 file fix → FileEdits.
- Agent runs tests / lint. On failure, trajectory carries the error; the model decides whether to roll back or continue.
- Generate PR description. Auto-write “what changed, why” from the think block and the diff.
Key decisions:
- Force Read first: skipping Read leaves the edit without reliable context and can widen the damage; this chapter has no local error-rate sample. Claude Code mandates this in its prompt; Codex enforces context anchors in V4A grammar.
- Add mtime to defeat races: if the user manually edits a file mid-5-file-change, the second edit should error immediately.
- Recovery boundary: run coordinated edits in an isolated worktree, temporary copy, or explicit snapshot. On test failure, retain the diff and error, then compensate or discard the isolated environment. A global stash can mix in unrelated user changes and is not a transaction.
Pitfalls:
- Measure the stable FileEdit size on the provider and model; split calls when they cross that boundary.
- If multi-file consistency matters, use an isolated worktree or snapshot. One ApplyPatch does not create a disk transaction.
- Don’t declare success before tests run (lazy verifier is a backstop).
Source: Codex’s goals.rs wires “run tests after edit” into the verifier. Claude Code uses stopHooks in query.ts for similar wiring.
Follow-up: “What about polyglot projects (Python backend + TS frontend)?” Each sub-project runs its own tests; failures merge into the trajectory. Codex’s run_tests auto-detects project languages.
Q7 · Architecture: Why doesn’t OpenClaw build an edit DSL? What are the consequences of staying generic?
OpenClaw is control-plane software. Its use case isn’t “a coding agent” but “an agent platform with many skills.” Skills may be coding, data analysis, customer support, scraping; building a DSL for one workload (coding edits) violates the generality goal.
Implementation: fs.read / fs.write / fs.edit live in the fs category of tool-catalog.ts as ordinary tools, with tool-fs-policy.ts providing a single boolean (workspaceOnly) for boundary.
No V4A, no str_replace protocol, no mtime check.
Consequences:
- Multi-file atomicity is fragile. No patch protocol means multi-file edits = multiple fs.write calls; intermediate failure leaves inconsistent state.
- Edit UX is weaker. The model must read then write; “write” usually means full-file overwrite (unless the tool supports diff-style edit, but default doesn’t).
- Coarser audit granularity. Trajectory shows “wrote file X,” not “changed which lines.”
- Forks can add it.
tool-policy-pipelineallowsbefore_tool_call/after_tool_callhooks for custom validation; V4A parsing can be plugged in.
Why this is acceptable: OpenClaw’s user is “agent platform user installing skills,” coding being one of many. A @coding-skill plugin can bring its own V4A protocol and str_replace tool; the OpenClaw kernel needn’t care.
Analogy:
- VSCode doesn’t bundle git; the git extension does. VSCode is control plane, git is skill.
- OpenClaw doesn’t bundle V4A; coding-skill does. Same design philosophy.
Practical: building an “agent platform”? Don’t bake DSLs for one workload into the kernel, make them skills/plugins. Building a “coding agent”? Go deep at V4A / str_replace level in the kernel.
Source: openclaw/src/agents/tool-fs-policy.ts:1-32 (the whole file is one boolean); openclaw/src/agents/tool-catalog.ts fs category.
Follow-up: “How does LangChain handle file edit?” LangChain has no built-in V4A; provides a file toolkit users hook themselves. Same control-plane philosophy as OpenClaw.
Q8 · Engineering: Hermes reuses V4A but “stuffs it as a function-call argument” instead of inlining. What’s the trade-off vs Codex?
Hermes: model puts a complete V4A string into the tool_use’s input.patch field; harness reads it and calls tools/patch_parser.py. Codex: model inlines V4A in assistant message text; harness scans for *** Begin Patch.
Trade-off:
| Dimension | Hermes (function arg) | Codex (inline) |
|---|---|---|
| Protocol complexity | Low (standard function call) | High (parses assistant text) |
| Size cap | Function-argument and output limits | Output limit |
| Model learning cost | Slightly lower (familiar function call) | Slightly higher (custom DSL) |
| Failure recovery | Standardized function-call errors | Needs custom “parse failed” return |
| Cross-model portability | All function-calling models | Depends on Anthropic / OpenAI allowing inline text |
Why Hermes picks function arg:
- Cross-model compatibility. Hermes targets OpenAI / Anthropic / Gemini. Every model supports function calling. Inline DSL in Gemini is awkward (its thinking mode mixes with text content).
- Simpler protocol handling. Python code does
result["patch"]in one line, easier than scanning assistant text.
Why Codex picks inline:
- Codex primarily runs on GPT, so hitting function-arg caps is normal.
- Rollout stores assistant message text, so inlined V4A lands in rollout directly. Replay needs no extra assembly.
- Codex needn’t be cross-model, locked to OpenAI, no portability concerns.
Practical:
- Multi-model agent → Hermes mode (function arg).
- Single-model + large-patch workload → Codex mode (inline).
- Unsure → start with function arg, switch when caps bite.
Source: hermes-agent/tools/patch_parser.py:1-29 (explicitly cites V4A reuse); hermes-agent/tools/file_tools.py (how the parser is called).
Follow-up: “What if Hermes hits the function-arg cap?” The model splits the patch into multiple tool_use calls (one per file). Atomicity goes, but it’s a working fallback.
Q9 · Concept: What is “the file-edit side-effect network”? Why does Claude Code make it so complete?
“Side-effect network” = the set of external systems an edit must update. Claude Code triggers four things after each edit:
- LSP diagnostics invalidate (
clearDeliveredDiagnosticsForFile). LSP server re-analyzes the file; next time the model asks for diagnostics, it gets fresh errors. - fileHistory tracks (
fileHistoryTrackEdit). Internal history table records “at time T, file X, diff Y.” The/diffcommand shows all in-session edits. - VS Code SDK notify (
notifyVscodeFileUpdated). If Claude Code runs as a VS Code extension, the editor refreshes the file (avoids stale display). - Transition reason write. On loop exit, transition gets
had_edits: trueso monitoring can distinguish “read-only session” from “modified session.”
Why so complete? Because agents aren’t islands. An edit isn’t just a file change; it impacts:
- Next turn’s context: stale LSP returns stale diagnostics.
- User’s visual perception: VS Code without notification shows old content while the agent thinks it’s new.
- Session-level retrieval: “What did you change earlier?” No fileHistory, no answer.
- CI / monitoring: without
had_edits, monitoring can’t categorize session types.
Codex / OpenClaw / Hermes do less:
- Codex: rollout write + execpolicy audit (~1.5 items).
- OpenClaw: tool event stream + session lane (~1 item).
- Hermes: memory commit + trajectory event (~1.5 items).
Claude Code does more because it positions as “IDE-native agent.” Deep editor integration mandates keeping IDE state consistent. Codex positions as “CLI / CI agent,” no editor to sync, only cares about rollout.
Practical: start with LSP-invalidate + fileHistory (2 items). Add VS Code notification when integrating an IDE. Add transition reason in production monitoring.
Source: claude-code/src/tools/FileEditTool/FileEditTool.ts:1-130 (everything after a successful edit).
Follow-up: “Shouldn’t LSP servers detect file changes themselves?” They usually do through file watchers. Delay varies by platform, load, and watcher configuration; measure active notification plus watcher fallback in the target IDE before claiming an ordering or latency guarantee.
Q10 · Open-ended: Design a layered protocol for file-editing tools.
A layered protocol borrowing selected boundaries from the four snapshots:
Layer 1 · Single-point edit (required)
interface SimpleEdit { file_path: string; old_string: string; // unique match required new_string: string; replace_all?: boolean;}Borrow Claude Code’s str_replace with uniqueness and mtime checks. It fits local, uniquely addressable edits; token count alone should not choose the tool.
Layer 2 · Large patch (optional)
interface BulkPatch { patch: string; // V4A format validate_only?: boolean; // dry-run}Borrow V4A. Lark grammar parser, phase 1 validate + phase 2 apply. When function-arg cap bites, fall back to SimpleEdit.
Layer 3 · Policy (required)
interface FsPolicy { workspace_root: string; forbidden_paths: string[]; allowed_extensions?: string[]; require_mtime_check: boolean; // default true}Borrow OpenClaw’s workspaceOnly + blacklist/whitelist. Run policy before every edit.
Layer 4 · Side effects (required in prod)
interface EditSideEffects { notify_lsp: boolean; track_history: boolean; notify_editor: boolean; // VS Code / Cursor / etc. emit_event: boolean; // for monitoring / audit}Borrow Claude Code’s side-effect network, each toggleable (small agents don’t need everything).
Layer 5 · Transition (required)
Every edit appends to trajectory:
interface EditOutcome { changed_files: string[]; diff: string; // unified diff bytes_changed: number; mtime_check_passed: boolean; side_effects_fired: string[]; error?: { code: string; message: string };}Monitoring aggregates on EditOutcome directly.
API example:
const editor = createFileEditor({ policy: { workspace_root: '/app', forbidden_paths: ['.env'] }, side_effects: { notify_lsp: true, track_history: true },});
await editor.simpleEdit({ file_path: 'src/foo.ts', old_string: '...', new_string: '...' });// orawait editor.bulkPatch({ patch: '*** Begin Patch ...' });Boundary of this composite: it puts V4A, mtime checks, policy, and edit side effects behind one interface. It does not reproduce any of the four runtimes, and this chapter has no benchmark showing that the result is lighter, deeper, or easier to maintain. Rollout storage and default side effects should follow recovery and audit requirements.
Effort depends on target languages, parser coverage, concurrent-write semantics, edit history, and the test matrix. There is no implementation record here from which to claim a fixed schedule.
Source: composite of the four implementations in this chapter. Follow-up: “Cross-language?” JSON Schema inputs and outputs plus a textual patch format let different languages implement the same protocol. Each implementation still needs its own parser and must pass the same success, ambiguity, and partial-failure cases.