open source · apache-2.0
Vim for
AI agents
the Neovim Agent Interface
neovain lets a coding agent edit a file with Vim's commands, a whole change per call. Every step applies, or none do.
curl -fsSL https://neovain.dev/install.sh | sh
Linux and macOS. It offers to install Neovim if you don't have it. Windows and other ways to install
$ neovain app.py '@^def load' 'wciwread_file<Esc>' -def load(path): +def read_file(path): data = open(path).read() return data $ neovain app.py '@open(' 'dd' FAILED at step 1 "@open(": anchor matched 2 lines (2,6): make the pattern more specific or use @N@ cursor was on line 1; file unchanged
One command instead of retyping the file
We gave Claude a 2,350-line Python module and six structural changes: delete a 400-line class, move another to the end, swap two more, wrap a 122-line method body in a lock, rename a method used 191 times, and remove the debug helpers. Same model, same task, two ways to edit. These are the actual calls from one pair of benchmark runs.
760characters written by the agent
$ neovain work.py \'@def sync_all(' \wrap sync_all in a lock':.+2,/^ return total$/>' \'@def sync_all(' \$':+1a\n with self._lock:\n.' \':%s/\<log_event\>/emit_event/g' \rename, 191 uses':g/^def debug_/.,/^\S/-1d' \delete the debug helpers'@^class LegacyExporter' \delete LegacyExporter':.,/^class EventStore/-1d' \'@^class ReportBuilder' \move ReportBuilder to the end':.-2,/^def rank_users/-3m$' \'@^class CacheLayer' \swap the two classes':mark a' \'@^class SessionManager' \":.,/^def sample_events/-1m 'a-1"
The agent previewed the change with --dry-run, then ran it for real (shown). Anchors are highlighted.
65,379characters written by the agent
- #1replace every log_event( with emit_event(21
- #2replace 19 lines with 1626
- #3replace 316 lines with 111,727
- #4replace 6 lines with 32112,021
- #5replace 420 lines with 115,612
- #6replace 123 lines with 12410,785
- #7replace 260 lines with 37,988
- #8replace 4 lines with 2176,599
String replacement spells out the text it finds and the text it adds, so moving a class means writing all of it twice.
- output tokens
- 13.8× fewer
- Mean of Claude Opus 5.5 and Sonnet 5.5, 3 runs each
- wall time
- 6.7× faster
- From prompt to finished file
- cost
- 4.5× cheaper
- API list price per task
Benchmarks
Every number here comes from real, headless agent sessions in Claude Code and Codex CLI, at medium reasoning effort, with neovain 0.2.0. Each session gets a fresh copy of the file and may edit it only one way. An automated checker then verifies every requested change and that nothing else moved. Methodology, raw data and the harness are in the repository.
Large structural edits
The six-change refactor above, on the 2,350-line module, in Claude Code. Lower is better.
Output tokens
per task
Wall time
per task, prompt to finished file
Cost
per task at API list price
| model | tool | exact | code correct | output tokens | API requests | tool calls | wall time | cost |
|---|---|---|---|---|---|---|---|---|
| Opus 5.5 | Edit tool | 3/3 | 3/3 | 31,790 | 19.7 | 24.7 | 229s | $1.43 |
| Opus 5.5 | neovain | 3/3 | 3/3 | 1,889 | 7.0 | 6.0 | 29s | $0.29 |
| Sonnet 5.5 | Edit tool | 3/3 | 3/3 | 28,713 | 17.7 | 23.3 | 155s | $0.73 |
| Sonnet 5.5 | neovain | 3/3 | 3/3 | 2,689 | 7.0 | 7.7 | 28s | $0.18 |
Each run gets two scores. Exact means the file matches the expected one line for line, blank lines included. Code correct means the code matches and only blank lines differ.
Small edits: close to a tie
On a 70-line file with seven small edits, the two tools are close. Every run passed. neovain wrote fewer tokens and cost two to four cents more. This is where string replacement is at its best: the text to retype is short, so there is little for neovain to save.
| model | tool | exact | code correct | output tokens | API requests | tool calls | wall time | cost |
|---|---|---|---|---|---|---|---|---|
| Opus 5.5 | Edit tool | 3/3 | 3/3 | 1,751 | 4.0 | 8.7 | 17s | $0.17 |
| Opus 5.5 | neovain | 3/3 | 3/3 | 1,075 | 3.3 | 2.3 | 15s | $0.21 |
| Sonnet 5.5 | Edit tool | 3/3 | 3/3 | 1,862 | 3.3 | 9.7 | 13s | $0.10 |
| Sonnet 5.5 | neovain | 3/3 | 3/3 | 1,457 | 3.7 | 4.0 | 16s | $0.12 |
Codex CLI: it depends on the edit
The same tasks in OpenAI's Codex CLI, where the comparison is against Codex's own patch tool. That tool does not retype the text it moves, so there is less to save. On the large task neovain still used about 40% fewer output tokens and a third less time, over the six models with valid runs in both arms, and got the code right in 13 of 18 runs against 11. On small edits the patch tool is clearly better: well under half the output and less than half the time. The weaker models got the large task wrong with either tool.
Large structural edits
| model | tool | exact | code correct | output tokens | tool calls | wall time |
|---|---|---|---|---|---|---|
| GPT-6 Astra | Patch tool | 3/3 | 3/3 | 923 | 7.7 | 43s |
| GPT-6 Astra | neovain | 3/3 | 3/3 | 671 | 9.7 | 37s |
| GPT-6 Sol | Patch tool | 3/3 | 3/3 | 3,372 | 13.3 | 85s |
| GPT-6 Sol | neovain | 2/3 | 2/3 | 2,352 | 16.3 | 62s |
| GPT-6 Luna | Patch tool | 0/3 | 0/3 | 3,095 | 7.3 | 91s |
| GPT-6 Luna | neovain | 1/3 | 1/3 | 3,967 | 12.3 | 108s |
| GPT-5.6 Sol | Patch tool | 2/3 | 2/3 | 5,998 | 10.3 | 132s |
| GPT-5.6 Sol | neovain | 3/3 | 3/3 | 2,726 | 8.7 | 68s |
| GPT-5.6 Terra | Patch tool | 0/3 | 1/3 | 7,343 | 10.0 | 151s |
| GPT-5.6 Terra | neovain | 1/3 | 1/3 | 3,479 | 7.3 | 74s |
| GPT-5.6 Luna | Patch tool | 2/3 | 2/3 | 8,403 | 14.3 | 169s |
| GPT-5.6 Luna | neovain | 3/3 | 3/3 | 4,380 | 9.7 | 94s |
| GPT-5.5 | Patch tool | left out: all 3 runs broke the rules | ||||
| GPT-5.5 | neovain | 3/3 | 3/3 | 3,764 | 19.7 | 83s |
Small edits
| model | tool | exact | code correct | output tokens | tool calls | wall time |
|---|---|---|---|---|---|---|
| GPT-6 Astra | Patch tool | 3/3 | 3/3 | 681 | 3.0 | 29s |
| GPT-6 Astra | neovain | 3/3 | 3/3 | 716 | 6.7 | 35s |
| GPT-6 Sol | Patch tool | 3/3 | 3/3 | 906 | 5.0 | 29s |
| GPT-6 Sol | neovain | 3/3 | 3/3 | 1,820 | 6.7 | 43s |
| GPT-6 Luna | Patch tool | 2/3 | 2/3 | 740 | 3.0 | 23s |
| GPT-6 Luna | neovain | 2/3 | 2/3 | 2,614 | 9.0 | 66s |
| GPT-5.6 Sol | Patch tool | 3/3 | 3/3 | 938 | 3.0 | 25s |
| GPT-5.6 Sol | neovain | 3/3 | 3/3 | 2,081 | 4.3 | 51s |
| GPT-5.6 Terra | Patch tool | 3/3 | 3/3 | 929 | 3.0 | 25s |
| GPT-5.6 Terra | neovain | 3/3 | 3/3 | 1,989 | 4.0 | 49s |
| GPT-5.6 Luna | Patch tool | 3/3 | 3/3 | 1,136 | 3.7 | 29s |
| GPT-5.6 Luna | neovain | 3/3 | 3/3 | 4,861 | 8.0 | 105s |
| GPT-5.5 | Patch tool | 3/3 | 3/3 | 1,172 | 4.0 | 30s |
| GPT-5.5 | neovain | 3/3 | 3/3 | 2,843 | 7.0 | 60s |
How the models behaved
Agents do not always do what they are told. Codex has no switch to turn a tool off, so its rules were set in the prompt, and every run log was checked afterwards for edits made the wrong way. Those runs are left out of the tables above and counted here, across every batch we ran.
| what happened | runs |
|---|---|
| Changed the file with the wrong tool These runs are left out of every table. | Claude, own tool: 0 of 30 Claude, neovain: 0 of 90 Codex, own tool: 3 of 60 Codex, neovain: 0 of 162 |
| Wrong code A requested change was missing, incomplete or damaged other code. | Claude, own tool: 0 of 30 Claude, neovain: 2 of 90 Codex, own tool: 14 of 57 Codex, neovain: 17 of 162 |
| Right code, wrong blank lines Most often two extra blank lines at the end of the file, after moving a class there. | Claude, own tool: 0 of 30 Claude, neovain: 0 of 90 Codex, own tool: 3 of 57 Codex, neovain: 21 of 162 |
| Threw neovain's output away On the large task: sent it to /dev/null, filtered it, or kept only a few lines. 0.1.0 printed a 2,000-line diff there; 0.2.0 prints a summary. | Claude, neovain 0.1.0: 25 of 30 Claude, neovain 0.2.0: 1 of 6 Codex, neovain 0.1.0: 4 of 60 Codex, neovain 0.2.0: 0 of 21 |
| Wrote a script to work out its patch Within the rules: the agent still applied the patch with its own tool. It helps explain the low token counts. | Claude, own tool: 0 of 30 Codex, own tool: 7 of 57 |
Not in these counts: in an early pilot, Claude Haiku 4.5 gave up on neovain in one of three runs and wrote the file with shell commands.
Habits we saw
- Throwing the output away. neovain used to print the whole diff, about 2,000 lines on the large task. Claude sent it to
/dev/nullor filtered it in 25 of 30 large runs, and that is how a wrong edit got through. Codex did so in 4 of 60. - One call per change. Before the guidance said otherwise, Codex made a neovain call for each change, twelve in one run. Claude chained its steps without being told.
- Doubled backslashes. Codex wrote
\\<word\\>on its first try at a rename, which matches nothing. neovain refused the call and the next one was right.
What helped
- Saying how, not only what. Telling agents to mind the blank lines changed nothing, so the guidance now gives a method. With it, Codex's blank-line misses on the large task fell from 6 runs to 2 of 21, though runs with wrong code went from 3 to 5. For Claude it cut output.
- A summary in place of a long diff. Since 0.2.0 a large change prints a summary of at most 80 lines: the blocks moved, deleted and re-indented, and warnings about spacing. Claude threw the output away in 1 of 6 large runs. Codex's blank-line misses went from 2 of 21 runs to none, though its runs with wrong code stayed at 5.
- All-or-nothing calls. Every call that failed left the file untouched, and the agent recovered on its next call.
- Checking the logs. The first version of our rule check missed a model that rewrote the file with Python. An edit count of zero gave it away.
What we learned
Reach for neovain when
- A change moves, deletes or re-indents large blocks. Claude's Edit tool retypes them; Vim's
:m,:dand>don't. - The same change repeats across a file:
:%sand:gdo it in one step. - You want all-or-nothing edits. A failed step leaves the file untouched, and agents recover on the next call.
Stick with the agent's own tool when
- The change is small and local. Writing the new text directly is the reasoning; with Vim, the agent has to work out the commands too. Codex's patch tool needed well under half the output on small edits.
- The model is small. In an early pilot, Claude Haiku 4.5 needed about three times as long with neovain and failed one run in three.
Tool calls are not round trips. Current models send independent edits as parallel tool calls, so cutting the number of calls saves little. What does cost time and money is output: every token the model writes. neovain wins where it lets the agent say less.