open source · apache-2.0

Vim for
AI agents

the Neovim Agent Interface

neovain lets a coding agent edit a file with Vim's commands, a whole change per call. Every step applies, or none do.

curl -fsSL https://neovain.dev/install.sh | sh

Linux and macOS. It offers to install Neovim if you don't have it. Windows and other ways to install

app.py
$ neovain app.py '@^def load' 'wciwread_file<Esc>'
--- app.py
+++ app.py
@@ -1,3 +1,3 @@
-def load(path):
+def read_file(path):
     data = open(path).read()
     return data

$ neovain app.py '@open(' 'dd'
FAILED at step 1 "@open(": anchor matched 2
lines (2,6): make the pattern more specific
or use @N@
cursor was on line 1; file unchanged

One command instead of retyping the file

We gave Claude a 2,350-line Python module and six structural changes: delete a 400-line class, move another to the end, swap two more, wrap a 122-line method body in a lock, rename a method used 191 times, and remove the debug helpers. Same model, same task, two ways to edit. These are the actual calls from one pair of benchmark runs.

neovain2 calls

760characters written by the agent

  1. $ neovain work.py \
  2. '@def sync_all(' \wrap sync_all in a lock
  3. ':.+2,/^ return total$/>' \
  4. '@def sync_all(' \
  5. $':+1a\n with self._lock:\n.' \
  6. ':%s/\<log_event\>/emit_event/g' \rename, 191 uses
  7. ':g/^def debug_/.,/^\S/-1d' \delete the debug helpers
  8. '@^class LegacyExporter' \delete LegacyExporter
  9. ':.,/^class EventStore/-1d' \
  10. '@^class ReportBuilder' \move ReportBuilder to the end
  11. ':.-2,/^def rank_users/-3m$' \
  12. '@^class CacheLayer' \swap the two classes
  13. ':mark a' \
  14. '@^class SessionManager' \
  15. ":.,/^def sample_events/-1m 'a-1"

The agent previewed the change with --dry-run, then ran it for real (shown). Anchors are highlighted.

Edit tool8 calls

65,379characters written by the agent

  1. #1replace every log_event( with emit_event(21
  2. #2replace 19 lines with 1626
  3. #3replace 316 lines with 111,727
  4. #4replace 6 lines with 32112,021
  5. #5replace 420 lines with 115,612
  6. #6replace 123 lines with 12410,785
  7. #7replace 260 lines with 37,988
  8. #8replace 4 lines with 2176,599

String replacement spells out the text it finds and the text it adds, so moving a class means writing all of it twice.

output tokens
13.8× fewer
Mean of Claude Opus 5.5 and Sonnet 5.5, 3 runs each
wall time
6.7× faster
From prompt to finished file
cost
4.5× cheaper
API list price per task

Benchmarks

Every number here comes from real, headless agent sessions in Claude Code and Codex CLI, at medium reasoning effort, with neovain 0.2.0. Each session gets a fresh copy of the file and may edit it only one way. An automated checker then verifies every requested change and that nothing else moved. Methodology, raw data and the harness are in the repository.

Large structural edits

The six-change refactor above, on the 2,350-line module, in Claude Code. Lower is better.

Output tokens

per task

Opus 5.5
Opus 5.5, neovain: 1,889 output tokens, mean of 3 runs1,889
Opus 5.5, Edit tool: 31,790 output tokens, mean of 3 runs31,790
Sonnet 5.5
Sonnet 5.5, neovain: 2,689 output tokens, mean of 3 runs2,689
Sonnet 5.5, Edit tool: 28,713 output tokens, mean of 3 runs28,713

Wall time

per task, prompt to finished file

Opus 5.5
Opus 5.5, neovain: 29s wall time, mean of 3 runs29s
Opus 5.5, Edit tool: 229s wall time, mean of 3 runs229s
Sonnet 5.5
Sonnet 5.5, neovain: 28s wall time, mean of 3 runs28s
Sonnet 5.5, Edit tool: 155s wall time, mean of 3 runs155s

Cost

per task at API list price

Opus 5.5
Opus 5.5, neovain: $0.29 cost, mean of 3 runs$0.29
Opus 5.5, Edit tool: $1.43 cost, mean of 3 runs$1.43
Sonnet 5.5
Sonnet 5.5, neovain: $0.18 cost, mean of 3 runs$0.18
Sonnet 5.5, Edit tool: $0.73 cost, mean of 3 runs$0.73
modeltoolexactcode correctoutput tokensAPI requeststool callswall timecost
Opus 5.5Edit tool3/33/331,79019.724.7229s$1.43
Opus 5.5neovain3/33/31,8897.06.029s$0.29
Sonnet 5.5Edit tool3/33/328,71317.723.3155s$0.73
Sonnet 5.5neovain3/33/32,6897.07.728s$0.18

Each run gets two scores. Exact means the file matches the expected one line for line, blank lines included. Code correct means the code matches and only blank lines differ.

Small edits: close to a tie

On a 70-line file with seven small edits, the two tools are close. Every run passed. neovain wrote fewer tokens and cost two to four cents more. This is where string replacement is at its best: the text to retype is short, so there is little for neovain to save.

modeltoolexactcode correctoutput tokensAPI requeststool callswall timecost
Opus 5.5Edit tool3/33/31,7514.08.717s$0.17
Opus 5.5neovain3/33/31,0753.32.315s$0.21
Sonnet 5.5Edit tool3/33/31,8623.39.713s$0.10
Sonnet 5.5neovain3/33/31,4573.74.016s$0.12

Codex CLI: it depends on the edit

The same tasks in OpenAI's Codex CLI, where the comparison is against Codex's own patch tool. That tool does not retype the text it moves, so there is less to save. On the large task neovain still used about 40% fewer output tokens and a third less time, over the six models with valid runs in both arms, and got the code right in 13 of 18 runs against 11. On small edits the patch tool is clearly better: well under half the output and less than half the time. The weaker models got the large task wrong with either tool.

Large structural edits

modeltoolexactcode correctoutput tokenstool callswall time
GPT-6 AstraPatch tool3/33/39237.743s
GPT-6 Astraneovain3/33/36719.737s
GPT-6 SolPatch tool3/33/33,37213.385s
GPT-6 Solneovain2/32/32,35216.362s
GPT-6 LunaPatch tool0/30/33,0957.391s
GPT-6 Lunaneovain1/31/33,96712.3108s
GPT-5.6 SolPatch tool2/32/35,99810.3132s
GPT-5.6 Solneovain3/33/32,7268.768s
GPT-5.6 TerraPatch tool0/31/37,34310.0151s
GPT-5.6 Terraneovain1/31/33,4797.374s
GPT-5.6 LunaPatch tool2/32/38,40314.3169s
GPT-5.6 Lunaneovain3/33/34,3809.794s
GPT-5.5Patch toolleft out: all 3 runs broke the rules
GPT-5.5neovain3/33/33,76419.783s

Small edits

modeltoolexactcode correctoutput tokenstool callswall time
GPT-6 AstraPatch tool3/33/36813.029s
GPT-6 Astraneovain3/33/37166.735s
GPT-6 SolPatch tool3/33/39065.029s
GPT-6 Solneovain3/33/31,8206.743s
GPT-6 LunaPatch tool2/32/37403.023s
GPT-6 Lunaneovain2/32/32,6149.066s
GPT-5.6 SolPatch tool3/33/39383.025s
GPT-5.6 Solneovain3/33/32,0814.351s
GPT-5.6 TerraPatch tool3/33/39293.025s
GPT-5.6 Terraneovain3/33/31,9894.049s
GPT-5.6 LunaPatch tool3/33/31,1363.729s
GPT-5.6 Lunaneovain3/33/34,8618.0105s
GPT-5.5Patch tool3/33/31,1724.030s
GPT-5.5neovain3/33/32,8437.060s

How the models behaved

Agents do not always do what they are told. Codex has no switch to turn a tool off, so its rules were set in the prompt, and every run log was checked afterwards for edits made the wrong way. Those runs are left out of the tables above and counted here, across every batch we ran.

what happenedruns
Changed the file with the wrong tool
These runs are left out of every table.
Claude, own tool: 0 of 30
Claude, neovain: 0 of 90
Codex, own tool: 3 of 60
Codex, neovain: 0 of 162
Wrong code
A requested change was missing, incomplete or damaged other code.
Claude, own tool: 0 of 30
Claude, neovain: 2 of 90
Codex, own tool: 14 of 57
Codex, neovain: 17 of 162
Right code, wrong blank lines
Most often two extra blank lines at the end of the file, after moving a class there.
Claude, own tool: 0 of 30
Claude, neovain: 0 of 90
Codex, own tool: 3 of 57
Codex, neovain: 21 of 162
Threw neovain's output away
On the large task: sent it to /dev/null, filtered it, or kept only a few lines. 0.1.0 printed a 2,000-line diff there; 0.2.0 prints a summary.
Claude, neovain 0.1.0: 25 of 30
Claude, neovain 0.2.0: 1 of 6
Codex, neovain 0.1.0: 4 of 60
Codex, neovain 0.2.0: 0 of 21
Wrote a script to work out its patch
Within the rules: the agent still applied the patch with its own tool. It helps explain the low token counts.
Claude, own tool: 0 of 30
Codex, own tool: 7 of 57

Not in these counts: in an early pilot, Claude Haiku 4.5 gave up on neovain in one of three runs and wrote the file with shell commands.

Habits we saw

  • Throwing the output away. neovain used to print the whole diff, about 2,000 lines on the large task. Claude sent it to /dev/null or filtered it in 25 of 30 large runs, and that is how a wrong edit got through. Codex did so in 4 of 60.
  • One call per change. Before the guidance said otherwise, Codex made a neovain call for each change, twelve in one run. Claude chained its steps without being told.
  • Doubled backslashes. Codex wrote \\<word\\> on its first try at a rename, which matches nothing. neovain refused the call and the next one was right.

What helped

  • Saying how, not only what. Telling agents to mind the blank lines changed nothing, so the guidance now gives a method. With it, Codex's blank-line misses on the large task fell from 6 runs to 2 of 21, though runs with wrong code went from 3 to 5. For Claude it cut output.
  • A summary in place of a long diff. Since 0.2.0 a large change prints a summary of at most 80 lines: the blocks moved, deleted and re-indented, and warnings about spacing. Claude threw the output away in 1 of 6 large runs. Codex's blank-line misses went from 2 of 21 runs to none, though its runs with wrong code stayed at 5.
  • All-or-nothing calls. Every call that failed left the file untouched, and the agent recovered on its next call.
  • Checking the logs. The first version of our rule check missed a model that rewrote the file with Python. An edit count of zero gave it away.

What we learned

Reach for neovain when

  • A change moves, deletes or re-indents large blocks. Claude's Edit tool retypes them; Vim's :m, :d and > don't.
  • The same change repeats across a file: :%s and :g do it in one step.
  • You want all-or-nothing edits. A failed step leaves the file untouched, and agents recover on the next call.

Stick with the agent's own tool when

  • The change is small and local. Writing the new text directly is the reasoning; with Vim, the agent has to work out the commands too. Codex's patch tool needed well under half the output on small edits.
  • The model is small. In an early pilot, Claude Haiku 4.5 needed about three times as long with neovain and failed one run in three.

Tool calls are not round trips. Current models send independent edits as parallel tool calls, so cutting the number of calls saves little. What does cost time and money is output: every token the model writes. neovain wins where it lets the agent say less.