Try the demo

How we measured DMN

We had Claude Code agents track down real bugs, with and without DMN, and wrote the rules down before any run. This page has the numbers, how we got them, where DMN didn't help, and what the numbers don't show.

21–25%fewer requests to find the code
16–22%fewer tokens to find the code
11 → 6tool calls in a median run

Finding the code behind a bug, on two code bases the model could rarely place without tools, and finding it about as often as without DMN. On a popular third one, where the model named the spot with no tools for 15 of 36 tasks, DMN saved a few requests and no tokens.

Results

Each bar is what an agent spent with DMN, as a share of what it spent without. Shorter is better; the line is "the same as without".

requests tokens
DMN's own codeprivate, 31 tasks
75% · −25%
78% · −22%
vortexpublic Rust, 30 tasks
79% · −21%
84% · −16%
ruffpopular, 36 tasks, 15 placed with no tools
92% · −8%
100% · ±0
0%50%100% = without DMN
Code base Tasks Found it (without / with) Median requests Median tokens Requests Tokens
DMN's own code 31 59 / 58 of 62 7.5 → 5 267k → 189k −25% −22%
vortex 30 58 / 58 of 60 8 → 5.5 301k → 210k −21% −16%
ruff 36 69 / 67 of 72 8 → 6 251k → 209k −8% ±0%

"Found it" counts runs that named the right function among their top three answers. The last two columns compare each task with itself, with and without DMN, averaged over tasks (a geometric mean), so they differ from the change in medians. The rule written before the runs let DMN's success rate fall by up to 10 points; it fell by at most 3. Runs on 2026-09-20 and 2026-09-21.

What a run looks like

On DMN's own code, the median run without DMN took 7.5 requests and 11 tool calls: search, read a file, search again, read another. With DMN it took 5 requests and 6 tool calls, and asked DMN once.

DMN doesn't hand over the answer. Its results point the agent at the right part of the code and the right words for it; the agent still reads the code and decides. Tasks where DMN's first results missed the right file gained the most on this code base, because being pointed at the right area was enough.

Where the gain is

Splitting the tasks shows where DMN helps and where it doesn't.

TasksCountRequestsTokens
DMN's code: no word in the report points grep at a few files23−29%−28%
DMN's code: a word in the report matches five files or fewer8−10%±0%
DMN's code: Rust and other21−29%−29%
DMN's code: the TypeScript GUI10−14%−4%
vortex: the fix touches several places10−32%−30%
vortex: the fix touches one place20−16%−8%

The pattern: DMN helps most when a bug report gives grep nothing to hold on to, and when the fix spans several places. When one word in the report already names the file, grep is as good.

The misses

Code the model can already place

We also gave the model each bug report with no tools at all and asked where the bug was. It could already name the spot for 3 of 31 tasks on DMN's code, 2 of 30 on vortex, and 15 of 36 on ruff, a popular project it has most likely seen in training. On ruff it often barely needs to search, so DMN had little to save.

DMN's own code10% guessable−22% tokens
vortex7% guessable−16% tokens
ruff42% guessable±0 tokens

Results full of test snapshots

On ruff, 44% of what DMN returned was test files, most of them snapshots that quote the same error text as the bug reports. That pushed the real source down the list. On DMN's own code, test files were 1%.

A bar we set, and missed

Before the runs we said a headline result needed both ratios at 0.75 or lower. DMN's own code came in at 0.754 and 0.784: clear and consistent, but just short. So we call it modest.

Speed

The counts above are about work. These are about time, each measured on its own, with where it came from.

WhatResultHow it was measured
A search, answered 5–12 ms Median over 17 developer questions each on Django (10 ms), Vite (5 ms) and ruff (12 ms), full-project indexes, on a desktop with a Ryzen 9 9950X3D and an RTX 4070 Ti SUPER. September 2026.
A search, CPU only 20 ms Median over 30 first-time questions and names on a 44-file sample project, on the CPU alone with the EmbeddingGemma model (26 ms at p90). Asking the same thing again takes about 1 ms.
From saving a file to finding it 2.2 s Median over 24 saves of a new function into the same sample, CPU only, found by its name (2.2 s at p90); 2.3 s over 12 more saves found by meaning alone. 2 s of it is a set pause that gathers a burst of saves into one update (DMN_CONTINUOUS_DEBOUNCE_MS); the update itself took about 0.15 s. September 2026.
One agent round trip ~3 s + reply Across 62 recorded coding runs, each request to the model took about 3.1 seconds plus 30 ms per token of reply.

Put together: finding the code took 21–25% fewer round trips, and each one skipped is a few seconds the agent isn't waiting. We haven't timed whole tasks end to end.

The race, run for real

The races on the home and engine pages replay real runs. Claude Code (Sonnet, headless) was asked "where do we rate-limit logins?" on a 44-file sample project: three times without DMN, three times with DMN as it ships (terse answers on), and three times with terse answers off, in turn. Every arm could grep, glob and read; the DMN arms could also run dmn search. The protocol was written down and committed before the first run, and all nine runs are reported.

Median of 3 runsCorrectRequestsTokensSeconds
Without DMN2/3316,4626.2
With DMN3/3217,3173.9
With DMN, terse answers off3/3215,7244.3

Tokens count cached input in full, as above. The terse-answer rule adds about 800 tokens to every request; on a two-request lookup that costs more than the shorter answers save, while on real bug reports it cut tokens by 23% (see what each part adds). One run without DMN named the right file and function but wrote the path with backslashes, which the check written down beforehand counts as wrong. One small project, so the benchmark above is the fuller measure.

What each part adds

DMN has a few parts that could save tokens: its search, a rule that makes the agent answer tersely, a budget on how much code a search returns, and serving it over MCP instead of as a skill. We measured each one on and off: real bug reports on DMN's own code (14 tasks, each twice), Claude Code on Opus 5.5, 196 runs, protocol written down first. Every setup found the right code in 26 of 28 runs.

SetupTokens95% CIAgainstWhat it means
DMN's search, as a skill0.85×0.71–1.01no DMN15% fewer tokens, not clear on its own here; 16% fewer requests (0.74–0.95), which is clear
+ terse answers (the compact rule, now on by default)0.77×0.61–0.97the skill alone23% fewer; 0.65× against no DMN (0.53–0.80), which is DMN as shipped now
Search budget of 5,000 characters0.92×0.80–1.06the skill aloneno clear effect: agents raised the budget themselves
Search budget of 3,000 characters0.99×0.86–1.14the skill aloneno clear effect, for the same reason
DMN over MCP instead of the skill1.12×0.97–1.29the skill aloneno clear effect: the agents never called the MCP tools

Terse answers are on by default for Claude Code; [skill] compact = "off" turns them off. A blinded check rated them about as useful as the full ones: 4.5 against 4.6 out of 5, a difference inside the margin set beforehand. They sometimes leave out a secondary detail, such as a suggested fix.

How we ran it

  • The task: the agent gets a real bug report and must name the function that needs the fix. Every task is a real bug, fixed in a real commit after the code snapshot.
  • The agent: Claude Code's search agent on Claude Opus 5, with read-only tools, a fresh agent for every run.
  • Without DMN: grep, glob and file reads. With DMN: the same, plus one command that asks DMN in plain words.
  • Fair tasks: two reviewers agreed the right answer, and the bug reports were written without looking at the fix, so they don't name it.
  • Every task, twice in each arm. Nothing was dropped for being too easy or too hard.
  • Counted from each transcript: requests to the model, and tokens (input, cached input and output).
  • Registered first: the tasks, the pass rule and the bar for a headline result were written down before any run, and not changed after.

Checks

  • No tools at all: tasks the model could answer with no tools were set aside in a second pass; the result held without them.
  • Home advantage: on tasks whose files had ever been used to tune DMN, the gain was about zero. All of it came from files never used for tuning.
  • One task carrying it: dropping any single task left the result about the same.
  • Noise: two runs of the same task often differ by a request or more, so we report averages over tasks, never single tasks.

What it doesn't show

  • It measured finding the code, not a whole coding session. Sessions also write and test code, where DMN doesn't change much. On one set of tasks we measured, finding code was about a fifth of a full session, which would make the session-wide saving closer to 5–10%. We haven't measured whole sessions on such code yet.
  • Cached tokens count in full. Our token numbers count cached input like any other input. Counting only uncached input and output, the gap on DMN's own code is about 9%.
  • One agent, one model. The runs used Claude Code with Claude Opus 5. Other agents and models may gain more or less.
  • On DMN's own code, embeddings ran on the CPU, the slowest setup. That affects speed, not these counts.

Two measures

The app's Insights view and dmn stats compare what DMN handed an agent with reading the same files whole. That runs higher than the head-to-head numbers on this page, which compare a whole task with DMN to the same task without it. The head-to-head numbers are the ones we stand behind.

The full report and its protocol are in DMN's repository, which isn't public yet.

Check it on your own code

dmn bench builds a search benchmark from your indexed code, with no labelling, runs it, and writes a report. The same seed on an unchanged index gives the same numbers. It measures how well DMN finds code, not how many tokens your agents spend; dmn stats shows that side.

dmn bench