Research

Owner Rules Written Into a Spreadsheet Roughly Doubled AI Answer Scores on 12 Synthetic Businesses

Report
Download the full reportPDF · Technical report · September 2026

The rules that make a spreadsheet's numbers right usually live in one person's head. Which code means a redo. Which six customers are really one. Which column changed meaning when the software did. Hand the file to an AI and those rules stay behind.

We built an open-source tool that asks the owner for them once and writes them into the file. Then we tested it the slow way: a written plan, a frozen tool, twelve businesses built after the freeze, two blind judges, and an independent recount of every number. This page walks through how the experiment was run and what it does and does not show. The full LaTeX technical report is attached above.

Short answer

We tested Sheet Geek, our open-source spreadsheet skill, with a plan written and frozen before any data. On 12 synthetic businesses nobody had seen, an AI handed the same workbook with the brain tab scored 10.32 of 15 against 5.25 without, higher on all 12 (Wilcoxon p = 0.00049). The gain came from the two stronger models, and the study shows that owner rules written into the file help, not that the tool's interview is what put them there.

What Is Sheet Geek?

Sheet Geek is an open-source skill we built (tested as spreadsheet-brain). It reads every row of a workbook with code, asks the owner only the few questions the data cannot answer, and writes the answers into the file as a tab named _brain. Any AI that opens the file later can read the owner's rules there.

Every spreadsheet has two layers: the cells, and the rules for reading them. The cells travel. The rules do not. In one of the businesses in this test, a landscaping company, the crew log writes hours as hours and minutes, so 7.30 means seven and a half hours. A blank crew size means that crew's usual size. An invoice re-issued with an A on the end replaces the original. Six rental houses count as one customer. The data shows the effect of every one of those rules, and no cell states any of them. An AI that sums the column gets a clean, confident, wrong number.

Research on data agents describes the same failure. Agents commit to a plausible reading of an underspecified task and hand back clean, runnable work that hides the wrong framing (Ambig-DS, 2026). Models often spot ambiguity when asked to judge it, yet rarely stop to ask on their own (Su and Cardie, 2026). Owner interviews already exist for data warehouses: Anthropic's data-context-extractor asks an analyst about BigQuery, Snowflake, Postgres and Databricks schemas, and Anthropic's analytics team found that metric definitions a model generated on its own were net-negative on their evals against a smaller layer curated by people (Anthropic, 2026). We did not find a tool that writes the owner's answers into the spreadsheet itself. That is the result of a search, not a proof.

The tool works in a loop. Code studies the file first: the tables, the keys, which columns match across tabs and files, what each formula feeds, and anything odd. Then it asks four questions, more in small rounds only when the data or the answers call for them, ten at most, and one closing question: what would someone new get wrong in this file? A Python script picks every question and counts every number. The model only reads them out and passes the answers back. The answers go into a visible _brain tab at the end of the workbook, each note marked as something the owner said, something code counted, or a guess. The data itself is never changed.

A practice file shows the difference in one line. Asked for the runway on a synthetic finance model, Claude Opus 5.5 without the brain found cash running out in May 2027 and flagged a $15M raise in the model as typed in and unconfirmed. With the brain, it found the same date and cited the owner's note that the round is not closed, so it left the raise out. Without the brain the AI has to ask. With it, the file already says. That workbook is one of the seven practice files the tool was developed on, so it shows the tool at its best (the runs). The test below uses businesses it never saw.

What Question Did the Experiment Ask?

One question, written down before the test existed. When an AI with no skill and no hint is handed a workbook, does a brain built by the frozen tool make its answers better than the same workbook without one, across businesses nobody has seen? The business, not the single answer, was the unit of analysis.

This was our second pre-registered study. The first, on version 0.1, found the brain raised judged answer quality on four held-out businesses: 7.00 of 15 against 5.06, better in 25 of 31 pairs (Wilcoxon p = 0.00009). But those pairs came from four businesses, and at the business level the effect was 3 of 4 positive and not significant (p = 0.11). A new user cares whether it helps on their business. So the second study tested at the level of the business, with twelve of them (Study 1 plan and results).

Pre-registration means the question, the measures, the primary test and the rules for dropping data are written before any data exists, and anything done differently is reported as a deviation. The plan was drafted on September 26, 2026, before version 0.2 of the tool existed. Its analysis was locked when the tool was frozen on September 29, before any test business existed. The plan is public as PREREG-confirmatory.md.

How Was the Experiment Run?

Twelve synthetic businesses were built after the tool was frozen. An agent playing each owner ran the tool's interview from a written brief. Four AI models answered the same request with and without the brain, each trial in a locked folder, and two blind judges from two model families scored every pair.

The study was designed and run by an AI agent, Claude Opus 5.5 operating in Claude Code, under Zach Kellman's direction. Before the freeze, version 0.2 was developed against the seven synthetic businesses used so far, with general fixes only, never a rule aimed at one practice fact. It took six practice runs and four rounds of fixes to clear every pre-set bar on one run. On the final run, the tool's own questions captured 73.7% of the owner facts per business, against 21.3% for version 0.1. The plan calls that number optimistic by construction, because the fixes were made while looking at those files. That is why this study measured capture again on new ones.

The confirmatory run itself took one morning: the tool frozen at 07:33, the analysis at 10:00, and an independent check after that. The figure is the whole run. The ten steps under it say what each box did.

  1. Plan written (September 26). The question, the measures, the primary test and the exclusion rules, before version 0.2 or any test business existed.
  2. Tool frozen (September 29, 07:33). Version 0.2.0's 46 tool files and the 12 procedure files for every later step (scripts, the builder spec and the read-block list) were fingerprinted with sha256, with 1,890 tests passing. No tool change was allowed until the analysis was done. The written freeze record was added to the plan at 07:58, after the brains existed. The fingerprint files date from 07:33, and the audit found the body of the plan unchanged.
  3. Twelve businesses built and verified (07:34 to 07:51). One builder agent per business, each in a different line of work and none like the practice files: landscaping, a dental practice, trucking, a nonprofit, a gym, music venue tickets, auto repair, a law firm, a home builder, a greenhouse grower, ad spend and leads, and a coffee roaster. Builders were told the business was for a research test of a tool, told never to look for the tool or read about it, and told not to design anything for an AI tool. Each made one or two workbooks (17 in all, main tables of 1,000 to 8,000 rows, five main workbooks written with xlsxwriter and seven with openpyxl), an owner brief with 6 to 8 owner facts (80 in all) that change correct answers but are stated nowhere in the file, one open request, and three plain questions whose answers were computed by code. A separate verifier recomputed every answer with its own code, confirmed the obvious answer came out different, searched every cell, note, header, tab name and file property for a stated fact, and checked that no question named its rule. All 12 passed on the first build.
  4. Brains built (07:52 to 07:54). One owner agent per business ran the tool's full interview, answering only from the brief: "not sure" where the brief was silent, and the brief's own words when no option said exactly what it says. It saved the brain as a _brain tab in a copy of each workbook. No arm had the brief pasted in as plain notes.
  5. Gate (07:55 to 07:58). Before any answer run, a code check confirmed each tab existed, held at least one owner note, and left the source data unchanged. Two checkers then read all 640 owner notes against the brief, one looking for contradictions and one for claims the brief never makes. All 12 brains passed on the first build, with no note flagged.
  6. Answer trials (07:59 to 09:58). Four models ran cold, with no skill and no hint that a notes tab might exist: Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5.5 and OpenAI GPT-5.6-Luna. Each business had one fixed message, the same for every model and both arms: the file names, the open request, the three questions, a pointer to Python, "Work numbers out from the files; don't guess," and a 500-word cap. Two repetitions per business, model and arm made 96 pairs and 192 trials, run in shuffled order. The only difference between the arms was the _brain tab.
  7. Isolation rule applied. Each trial ran in a new folder holding only its workbooks, with a scrubbed environment. Claude runs had 94 operating-system read-block rules; Codex runs could not write and had plugins and apps off. Two trials of one business never ran at the same time. Code scanned every command. A trial that searched the whole disk or a shared temp folder, used Spotlight, went above its own folder, or touched another trial's or a judge's folder was re-run once, and a second flag dropped the pair.
  8. Blind judging (09:16 to 09:59). Two judges from two model families, Claude Opus 5.5 and OpenAI GPT-6-Astra, saw each pair's answers in coin-flip order, with the owner's brief and the reference answers and grading rules. They were told not to open the workbooks, that length, formatting and tone earn nothing, and that the workbook may hold a tab of notes written by its owner, so an answer citing it is not inventing a source. That last sentence was the one change from Study 1. Each judge scored both answers 0 to 5 on target, right and useful (0 to 15 in all), graded the three plain questions, listed the mistakes the owner would catch, marked any mention of a notes tab, and picked a winner. A pair's score is the mean of the two judges.
  9. Analysis (10:00). The frozen analysis script took, for each business, the mean of brain score minus no-brain score over its pairs. It then ran a two-sided Wilcoxon signed-rank test on the 12 business differences at alpha 0.05, with a 95% bootstrap interval over businesses (10,000 resamples). Per-model results were planned as description, not tests, and pair-level tests as descriptive only, since pairs within a business are not independent.
  10. Independent verification (after 10:00). A separate pass recomputed every reported number from the raw trial and verdict files with its own code, not the project's script. All matched, with the bootstrap interval inside random-seed noise. OpenAI GPT-6-Astra, given only the per-pair judge totals, reproduced the primary result. An integrity audit checked the fingerprints, the timeline and the isolation, and its corrections became deviations 1, 2 and 4.

What Did the Experiment Find?

The pre-registered test passed. Open-request scores averaged 5.25 of 15 without the brain and 10.32 with it. The brain answer scored higher on all 12 businesses, a business-level gain of +5.05 points (95% interval +4.19 to +5.78, Wilcoxon p = 0.00049), the smallest p that twelve businesses can give.

Every business gained. Ten of the twelve gained four points or more, and the home builder gained least, +1.42. If the brain did nothing, each business's gain would be as likely to be negative as positive. Of the 4,096 ways twelve signs can fall, only 2 are as lopsided as this one, all twelve the same way. That is p = 0.00049. By pair, described rather than tested, the brain answer scored higher in 61 of 73.

Two averages appear, and both are right. The figure's average row is the mean of the 12 business means, 5.27 without the brain. The 5.25 is the mean over all 73 pairs. Either way the score roughly doubled, and both averages include Claude Haiku 4.5, which gained almost nothing.

Open-request score by business, 0 to 15, mean of two blind judges
BusinessPairsWithoutWithGainBrain higher
Gym64.6711.08+6.425 of 6
Trucking fleet74.2910.64+6.367 of 7
Law firm64.5810.83+6.255 of 6
Ad spend and leads64.7510.92+6.175 of 6
Auto repair shop64.5810.67+6.085 of 6
Nonprofit65.9211.58+5.675 of 6
Coffee roaster65.3310.67+5.335 of 6
Dental practice65.0810.00+4.926 of 6
Landscaping65.089.42+4.336 of 6
Greenhouse grower65.759.75+4.004 of 6
Music venue tickets66.259.92+3.674 of 6
Home builder66.928.33+1.424 of 6
Mean of 12 businesses735.2710.32+5.0512 of 12 businesses
All 73 pairs, pooled735.2510.32+5.0761 of 73 pairs
Source: Study pairs file and the frozen analysis, recomputed and matched to results.json. Pooled over all 73 pairs, the means are 5.25 and 10.32.

The secondary measures point the same way. Judged against the key, plain questions came back right 133 of 219 times with the brain (60.7%) and 23 of 219 without (10.5%), and the two judges differed on only 2 of 438 grades. Mistakes the owner would catch fell from 7.6 per answer to 3.7, on all 12 businesses. The judges preferred the brain answer on every business, 58 pairs to 8 with 7 ties. Each business-level secondary test gave p = 0.00049.

The judges are models, so the independent verifier also graded the plain questions by code. On the 18 questions whose key is a single number, it counted an answer right if it contained a number within 1% of the key: 67.9% with the brain and 25.7% without (business-level p = 0.0033). That check was not in the plan, and it is more lenient than the judges. All 21 of its disagreements with them were numbers it accepted and the judges marked wrong. Even with the brain, 86 of 219 plain questions were still answered wrong. The brain moved answers from mostly wrong to mostly right, not to always right.

Which AI Models Gained From the Brain?

Two of them. Claude Opus 5.5 gained +6.98 points and GPT-5.6-Luna +7.96. Claude Haiku 4.5 gained +0.13, got no plain question right in either arm, and read the notes tab in only 5 of its 24 brain trials. Claude Sonnet 5 was almost entirely dropped, so the result covers three models, not four.

Results by answering model (described, not tested)
ModelPairsWithoutWithGainPlain questions rightMistakes per answer
Claude Opus 5.5247.6914.67+6.9832% to 100%6.5 to 0.4
GPT-5.6-Luna244.6312.58+7.960% to 81%7.3 to 1.7
Claude Haiku 4.5243.483.60+0.130% to 0%9.0 to 9.1
Claude Sonnet 514.5013.00+8.500% to 100%10.5 to 3.0
Source: Study pairs file by model, matched to results.json. Values round half up, so GPT-5.6-Luna's 4.625 shows as 4.63 where the plan's results text rounds it to 4.62. Sonnet 5 kept one pair and says nothing on its own.

With the brain, Opus 5.5 got every plain question right, against 32% without, and its mistakes per answer fell from 6.5 to 0.4. GPT-5.6-Luna went from no plain question right to 81%. Haiku 4.5 is the null result, and the reason is simple: in 19 of its 24 brain trials it never opened the tab, and the 5 trials that did read it gained only +0.2. A brain helps a model that reads it.

Sonnet 5 is missing for a different reason. Under the frozen isolation rule, carried over unchanged from Study 1, it searched the whole disk or the shared temp folder for its files in 43 of its 48 trials. Each was re-run once and 39 were flagged again, so 23 of its 24 pairs were dropped and the plan's 96 pairs became 73. The audit confirmed all 39 were real searches. By plan, every number in this section is described, not tested.

Did the Result Survive the Checks?

Yes, in every check the verifier ran, though none of them was pre-registered. Each judge alone found a gain on all 12 businesses. Leaving out the pairs where an answer named the notes tab cut the gain to +3.90, still 12 of 12. Leaving out each model in turn gave +3.62 to +7.46.

The largest threat is the blinding. Judges were never told which answer had the brain, but a brain answer often said where its rules came from. Depending on who counted, 20 to 25 of the 73 brain answers named the notes tab: 20 by the audit's text search, 24 by judge 1's marks and 25 by judge 2's. (A stricter text rule used in the sensitivity check finds 17, all Opus 5.5.) The results file also records one no-brain answer naming a notes tab, one that answer created itself, where the plan's results text says none. So a judge could tell the arms apart in about a third of pairs. Dropping the 25 pairs either judge flagged left +3.90, still higher on all 12 businesses (p = 0.00049). But 23 of those 25 were Opus 5.5 pairs, so that check mostly tests GPT-5.6-Luna and Haiku 4.5. Applying the plan's half-pairs rule, which drops the gym, leaves 11 businesses and p = 0.00098.

The judges agreed closely on scores (Spearman 0.91) and on plain-question grades (99.5%), and less on which answer won: 82% of pairs. Judge 1 alone gave +5.42 and judge 2 alone +4.68, each on all 12 businesses. The gap is not specific to Opus 5.5 answers: judge 1's gain is also larger on GPT-5.6-Luna answers (+8.46 against +7.46), mostly because judge 1 scored no-brain answers lower. Judge 2 still matters most here. Judge 1 was the same model as one of the answering models and as every agent that built the test, and model judges carry known position, verbosity and self-enhancement biases (Zheng et al., 2023).

Where Did the Owner Knowledge in the Brains Come From?

Mostly from one question. All 80 owner facts ended up in the brains, but only 34 of them (42.5%) came from a question the tool asked about that fact. The rest came mostly from the owner agents' reply to the tool's closing question, what someone new would get wrong, where they pasted their brief nearly word for word.

Capture on these new businesses averaged 41.9% per business, with a low of 16.7%. That is below all three bars version 0.2 had to clear on its practice files: a 60% per-business mean, 55% pooled and 35% for the lowest business. It is the optimism the plan warned about. In this study capture was a secondary measure, reported and not tested.

It changes what the headline means. The brains carry the brief's rules nearly word for word, including definitions that match how the answer key was computed. A numeric scan found no note that simply states a key answer. Still, the gain rests on owner agents that knew exactly what to type when asked what someone new would get wrong, and a busy real owner may say less. The tool's own faults also stayed in, because it was frozen. The owner agents logged them as they went, and on the home builder the scan read only half of one tab. The home builder also had the smallest gain. The study does not test whether the two are linked.

What Does This Study Not Show?

It does not show that the tool's interview is what helps, because no arm had the owner's brief pasted in as plain notes. It did not use real owners or human judges, since every owner and judge was a model. And it covers three answering models, one of which gained nothing.

What Went Differently From the Plan?

Four things, numbered as in the pre-registration. The agent running the study saw the owner interview logs before the gate. The gate's wiring script was written after the freeze. Sonnet 5 was almost entirely dropped by the isolation rule. And the audit found limits in the leak detector. Only the third changed what the result covers.

  1. The operator saw the owner interview logs before the gate. The operator was the AI agent running the study, which was not supposed to read any brief, key or workbook before the brains were gated. The brain-build workflow returned all 12 owner question-and-reply logs, which repeat the brief's facts, to its session, and part of the landscaping log was displayed. This deviation first named only the landscaping log, and the audit corrected it. The tool was already frozen, the gate and everything after it ran by script, and nothing was changed because of it.
  2. The gate's wiring script was written after the freeze. confirm_gate.sh, written at 07:52, chains the frozen code check and the frozen two-checker gate and passes a business only when both pass. It was not among the fingerprinted procedures. The audit found its logic matches the plan.
  3. Sonnet 5 is almost absent. The frozen isolation rule dropped 23 of its 24 pairs, as described above. Every business kept at least 6 of its 8 pairs, so none left the primary test, but the result covers three answering models, not four.
  4. The leak detector had limits. It did not expand $TMPDIR, caught /private/tmp only as a search's first root, and by design did not flag Claude trials that searched the home folder, which the operating system blocks. Five Haiku trials and the one kept Sonnet pair ran such searches. Every read outside the trial's own folder came back denied, and the audit found no trial that read another trial's files, a brief, a key or a judge folder. Dropping the Sonnet pair would move trucking from +6.36 to +6.00.

Two more details belong next to these. The written freeze record, the plan's Freeze section with the first deviation, was added to the plan at 07:58, after the brains existed, while the fingerprint files date from 07:33, before any business existed. And judge 1 leaves no record of which files it read, so its isolation rests on the operating system's read blocks. Judge 2's logs show it read only its own staged files.

Where Can You Check the Code, Data and Full Report?

In public. The tool, the pre-registration, the results file, every pair's scores, the questions and answer keys, the trial and judge scripts, and the verification reports are in the Sheet Geek repository, release v0.2.0, under Apache-2.0. The LaTeX report at the top of this page carries the full method, deviations and limitations; per-business and sensitivity tables are in the repository.

Start with the plan, PREREG-confirmatory.md, which holds the design, the freeze, the four deviations and the results as the plan reported them. The study folder, evals/rigor-confirmatory, holds the builder spec, results.json, the per-pair file pairs.csv, the capture scores, the scripts that ran the trials and the judges, and the independent verification reports. The code is in the Sheet Geek repository at release v0.2.0. The twelve test workbooks, the owner briefs and the answer texts are not in the repository.

To try it, add the plugin to Claude Code from the repository and open any workbook. Answer its questions like you are training your replacement, then read what it wrote before you send the file. A brain is only as good as what the owner told it. Writing down what a business knows so its data can be read correctly is also the core of our operations intelligence work.

Common Questions

Does Sheet Geek make AI answers more accurate?
Yes. In the pre-registered test, pooled over the models that stayed in, judged scores rose from 5.25 to 10.32 of 15 and were higher on all 12 synthetic businesses. Plain questions were right 60.7% of the time with the brain and 10.5% without. The gain came from Claude Opus 5.5 and GPT-5.6-Luna. Claude Haiku 4.5 gained nothing, mostly because it rarely opened the tab (described, not tested).
Is it the tool's interview or the notes themselves that help?
The study cannot say. There was no arm with the owner's brief pasted in as plain notes, so it shows that owner rules written into the file help. Only 42.5% of the owner facts came from the tool's own targeted questions. The rest came mostly from the owner agents' reply to its closing question.
Was the judging really blind?
By design: coin-flip answer order, and no judge was told which answer had the brain. In practice, 20 to 25 of 73 brain answers named the notes tab, so a judge could tell the arms apart in about a third of pairs. Dropping those pairs left a gain of +3.90, still on all 12 businesses, though that check mostly tests two of the three models.
Why is Claude Sonnet 5 missing from the results?
The frozen isolation rule removed it. Sonnet 5 searched the whole disk or a shared temp folder for its files in 43 of 48 trials and did it again in 39 re-runs, so 23 of its 24 pairs were dropped. The result covers Claude Opus 5.5, GPT-5.6-Luna and Claude Haiku 4.5.
Who ran the experiment, and is there a conflict of interest?
There is one. Actual Intelligence Labs built the tool. The experiment was designed and run by an AI agent, Claude Opus 5.5 in Claude Code, under Zach Kellman's direction, and the same model built the test businesses, played the owners and served as one judge. The plan was frozen first, and the scores, scripts and verification reports are public.
Can I check the numbers myself?
Yes. The pre-registration, the results file, every pair's scores, the analysis and trial scripts, and the independent verification reports are in the public Sheet Geek repository, release v0.2.0. The LaTeX technical report attached to this page has the full method, the deviations and the limitations; the per-business and sensitivity tables are in the repository.

Sources

SourcePublisherLink
Sheet Geek (released as spreadsheet-brain v0.2.0): open-source skill, studies and verificationActual Intelligence Labs, GitHubgithub.com
Pre-registered confirmatory test: does a spreadsheet brain help on businesses nobody has seen?Actual Intelligence Labs, 2026github.com
Study 2 data, scripts and independent verification (evals/rigor-confirmatory)Actual Intelligence Labs, 2026github.com
Pre-registered test: does a spreadsheet brain make an AI better at a workbook? (Study 1)Actual Intelligence Labs, 2026github.com
Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science AgentsStoisser et al., arXiv 2026arxiv.org
Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying QuestionsSu and Cardie, arXiv 2026arxiv.org
data-context-extractor skillAnthropic, GitHubgithub.com
How Anthropic enables self-service data analytics with ClaudeAnthropic, 2026claude.com
Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaZheng et al., arXiv 2023arxiv.org
Self-Preference Bias in LLM-as-a-JudgeWataoka et al., arXiv 2024arxiv.org
Individual Comparisons by Ranking MethodsWilcoxon, Biometrics Bulletin 1945doi.org

Questions to explore next

Keep Exploring

Chalk stick figure in a hard hat presenting a little machine of blue gears it just built

Bring us the bottleneck.
We’ll build the system.

Building What's Next.