What Is Sheet Geek?
Sheet Geek is an open-source skill we built (tested as spreadsheet-brain). It reads every row of a workbook with code, asks the owner only the few questions the data cannot answer, and writes the answers into the file as a tab named _brain. Any AI that opens the file later can read the owner's rules there.
Every spreadsheet has two layers: the cells, and the rules for reading them. The cells travel. The rules do not. In one of the businesses in this test, a landscaping company, the crew log writes hours as hours and minutes, so 7.30 means seven and a half hours. A blank crew size means that crew's usual size. An invoice re-issued with an A on the end replaces the original. Six rental houses count as one customer. The data shows the effect of every one of those rules, and no cell states any of them. An AI that sums the column gets a clean, confident, wrong number.
Research on data agents describes the same failure. Agents commit to a plausible reading of an underspecified task and hand back clean, runnable work that hides the wrong framing (Ambig-DS, 2026). Models often spot ambiguity when asked to judge it, yet rarely stop to ask on their own (Su and Cardie, 2026). Owner interviews already exist for data warehouses: Anthropic's data-context-extractor asks an analyst about BigQuery, Snowflake, Postgres and Databricks schemas, and Anthropic's analytics team found that metric definitions a model generated on its own were net-negative on their evals against a smaller layer curated by people (Anthropic, 2026). We did not find a tool that writes the owner's answers into the spreadsheet itself. That is the result of a search, not a proof.
The tool works in a loop. Code studies the file first: the tables, the keys, which columns match across tabs and files, what each formula feeds, and anything odd. Then it asks four questions, more in small rounds only when the data or the answers call for them, ten at most, and one closing question: what would someone new get wrong in this file? A Python script picks every question and counts every number. The model only reads them out and passes the answers back. The answers go into a visible _brain tab at the end of the workbook, each note marked as something the owner said, something code counted, or a guess. The data itself is never changed.
A practice file shows the difference in one line. Asked for the runway on a synthetic finance model, Claude Opus 5.5 without the brain found cash running out in May 2027 and flagged a $15M raise in the model as typed in and unconfirmed. With the brain, it found the same date and cited the owner's note that the round is not closed, so it left the raise out. Without the brain the AI has to ask. With it, the file already says. That workbook is one of the seven practice files the tool was developed on, so it shows the tool at its best (the runs). The test below uses businesses it never saw.
What Question Did the Experiment Ask?
One question, written down before the test existed. When an AI with no skill and no hint is handed a workbook, does a brain built by the frozen tool make its answers better than the same workbook without one, across businesses nobody has seen? The business, not the single answer, was the unit of analysis.
This was our second pre-registered study. The first, on version 0.1, found the brain raised judged answer quality on four held-out businesses: 7.00 of 15 against 5.06, better in 25 of 31 pairs (Wilcoxon p = 0.00009). But those pairs came from four businesses, and at the business level the effect was 3 of 4 positive and not significant (p = 0.11). A new user cares whether it helps on their business. So the second study tested at the level of the business, with twelve of them (Study 1 plan and results).
Pre-registration means the question, the measures, the primary test and the rules for dropping data are written before any data exists, and anything done differently is reported as a deviation. The plan was drafted on September 26, 2026, before version 0.2 of the tool existed. Its analysis was locked when the tool was frozen on September 29, before any test business existed. The plan is public as PREREG-confirmatory.md.
How Was the Experiment Run?
Twelve synthetic businesses were built after the tool was frozen. An agent playing each owner ran the tool's interview from a written brief. Four AI models answered the same request with and without the brain, each trial in a locked folder, and two blind judges from two model families scored every pair.
The study was designed and run by an AI agent, Claude Opus 5.5 operating in Claude Code, under Zach Kellman's direction. Before the freeze, version 0.2 was developed against the seven synthetic businesses used so far, with general fixes only, never a rule aimed at one practice fact. It took six practice runs and four rounds of fixes to clear every pre-set bar on one run. On the final run, the tool's own questions captured 73.7% of the owner facts per business, against 21.3% for version 0.1. The plan calls that number optimistic by construction, because the fixes were made while looking at those files. That is why this study measured capture again on new ones.
The confirmatory run itself took one morning: the tool frozen at 07:33, the analysis at 10:00, and an independent check after that. The figure is the whole run. The ten steps under it say what each box did.
- Plan written (September 26). The question, the measures, the primary test and the exclusion rules, before version 0.2 or any test business existed.
- Tool frozen (September 29, 07:33). Version 0.2.0's 46 tool files and the 12 procedure files for every later step (scripts, the builder spec and the read-block list) were fingerprinted with sha256, with 1,890 tests passing. No tool change was allowed until the analysis was done. The written freeze record was added to the plan at 07:58, after the brains existed. The fingerprint files date from 07:33, and the audit found the body of the plan unchanged.
- Twelve businesses built and verified (07:34 to 07:51). One builder agent per business, each in a different line of work and none like the practice files: landscaping, a dental practice, trucking, a nonprofit, a gym, music venue tickets, auto repair, a law firm, a home builder, a greenhouse grower, ad spend and leads, and a coffee roaster. Builders were told the business was for a research test of a tool, told never to look for the tool or read about it, and told not to design anything for an AI tool. Each made one or two workbooks (17 in all, main tables of 1,000 to 8,000 rows, five main workbooks written with xlsxwriter and seven with openpyxl), an owner brief with 6 to 8 owner facts (80 in all) that change correct answers but are stated nowhere in the file, one open request, and three plain questions whose answers were computed by code. A separate verifier recomputed every answer with its own code, confirmed the obvious answer came out different, searched every cell, note, header, tab name and file property for a stated fact, and checked that no question named its rule. All 12 passed on the first build.
- Brains built (07:52 to 07:54). One owner agent per business ran the tool's full interview, answering only from the brief: "not sure" where the brief was silent, and the brief's own words when no option said exactly what it says. It saved the brain as a _brain tab in a copy of each workbook. No arm had the brief pasted in as plain notes.
- Gate (07:55 to 07:58). Before any answer run, a code check confirmed each tab existed, held at least one owner note, and left the source data unchanged. Two checkers then read all 640 owner notes against the brief, one looking for contradictions and one for claims the brief never makes. All 12 brains passed on the first build, with no note flagged.
- Answer trials (07:59 to 09:58). Four models ran cold, with no skill and no hint that a notes tab might exist: Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5.5 and OpenAI GPT-5.6-Luna. Each business had one fixed message, the same for every model and both arms: the file names, the open request, the three questions, a pointer to Python, "Work numbers out from the files; don't guess," and a 500-word cap. Two repetitions per business, model and arm made 96 pairs and 192 trials, run in shuffled order. The only difference between the arms was the _brain tab.
- Isolation rule applied. Each trial ran in a new folder holding only its workbooks, with a scrubbed environment. Claude runs had 94 operating-system read-block rules; Codex runs could not write and had plugins and apps off. Two trials of one business never ran at the same time. Code scanned every command. A trial that searched the whole disk or a shared temp folder, used Spotlight, went above its own folder, or touched another trial's or a judge's folder was re-run once, and a second flag dropped the pair.
- Blind judging (09:16 to 09:59). Two judges from two model families, Claude Opus 5.5 and OpenAI GPT-6-Astra, saw each pair's answers in coin-flip order, with the owner's brief and the reference answers and grading rules. They were told not to open the workbooks, that length, formatting and tone earn nothing, and that the workbook may hold a tab of notes written by its owner, so an answer citing it is not inventing a source. That last sentence was the one change from Study 1. Each judge scored both answers 0 to 5 on target, right and useful (0 to 15 in all), graded the three plain questions, listed the mistakes the owner would catch, marked any mention of a notes tab, and picked a winner. A pair's score is the mean of the two judges.
- Analysis (10:00). The frozen analysis script took, for each business, the mean of brain score minus no-brain score over its pairs. It then ran a two-sided Wilcoxon signed-rank test on the 12 business differences at alpha 0.05, with a 95% bootstrap interval over businesses (10,000 resamples). Per-model results were planned as description, not tests, and pair-level tests as descriptive only, since pairs within a business are not independent.
- Independent verification (after 10:00). A separate pass recomputed every reported number from the raw trial and verdict files with its own code, not the project's script. All matched, with the bootstrap interval inside random-seed noise. OpenAI GPT-6-Astra, given only the per-pair judge totals, reproduced the primary result. An integrity audit checked the fingerprints, the timeline and the isolation, and its corrections became deviations 1, 2 and 4.
What Did the Experiment Find?
The pre-registered test passed. Open-request scores averaged 5.25 of 15 without the brain and 10.32 with it. The brain answer scored higher on all 12 businesses, a business-level gain of +5.05 points (95% interval +4.19 to +5.78, Wilcoxon p = 0.00049), the smallest p that twelve businesses can give.
Every business gained. Ten of the twelve gained four points or more, and the home builder gained least, +1.42. If the brain did nothing, each business's gain would be as likely to be negative as positive. Of the 4,096 ways twelve signs can fall, only 2 are as lopsided as this one, all twelve the same way. That is p = 0.00049. By pair, described rather than tested, the brain answer scored higher in 61 of 73.
Two averages appear, and both are right. The figure's average row is the mean of the 12 business means, 5.27 without the brain. The 5.25 is the mean over all 73 pairs. Either way the score roughly doubled, and both averages include Claude Haiku 4.5, which gained almost nothing.
| Business | Pairs | Without | With | Gain | Brain higher |
|---|---|---|---|---|---|
| Gym | 6 | 4.67 | 11.08 | +6.42 | 5 of 6 |
| Trucking fleet | 7 | 4.29 | 10.64 | +6.36 | 7 of 7 |
| Law firm | 6 | 4.58 | 10.83 | +6.25 | 5 of 6 |
| Ad spend and leads | 6 | 4.75 | 10.92 | +6.17 | 5 of 6 |
| Auto repair shop | 6 | 4.58 | 10.67 | +6.08 | 5 of 6 |
| Nonprofit | 6 | 5.92 | 11.58 | +5.67 | 5 of 6 |
| Coffee roaster | 6 | 5.33 | 10.67 | +5.33 | 5 of 6 |
| Dental practice | 6 | 5.08 | 10.00 | +4.92 | 6 of 6 |
| Landscaping | 6 | 5.08 | 9.42 | +4.33 | 6 of 6 |
| Greenhouse grower | 6 | 5.75 | 9.75 | +4.00 | 4 of 6 |
| Music venue tickets | 6 | 6.25 | 9.92 | +3.67 | 4 of 6 |
| Home builder | 6 | 6.92 | 8.33 | +1.42 | 4 of 6 |
| Mean of 12 businesses | 73 | 5.27 | 10.32 | +5.05 | 12 of 12 businesses |
| All 73 pairs, pooled | 73 | 5.25 | 10.32 | +5.07 | 61 of 73 pairs |
The secondary measures point the same way. Judged against the key, plain questions came back right 133 of 219 times with the brain (60.7%) and 23 of 219 without (10.5%), and the two judges differed on only 2 of 438 grades. Mistakes the owner would catch fell from 7.6 per answer to 3.7, on all 12 businesses. The judges preferred the brain answer on every business, 58 pairs to 8 with 7 ties. Each business-level secondary test gave p = 0.00049.
The judges are models, so the independent verifier also graded the plain questions by code. On the 18 questions whose key is a single number, it counted an answer right if it contained a number within 1% of the key: 67.9% with the brain and 25.7% without (business-level p = 0.0033). That check was not in the plan, and it is more lenient than the judges. All 21 of its disagreements with them were numbers it accepted and the judges marked wrong. Even with the brain, 86 of 219 plain questions were still answered wrong. The brain moved answers from mostly wrong to mostly right, not to always right.
Which AI Models Gained From the Brain?
Two of them. Claude Opus 5.5 gained +6.98 points and GPT-5.6-Luna +7.96. Claude Haiku 4.5 gained +0.13, got no plain question right in either arm, and read the notes tab in only 5 of its 24 brain trials. Claude Sonnet 5 was almost entirely dropped, so the result covers three models, not four.
| Model | Pairs | Without | With | Gain | Plain questions right | Mistakes per answer |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 24 | 7.69 | 14.67 | +6.98 | 32% to 100% | 6.5 to 0.4 |
| GPT-5.6-Luna | 24 | 4.63 | 12.58 | +7.96 | 0% to 81% | 7.3 to 1.7 |
| Claude Haiku 4.5 | 24 | 3.48 | 3.60 | +0.13 | 0% to 0% | 9.0 to 9.1 |
| Claude Sonnet 5 | 1 | 4.50 | 13.00 | +8.50 | 0% to 100% | 10.5 to 3.0 |
With the brain, Opus 5.5 got every plain question right, against 32% without, and its mistakes per answer fell from 6.5 to 0.4. GPT-5.6-Luna went from no plain question right to 81%. Haiku 4.5 is the null result, and the reason is simple: in 19 of its 24 brain trials it never opened the tab, and the 5 trials that did read it gained only +0.2. A brain helps a model that reads it.
Sonnet 5 is missing for a different reason. Under the frozen isolation rule, carried over unchanged from Study 1, it searched the whole disk or the shared temp folder for its files in 43 of its 48 trials. Each was re-run once and 39 were flagged again, so 23 of its 24 pairs were dropped and the plan's 96 pairs became 73. The audit confirmed all 39 were real searches. By plan, every number in this section is described, not tested.
Did the Result Survive the Checks?
Yes, in every check the verifier ran, though none of them was pre-registered. Each judge alone found a gain on all 12 businesses. Leaving out the pairs where an answer named the notes tab cut the gain to +3.90, still 12 of 12. Leaving out each model in turn gave +3.62 to +7.46.
The largest threat is the blinding. Judges were never told which answer had the brain, but a brain answer often said where its rules came from. Depending on who counted, 20 to 25 of the 73 brain answers named the notes tab: 20 by the audit's text search, 24 by judge 1's marks and 25 by judge 2's. (A stricter text rule used in the sensitivity check finds 17, all Opus 5.5.) The results file also records one no-brain answer naming a notes tab, one that answer created itself, where the plan's results text says none. So a judge could tell the arms apart in about a third of pairs. Dropping the 25 pairs either judge flagged left +3.90, still higher on all 12 businesses (p = 0.00049). But 23 of those 25 were Opus 5.5 pairs, so that check mostly tests GPT-5.6-Luna and Haiku 4.5. Applying the plan's half-pairs rule, which drops the gym, leaves 11 businesses and p = 0.00098.
The judges agreed closely on scores (Spearman 0.91) and on plain-question grades (99.5%), and less on which answer won: 82% of pairs. Judge 1 alone gave +5.42 and judge 2 alone +4.68, each on all 12 businesses. The gap is not specific to Opus 5.5 answers: judge 1's gain is also larger on GPT-5.6-Luna answers (+8.46 against +7.46), mostly because judge 1 scored no-brain answers lower. Judge 2 still matters most here. Judge 1 was the same model as one of the answering models and as every agent that built the test, and model judges carry known position, verbosity and self-enhancement biases (Zheng et al., 2023).
Where Did the Owner Knowledge in the Brains Come From?
Mostly from one question. All 80 owner facts ended up in the brains, but only 34 of them (42.5%) came from a question the tool asked about that fact. The rest came mostly from the owner agents' reply to the tool's closing question, what someone new would get wrong, where they pasted their brief nearly word for word.
Capture on these new businesses averaged 41.9% per business, with a low of 16.7%. That is below all three bars version 0.2 had to clear on its practice files: a 60% per-business mean, 55% pooled and 35% for the lowest business. It is the optimism the plan warned about. In this study capture was a secondary measure, reported and not tested.
It changes what the headline means. The brains carry the brief's rules nearly word for word, including definitions that match how the answer key was computed. A numeric scan found no note that simply states a key answer. Still, the gain rests on owner agents that knew exactly what to type when asked what someone new would get wrong, and a busy real owner may say less. The tool's own faults also stayed in, because it was frozen. The owner agents logged them as they went, and on the home builder the scan read only half of one tab. The home builder also had the smallest gain. The study does not test whether the two are linked.
What Does This Study Not Show?
It does not show that the tool's interview is what helps, because no arm had the owner's brief pasted in as plain notes. It did not use real owners or human judges, since every owner and judge was a model. And it covers three answering models, one of which gained nothing.
- No control arm for the interview. There was no third arm with the owner's brief pasted into the workbook as plain notes. The study shows that owner rules written into the file help an AI answer. It does not show that the tool's interview is what put them there, and 46 of the 80 facts did not come from its targeted questions, mostly arriving through the closing question.
- One model held most of the roles. Claude Opus 5.5 built the businesses, verified them, played the owners, gated the brains, served as judge 1, and was one of the four answering models. Model judges can favor text that reads like their own (Wataoka et al., 2024). The judge from the other model family found a smaller gain, +4.68 against +5.42, though judge 1's gain was just as much larger on GPT-5.6-Luna answers, so the gap is not specific to Opus 5.5 text.
- Judges were told a notes tab might exist. The judge prompt said the workbook may hold an owner's notes tab, so citing one is not inventing a source. That removed a bias against the brain found in Study 1. Combined with the 20 to 25 brain answers that named the tab, it also let a judge infer which arm it was reading in about a third of pairs.
- Everything is synthetic. The businesses were invented, the owners were agents reading a complete brief, and every judge was a model. No real owner sat through the questions and no person graded an answer. Businesses built to hold owner-only rules may make a brain look more useful than it would be on a real file where most answers sit in the data.
- Three models, two command-line tools. Sonnet 5 kept 1 of its 24 pairs, and Haiku 4.5 gained nothing. Every run went through Claude Code or Codex with a Python shell. Chat apps such as ChatGPT and claude.ai were not tested.
- The frozen tool's faults stayed in. Faults found during the run, such as the half-read tab on the home builder, were logged and left alone, as a freeze requires.
- Answers were still far from perfect. With the brain, the mean score was 10.32 of 15, and 86 of 219 plain questions were still answered wrong.
- We built it. Actual Intelligence Labs built the tool, designed the test and wrote this report, and we build systems like it for clients. The experiment was designed and run by an AI agent, Claude Opus 5.5 operating in Claude Code, under Zach Kellman's direction. That conflict of interest is why the plan was written first, the deviations are listed below, and the plan, scores, scripts and verification reports are public.
What Went Differently From the Plan?
Four things, numbered as in the pre-registration. The agent running the study saw the owner interview logs before the gate. The gate's wiring script was written after the freeze. Sonnet 5 was almost entirely dropped by the isolation rule. And the audit found limits in the leak detector. Only the third changed what the result covers.
- The operator saw the owner interview logs before the gate. The operator was the AI agent running the study, which was not supposed to read any brief, key or workbook before the brains were gated. The brain-build workflow returned all 12 owner question-and-reply logs, which repeat the brief's facts, to its session, and part of the landscaping log was displayed. This deviation first named only the landscaping log, and the audit corrected it. The tool was already frozen, the gate and everything after it ran by script, and nothing was changed because of it.
- The gate's wiring script was written after the freeze. confirm_gate.sh, written at 07:52, chains the frozen code check and the frozen two-checker gate and passes a business only when both pass. It was not among the fingerprinted procedures. The audit found its logic matches the plan.
- Sonnet 5 is almost absent. The frozen isolation rule dropped 23 of its 24 pairs, as described above. Every business kept at least 6 of its 8 pairs, so none left the primary test, but the result covers three answering models, not four.
- The leak detector had limits. It did not expand $TMPDIR, caught /private/tmp only as a search's first root, and by design did not flag Claude trials that searched the home folder, which the operating system blocks. Five Haiku trials and the one kept Sonnet pair ran such searches. Every read outside the trial's own folder came back denied, and the audit found no trial that read another trial's files, a brief, a key or a judge folder. Dropping the Sonnet pair would move trucking from +6.36 to +6.00.
Two more details belong next to these. The written freeze record, the plan's Freeze section with the first deviation, was added to the plan at 07:58, after the brains existed, while the fingerprint files date from 07:33, before any business existed. And judge 1 leaves no record of which files it read, so its isolation rests on the operating system's read blocks. Judge 2's logs show it read only its own staged files.
Where Can You Check the Code, Data and Full Report?
In public. The tool, the pre-registration, the results file, every pair's scores, the questions and answer keys, the trial and judge scripts, and the verification reports are in the Sheet Geek repository, release v0.2.0, under Apache-2.0. The LaTeX report at the top of this page carries the full method, deviations and limitations; per-business and sensitivity tables are in the repository.
Start with the plan, PREREG-confirmatory.md, which holds the design, the freeze, the four deviations and the results as the plan reported them. The study folder, evals/rigor-confirmatory, holds the builder spec, results.json, the per-pair file pairs.csv, the capture scores, the scripts that ran the trials and the judges, and the independent verification reports. The code is in the Sheet Geek repository at release v0.2.0. The twelve test workbooks, the owner briefs and the answer texts are not in the repository.
To try it, add the plugin to Claude Code from the repository and open any workbook. Answer its questions like you are training your replacement, then read what it wrote before you send the file. A brain is only as good as what the owner told it. Writing down what a business knows so its data can be read correctly is also the core of our operations intelligence work.

