Automate Reporting: stop rebuilding the same spreadsheet every month | NexBDM Blog
NexBDM

NexBDM Blog

Automate Reporting: stop rebuilding the same spreadsheet every month

By NexBDM Team · 2026-08-24

Key takeaways

  • The monthly rebuild is where reporting errors enter. We traced the spreadsheet error statistic everybody quotes, found it reports the sample size as the finding, and set out which reports to automate first.

The monthly rebuild is where reporting errors enter. We traced the spreadsheet error statistic everybody quotes, found it reports the sample size as the finding, and set out which reports to automate first.

To automate reporting, you replace the monthly rebuild with a fixed pipeline: one source of record per number, a template that does not change, and a scheduled run that produces the report without anyone retyping it. The rebuild is where errors enter, because every rebuild is a fresh chance to introduce one.

Most South African small businesses do not have a reporting problem. They have a rebuilding problem. The sales summary, the debtors age analysis, the stock movement sheet: each one gets assembled again from scratch every month, from the same exports, in the same order, by the same person, into a workbook that already exists.

Before we get to how to stop doing that, it is worth dealing with the statistic that gets quoted at you whenever spreadsheets come up, because we traced it this week and it does not say what the quotes say it says.

The claim everybody repeats, and what the source table actually says

You will have seen some version of this: "88 percent of spreadsheets contain errors." It appears in vendor blogs, LinkedIn posts and conference slides, usually with no citation at all, occasionally attributed to academic research on spreadsheet auditing.

We pulled the underlying papers. Here is what happened to that number.

It traces to a summary table in Errors in Operational Spreadsheets, by Stephen Powell, Kenneth Baker and Barry Lawson of Dartmouth College, published in the Journal of Organizational and End User Computing, volume 21 issue 3, July to September 2009. Their sentence reads:

"In total, 88 spreadsheets are represented in the table. For all 88, the weighted average percentage of spreadsheets with errors is 94%."

88 is the sample size. The finding is 94 percent. The circulating statistic has taken the number of spreadsheets studied and reported it as the proportion that were wrong. It is not an exaggeration and it is not a fabrication. It is a transposition, repeated often enough that the corrected figure now looks like the error.

So what is the real number

There is not one. That is the honest answer, and it is more useful than a headline percentage.

Three separate figures come out of this literature, and they measure different things on different samples.

FigureWhat it actually measuresSample
94 percent of spreadsheets have errorsWeighted average across seven field audits summarised by Panko88 spreadsheets
5.2 percent of formula cells contain errorsWeighted average cell error rate43 spreadsheets, not 88
0.9 to 1.8 percent of formula cellsErrors found under a purpose built audit protocol, depending on how an error is defined50 operational spreadsheets

Three things follow, and none of them survive compression into a single statistic.

First, the 5.2 percent rate is not drawn from the same sample as the 94 percent. The Dartmouth authors state it plainly: cell error rate data was available for only 43 of the 88 spreadsheets. Quoting both figures in one breath, as though 94 percent of spreadsheets are wrong and 5.2 percent of every workbook's formulas are the reason, splices two samples together.

Second, the authors reporting the 94 percent list their own reasons to doubt it. They write that three of the seven sources are unpublished, that the majority gave little or no information on how they defined an error or on the methods used to find one, and that a single source accounts for around 70 percent of the sample behind the cell error rate. Those caveats sit in the same paragraph as the number. They never travel with it.

Third, Panko's own paper reports different figures again. In Spreadsheet Errors: What We Know. What We Think We Can Do, his table of seven field audits covers 367 spreadsheets and gives a weighted average of 24 percent with errors, rising to 91 percent across the 54 audited from 1997 onwards. Same researcher, same method, different denominator, wildly different headline. He also notes that most of those audits reported only substantive errors rather than all errors, and that the older audits used techniques unlikely to catch a majority of them.

So the same body of research supports 24 percent, 91 percent and 94 percent depending on which rows you count. Anyone quoting one of them as the spreadsheet error rate has picked a row.

The finding that actually matters, and almost nobody quotes it

The most rigorous study in the set is also the quietest. Powell, Baker and Lawson built a spreadsheet auditing protocol and applied it to 50 diverse operational spreadsheets, real workbooks in real use. They found errors in 0.9 to 1.8 percent of all formula cells, depending on how an error is defined. That is roughly three to five times lower than the received wisdom of about 5 percent they set out to test.

Their second finding is the one to build on: the error rate differed widely from spreadsheet to spreadsheet.

That sentence changes the operational question completely. If errors were spread evenly at 5 percent of formulas everywhere, spreadsheets would simply be a bad tool and the answer would be to stop using them. They are not spread evenly. Most workbooks are broadly fine and a few are badly wrong, which means the useful question is not "are spreadsheets risky" but "which of mine".

The answer to that is fairly predictable. Risk concentrates in the workbook that gets rebuilt, because a rebuild is the moment a human touches structure rather than data. A file that is opened, filtered and printed has one exposure. A file that is copied, re-pointed at a new export, has a row inserted and gets its formula ranges dragged down has a new one every single month.

A scope note, stated plainly: this research is from the United States and the United Kingdom, and it is old. The Dartmouth study is from 2009 and the audits it summarises run from 1987 to 2000. We have not found a South African equivalent and we are not going to relabel American and British data as local. What carries across is the mechanism, not the percentages.

Which reports to automate first

Rank them by frequency multiplied by rebuild depth, not by how much anyone complains about them.

  1. Anything rebuilt monthly from the same two or three exports. Debtors ageing, sales by rep, stock movement. High frequency, and the assembly steps are identical every time, which is exactly what a machine is good at and a tired person is not.
  2. Anything where last month's file is the starting point. "Save as, then update" is the highest risk pattern in this whole article, because errors compound forward and nobody re-checks the parts that were already right.
  3. Anything a third party sees. Reports going to a bank, a funder, an auditor or a board carry a cost of being wrong that has nothing to do with how long they took to build.
  4. Anything only one person can produce. If the month end pack has a single point of failure who is also the only person who knows which tab feeds which, that is a continuity problem wearing a reporting costume.

Notice what is not on that list: the report somebody finds annoying but runs twice a year. Annoyance is a poor ranking signal. We made the same argument about ranking processes by frequency times error cost in our guide to workflow automation for small business in South Africa, and reporting is the clearest case of it.

What changes when this is built properly

Four mechanisms, in the order they matter.

One source of record per number. Every figure resolves to exactly one system that owns it. Revenue comes from the accounting system, not from the accounting system for eleven months and a corrected spreadsheet in the twelfth. When two sources disagree, the pipeline fails loudly instead of quietly preferring one.

The template stops moving. The structure is defined once and the data flows into it. This is the mechanism that actually removes the errors the research describes, because a dragged formula range and an inserted row are structural edits, and a fixed template has no month in which structure is edited.

The run is scheduled, not remembered. The report is produced on a date whether or not anyone thinks of it. Reports that depend on somebody remembering are late in exactly the months when everyone is busiest, which are the months you most want to be looking at them.

Checks run before the report is trusted, not after it is questioned. Totals reconcile to the source, row counts match, dates fall inside the period, and the report says so on its face. The most valuable thing a pipeline produces is not the report. It is the assertion, attached to the report, that a specific set of checks passed.

What does not change

Automation moves the assembly. It does not move the judgement, and the gap between those two is where most disappointment with reporting tools comes from.

A pipeline can tell you that debtor days moved from 47 to 63. It cannot tell you that the move is one large customer who is about to place a bigger order, so pushing them for payment this week is the wrong call. It cannot tell you which of two divisions is worth the next hire. It cannot tell you that a number looks wrong because you know something about the month that is not in any system.

It will also not fix a report that was measuring the wrong thing. A monthly pack that has quietly stopped answering any question anybody has does not improve by arriving faster and more reliably. It becomes a punctual document nobody reads. Automating a report is a good moment to ask whether it should exist, and that is not a question software can answer for you.

And the errors do not go to zero. They move. Instead of a dragged formula, the failure mode becomes a source system that changed a column name, or a mapping written once and never revisited. That is a genuinely better trade, because those failures are findable and fixable in one place rather than scattered across twelve monthly copies. It is a trade, not a cure.

Frequently Asked Questions

Is it true that 88 percent of spreadsheets contain errors?

No. In the source table, 88 is the number of spreadsheets studied and 94 percent is the proportion found to contain errors. The circulating statistic reports the sample size as the finding. The authors who publish the 94 percent also list several reasons to treat it cautiously.

What is the actual spreadsheet error rate?

The research supports 24, 91 or 94 percent of spreadsheets depending on which audits are counted. The best controlled study, 50 operational spreadsheets audited under a purpose built protocol, found errors in 0.9 to 1.8 percent of formula cells, and found the rate varied widely between workbooks.

Which reports should a small business automate first?

Rank by frequency multiplied by rebuild depth. Anything assembled monthly from the same exports, anything built by copying last month's file, anything a bank or auditor sees, and anything only one person can produce. Rarely run reports come last regardless of how irritating they are.

Does automating reporting mean replacing Excel?

Usually not. The problem is the monthly rebuild, not the spreadsheet. A workbook that receives data into a fixed structure on a schedule has removed the risky step. Replacing the tool without removing the rebuild moves the same fragility somewhere less familiar.

How do you know the automated report is right?

By making it prove itself. Totals reconcile against the source system, row counts match, dates fall inside the period, and the report states which checks passed. A report that cannot show its checks is asking for the same trust as the manual version.

Where to start

Pick the single report that gets rebuilt most often and, before changing anything, write down every step somebody currently performs to produce it. Most teams find between nine and twenty steps for a report they describe as "just an export". That list is the specification, and it is usually also the moment the decision makes itself.

If you want that mapped across the whole business rather than one report, that is what a Business Autopsy does: it traces where the hours actually go before anybody proposes software. You can also read our guide to increasing team output without hiring, or start a discovery conversation about your own month end.

Book a free strategy call →