NexBDM Blog
AI for Construction: what it counts reliably, and where it quietly bids you short
By NexBDM Team · 2026-09-06
Key takeaways
- Peer reviewed testing put an AI quantity takeoff up against a contractor's own. Counts matched. Areas came back low, and worst on facades. On a tender, an error that runs low is the one you pay for on site.
Peer reviewed testing put an AI quantity takeoff up against a contractor's own. Counts matched. Areas came back low, and worst on facades. On a tender, an error that runs low is the one you pay for on site.
AI for construction is dependable at counting. Doors, windows and fittings come back matching a human takeoff almost exactly. Area measurement is where it slips, and peer reviewed testing found the error runs one way, downward, and is worst on facades. On a tender, a quiet downward bias is the expensive kind.
That paragraph is the whole post, but it is worth showing the work, because the two halves come from different places and neither one is guesswork.
Why this matters more in 2026 than it did last year
South African contractors are bidding into a harder market, and there is a quarterly measurement of exactly how much harder.
The cidb SME Business Conditions Survey is run by the Bureau for Economic Research at Stellenbosch University on behalf of the cidb. In the Q4 2025 edition, published February 2026:
| Indicator | 2025Q3 | 2025Q4 |
|---|---|---|
| cidb SME Business Conditions Index | 37 | 42 |
| Insufficient demand as a constraint, all contractors | 69% | 74% |
| Insufficient demand, general building | 66% | 72% |
| Insufficient demand, civil engineering | 71% | 76% |
| Tendering competition, net %, civil engineering | 24 | 47 |
Read those together. Confidence recovered, and at the same time roughly three quarters of surveyed contractors named insufficient demand as a constraint, while tendering competition among civil engineering firms nearly doubled in a single quarter.
That is a market where the same work is being chased by more bidders. The rational response is to bid on more jobs, which means producing more takeoffs and more priced documents in the same number of hours. That pressure is precisely what makes automated takeoff attractive, and it is also what makes the direction of its error worth knowing before you lean on it.
What the research actually measured
Most pages answering "AI for construction" quote accuracy percentages with no study behind them. There is a real one, and it is recent.
Hooman Sadeh, Dominick Geloso and Dimitar Todorov of Utica University published Testing AI Accuracy in Quantity Takeoff: A Methodological Case Study in Commercial Construction in EPiC Series in Built Environment, Volume 7, 2026, pages 753 to 762, in the proceedings of the Associated Schools of Construction 62nd Annual International Conference.
The method: take a live commercial construction project, run an AI quantity takeoff platform over the drawings, and compare its output line by line against the takeoff the contractor produced. Twenty nine line items covering exterior, floor and ceiling finishes, plus windows and doors. Units were square feet, each, and linear feet. The comparison used Wilcoxon signed rank and Kruskal Wallis tests rather than a headline accuracy claim.
Counting is solved. Measuring is not.
| What was tested | Result | What it means |
|---|---|---|
| All items together | Z = -1.98, p = .048 | Statistically significant overall underestimation |
| Area quantities (square feet) | Z = -2.22, p = .027, 12 of 15 items lower | Significant, and consistently low |
| Count quantities (each) | Z = -0.45, p = .655, 11 of 13 identical | No significant difference. Counting is reliable |
| Deviation by finish type | H(2) = 8.42, p = .015 | Accuracy depends on what is being measured |
The finish type breakdown is the part worth memorising. Ranked by size of deviation, exterior finishes came out worst at a mean rank of 12.20, floor finishes sat in the middle at 7.80, and ceiling finishes were the most accurate at 4.00. The gap between exterior and ceiling was statistically significant at p = .022.
The authors put the reason plainly: accuracy holds for horizontal elements such as ceilings and floors and diverges for vertical facade components, likely because of geometric and visual complexity in delineating walls. Repetitive, clearly bounded, orthogonal shapes are read well. Irregular geometry and complex facades are not.
Note also that the median deviation across all items was 0.0%. If you had checked only the middle of the distribution you would have concluded there was no problem at all. The bias shows up in the ranks, not the median, which is exactly the kind of error a spot check of three lines will miss.
Why a downward bias is the dangerous direction
An estimating error that runs high loses you the tender. You are underbid, you find out immediately, and the cost is an opportunity.
An error that runs low wins you the tender. You find out on site, months later, when the material order does not cover the elevation. The cost is real money, and it comes out of a margin that the cidb survey already describes as being competed down.
So the two errors are not symmetrical, and a tool whose residual bias points downward is a tool to check in one specific place rather than a tool to distrust generally. The place is any measured area on a vertical or irregular surface: cladding, plaster, paint, glazing, curtain walling, anything on an elevation with returns and reveals in it.
The check that follows from the finding
- Accept the counts. Doors, windows, fittings, sanitaryware, luminaires. The evidence says these match. Reviewing them by hand spends your attention where there is no error to find.
- Re-measure the elevations by hand. Not all of them, and not the whole drawing. The facade areas, because that is where the deviation concentrates.
- Compare the total, not a sample. A systematic bias of a few percent spread across fifteen line items will pass any three line spot check and still be sitting in the total.
- Refuse scanned drawings. Every published measurement result assumes a native vector drawing. A photographed or scanned set is a different problem, and these numbers do not transfer to it.
One more caveat worth carrying, because it is usually dropped. The widely repeated figures of a 51.3% reduction in takeoff time and a 20.4% improvement in measurement accuracy come from Zhao and colleagues in 2025, and that study was run in an undergraduate estimating course, not on a live commercial bid. Students improving against their own baseline is a real result about teaching. It is not a claim about your estimator.
The part of a construction business AI is genuinely good at, and nobody writes about
Takeoff gets the attention because it is visible and technical. It is not where most contractors lose most hours.
The hours go into the paperwork around the job: the tender pack, the compliance documents that have to be current on the day of submission, the variation orders, the progress claims, the retention releases, and the supplier invoices that must be matched to a job before anyone can say whether the job made money.
All of that is document handling, and none of it requires judgement about geometry. It requires the same fact, captured once, to appear correctly in eleven places.
Where the manual work actually stops
Five mechanisms, in the order they repay the effort. Not one of them needs a takeoff tool.
- Capture the job record once, at award. Client, site, contract sum, retention percentage, payment terms, practical completion date, responsible person. Everything downstream reads from that record. Nothing is re-typed onto a claim, a variation or an invoice, which is where the transposed digit lives.
- Hold the compliance pack as dated fields, not as a folder. Tax clearance, cidb registration, B-BBEE affidavit, letter of good standing and CSD registration all carry expiry dates. Held as dates, they can raise a reminder weeks before a bid is disqualified for a lapsed certificate. Held as PDFs in a folder, they can only be discovered stale on the morning of the deadline. Our guide to Central Supplier Database registration covers what has to stay current, and document management covers where it should live.
- Number and log every variation as it is instructed, not at month end. A variation captured on the day carries the instruction, the date and the person who gave it. Reconstructed six weeks later it carries an argument.
- Drive the progress claim from the job record. The claim is a calculation over things you already hold: work measured to date, previous claims, retention, less what has been certified. Rebuilding that spreadsheet every month is the definition of work that should be a review rather than a rebuild.
- Put the certified date on a clock. Payment cycles in this sector run long. A claim that has passed its certification date without payment should surface itself as a task with an owner and a due date, which is the same mechanic as invoice follow up, pointed at a certificate instead of an invoice.
Read together, those five move the work from reconstruction to capture. Reconstruction is what makes construction admin expensive: not the entry itself, but the weeks of distance between the thing happening on site and the person in the office trying to work out what it was.
Before you buy anything
The common failure is not bad software. It is buying a takeoff engine to solve an administration problem, a mismatch we see often enough that we wrote why AI projects fail about it and turned the supplier questions into an AI vendor checklist. If your hours are going into claims and compliance rather than measurement, a takeoff tool will not give them back. General business process automation and admin automation are the nearer neighbours.
If you want the honest version of which one it is for your business, a Business Autopsy maps where the hours actually go before anyone recommends a tool.
Frequently Asked Questions
Is AI accurate enough for construction estimating?
For counted items, yes. Peer reviewed testing found no significant difference against a contractor takeoff, with 11 of 13 count items identical. For measured areas it showed a significant downward bias, worst on exterior facades, so elevations still need human re-measurement.
What does AI for construction do best?
Repetitive, clearly bounded, orthogonal work: counting doors and fittings, measuring ceilings and floors, extracting fields from documents, matching invoices to jobs. It is least reliable on irregular geometry and complex facade areas.
Can AI read scanned or photographed drawings?
Far less reliably. Published accuracy results assume native vector drawing files. A scanned or photographed set is an image recognition problem rather than a geometry problem, and measurement results from clean drawings do not transfer to it.
Where should a small South African contractor start?
With the paperwork, not the drawings. Compliance certificate expiry dates, variation logging, and progress claims driven from a single job record return more hours than takeoff automation for most contractors in the lower cidb grades.
Does AI replace an estimator?
No. The Utica University study concluded that estimator oversight remains essential for accuracy assurance in complex building elements. The realistic gain is moving estimator time from tracing measurements to pricing, constructability and risk.
Sources
- cidb SME Business Conditions Survey, Q4 2025, conducted by the Bureau for Economic Research, Stellenbosch University, on behalf of the Construction Industry Development Board. Published February 2026.
- Sadeh, H., Geloso, D. and Todorov, D. (2026). Testing AI Accuracy in Quantity Takeoff: A Methodological Case Study in Commercial Construction. EPiC Series in Built Environment, Volume 7, pages 753 to 762. Proceedings of the Associated Schools of Construction 62nd Annual International Conference.
- Zhao et al. (2025), cited within the above, for the undergraduate estimating course results.