All whitepapers

What clean data in the field actually means

The message you will not get

A field worker with muddy hands, gloves on, in the rain, will send one message. Not an app, not a login, not a form. We designed for that sentence, and everything in this paper follows from taking it literally.

The message might say: “we hit solid rock at the 200 mark, cant get through with the 3t.” No units, no cost code, no site reference, missing apostrophe and all. That is how crews actually write, and it matters that they do. Somewhere inside that message is a fact that will move a five-figure number: the ground has changed, the rate that was bid is no longer the rate being achieved, and every remaining metre on that stretch now costs something nobody has priced.

The person who needs that fact is a project manager who will otherwise meet it at invoicing. In field construction, the facts that explain why a site drifted arrive as voice notes and photos in message threads, and then get retyped, partially, into spreadsheet side-files. The loss happens at the retype. It is invisible because nobody logs what they did not retype.

How large is the gap between the number a job was priced at and the number it closes at? The largest published measurement is Flyvbjerg, Skamris Holm and Buhl (2002): 258 transportation infrastructure projects worth about US$90 billion, across 20 nations, comparing actual construction cost at completion against the budget at the decision to build. In 9 out of 10 projects, costs were underestimated. The average escalation was 28% across all project types, and the authors found no improvement over 70 years of data: “no learning that would improve cost estimate accuracy seems to take place.” Two honest notes on this source. Their measurement is public transport infrastructure, and a competitive fibre bid is a different measurement, so we use the study as evidence that the gap is real, persistent, and global rather than as a rate to apply to anyone’s bid. And the authors’ own preferred explanation is strategic misrepresentation at the approval stage, deliberate lowballing to get projects built, which no field-data tool addresses. What field data does address is the part of the gap that opens after the decision: the ground that changed, the standing time nobody recorded, the week that went slow for a reason that never made it into a record.

Actual cost against the estimate, measured across 258 projects

and in 9 out of 10 projects, costs are underestimated: an 86% likelihood for a randomly selected project

Average escalation of actual construction cost over the budget at the decision to build. 258 transportation infrastructure projects worth about US$90bn, across 20 nations. Flyvbjerg, Skamris Holm & Buhl (2002), Journal of the American Planning Association 68(3), open access at arXiv:1303.6604.

This paper is about what “clean data” means under those conditions. Our argument, compressed to a sentence: data quality in the field is decided at acquisition, by protocol, before any database, dashboard, or model exists. The most expensive data defect is the message that was never sent. And “clean” itself turns out to be a relation between a dataset and the question you intend to ask of it, a relation you can still fix at the moment of capture and rarely afterwards.

There is a measured cost literature for what missing field data does to construction businesses, and there is a published machine-learning study whose method demonstrates the relation argument better than any theory could. This paper reads both closely and takes positions on them. We start with the money.

What the missing message costs

The strongest available number comes from the one study that had a contractor’s own books rather than a survey. Love, Smith, Ackermann, Irani and Teo (2018) analysed 19,605 rework events across 346 construction projects delivered by a single contractor between 2009 and 2015, in what its authors describe as the first longitudinal, in-depth study of rework costs in construction. Over the period, rework reduced the contractor’s mean yearly profit by 28%.

28%of mean yearly profit lost to rework19,605 rework events · 346 projects · one contractor's audited books, 2009–2015
Love, Smith, Ackermann, Irani & Teo (2018), Production Planning & Control, 29(13), described by its authors as the first longitudinal, in-depth study of rework costs in construction.

One scoping note before we lean on this number, because precision here is what separates a meta-analysis from a pitch. Rework has many causes: design errors, scope changes, workmanship, and information failures together. No published source decomposes rework cost by cause, so the 28% is the ceiling on what information failure could be worth, and the true share belongs to the field measurements this industry has not yet made. What the rework literature does establish is the shape of the cost and where it hides, and both points carry.

The shape first. 88 of those 19,605 events, 0.45%, accounted for 34% of the total rework cost. The cost is a tail risk rather than a steady tax. A reporting system that captures the average event and misses the exceptional one has captured almost nothing, because the money sits in the exceptions: the rock nobody reported, the redesign nobody logged, the asphalt that was priced as soil.

The cost of rework is a tail, not a tax

The same 88 events, as a share of all events and as a share of all cost. 19,605 rework events across 346 projects, one contractor's audited books, 2009–2015. Love, Smith, Ackermann, Irani & Teo (2018), Production Planning & Control.

Three more findings frame the problem:

Outside construction, Haug, Zachariassen and van Liempd (2011) give the cost of poor data quality a structure worth borrowing: costs split along two axes, direct versus hidden and operational versus strategic. Hidden costs are the ones not normally included in any purchase price, so nobody budgets them, and what is not budgeted is not managed. A missing progress report lands in the hidden-strategic cell, the corner of the grid where money disappears without an invoice ever saying so. Their case study puts a figure on the most strategic hidden cost of all: the company’s inability to know its own production costs at the point of quoting was estimated, in that single case company, at 5–7% of total fixed costs. One case study, carefully done, not an industry statistic; we cite it for the mechanism. Sales staff could not set the right price because the cost data was not there. A contractor pricing the next site off gut feel, because the last site’s achieved rates were never recorded, is the same failure wearing high-vis.

Where the cost of a missing report lands

The two axes and the cell definitions are Haug, Zachariassen & van Liempd (2011); the field-construction examples in each cell are ours. Hidden costs are the ones "not normally included in the purchase price": unbudgeted, therefore unmanaged.

One finding from the 2018 study matters more to us than its numbers. The authors observe that rework costs have “a proclivity … to be largely ignored, concealed or considered to be normal function of operations”. Of their 346 projects, complete rework cost data existed for 98. The industry’s canonical cost-of-bad-data study had to be a bespoke six-year research effort because the data is not routinely captured, and even then, capture was complete on 28% of the projects of a contractor that was actively cooperating. The literature that measures the cost of missing field data is thin for the same reason the cost exists.

A boundary we state plainly: all of the measured evidence above comes from buildings and general infrastructure. We found no published study measuring any of this for telecom civils specifically, and no measured completion-rate statistic for site reporting anywhere in the open literature. This paper states that gap rather than borrowing adjacent numbers to fill it.

The conversation that stays open

If the dominant failure is omission, the useful design question is not “how do we clean the data?” It is: why was the message never sent, and what would have to change for it to be sent?

The answer is not that field workers are careless. The worker knew about the rock; they were standing on it. The information existed all day, in a person’s head and in the ground. What never existed was a cheap enough way to move it into a record. The crew lead is driving back from site at the end of a ten-hour day. The project manager, the one person with a reason to want the number, has no time to chat with every crew, transcribe voice notes, and copy figures out of message threads into a spreadsheet. The thread scrolls away; the retype never happens. Both people are behaving rationally. The record loses.

That is the gap we are building for, and the mechanism deserves a precise statement, because it is neither a form nor a data-cleaning pipeline.

The acquisition is a conversation, and the conversation stays open until the record is complete. A worker sends whatever they can, however they can: a voice note, a photo, a few words in their own language, over the messaging thread already on their phone. The system does not need a complete, structured record out of that first message, because it does not have to stop there. Where a form fails silently and a dashboard waits passively, a conversation answers. It confirms what it understood, in the worker’s own terms. It asks for what is missing, at the worker’s pace, resuming whenever the worker can reply. It carries the burden of the follow-up itself, so that neither the worker nor the project manager has to. The target is completeness reached inside the conversation, minutes or hours after the work happened, instead of reconstructed weeks later from whatever survived the retype.

The conversation stays open until the record is complete

The mechanism: never block on a parse, ask for what is missing at the worker's pace, and close the record while there is still someone to ask.

Two practical questions belong here, because every practitioner asks them.

Why would the worker answer? For the same reason helmets get worn on a site where helmets are checked: reporting is part of the job, and the consequences of not reporting already exist today, for the worker and for the company. What decides whether compliance actually happens is friction. A crew lead who will never fill a 150-line form at the end of a shift will say, in one sentence, “all on plan, just spent two hours longer on the roof iron because of the rain”, because saying it costs nothing. The conversation wins not by inventing a new incentive but by dropping the cost of complying below the cost of ignoring it. Answering one specific question is easier than composing a report, and easier than explaining next week why the record is empty.

What about a trench with no signal? Nothing about a conversation requires it to be live. Crews already write notes in whatever app they like when they are offline and send them when coverage returns; the thread picks up where it left off, and the follow-up questions wait. The shift this design makes is from forms to messages, and messages already tolerate every connectivity pattern a site produces. Where the work needs more than text, photos or structured QA capture, a purpose-built surface can serve that need; the spine of the protocol stays the conversation.

What this replaces is the assumption underneath cleaning: that data arrives once, in whatever state, and quality is manufactured afterwards. In a conversational protocol the record is not considered acquired until it is complete, so the qualities everyone wants, attributed, unit-carrying, traceable, are established while there is still someone on the other end to ask. The rest of this paper argues that this ordering is a constraint rather than a preference: cleaning operates on what arrived, and in the field, the cost that took 28% of a contractor’s profit is dominated by what never arrived at all.

“Clean” is not a property

The data-quality literature settled long ago on “fitness for use” as the definition of quality: Juran’s phrase, made canonical for data by Wang and Strong (1996), whose framework treats quality as fitness for the data consumer’s purpose. We are not claiming that insight. Our claim is about when the relation must be fixed: before ingestion, at the capture protocol, because everything downstream of capture inherits what capture discarded. A published machine-learning study demonstrates the point with unusual clarity, precisely because its authors did everything right.

Deza, Ihshaish and Mahdjoubi set out to classify construction cost descriptions into the ICMS cost standard using machine learning (arXiv:2211.07705, a preprint, submitted to Engineering Applications of Artificial Intelligence). A naming note their own text forces on us: the paper’s title expands ICMS as International Construction Measurement Standard, its body as International Cost Management Standard, and the standard’s own coalition writes International Construction Measurement Standards. The source is inconsistent, so we use the title’s expansion once and the acronym hereafter. They retrieved 123,210 materials-and-costs line items, natural-language text manually labelled by quantity-surveyor experts with the support of the Royal Institution of Chartered Surveyors, from 24 UK infrastructure projects, and modelled 51,906 of them.

The closest published corpus to field text, and it is the tidy end

Bills-of-Quantities descriptions, written deliberately, for money, by subcontractors, and still this inconsistent. Deza, Ihshaish & Mahdjoubi (2022), arXiv:2211.07705 (preprint), §2.1.

To prepare that text for their models, they cleaned it. Special characters, punctuation and numbers were removed. The price breakdown was excluded from the study entirely.

For their question, which of the standard’s cost categories does this line belong to, that is correct preprocessing: a quantity carries no signal about a category, and stripping it reduces noise. Their results vindicate the choice.

For our question, what did this metre actually cost, the identical cleaning step is total data loss. Every number is gone and the price column was never in scope. The corpus, cleaned correctly by a competent team for their purpose, contains no answer to ours at all.

One corpus, one cleaning pipeline: correct for one question, fatal for the other. “Is this data clean?” has no answer. “Is this data clean for this question?” does. The consequence is practical, not philosophical: the cleaning question has to be answered before ingestion, because by the time a model exists, the pipeline has already discarded whatever it discarded. Most teams answer it after. They collect, then clean, then discover what the cleaning cost them. The capture protocol is the last place where the answer is still cheap.

We should be honest about what this comparison can carry. A Bill of Quantities is written deliberately, in a billing document, by subcontractors, for money. It is the tidy end of construction free text, and it still needed cleansing; still recorded the same fact four ways (the source’s own examples: ‘cable 10 m’, ‘cable 10 meter’, ‘cable 10m’, ‘ten metre cable’); still carried misspellings; still had labels that, in its own authors’ words, come down to subjective judgment. If the deliberate, invoiced end of the industry’s text is that inconsistent, a voice note from a trench is strictly harder. We use their corpus only as an argument from the easier case, in that direction, and we transfer no number from their dataset to ours.

Why entry validation is a data-loss mechanism

The standard reflex, on hearing “the input is a mess”, is to validate at the door: required fields, format checks, “invalid entry, please resend.”

In the field, every validation placed in front of the message is a message that never arrives. The worker sending one message from the rain does not have a second attempt in them, so the rule we designed to is blunt: never block on a parse. A message the system fails to understand is still accepted, flagged for a human to look at, and answered honestly with “a person will check this”, because a sender who gets “invalid format” back will not resend, and the fact they were carrying dies in the doorway. Not rejecting extends even to senders the system does not recognise. A message from an unknown number is accepted and marked as exactly that, because “we could not attribute this” is recoverable later and “we discarded it” never is.

Does that mean accepting garbage into the books? No, and the boundary between those two things is the design rule we would defend above all the others.

Accept any message; be strict about every row. The conversation never rejects; the data model always may. A cost that cannot be attributed to a site is refused as a row, because an unattributed cost is exactly the failure mode this design exists to kill. But the worker’s message reporting that cost has already been accepted, and the resolution is a question asked in the conversation, not an error returned to a person standing in a trench. Blocking happens at the boundary of the data model and never at the boundary of the conversation. The same split resolves the permissions case: a worker not authorised to report costs still gets their message accepted; it simply creates no cost entry, and the matter passes to the crew lead. Accepted as a message; refused as a row.

Validation is not removed; it is relocated. It becomes conversational, asynchronous, and never a precondition of acceptance. That single relocation is most of what “protocol, not cleaning” means. It is also the part a form-based system cannot retrofit: when the product surface is the form, rejection is the interface, and a vendor whose system of record is built on validated submission can bolt a chat window onto the front without changing the fact that an incomplete record bounces. Non-rejection is a data-model decision, and it has to be made at the foundation.

Eight properties you can write a check for

“Clean data” earns the adjective only if it is testable. Below are the properties we hold the data model to, each stated as the check you could write for it. To be precise about epistemic status: these are the design’s commitments, what the system is built to guarantee, not measurements of a shipped product’s behaviour, and this paper makes no claim of that second kind anywhere. Readers who know event-sourcing patterns, provenance standards, or double-entry bookkeeping will recognise their ancestors here; the properties are proven patterns applied to a domain that mostly lacks them, not our inventions.

# Property The check you could write
1 Provenance. Every stored fact carries the identity of the message that produced it. Every number opens the message it came from. No progress figure, cost entry, or blocker can exist without one.
2 Non-destructive correction. A correction is an append, never an edit. Nothing already recorded is ever rewritten in place; a worker’s undo works by adding a countering entry, never by deleting one.
3 Recomputability. Nothing is displayed that cannot be rebuilt from what was reported. Recompute every derived figure from the underlying reports; compare with what was served. Any non-zero difference is a defect.
4 Explicit units and money. No bare number, anywhere. No quantity travels without its unit; money is whole minor units against an explicit ISO-4217 currency, never a floating-point value.
5 Mandatory attribution. A cost that belongs to no site cannot exist as a row. Recording an unattributed cost fails; the conversation resolves the attribution instead.
6 No silent loss. A message the system failed to understand is visible unfinished business, never a gap. Any message still awaiting human review after four working hours has raised an escalation, because a silent parse failure is worse than no system at all: the worker believes they reported it.
7 Absence is a value. “Nothing happened today” is stored as a zero-progress report, not left blank. A day reported as no-progress is distinguishable from a day nobody reported at all. “No work happened” is information.
8 Determinacy. A derived figure’s definition admits exactly one reading. Two independent implementers, given only the definitions, compute the same number on the same test data.

Where a careful reader should push on this table:

Property 5 looks like a contradiction of the previous section. We spent a page arguing against rejection, and here is a refused write. The resolution is the boundary from that same page. The message carrying the unattributed cost is accepted and safe; it is the row that waits until the conversation has established which site the money belongs to. Strictness about rows is what makes tolerance about messages affordable.

Properties 1 and 2 are what keep repair possible. Cleaning still happens downstream of a protocol like this: units get normalised, duplicates reconciled, a wrong quantity superseded. What the protocol decides is whether cleaning remains possible. When every transformation is non-destructive and every fact names its source, a defect found in month six is fixable in month six. A protocol failure, the message never captured or the source never linked, is the one kind of data-quality failure no downstream effort reaches.

A tempting ninth property is wrong, instructively. “The same quantity is never stored twice” sounds like obvious hygiene. In an append-only model it is false by design: a correction is the same quantity stored again, on purpose, next to what it replaces. The property that actually holds is that the same quantity is never counted twice. A correction names exactly what it supersedes, and every inbound message carries an identity that makes a redelivery harmless. Storage de-duplication and accounting de-duplication are different properties, and conflating them is how an append-only record gets “tidied” into a mutable one by someone trying to help.

Property 8 is listed last because it is the one we have tested hardest. The full story is in the trade-offs section. The short version: the first seven properties check a value against a definition, and running the determinacy check against our own definitions proved that a definition admitting three readings cannot be checked at all.

What cleaning cannot reach

Back to the preprint, for the finding that sets the ceiling on every downstream repair.

Deza et al.‘s models confused two of the standard’s categories in a specific, diagnosable way. The distinguishing property was whether an installation was permanent or temporary, and, in the authors’ words, neither the word “permanent” nor “temporary” was necessarily present in the description. The classifier was not too simple; a bigger one would not have helped. The information required to be right was not in the input. Their stated remedy is more training data “of diverse contextual nature”, which is a capture change prescribed by a modelling paper.

That is the hard ceiling. Cleaning a CSV recovers structure that was lost in transit: the value is there, mangled, and effort restores it. Field data’s usual problem is worse. The structure existed, in the worker’s head and in the physical world where the trench either got dug or did not, but no machine-readable encoding of it was ever made. A fact the message never carried cannot be recovered by any model, pipeline, or budget applied downstream, because there is nothing to recover from. It can only be asked for while there is still someone to ask, and that is precisely what a conversation that stays open until the record is complete is for.

Why the model is not the problem

The same study also answers the instinct that a hard data problem deserves a sophisticated model. Three of its findings, then one piece of our own arithmetic.

First: the capture decision bounded the model before the model existed. The cost side of the ICMS standard has 109 categories. In a 24-project corpus, 37 had no entry at all, and only 32 cleared the 250-samples-per-category minimum the authors required, a cut-off they themselves describe as set “relatively arbitrarily”. A further large share of the retrieved corpus went unmodelled for thin coverage and duplicate removal together; the source does not separate the two causes, so we do not lean on the combined figure. The category coverage alone carries the point. No architecture recovers a category nobody recorded. The schema, what got captured, at what granularity, distributed how, decided the model’s reach years before anyone trained anything.

The schema decided the model's reach before anyone trained anything

ICMS cost-side categories covered by a manually labelled corpus of 24 UK infrastructure projects, at the authors' 250-samples-per-category minimum. Deza, Ihshaish & Mahdjoubi (2022), arXiv:2211.07705 (preprint).

Second: sophistication was not the binding constraint. Five model families competed. The best was a multilayer perceptron with a single hidden layer over bag-of-words, at 0.930 macro F1, above the bidirectional GRU at 0.913 and the bidirectional LSTM at 0.907. The recurrent architectures, in the authors’ analysis, overfitted text too simple to need them: the median description is 14 words, and the signal sits in local key features, making long-range memory “largely inconsequential at this task”. One task, one corpus. This does not generalise to a law, the authors decline to over-conclude from it, and we match their restraint. It does generalise to a question worth asking in every planning meeting: is the model the constraint here, or is the input? If you have ever suspected that a team reaching for a bigger model is avoiding a data conversation, this is that suspicion with a results table attached.

The simplest model won

Five model families classifying 51,906 construction cost descriptions into 32 ICMS categories. The recurrent networks overfitted text whose median length is 14 words. Deza, Ihshaish & Mahdjoubi (2022), arXiv:2211.07705 (preprint), Table 1. Scale cropped to 0.90–0.94; the values are printed.

Third: the errors that remained were input errors. The permanent-versus-temporary confusion of the previous section is where their accuracy stopped, against absent information, with a capture change as the named remedy.

Now our own arithmetic, pointed at ourselves. A single deployment at the scale our worked scenario plans for, 4,850 metres of trench across six sites, yields single-digit daily observations for most surface types. At six observations, and assuming the day-to-day variability that ground-driven work exhibits (coefficients of variation between roughly 0.5 and 0.8), the 95% confidence half-width on a unit rate runs about ±40% to ±64% of the estimate itself. No architecture fixes n = 6, and publishing a learned conclusion on n = 6 would be a confident number with nothing underneath it, the exact defect this paper is about. So the sequence is fixed, and each stage exists to make the next stage’s inputs real: clean data model → descriptive statistics → reference-class distributions → and only then anything learned. Skipping ahead does not skip the work. It relocates the work downstream, to a place where it can no longer be done.

That sequence is also where the value compounds. Every completed record adds to a provenance-linked history of achieved rates by ground type, crew, and season, and the reference classes built from that history are what let the next bid be priced against measured reality instead of averages and gut feel. Reference-class forecasting is the one method in this literature with a track record against optimism, and it runs on exactly the kind of data this protocol is designed to capture. A record like that cannot be bought or backfilled; the overwrites and the unsent messages of the past are gone. It can only be accumulated, deployment by deployment, by whoever starts capturing correctly first.

Build these six things now, defer the rest

A data model for field acquisition scales or fails on a short list of early decisions, and there is a one-question test for which list any decision belongs to: if this is missing in a year, is the missing thing information that was never written down? If yes, no migration recovers it, because there is nothing to migrate. If no, it is a feature, and features can wait.

Six things pass the test. All six fail the same way when skipped, and all six are cheap on day one:

Nearly everything else defers safely: localisation (a per-worker language preference is a field from day one, so translating the product later is a feature rather than a migration), finer-grained roles, tunable escalation thresholds, per-activity cost allocation, mapping, invoicing, cross-deployment benchmarking. The pattern: leave the column, defer the feature. A nullable field costs nothing today and converts a future migration into a future afternoon. That is what “the data model scales” actually means. Not that it does everything now, but that nothing done now forecloses anything later, because the six expensive-to-reverse decisions are the ones you can never buy back.

Seven positions we take, and what each one costs

A design is a set of trade-offs, and a paper that will not name its own trade-offs is a brochure. Here are the positions this design takes, each with the cost we accept for it, because a reader who can see what a design gives up can trust what it claims.

  1. The headline unit rate is an approximation, by choice. It divides all trenching-phase cost by metres trenched. The cleaner method, per-activity allocation, requires workers to attribute every cost to a task, and they will not reliably do that. We take the approximation and label it; the error it admits grows with however much non-trenching cost lands during the trenching phase. A labelled approximation beats an unlabelled precision that depends on behaviour no crew exhibits.
  2. Forecasts fail loud, not quiet. Every forecasting rule has inputs that produce nonsense, especially in a deployment’s first days when history is thin. Our position is that the failure modes get documented and bounded rather than hidden, because a reader who knows where a forecast breaks can calibrate against it, and one who discovers it alone stops trusting every number the system serves.
  3. An idle site defaults optimistic, and a separate rule exists because of it. With no recent reported metres, remaining work is forecast at the bid rate, so a stalled site looks fine on rate. That is why silence is its own first-class signal with its own escalation rule, rather than something the rate calculation is asked to detect. Defaults that err optimistic deserve to be said out loud.
  4. The worker’s judgement is recorded and quarantined. The worker’s own percentage-complete estimate is captured and deliberately excluded from every derived figure; the gut number and the computed number never mix. A judgement is data about the judge, not about the trench, and keeping it separate preserves both.
  5. Transcription is treated as inside the trust boundary, and provenance is the answer. A voice note’s transcript is a derived artefact, and a wrong transcript parses into a plausible wrong row. This is precisely why every number opens the message it came from: the design assumes some derived values will be wrong and makes each one auditable in seconds, instead of assuming a pipeline can be made incapable of error.
  6. A committed cost stays visible until it resolves. Committed cost enters the forecast before anything is billed, correctly, since the money is already promised. Site acceptance then requires every committed cost to have become an incurred one or been explicitly written off, so nothing exits the books silently. The residual cost we accept: a committed cost that never converts sits in the forecast until closeout, and we prefer that conservatism to aging it out by rule.
  7. Determinacy has to be tested, not assumed. We know because we ran the test. Property 8 says two independent implementers, given only the definitions, must compute the same number on the same data. We ran exactly that check against our own definitions, and it earned its place on the list: the phrase “the last 14 days” admitted three defensible readings, which on the same worked scenario produce forecasts $10,352 apart, about 30% of the $35,000 bid. Both implementations were correct. The sentence was the defect, and the fix was one clarifying clause naming calendar days and the empty-window default. This is why determinacy is the property that makes the other seven meaningful: they all check a value against a definition, and a definition admitting three readings cannot be checked at all. No test of outputs could have caught it; the code was right on every reading. A check on definitions is the only net this class of defect cannot slip, and we now run it on every derived figure we define.

One sentence, three forecasts

widest − narrowest = $10,352, about 30% of the $35,000 bid

Three good-faith readings of "the last 14 days" in our own specification, each computing the forecast on the same worked scenario. Both implementations that surfaced this were correct; the sentence was the defect. The fix was one clarifying clause.

What would falsify this

The properties are stated as checks, so they can fail, and the argument should be attackable at its joints. Here are the joints, including the one that would hurt most.

“The conversation closes the record.” The claim that carries the most weight rests on worker behaviour: crews answering follow-up questions at their own pace instead of ignoring them. If they mute the conversation the way they abandon forms, the record stays open and this design has failed at its center. We think the friction argument holds, answering one question is cheaper than composing a report or explaining a gap later, but the honest test is quantitative: completion rates of records acquired conversationally versus records reconstructed from passive capture, measured on deployments, published either way. The literature supplies the counterweight we expect to be tested against. In Haug et al.’s case company, management’s own view was that perfectly correct data “would never be possible nor expedient”: the optimal data quality level sits below 100%, because past some point the next unit of quality costs more than it returns. We agree, and the design reflects it. The conversation pursues the data the record needs, the attribution, the quantity, the unit, and not all the data there is.

“Protocol, not cleaning.” If everything we file under protocol turns out to be recoverable after the fact with enough effort, the paper collapses into “store your raw data”, which is not novel. Our defence is the non-rejection half, which is not in the standard advice: no downstream tooling recovers a message that bounced off a validation gate, and no cleaning pass reopens a conversation with a crew that finished the job in March.

“Clean is a relation.” The relation itself is fitness-for-use, established doctrine since Wang and Strong. What is attackable is our timing claim, that the relation must be fixed at capture. A reader who accepts the relation but locates the fix downstream owes an account of where Deza et al.’s deleted numbers come back from.

Method note

Every factual claim in this paper traces to a source read at full text: a peer-reviewed publication, one preprint (arXiv:2211.07705, flagged as a preprint throughout), or our own design documents, with derived arithmetic and its assumptions shown where we computed it. Where the preprint’s abstract and body disagree, we cite the body’s precise figures. Claims about our mechanism describe design commitments rather than a shipped product’s measured behaviour, and the worked example running through the paper (a fibre build of 412 homes, 4,850 m of trench, six sites, $208,000 bid total, with the $35,000 site quoted throughout) is our own worked scenario. The cost-escalation figures in the opening are Flyvbjerg, Skamris Holm and Buhl’s measurement of public transport infrastructure, and we state rather than blur that a competitive fibre bid is a different measurement. Several widely circulated cost-of-bad-data headline figures were considered and deliberately excluded, and since other papers build their case on them, we name them. The claim that poor data costs the US economy “$611 billion a year” traces through secondary citations to vendor material from around 2005, with no published method or dataset behind it. The “$3.1 trillion a year” figure and the “15–25% of revenue” range circulate just as widely and trace the same way: to marketing collateral and opinion surveys, not to a measurement anyone can check. The often-quoted “8–12% of revenue” range goes back to Redman (1998), who attributed it to proprietary studies that were never published. These are the biggest numbers in the field, they would have made our opening more dramatic, and none of them survives the test we apply to every other figure in this paper. If a number is not in the references below, we do not stand behind it.

References