What is “bad data” in a business context? Usually something far less dramatic than a corrupted database. A product weight that says 2.4 in one system and 2400 in another, with nobody sure whether the unit is kilograms or grams. Or a customer that exists three times in the CRM under slightly different spellings. Or a supplier record with a payment term that was renegotiated a year ago but updated in only one of four systems. Bad data looks mundane. That is exactly why it survives and persists in your systems for years.
The data involved is the everyday operational core of a business. Product data: descriptions, attributes, prices, images, technical specifications, packaging details. Customer data: contacts, addresses, terms, history. Supplier data: certifications, lead times, conditions. And the transactional data built on top of all three: orders, invoices, shipments. When the core records are wrong, every transaction that references them inherits the error.
This raises an obvious question: with so many B2B tools built for data management, why is the data still wrong? A typical mid-sized manufacturer runs an ERP, a CRM, a webshop, a marketing platform, spreadsheets for everything the other systems don’t cover, and often some shadow tools that individual departments signed up for on their own, without IT department involvement. Each tool manages its slice of data reasonably well. The problem sits between them. Every system holds its own copy of the same product, the same customer, the same supplier. The copies drift apart with every manual update, partial import, quick fix made under deadline pressure. There is no single place where the correct version lives, so “correct” becomes a matter of which system you happened to open.
So bad data is fundamentally an alignment problem. Buying another tool adds another copy to keep in sync. Errors are born in the gaps between systems, and money quietly leaves through the manual work of bridging those gaps.
The scale of that leak shows up from two directions. At the record level, a Harvard Business Review study had 75 executives check the last 100 records their own departments created and mark the errors. On average, 47% of newly created records contained at least one critical, work-impacting error, and only 3% of the resulting quality scores rated acceptable by the loosest standard. At the company level, the same problem translates into money. In a Gartner survey of 154 enterprise customers of data quality vendors, respondents estimated that poor data quality costs them an average of $12.9 million per year. The sample skews toward large enterprises, so the absolute number matters less than what it signals: companies that measure the problem find it expensive.
Where The Money Goes
Bad data tends to cause friction rather than visible disasters, and friction compounds. Thomas C. Redman, a data quality researcher and author, calls the mechanism the “hidden data factory“. When a salesperson receives a flawed prospect record, they fix it themselves, and when finance gets an order with a wrong price, someone corrects it by hand. When two systems disagree, an analyst spends an afternoon reconciling them. Each of these fixes takes only minutes, which is exactly why they rarely get logged, budgeted, or noticed by anyone tracking costs.
The second mechanism does more damage: decisions made on numbers that should not have been trusted. For example, a manager who doubts the forecast will hedge it, and a planner who has been burned before adds extra stock just in case. The same happens in marketing, where campaigns often run on customer records that went stale a year ago. On the balance sheet, all of this simply shows up as slower growth and thinner margins.
The costs also escalate with time. A widely used rule of thumb says an error costs roughly 1x to prevent entry, 10x to correct later, and 100x once it has driven a bad decision or reached a customer. The exact multipliers are folklore more than science, but the direction they point in is well documented: the later you catch an error, the more it costs.
Why Companies Don’t Notice
According to Gartner surveys, 59% of organizations do not measure data quality at all. That single fact explains most of the problem. You cannot manage a cost you never see, and data errors are structured to stay unseen. They are distributed across departments, absorbed as individual workarounds, and blamed on anything but the data itself. A missed delivery gets blamed on logistics, while a failed campaign is written off as weak creative. When the forecast turns out wrong, the market takes the blame. The data behind all three failures rarely comes up.
There is also an ownership gap. Data quality sits between IT and the business, so it often belongs to neither. IT maintains the systems but does not know whether a supplier record is correct. The business knows the record is wrong but has no mandate or tool to fix it at the source.
What Good Data Quality Means In Practice
“Good data” is not one thing. It is a set of measurable dimensions, and different failures cost you in different ways:
| Dimension | The question it answers | Typical business impact when it fails |
| Accuracy | Does the value match reality? | Wrong prices, wrong shipments, wrong invoices |
| Completeness | Are required fields filled? | Products that cannot be listed, orders that stall |
| Consistency | Do systems agree with each other? | Reconciliation work, conflicting reports |
| Timeliness | Is the data current? | Decisions based on stale customer or stock data |
| Uniqueness | Does each entity exist once? | Duplicate customers, double outreach, skewed analytics |
| Validity | Does the value follow the defined format and rules? | Failed imports, broken integrations, compliance gaps |
The table matters because it turns a vague complaint (“our data is bad”) into something you can measure, assign, and improve field by field.
How To Build A Process That Holds
Fixing data quality means changing how records get created, which is why one-time cleanup projects keep disappointing. Six months after a heroic deduplication effort, the same error rates return, because the processes that produced the errors never changed. Lasting improvement comes from a continuous loop: define rules, measure against them, catch violations early, and fix causes rather than symptoms.
The practical sequence looks like this:
Define quality rules per field and per entity. “Every sellable product needs a net weight in grams and at least one image” is a rule you can enforce; “product data should be complete” is not.
Assign ownership. Every critical data domain needs a named person who can approve and correct records, a person, not a committee, and not “IT”.
Measure continuously, not annually. Automated checks on new and changed records catch errors at entry, where they are cheapest to fix, before they reach an invoice, a shipment, or a forecast.
Fix at the source. If a field keeps arriving empty, change the intake form or the import mapping, not just the records.
A common pattern at mid-sized manufacturers illustrates the loop in practice. Product and supplier data lives in ERP exports, spreadsheets, and inboxes, and three departments maintain three versions of the same record. Most fixes at this stage are procedural: audits, naming standards, required fields, and defined ownership per attribute, as outlined in guides on product data quality. Once critical records sit in one governed system with validation at entry, most of the recurring correction work disappears within months, and that result is achievable on many platforms.
Measuring The Return
Data quality initiatives compete for budget with projects that promise revenue, so they need their own business case. The baseline numbers already exist: correction hours, order error rates, duplicates, time-to-market. Tracked before and after, they make the return visible. Reduced manual work shows up first and fewer downstream errors follow, while the largest gain arrives last, when forecasts, pricing, and planning finally rest on numbers people trust. Analyses of the ROI of master data management break these value drivers down and show how to put numbers on each.
One more reason this matters now: AI. Every model, from demand forecasting to product content generation, amplifies the quality of its inputs. Feed it flawed records, and you get confident, automated, scaled-up mistakes. Companies that invest in data quality first get compounding returns on every analytics and AI initiative that follows. Companies that skip it pay the hidden data factory tax, plus interest.
Start small. Pick one critical data domain, run the 100-record check from the HBR study on your own data, and put a number on what you find. The number will be worse than you expect. That is the point. Once the cost is visible, fixing it stops being an IT chore and becomes a business decision with a clear payback.