A table with data sheets and sticky notes – representing the look into your own data before an optimisation project
Back to blog
Data QualityProcess OptimisationDigitalisation

Is Your Data Good Enough? What It Could Already Tell You

Sven HennessenProcesses

Before every optimisation, before every automation, there's a question most skip: how good is our data actually? Automation doesn't correct data errors — it speeds them up. And the data you already have reveals more than you think, if you know where to look.

Imagine you automate your order process tomorrow. Everything is faster, everything runs digitally. And three times as quickly as before, incorrect orders land with the supplier — because the address in your system has been outdated for two years. Automation doesn't correct data errors. It speeds them up.

What poor data quality actually costs

Gartner estimates the average annual cost of poor data quality at around $12.9 million (per organisation). For small and mid-sized businesses, that figure sounds abstract. It becomes more concrete when you break it down: every incorrect master data address generates a return. Every wrong stock level triggers an over-order or a delivery delay. Every duplicated customer record multiplies errors into downstream systems.

What makes it insidious: most of these costs don't appear in a single line of your profit and loss statement. They hide in the time experienced staff spend manually fixing exceptions. In decisions made on wrong numbers. In systems nobody trusts anymore, and around which Excel workarounds emerge that end up carrying the real knowledge of the business.

Garbage in, garbage out: why maturity comes before speed

"Garbage in, garbage out" is the oldest principle in data processing — and the most frequently ignored. It means: the quality of your output is never better than the quality of your input. Automate a process with bad data and you produce bad results faster. Train an AI model on inconsistent legacy data and it makes worse decisions than an experienced dispatcher — just a hundred times quicker.

That's why data maturity isn't a technical footnote; it's a project prerequisite. Broadly, five maturity levels can be distinguished: from raw, unsystematically captured data through consolidated and cleansed records to reliable, analysable foundations. Most mid-sized businesses don't start at level five. Many underestimate which level they're actually on.

That's not a criticism — it's the starting point. Knowing your data maturity before starting a project allows for realistic planning: what needs to be cleansed first, which automation is immediately possible, and where groundwork is still needed.

Improving data quality: a cycle, not a one-off effort

The good news: data quality isn't a large project you tackle once and then tick off. It's a cycle you work through step by step — first for the most important data domain, then the next. This iterative approach has proven itself in practice because it keeps risk small and delivers early results, rather than spending months "cleaning up" before anything becomes usable. The methodological foundation is the Control logic from Six Sigma / DMAIC: not clean once, but stay clean.

A four-step cycle – Profile, Cleanse, Establish rules, Monitor – which after monitoring restarts at profiling the next data domain

The four steps are deliberately straightforward:

  • Profile: First measure where the errors actually sit — duplicates, gaps, conflicting formats. No gut feeling, but an honest current state of the data domain in question.
  • Cleanse: Fix the errors found, merge duplicates, close gaps. One-off cleanup work — the part most people associate with "data quality."
  • Establish rules: The critical step most skip. Mandatory fields, validations, and a clear single source of truth ensure the same errors don't reappear tomorrow. Without this step, you'll be cleansing again in six months.
  • Monitor: A few simple metrics — duplicate rate, share of incomplete records — show continuously whether quality holds. If a value tips, you intervene early instead of being unpleasantly surprised at the next project.

Then the cycle starts again — with the next data domain. Data quality grows with the business rather than having to be "finished" upfront in one large, risky effort. For mid-sized businesses, this is the pragmatic path: small, verifiable steps instead of big bang.

Your data is already talking — are you listening?

The positive counterpart to the warning: the data you already have usually contains more answers than you realise. Not as an abstract big-data promise, but very concretely.

How often are orders manually corrected? What's the rework rate on incoming invoices? At which step do most queries arise? These aren't feelings — they're measurements that already exist in your systems. They show where the process is stuck, what data quality actually looks like, and where the biggest levers for improvement lie.

If 30 percent of tickets in a ticketing system carry the status "query to creator", that's not a coincidence. It's a data point saying: required fields are missing during capture. Or the right people are seeing the wrong information. Both are solvable — but only if you look.

The honest maturity check before the project

Before any optimisation project, five honest questions are worth asking:

  • Where does our master data come from — and who maintains it, by what rules?
  • What's the manual correction rate in the affected processes?
  • How many systems contain the same data in different versions?
  • Who is the reliable single source of truth for the most important decision foundations?
  • Which decisions do we make today on gut feeling because the data picture is unclear?

The more of these questions lead to "not entirely sure", the more valuable it is to clarify them before investing. Not to slow the project down — but to build it on a foundation that holds.

We look at this state together with our clients before developing any technical concepts. What regularly happens: the usable data is often better than feared. And the problematic areas are smaller and more precisely addressable than expected. That's the difference between a project that delivers and one that's flanked by Excel workarounds a year after go-live.

The next part of this series addresses a prerequisite that's even less often explicitly clarified than the data situation: who actually carries the responsibility for this process running well in the long term?

Sources

Need support?

Wondering what your existing data can already deliver — or whether it's ready for automation at all? These are exactly the questions we're happy to work through together, before a single line of code is written.

Get in touch