Data Quality for AI: Stop Blaming the LLM
    AI ImplementationCTO

    Data Quality for AI: Stop Blaming the LLM

    Your AI/LLM project is failing because of poor data quality, not the model. A CTO's guide to fixing the foundation before building the skyscraper.

    Executive Summary

    Pain

    You've put a good deal of effort into a new AI analytics tool, but the answers it gives are a bit off and don't quite match what you know to be true.

    Risk

    Spending a lot of time and money trying to fix the model's 'hallucinations', while the business starts to lose faith in both the new tool and the data team.

    Fix

    In my experience, the fix isn't to keep tweaking the model. The issue usually lies in the data it's being fed. It's an engineering problem before it's a data science one.


    When the new AI tool meets your business

    It's a pattern I've seen quite a few times, especially in fast-growing tech companies. You've scaled up, you have a decent data stack, and the board is, quite reasonably, asking about AI. You hire good people, get the right tools, and build a pilot, maybe an internal 'chat with your data' tool. The first demos often look very impressive.

    Then it gets used by the rest of the business. The Head of Sales asks for 'Monthly Recurring Revenue', and the number doesn't quite line up with the board pack. The COO asks about 'customer churn rate', and you realise the definition it's using comes from an old, forgotten data source. The tool is fast, confident, and unfortunately, incorrect.

    This is often the point where technical ambition runs up against years of accumulated architectural debt. The natural reaction is to blame the model, start tweaking prompts, or assign more data scientists to the problem. In my experience, this approach doesn't usually lead to a good outcome.

    The problem is usually the data, not the model

    You've probably moved to the cloud and hired some very capable engineers. The difficulty is that, without a pause to tidy up, you can end up just migrating existing issues. An automated but messy process simply produces confusing data more quickly. Without some thought given to how it's managed, a modern stack can sometimes just amplify the noise.

    The problem isn't usually the AI itself, but the data it's trying to make sense of. It's the old 'garbage in, garbage out' problem, only now the output is delivered with remarkable confidence and clarity.

    I was working with a company recently that had to put a large AI project on hold. Their engineers found the model was attempting to learn from over 300 different, poorly documented tables. They had three separate columns for 'customer ID' and five competing definitions of an 'active user'. The project was on shaky ground before the AI-specific work had even begun. This is where a clear CTO Data Strategy that focuses on the foundations becomes so important.

    Data Quality for AI: Fix your data, not just blame the LLM. Improve AI performance with better data.

    How to prepare your data for AI

    To get AI to work reliably, it helps to shift perspective a bit. Instead of thinking about individual reports, it's more like managing a production line. The 'product' is trustworthy, consistent data, and the AI is just one of the customers using it.

    Getting this right involves a deliberate, and I'll admit, rather unglamorous shift in focus. But it is important.

  1. Curate your data sources: The first step isn't to add more data, but to reduce what's there. It's worth doing a proper audit of your data. Which tables are actually being used? Which reports do people trust? In my experience, a large portion of the data in a company's warehouse, perhaps 60-70%, is effectively noise. The goal is to identify and consolidate this down to a core set of well-documented, reliable assets. This is fundamental to building a Single Source of Truth.
  2. Define your business logic as code: One of the most common reasons for inconsistent data is that key business logic exists only in people's heads, or it's hidden away in thousands of lines of SQL. For an AI to use it reliably, this logic needs to be defined explicitly as code, usually in a central Semantic Layer. That way, when the tool asks 'What is our net margin?', the definition is unambiguous.
  3. Introduce some light data governance: This doesn't mean creating a slow, bureaucratic committee. It's more about having some sensible checks and balances in place. Light Data Governance can be as simple as having automated data quality checks, using CI/CD pipelines for your metric definitions, and making sure every core data set has a clear owner. It just helps prevent surprises for the leadership team and saves the engineering team from chasing down preventable errors.
  4. This work is about people, not just technology

    It's sensible to be prepared for some resistance. This kind of project is often less about the technology and more about people and process. You might be asking someone to give up a departmental dashboard they've used for years, even if it's based on incorrect data. You'll likely need to get the Head of Product and the Head of Finance to sit down and finally agree, in writing, on the official churn definition. It can be slow, patient work that requires a bit of negotiation. The trade-off is usually that things might move a little slower for one quarter, so they can move much faster for the next three years.

    What happens when you get the foundations right

    Once this foundation is in place, you can have another go at the AI project. This time, when it uses your data, it's drawing from a clean, curated, and well-understood source. The answers it gives are not just quick, they're also defensible. The business can start to trust the outputs, not because the AI is magic, but because the data it's built on is sound.

    As a CTO, your time shifts from debugging strange outputs to providing clarity. In my view, that's a much better use of everyone's time.

    Ready to Transform Your Data?

    Book your free clarity call today and discover how NorthStar Analytics can help you build a single source of truth.