A/B Testing: Why Your Results Are a Liability
    Product LeadGaming & Mobile Apps

    A/B Testing: Why Your Results Are a Liability

    For Product Leads: Stop making decisions on bad data. If your A/B test results are contaminated, it's not a tool problem—it's an architectural failure. Here's the fix.

    Executive Summary

    Pain

    You're running lots of A/B tests, but the results often seem to contradict each other, or don't line up with your main business metrics. This can undermine trust and means engineering time gets wasted.

    Risk

    Shipping a feature based on a flawed test can be worse than shipping nothing. You might be spending money on development that ends up harming user retention or LTV, making your product roadmap a bit of a gamble.

    Fix

    The answer isn't another experimentation tool. It's a proper fix to your event tracking: a clear, governed framework to make sure the data going into your tests is clean, consistent, and reliable.


    Why a 'winning' test result might not be a win

    It's a familiar scene. The Product Manager presents a 'winning' A/B test to the leadership team. The chart shows a 15% uplift in user engagement for the new feature, with a p-value of 0.01. It looks like a clear victory. The Head of Engineering is praised for a quick delivery.

    Then the CFO clears their throat. "If engagement is up 15%, why is our daily active users metric flat? And why does the finance dashboard show a dip in in-app purchases for that cohort?"

    Silence. The confidence in the room drains away. The 'statistically significant' result doesn't mean much because it contradicts what the business sees as its source of truth. The data is unreliable, the decision is questionable, and the credibility of your Product Analytics team has taken a knock. It's more than a wasted sprint; it's a loss of trust.

    The problem is usually in the data foundation

    This sort of thing isn't usually a one-off analyst error. It's often the result of a system where the pressure to move quickly has come at the cost of data integrity. In my experience, this is a common problem for scale-ups, particularly in gaming and mobile apps. It's a common trap. You hire good engineers and invest in the best tools: Amplitude, Mixpanel, Snowflake: but the underlying tracking issues just get carried over. Automating a messy process just means you produce unreliable data more quickly.

    The root cause is technical, not a fault of the people involved. It's what you might call 'instrumentation debt'.

  1. Inconsistent events: Your event tracking can become a bit of a free-for-all. Events are named inconsistently (`user_signup` vs. `UserSignedUp`), triggered from different places (client-side vs. server-side), and it's not always clear which ones are still in use. When engineers are under pressure to ship features, tracking can become an afterthought.
  2. No single source of truth: The way a 'retained user' is defined in your A/B testing tool might differ from the definition used in the main company dashboards. It's no surprise the numbers don't match. Without a central Semantic Layer, each dashboard can end up telling a slightly different story.
  3. Unreliable inputs: You're running tests on a shaky foundation. If the underlying user behaviour analytics data isn't reliable, the experiment is flawed from the start. You can't trust the results if you can't trust the data going in.
  4. The problem isn't the A/B test itself, but the data architecture that underpins it.

    A/B testing pitfalls: Why your results might be misleading. Avoid common mistakes!

    A structured approach to fixing event tracking

    The fix isn't to buy another tool or hire more analysts to stitch data together in spreadsheets. It's about fixing the underlying system, not just patching up the reports. It helps to think of your data tracking less like a one-off bit of code, and more like a production line for making decisions.

  5. Audit and consolidate. First, you need to get things under control. This usually starts with a full audit of your event schema. We find the events that are still firing away but haven't been used in months, which are often costing money. This leads to the necessary, sometimes difficult, conversations to agree on a single, unified tracking plan that everyone in Product and Engineering can work from.
  6. Govern and document. This new plan is then put into practice with some light governance. This doesn't mean writing a 100-page document that gathers dust. Instead, the rules are built directly into the workflow, using CI/CD checks and tools like Avo. This makes sure that no new event can be added unless it meets the agreed standard, which is how you maintain Data Integrity as you grow.
  7. Train and support the teams. Finally, we help the teams get up to speed. This isn't generic classroom training. A good way to do it is by working with them to fix their own dashboards. We often help establish 'champions' within the product teams, who become the first point of contact for data quality. The aim, for us, is to get to a point where we're no longer needed.
  8. Accepting the initial slow-down

    Let's be honest. Putting a proper A/B Testing Analytics framework in place will probably feel like you're slowing down for a quarter. Your product managers might resist what feels like bureaucracy in a formal tracking plan. Your engineers might be tempted to take shortcuts. This is the difficult part that's easy to put off. It's as much a challenge of persuasion as it is a technical one.

    But in my experience, it's the only way to move from a culture of guesswork to one where decisions are genuinely based on reliable data. You might have to move a bit slower for a few weeks to be able to move much faster for the next few years. The alternative is to carry on building your product roadmap on a foundation of unreliable data.

    Ready to Transform Your Data?

    Book your free clarity call today and discover how NorthStar Analytics can help you build a single source of truth.