The most political meeting in any company is the one where investment is allocated. Ideas arrive attached to people, and people arrive attached to seniority. In the absence of evidence, the outcome is decided by whoever argues best, and argument quality correlates only loosely with commercial merit. Experimentation is the antidote, and it is widely misunderstood. In most organisations it is treated as a digital marketing technique: a way to choose between two button colours, owned by a growth team, reported in a channel dashboard. That framing is why it never reaches the boardroom. I think of experimentation differently, as a management system for reducing uncertainty before capital is committed. Its purpose is not to optimise a landing page. It is to give a leadership team a fair, credible and repeatable method for determining which of its beliefs are true, and to do so before those beliefs are funded at scale. Organisations rarely struggle to produce opinions. They struggle to produce evidence that anyone accepts.
Why intuition remains dominant
Intuition is not the enemy. Experienced judgement is a legitimate and often superior input where data is thin. The problem is that intuition tends to win arguments it should have lost, for reasons that are structural rather than personal.
- Seniority outweighs evidence, because disagreeing with a senior view carries a career cost that being wrong quietly does not.
- Attachment to ideas. People defend proposals they authored more vigorously than proposals they merely support.
- Confirmation bias in analysis design, where the question is framed in a way that makes the desired answer likely.
- Post-launch narrative, where a result is explained after the fact in whatever terms make it look intentional.
- Market movement mistaken for impact. A rising category makes almost every initiative look effective.
- Failed concepts reframed rather than learned from, so the same idea returns in eighteen months with new packaging.
None of this is corrected by exhortation. It is corrected by making the evidence standard explicit and agreeing it before anyone knows the answer.
What a good experiment accomplishes
A good experiment is not primarily a statistical artefact. It is a decision-making device, and it should do six things.
- 01State a hypothesis that could be false, in commercial rather than statistical language.
- 02Predict the behaviour that should change, for whom, and through what mechanism.
- 03Establish a credible comparison, so that the counterfactual is genuinely plausible.
- 04Measure incremental effect rather than response, netting off what would have happened anyway.
- 05Examine heterogeneous effects, because an average effect of zero often hides a strong effect in one segment and a negative one in another.
- 06Produce a decision. Scale, refine or stop, agreed against thresholds set in advance.
The sixth point is where most experimentation programmes fail. A test that concludes with 'interesting, let us investigate further' has produced analysis, not evidence.
The academic evidence on experimentation's value is encouraging and should be read carefully. Research published in Management Science by Koning, Hasan and Chatterji, examining thousands of startups, found that firms adopting A/B testing tools saw meaningful performance improvements in the following year, with the strongest effects among firms running more tests. The authors note that the benefits accrued disproportionately to firms already positioned to act on results, and that the setting was technology startups. It would be a misreading to assume the same magnitude transfers automatically to a large bank or retailer with longer feedback cycles and heavier operational constraints.
Higher growth after adoption
Experimentation is bigger than digital A/B testing
The most valuable experiments I have been involved in were not digital. They were commercial and operational, and they were designed by people who understood the business rather than by specialists in test infrastructure.
Customer holdouts
The single most useful tool in a customer-marketing organisation. A permanent, randomly selected group excluded from a programme, giving a continuous read on incrementality.
Geographic tests
Where individual randomisation is impossible, matched markets or regions provide a workable comparison, provided the matching is done on pre-period behaviour rather than on convenience.
Store or branch level
Physical estates make excellent experimental units for pricing, layout, staffing and service interventions.
Pricing experiments
Willingness to pay is an empirical question that most organisations answer with a committee. Careful, ethically bounded price testing usually produces the largest single margin insight available.
Phased rollouts
Sequencing a launch across regions or cohorts converts an operational necessity into a measurement opportunity, at almost no extra cost.
Operational interventions
Servicing scripts, delivery windows, onboarding steps and collections treatments are all testable, and often have larger effects than the marketing layered on top of them.
The LEARN framework
Framework
The LEARN framework
Five stages that convert an opinion into a decision. The order matters: three of the five happen before anything launches.
- LLocateIdentify the specific uncertainty that is holding up or distorting the decision. If resolving the uncertainty would not change the choice, do not test it.
- EEstablishDefine the hypothesis and the predicted mechanism. What should change, for whom, and why should it change?
- AAssignSet the success metric, the minimum detectable effect and the decision thresholds before launch. Pre-registration is the discipline that makes the result believable.
- RRunExecute a credible, adequately powered test for a period long enough to capture the behaviour, including any novelty effect and its decay.
- NNormaliseRecord the learning in a shared repository and act on the threshold agreed in stage three: scale, refine or stop.
How leaders actually create an experimentation culture
Culture here is a function of what leadership visibly rewards, and there are eight practices that do most of the work.
- Reward learning explicitly, including the learning that came from a null result.
- Publish negative results with the same prominence as positive ones. This single practice changes behaviour faster than any training programme.
- Separate the quality of an idea from the status of the person who proposed it, structurally, by pre-registering measures.
- Build a searchable knowledge repository, so nobody pays twice to learn the same thing.
- Track test velocity as an operating metric, because learning rate is a genuine competitive variable.
- Quantify value captured from experimentation annually, in currency, and report it to the board.
- Prevent repeated testing of the same question by requiring a repository search before a test is approved.
- Make stopping safe. If cancelling an initiative damages a career, initiatives will not be cancelled.
The cheapest thing an organisation can buy is the knowledge that an expensive idea does not work.
From isolated tests to an operating system
Experimentation becomes strategically material when it is connected to planning. That requires four things that are organisational rather than technical.
Governance. A single forum that reviews test design before launch and results after, with authority to release or withhold funding based on the outcome.
Prioritisation. Tests are a scarce resource, constrained by traffic, customer contact tolerance and operational capacity. They should be prioritised by value at stake and by how much the answer would change the decision, not by which team asked first.
Capability. A small central group that designs and validates tests, embedded with commercial teams rather than isolated from them. Central expertise, distributed application.
Integration with annual planning. The most powerful change is to require that any initiative above a defined investment threshold arrives at the planning cycle with either experimental evidence or an explicit statement of why experimentation was not possible.
Risks, limits and where experimentation is the wrong tool
Underpowered tests are worse than no tests, because they produce confident conclusions from noise. A programme that runs many small tests and declares winners from them will systematically scale false positives, and it will do so while believing it is being rigorous.
Short-term optimisation is a genuine hazard. Experiments measure what can be observed within the test window, which biases organisations towards interventions with fast, visible effects and away from investments in brand, trust and product quality that compound over years.
There are ethical boundaries that should not be crossed for the sake of a clean measurement. Testing that exploits vulnerability, withholds a material benefit from customers who need it, or manipulates rather than informs, is not acceptable regardless of the statistical elegance. Financial services, healthcare and any context involving financial hardship require particular care.
Some decisions cannot be experimented on at all: acquisitions, market entries, platform replacements, brand repositioning. Here the discipline transfers rather than disappears, through explicit assumptions, sensitivity analysis, staged commitment and pre-agreed evidence thresholds for continuing.
Finally, an honest counterargument. Excessive experimentation can slow an organisation and create a culture in which nobody will commit without proof. Judgement remains essential, particularly where the window of opportunity is short. The aim is not to test everything, it is to know which of your beliefs are load-bearing and to test those.
Questions for leadership teams
- 01Which major investments are currently being scaled without a credible comparison?
- 02What is our annual spend on decisions that were never tested and could have been?
- 03Can a team stop an initiative without reputational damage to the people who proposed it?
- 04Do we know what this organisation has already tested, and where is that recorded?
- 05Are our experiments designed to produce decisions, or to produce more analysis?
Experimentation does not remove judgement from leadership. It gives leaders better evidence on which to apply that judgement, and it changes the character of the argument in the room from who is most persuasive to what is most likely to be true. That shift, more than any individual test result, is what makes the investment worthwhile.
Sources and further reading
- Experimentation and Start-Up Performance: Evidence from A/B TestingKoning, Hasan & Chatterji, Management Science, 2022 — Peer-reviewed study of technology startups adopting A/B testing tools. Effects concentrated among firms testing at scale.
- The Surprising Power of Online ExperimentsHarvard Business Review, 2017 — Practitioner account from Microsoft's experimentation platform on hit rates and the prevalence of null results.
- Building a culture of experimentationHarvard Business Review, 2020 — Analysis of the organisational conditions required for experimentation to influence decisions.
The views expressed in this article are personal and do not necessarily represent the views of Claudio's current or former employers. Company and client examples are based solely on publicly available information.