Experimentation · Buyer’s Guide
Google Optimize died in 2023 and was never replaced. The market that filled the gap is now consolidating hard — and most sites buying into it don’t have the traffic to use it.
Google shut down Optimize on 30 September 2023 and pointed everyone at third-party tools. It had been the free on-ramp for a generation of marketers, and its removal did something predictable: it pushed a large population of small and mid-sized sites toward paid platforms, most of which are priced for organisations with far more traffic.
The market has moved a great deal since. VWO and AB Tasty have merged into one company. OpenAI acquired Statsig. Webflow absorbed Intellimize. If your shortlist is more than two years old, it describes a market that no longer exists.
But before any of that matters, there’s an arithmetic question that determines whether you should buy anything at all, and it’s the one section of this article I’d insist on.
Do You Have Enough Traffic? Answer this before reading further
A/B testing is a statistical procedure, and statistical procedures have sample size requirements. Below them, you don’t get weak results — you get results indistinguishable from coin flips, presented with the same confident interface as real ones.
The rough shape: detecting a genuine relative improvement of around 10% on a baseline conversion rate of 3% requires something on the order of tens of thousands of visitors per variant. Smaller effects need dramatically more. A 5% relative lift — still commercially meaningful — can require several times that.
Which produces an uncomfortable conclusion for most sites. If you receive 5,000 visitors a month to the page you want to test, a properly powered test on a modest effect will take longer than a year, during which your site, traffic mix and market will all have changed enough to invalidate it.
The honest thresholds, roughly:
Under 10,000 monthly visitors to the tested page: don’t A/B test. Make changes based on judgment, qualitative research and obvious problems, and measure at the level of the whole business rather than the individual change.
10,000 to 50,000: test rarely, only large changes, and only where you expect big effects. Redesigns, not button colours.
50,000 to 500,000: a genuine testing programme becomes viable. This is where the mid-market tools earn their fees.
Above 500,000: testing should be continuous and the enterprise platforms start to justify themselves.
The Market, After Consolidation Who’s left and what they cost
Pricing in this category is unusually opaque — several vendors gate it entirely behind a sales conversation, which is itself information about who they want as customers.
Free and open-source. GrowthBook, self-hosted, is the most credible free option and assumes engineering involvement. PostHog offers a free tier including feature flags and experimentation. Both are developer-oriented; neither is a no-code tool for marketers. If you have a developer and no budget, these are the answer.
Entry tier, roughly $29–$70 a month. A cluster of lighter tools aimed at marketers on site builders, generally combining basic testing with heatmaps or analytics. Adequate for occasional tests on a single page, and honest about being that.
Mid-market, roughly $200–$700 a month. VWO and Convert are the established names. VWO offers a free Starter tier under 10,000 monthly tracked users, with Growth pricing quoted in the region of $200–$315 a month depending on source and billing terms, and higher tiers well above that. Convert sits toward the upper end of the band. Pricing in this tier is driven by monthly tracked users, so costs scale directly with traffic — which means a successful site’s testing bill rises exactly as fast as its traffic does.
Enterprise. Optimizely and Adobe Target, with published estimates for Optimizely running from tens of thousands into six figures annually. These are procurement exercises rather than purchases, and everything is negotiable.
Developer-led feature flagging. LaunchDarkly, Split, Statsig. Different product category solving an adjacent problem — server-side experimentation and controlled rollout rather than marketing page tests. If your experiments are about product behaviour rather than page layout, this is the right shelf.
One forward-looking note worth factoring into a decision now: the merged VWO/AB Tasty entity is widely expected to move upmarket, which typically means pricing pressure at the lower tiers over the following year or two. If you’re evaluating and the terms are good, locking in is a reasonable hedge.
What to Test, In Order Effect size descending
Given that every test costs weeks of traffic, what you choose to test matters more than the tool you test with.
The offer itself. Price, guarantee, what’s included, the risk the customer is asked to take. Larger effects than any design change, and almost nobody tests it because it feels like a business decision rather than a marketing one. It is a business decision. Test it anyway.
The headline and the promise. What the page claims it will do for the visitor. Consistently among the highest-leverage single elements.
Page structure and length. What order information appears in, and how much of it there is. Large effects, particularly on considered purchases.
Form length and friction. Removing fields is one of the few interventions that reliably improves conversion across contexts. Every field is a chance to leave.
Trust and proof placement. Where reviews, guarantees and security signals appear relative to the decision point.
Then, distantly, everything else. Button colour, microcopy, image choice, spacing. Real effects exist here and they are small enough that most sites cannot detect them, which means testing them consumes your entire traffic budget to learn nothing.
“The button-colour test is famous because it’s easy to run, not because it’s ever moved a business.”
Reading a Result Honestly Where most programmes go wrong
Running a test is easy. Interpreting one is where the discipline lives, and a handful of errors account for most of the bad decisions made from good data.
Statistical significance is not business significance. A result can be statistically solid and commercially irrelevant — a 0.4% improvement that’s real and reliably measured may not justify the engineering to ship it. Decide before the test what size of effect would actually change what you do.
A flat result is a result. Most tests don’t produce winners, and a well-run programme has a majority of inconclusive outcomes. That isn’t failure. Learning that a change you were confident about does nothing prevents you from rolling it out everywhere, which is worth real money.
Watch the metric downstream of the one you tested. A checkout change that raises conversion and reduces average order value may have made you poorer. A headline that raises click-through and lowers purchase intent has moved the wrong number. Always instrument at least one step beyond the thing you’re optimising.
Novelty effects fade. Regular visitors react to change as change. An improvement visible in week one that shrinks by week four was partly the novelty, and tests that run for a single week are especially prone to this.
Segment the result before shipping. A variant that wins overall can lose badly on mobile, or among returning customers, or in one market. Aggregate wins that hide segment losses are common and shipping them makes things worse for a group you didn’t examine.
Re-test the important ones eventually. Results decay. Audience, competition, seasonality and site context all change, and a conclusion from three years ago describes a site and a market that no longer exist.
Where the Ideas Should Come From Not from a list of best practices
The testing tool is the cheap part. Knowing what to test is the expensive part, and it doesn’t come from articles listing tactics.
Session recordings and heatmaps. Watch twenty recordings of people failing to convert. It’s tedious and it’s the fastest route to a genuine hypothesis. You will see people miss a button you thought was obvious, or scroll past the thing you built the page around.
Ask people who didn’t buy. An exit survey with one question, or an email to people who abandoned. The answers are frequently mundane — unclear delivery cost, uncertainty about returns, an unanswered question — and mundane is exactly what you can fix.
Read support tickets and pre-sale questions. Every question a customer has to ask is something the page failed to answer. This is a free, continuously-updated list of page defects and almost no marketing team reads it.
Look at your own funnel data. Where the largest proportional drop-off occurs is where a test has the most room to move. Testing the step that already converts at 85% is bounded by arithmetic.
A good hypothesis has a shape: because [observed evidence], we believe [specific change] will [predicted effect] for [defined audience], measured by [metric]. If you can’t write that sentence, you have an idea rather than a test, and running it will produce a number you can’t interpret either way.
Beyond A/B: The Rest of the CRO Toolkit Often better value
Split testing dominates the conversation and is one instrument among several. The others have no sample size requirement, which for most sites makes them strictly more useful.
Session recording and heatmaps. Watching real people fail is the highest-yield diagnostic available, and twenty sessions is enough to find something. Several tools bundle this at entry-tier pricing. The discipline is to watch failures rather than successes — successful sessions teach you nothing you didn’t already assume.
Form analytics. Field-level drop-off telling you exactly which question people abandon on. Frequently a single field — a phone number, a company size, an unexplained requirement — is responsible for a large share of abandonment, and removing it is a decision that needs no test.
On-site surveys. One question at the right moment. “What almost stopped you buying today?” on the confirmation page, or “What were you looking for?” on an exit. Cheap, fast, and it surfaces objections you’d never have hypothesised.
Moderated user testing. Five people attempting a task while narrating. It’s the most information-dense hour available in this whole discipline and it costs very little.
Funnel analysis in tools you already own. GA4’s funnel explorations are free and show you where the largest proportional drop-off is, which is where any intervention has the most room to work.
A budget observation worth stating plainly: for a site under the traffic threshold, spending $50 a month on session recording and an hour a week watching it will produce more improvement than $300 a month on a testing platform you cannot power. The testing platform feels more rigorous. The recordings are more useful.
The Options, Compared
| Option | Rough cost | Needs a developer | Best for |
|---|---|---|---|
| GrowthBook (self-hosted) | Free + hosting | Yes | Teams with engineering and no budget |
| PostHog | Free tier, then usage | Yes | Product teams wanting flags and analytics together |
| Entry-tier marketer tools | ~$29–$70/mo | No | Occasional tests on site builders |
| VWO | Free under 10k MTU; ~$200–$315/mo Growth | No | The default mid-market choice |
| Convert | ~$600–$700/mo | Partially | Privacy-sensitive testing at mid-market scale |
| Optimizely / Adobe Target | Five to six figures a year | Yes | Enterprise programmes with dedicated teams |
| Nothing | Free | No | Under 10k monthly visitors |
Running a Programme Rather Than Occasional Tests If you’re above the threshold
Sites with enough traffic face a different problem: not whether to test, but how to make testing a repeatable capability rather than a series of one-offs.
Keep a prioritised backlog. Every hypothesis, scored on expected effect, confidence, and effort to build. The scoring is imprecise and its purpose is forcing the comparison — without it, you test whatever was suggested most recently or most loudly.
Run one test per page at a time. Concurrent tests on overlapping traffic interfere with each other in ways that are difficult to untangle afterwards. Sequential is slower and interpretable.
Document every result, including the failures. A written record of what didn’t work is the most valuable asset a testing programme accumulates, and the one most reliably lost to staff turnover. Teams without it re-run the same losing test every eighteen months.
Set the cadence by your traffic, not by ambition. If a properly powered test takes three weeks, you get roughly seventeen a year. Planning fifty is planning to stop tests early, which is planning to make decisions from noise.
Ship the winners properly. A surprising number of validated improvements never make it into the production codebase, living on indefinitely as a test variant served by a JavaScript overlay — which slows the page and creates a dependency on the testing tool that nobody intended.
What to Do Instead, If You’re Below the Threshold
Most readers of this article will be in the bottom row of that table, so this section matters more than the rest.
Fix the obvious things without testing them. A page that loads in six seconds, a form with fourteen fields, no visible pricing, no delivery information — these don’t need a test. They need fixing. Testing whether a broken thing should be fixed is a way of postponing the fix.
Use qualitative research, which has no sample size requirement. Five user sessions will tell you more about why your page fails than a badly-powered test will, and you can run them this week.
Measure at business level over longer periods. Change the page, then compare the following quarter against the previous one. Confounded and imprecise, and vastly better than a test you don’t have the traffic to run.
Test on aggregate traffic rather than per-page. Site-wide changes — navigation, checkout flow, trust signals in a global template — accumulate sample across every page, which sometimes brings a test into range that a single landing page never would.
Market consolidation events and pricing ranges verified against published 2026 reporting and vendor pricing pages where available; several vendors do not publish pricing. Sample size guidance is directional — use a proper calculator with your own baseline conversion rate and target effect. This article contains no affiliate links.