Skip to content

Services

CRO & A/B Testing

Test design, sample size, and statistical significance done properly — including honesty about how often a test finds nothing at all.

of tests
74.2%
Clients served
60+
Ad spend managed
$20M+
Years operating
3

found no statistically significant difference at all. Only 17.4% found a significant winner. That's the honest base rate for A/B testing, and any agency promising a win on every test is setting an expectation the data doesn't support.

2,408-test study, Visionary Marketing, 2026

Across a 2026 study of 2,408 A/B tests, 8.4% reached statistical significance with the variant performing *worse* than the control (Visionary Marketing, 2026 dataset). A test that reaches significance is not the same as a test that won, and the difference is money.

Most CRO programs fail at the setup stage, long before a test ever launches.

ThinkMedia runs CRO and A/B testing for companies spending $5,000 or more a month on marketing. We design tests properly, report the ones that fail exactly as clearly as the ones that win, and won't call a result before the sample size supports it.

Four mistakes show up constantly:

  1. 01

    Peeking. A test shows a promising result on day three, someone stops it and calls a winner, and the result is statistically meaningless — it was never run to the sample size the test needed to reach a trustworthy conclusion.

  2. 02

    No pre-defined sample size or minimum detectable effect. Without deciding in advance how big an effect matters and how many visitors are needed to detect it, any “significant” result midway through a test is just noise that happened to look like a pattern.

  3. 03

    Testing too many things in one variant. Headline, image, and button color all change at once, so a win or a loss can't be traced to a specific cause, and the next test starts from zero learning instead of building on the last one.

  4. 04

    Declaring victory after a partial business cycle. A test run for three days misses weekday-versus-weekend behavior entirely, and a real result needs at least one full business cycle, typically seven to fourteen days, to hold up.

How this runs

What you get.

Reporting
a written result for every test, including the ones that didn't win, with the effect size and confidence level stated plainly rather than summarized as a vague “improvement.”
Access
your testing tool and results data live under your own account from day one. On exit, you keep every test result and the full historical log.
Communication
a named strategist who designs and reads your tests directly, reachable with a one-business-day response commitment.
Terms
thirty days' notice, no annual minimum, flat fee — never billed per “win,” since that would create an incentive to call tests early.

What's included

The actual thing we do, and how often.

  • monthly planning cycle

    Test prioritization and hypothesis design

    each test starts with a written hypothesis and a stated reason to believe it will move the metric, not a guess dressed up as a test.

  • per test

    Sample size and Minimum Detectable Effect calculation

    calculated before launch, so everyone agrees in advance how long the test needs to run and what size of effect would actually be worth shipping.

  • per test

    Test implementation and QA

    variants built and checked in a staging environment before launch, including a validation pass that tracking fires correctly for both control and variant.

  • ongoing per test

    Statistical monitoring without peeking

    results tracked against the pre-defined sample size, with no early calls and no stopping the moment a result looks promising.

  • at test conclusion

    Test result reporting

    a written report stating whether the result was significant, what the effect size was, and what we're testing next because of it — including tests that found nothing, reported with the same clarity as the ones that won.

How we work

Four phases, always in this order.

  1. Funnel Teardown

    Days 1–10

    We map your funnel and identify where the biggest, most testable drop-offs actually are, using existing analytics data rather than guessing which page deserves attention first.

  2. Test Blueprint

    Days 7–17

    A prioritized test roadmap, each test with a written hypothesis, target metric, and calculated sample size before anything gets built. This is where we agree on what a real win would look like, before we're looking at results that could bias that judgment.

  3. Build

    Ongoing, per test

    Each test is built, QA'd, and launched on its own schedule. What we deliberately do not do: call a test early because the dashboard looks good on day four. A test ends when it hits its pre-calculated sample size, not when the number we want to see shows up.

  4. Compounding

    Ongoing

    A continuous testing calendar, each test informed by what the last one taught, whether it won, lost, or found nothing — since even a null result usually rules something out and narrows what's worth testing next.

Honest scoping

Who this is for, and who it isn't.

You're a good fit if you have enough traffic to reach statistical significance within a reasonable timeframe — generally a few thousand visitors per variant per month at typical conversion rates — and an existing funnel worth optimizing rather than building from scratch.

You're not a good fit yet, and we'll say so before taking the engagement, if your traffic volume is too low to reach significance within a normal testing cycle, since a test that would take a year to reach sample size isn't a good use of your budget yet. If your conversion funnel doesn't exist yet, or converts so rarely that only a handful of conversions happen a month, build the funnel first — testing optimizes something that's already working, it doesn't create it.

**A/B testing vs. redesigning based on opinion**

Asked and answered

Common Questions

  • ThinkMedia charges a flat monthly fee for ongoing testing programs, not a per-test or per-win fee, starting for companies spending $5,000 or more a month on marketing overall. Pricing is scoped after the funnel teardown.

  • It depends on your traffic and baseline conversion rate, calculated per test before launch — typically two to six weeks to reach a pre-defined sample size, run across at least one full business cycle to account for weekday and weekend behavior.

  • Not every test wins, and any agency promising that isn't being honest with you. Across a large 2026 study, only about 17% of tests found a significant winning variant, while roughly 74% found no significant difference at all — that's the normal, expected outcome distribution for rigorous testing.

  • You own your test results and data at all times. Testing tools are set up under your own account from day one, and the full historical test log stays with you if you leave.

  • There's no fixed number, since it depends on your baseline conversion rate and the effect size worth detecting, but generally a few thousand visitors per variant per month is the practical floor for reaching significance within a normal testing cycle.

  • It gets reported exactly as clearly as a win would be, with the effect size and confidence level stated plainly, and it usually narrows what's worth testing next — a null result rules something out, which is real information, not a wasted test.

  • A named strategist assigned at kickoff, the same person who calculates sample sizes and reads results, reachable directly rather than through a rotating analytics team.

  • ThinkMedia has no minimum contract length. Thirty days' notice ends the engagement in either direction, with no annual minimum.

Proof

Results.

Every case study

Get in touch

Tell us what you’re running.

  1. You send the details

    Channels, monthly spend, and the part that is not working.

  2. We look at the accounts

    Sixty to ninety minutes inside them, before we say anything.

  3. You get the findings

    A 45-minute call covering everything — including what you can fix yourself.

marks a required field.

Including https://

As much or as little as you like.

A person replies from a real address. We never sell or share your details.