This week I’m trying something different: one longer look at a paper instead of the usual roundup. Let me know what you think. As always, you’ll also find 1k+ open roles, upcoming events, and a cartoon that made me smile.
Brought to you by
Convert—A/B tests & personalization for growth teams
Experimentation Jobs—Where experimenters find their next role
🔎 Interesting things you might have missed
[Paper] Statistical foundations of LLM-based A/B testing
Can a LLM replace humans in an A/B test? Well, not really. And now there’s a rigorous statistical argument for why.
A new paper from Spotify’s experimentation team frames the question properly: an LLM’s response is a surrogate for a human response, not a replacement for it. Surrogate endpoints have been studied in clinical trials for decades, and the same logic applies here. A surrogate is only useful if it’s calibrated against the real outcome you actually care about, and if that calibration holds up when you apply it to something new.
The empirical test on the Upworthy Research Archive makes the point clear. Looking at 417 tests where headlines were phrased as a question or not, raw LLM predictions recovered only 39% of the real human treatment effect. Calibrating those predictions against historical human data closed most of the gap, but that calibration step is doing all the work, and it only works if the new test resembles the ones you calibrated on.
And that is the important part. The paper shows you can build diagnostics to check whether your calibration is holding, but those diagnostics can only be run on historical treatments. They can tell you the assumption failed. They can never tell you it holds for a truly new idea, because you have no human data yet to check it against. So the paper’s conclusion is that LLM-based A/B testing is correct by assumption, while testing on humans is correct by design.
What does that mean in practice? That maps to how I’ve already been using LLMs in this space: useful for checking, prioritizing, and pressure-testing variants when you’re working within a pattern you’ve already validated with real people. Not a substitute for running the experiment when you’re testing something truly new. The paper even gives you a way to think about it economically: run a small human pilot when the treatment is novel or the stakes are high, and lean on the LLM more when you’re iterating on something familiar and low risk.
____
Variance reduction below the randomization grain
Marketplace experiments must randomize at region level to avoid spillover, often killing statistical power. Instacart’s team has a fix: train prediction models at the order level, aggregate up, and use as CUPED covariates. Result: 18–40% variance reduction and experiment runtimes cut by a third. READ
____
Last week’s favourite:
[Paper] Trustworthy A/B patterns
Ronny Kohavi et al. performed eight large-scale replications of popular CRO patterns like rounded buttons, page speed, coupon fields and sticky CTAs. Results show that previously reported effects are wildly exaggerated. Only two replications showed significant effects in the expected direction. READ
🚀 Job opportunities
Find 1k+ open roles on ExperimentationJobs.com. This week’s featured roles:
Lead UX Consultant / Team Lead at Creative CX (London, United Kingdom)
Manager, Digital Optimization at Allstate (Canada)
Conversion Rate Optimization (CRO) Specialist at Herold (Wien, Austria)
CRO Manager at Sunny Cars (Haarlem, Netherlands)
Lead Data Scientist, Experimentation Platform at Strava (San Francisco, USA)
📅 Upcoming events
A running list of upcoming events. Subscribe here. (👋= join me, 🎁= discount)
👋16 Jul: Experimentation Culture Awards (online)
29 Jul: Experimentation London #12 (London, UK)
31 Jul: TLC: 5 Stories to Rewire Your Statistical Intuition with Ishan Goel (online)
🆕14 Aug: TLC: Bridging The Great Disconnect with Ole Gregersen (online)
9-10 Sep: Women in Experimentation Summit (online)
👋24 Sep: Experimentation Elite Awards (London, UK)
👋5 Nov: Experimentation Heroes (Amsterdam, Netherlands)
😃 Something that made me smile

When you’re ready, here’s how I can help
If this newsletter sparked something and you want to talk it through, a few ways I help:
Scale experimentation
Strategy, setup, metrics, and ways of working for teams that want more impact.Careers and hiring
Support with your next role, or help finding and hiring strong experimenters.Quick sparring
A fresh outside perspective on your ideas, roadmap, or experimentation setup.
Interested?
Just hit reply and tell me what you’re thinking about.
Have a great week and keep experimenting.
Thanks, Kevin


