- AI for Ecommerce and Amazon Sellers
- Posts
- Guide: Seven Techniques for Smarter A/B Testing with Claude and AI
Guide: Seven Techniques for Smarter A/B Testing with Claude and AI
A Full SOP Inside

From Our Sponsor:
Delete Your Negative Reviews to 2x Revenue
Every negative review on your listing is quietly killing your conversion rate. And when conversion drops, PPC gets more expensive, organic ranking slips, and traffic you're paying for stops converting.
Most sellers treat bad reviews like weather. Something you just live with.
But there's a compliant process for getting eligible negative reviews removed, and sellers doing it are seeing real lifts across the board.
Seth Stevens from Seller Growth Network is running a free session breaking down exactly how:
How to get reviews removed while staying within Amazon's TOS
The direct impact on conversion, ranking, and profitability
Why the current window for this matters
If you've got negative reviews dragging down your listings, this one's worth your time.
Guide: Seven Techniques for Smarter A/B Testing with Claude and AI

A/B testing, showing one group of visitors your existing page and another group a changed version, then measuring real behaviour, has been mechanically solved for years. The tools work. What hasn't been solved is the thinking around them: deciding what deserves a test, resisting the pull to see your own idea win, and connecting one result to the next so that a year of testing produces knowledge rather than a spreadsheet of disconnected wins and losses.
This is where AI earns its place, and it's worth being clear about where it doesn't. The tempting use of a tool like Claude is to generate variations faster. That's the cheap half of the work, and speeding it up mostly lets you run more disconnected tests — the opposite of progress. The valuable half is the connective thinking: classifying tests so patterns emerge, stripping human bias out of how tests are framed, and running the kind of joined-up analysis that makes test forty smarter than test four. That half has always been neglected because it's slow, effortful, and easy to skip. It's also the half an AI can genuinely help you sustain.
The techniques below are built on a simple classification system that sorts every ecommerce A/B test into one of eight purposes: Brand (trust and credibility), Discovery (helping people find the right product), Product Appeal (making a product more desirable through messaging or imagery), Product Detail (giving the specifics people need to choose), Price and Value (improving the price-to-value perception), Usability (removing friction), Quantity (increasing cart size), and Scarcity (creating urgency). Once every test carries a purpose tag, individual tests stop being isolated events and start forming a picture of what your particular customers actually care about. Each technique pairs a principle from that system with a practical way to run it using Claude, including Cowork, Anthropic's agentic assistant that can work directly with your files and data rather than only holding a conversation.
What You'll Need
Tools
Claude (for analysis, hypothesis framing, and connective work across your test programme)
Claude Cowork (for working directly with exported test data, reviews, and support tickets)
Data and Assets
Export of your past 12–24 A/B test results (format doesn't matter — spreadsheet, CSV, or even a summary document)
Product reviews, post-purchase survey responses, and/or support tickets
Access to your analytics platform for funnel metrics (transactions, revenue per visitor, add-to-cart rate, cart views, checkout step views)
Step 1: Audit Your Test Back-Catalogue
The most revealing test in your programme is the audit of tests you've already run. Take your last twelve to twenty-four tests and tag each one with its purpose using the eight-category system above. This almost always exposes two things: heavy clustering into one or two purposes, and purposes with no tests at all. The empty purposes are where your blind spots live. If you've never run a Price and Value test, you don't know it doesn't work for your customers. You only know you've never looked.
This is a natural first job for Cowork. Export your test history, however messy, and give Claude the eight purpose definitions. It will tag each test, count how often each purpose has been tested and won, and show you the shape of your programme at a glance. Ask it to run the tags separately for mobile and desktop, because the winning purposes often differ between the two and a combined view averages the difference away. Mobile matters here: most stores now take more traffic on mobile than desktop, yet convert at roughly half the rate, making it the largest under-examined area in most programmes.
The output is a diagnosis rather than an answer. If your history shows nine Usability tests and nothing on Brand, Price, or Quantity, Claude hasn't told you those purposes will win. It's told you where you've stopped looking — which is the more useful thing to know before you spend another design cycle.
Step 2: Replace the Hypothesis with a Question Set
Every conventional framework tells you to write a hypothesis — something like "adding a product video will lift conversion by eight per cent." On paper this is harmless. In practice it introduces a specific and expensive problem: a prediction comes from a person, people become attached to their own predictions, and once a test is privately "yours" you start rooting for it. That's how tests get stopped early on four days of favourable data, and how a losing result gets quietly filed away instead of examined.
The alternative is to frame each test as a set of neutral questions rather than a single prediction. For that same video example, the questions might be: do people care about watching a video at all, will they actually watch it and how far, does it change the add-to-cart rate, does it change how much of the rest of the page they read, and does it affect returning visitors differently from new ones. Questions have no author to embarrass, they can't be confirmed by a few good days, and they teach you something even when the test loses — because a loss now tells you people didn't care about the thing you added.
This is one of the cleanest uses of AI in the process, precisely because a model has no reputation staked on the outcome. Ask Claude to convert any proposed test into its underlying question set and it will produce a genuinely disinterested list, free of the quiet lobbying a human proposer brings. The tool doesn't just help you run the technique — it embodies the psychology the technique is trying to protect.
Step 3: Generate Test Ideas from Real Customer Language
The audit in Step 1 tells you which purposes you've neglected. The obvious next step is to generate ideas to fill them, and this is where most teams go wrong by reaching for generic best practices. A best practice is simply someone else's test result applied to your customers, and it frequently fails to transfer. Free-shipping messaging does nothing for a store whose customers already know shipping is free.
The better source of ideas is your own customers' words. Point Claude at your product reviews, post-purchase survey responses, and support tickets, and ask it to surface the recurring objections, hesitations, and points of confusion — then map each to the purpose it belongs to. A cluster of reviews mentioning uncertainty about fit points to Product Detail. Repeated remarks about not trusting the brand at first points to Brand. This grounds your ideas in evidence about why people hesitate rather than in a list of tactics that worked for a different shop. First understand why visitors aren't converting; only then decide what to show them.
Step 4: Pre-Screen Variations with Synthetic Customers
You'll usually have more ideas than traffic. Recent research suggests that AI models can approximate human purchase-intent survey responses with reasonable fidelity when prompted carefully, which opens a useful step: building synthetic customer personas from your real review data and running candidate variations past them before any live traffic is spent. Claude can play these personas, react to two or three versions of a page, and help you rank which are worth the cost of a real test.
The caution here is the whole point, so treat it as part of the technique rather than a footnote. A synthetic customer is a borrowed prior — exactly the thing this discipline warns against when it cautions that best practices are someone else's results. An AI persona is a well-informed guess about your customer, not your customer. Use it to narrow the field and kill the obvious non-starters, never to declare a winner. The live split test remains the only thing that decides. Used this way, pre-screening lets you spend your limited traffic on the variations most likely to teach you something, without pretending the traffic was unnecessary.
Step 5: Brief Multiple Concepts for Important Tests
A common and costly mistake is building a single design concept for an important test. Design details can decide an outcome, so one mediocre execution can bury a genuinely good idea — and you'll tend to write off the idea rather than the execution. For a test that matters, you want more than one honest attempt at the change.
Claude is well suited to producing distinct concepts rather than variations on a single theme. Ask it for three separate approaches to the same purpose, each with its own reasoning about why a different customer might respond, then choose the strongest to build. The value isn't that the AI writes faster — it's that it widens the range of ideas on the table before you commit design and development time, so the test measures the idea properly rather than accidentally measuring a weak first draft.
Step 6: Standardise Measurement and Read the Whole Funnel
Most testing advice obsesses over statistical significance — which only tells you whether a difference is real — and neglects what to measure, which tells you what actually happened. Fix this by using the same set of success metrics on every test so results are comparable across the whole programme: transactions, revenue per visitor, add-to-cart clicks, cart page views, and views of each step in the checkout. Then add a goal on the specific thing you changed. If you added an accordion of product detail, measure whether anyone opened it. A test where nobody opened it means something entirely different from one where a third of people opened it and still didn't buy — and only the second result tells you where to look next.
Claude's contribution here is at the analysis stage. Feed it the exported results — including every funnel metric rather than the headline conversion rate alone — and ask it to locate where in the funnel behaviour actually changed. A test that lifted add-to-cart clicks but not transactions worked and then lost the gain somewhere downstream, which points to a completely different follow-up than a test that moved nothing at all. Weight the bottom of the funnel most heavily when you plan, because a lift at the payment step flows straight to revenue, whereas a lift higher up is diluted by everyone who drops off before checkout.
Step 7: Build the Purpose Ledger and End Every Test with a Named Next Test
This is the step almost every programme skips, and it's the reason month six of testing so often looks exactly like month one. Someone raises an idea, you test it, you move on, and nothing accumulates. This pattern is common enough to have a name — tunnel-vision testing — where each element is tested in isolation and never connected into a larger understanding of what your customers want. The cure is a single living document, sometimes called a purpose ledger: a running record for each of the eight purposes showing tests run, won, lost, and inconclusive, plus what each result changed about your understanding of that purpose. The individual test results are raw material. The ledger is the actual product of a testing programme.
This is the most demanding use of Claude and the one that best justifies putting AI near your testing at all. After each test, have it produce a segmented read rather than a single overall number — at minimum splitting new visitors from returning ones and separating traffic sources, because a flat overall result often hides a real win in one segment and a real loss in another. Brand tests, for instance, frequently do little for returning visitors who already know you and a great deal for newcomers. Then have it update the ledger with what the result changes about that purpose, and crucially, end with a specific named follow-up test rather than a vague note to explore further. That single discipline — an actual next test at the end of every analysis — is what keeps a programme moving instead of restarting the idea hunt every fortnight.
For the complete SOP including the eight-purpose classification reference, purpose ledger template, and worked examples of each technique, click here — a free gift from me :)
Do You Love The AI For Ecommerce Sellers Newsletter?
You can help us!
Spread the word to your colleagues or friends who you think would benefit from our weekly insights 🙂 Simply forward this issue.
In addition, we are open to sponsorships. We have more than 66,000 subscribers with 75% of our readers based in the US. To get our rate card and more info, email us at [email protected]
The Quick Read:
OpenAI builds a dedicated SMB ads unit, per six job postings spanning growth, data science and sales ops. It recruits from Meta, Google and TikTok, and outsources SMB sales to vendors it polices. The Google-Meta playbook, at speed.
Time runs a hidden version of its site that feeds sponsored content to AI crawlers from Anthropic, OpenAI and Perplexity. Ally Bank goes first. The pitch: shape what assistants say about a brand, with readers none the wiser.
xAI ships Imagine Image 2.0 as Quality Mode on Grok, with sharp text, precise region editing, multi-ref inputs and smart resize. It ranks second on Arena for both image generation and editing.
Meta is allegedly building its own web search index so its AI avoids routing queries through Google. A developer reports heavy Meta scraping of text and images across sites, hinting at image, video or world-model ambitions.
TikTok Shop US GMV grows 103% to $11.8B in H1 2026, reclaiming its top market. Shop overtakes video and live as where sales close at 51% of GMV. Prices fall in 21 of 27 categories as competition tightens.
Perplexity blocks Time's markdown ads from its index, warning publishers who serve AI-crawler-only sponsored content risk a trust-score downgrade. It calls the practice deceptive cloaking, label or not.
The Tools List:
🎥 Arcads - Generate UGC-style video ads from a script using 1,000+ AI actors.
🌪️ Relay.app - Workflow automation with built-in human-in-the-loop approval steps.
🗣️ Cartesia - Ultra-low-latency text-to-speech and voice cloning for real-time apps.
📋 Beautiful.ai - AI presentation maker with auto-designed, on-brand slides.
👩🏽💼 Clay - Enrich and scrape web data to build hyper-personalized prospecting lists.
About The Writer:

Jo Lambadjieva is an entrepreneur and AI expert in the e-commerce industry. She is the founder and CEO of Amazing Wave, an agency specializing in AI-driven solutions for e-commerce businesses. With over 13 years of experience in digital marketing, agency work, and e-commerce, Joanna has established herself as a thought leader in integrating AI technologies for business growth.
For Team and Agency AI training book an intro call here.