Listed players
SST2.44▼ -7.58%TIG40.00▲ +3.90%TEAD0.56▲ +3.77%PERI8.50▼ -2.97%TBLA3.23▼ -2.71%INUV0.57▼ -1.74%AV10.06▼ -1.59%GOOGL343.50▲ +1.56%SNAP5.58▼ -1.24%PINS19.26▼ -1.03%MSFT517.53▲ +0.92%PPLI41.28▲ +0.81%IOS32.24▲ +0.44%META728.08▲ +0.30%GDDY97.21▲ +0.24%DV13.49▲ 0.00%MCHX1.29▲ 0.00%
Ticker byClearTrust

Lesson 5 of 5 · 9 min read · advanced

This lesson counts towards the ClearTrust Media Buying & Compliance certificate. Enrol with your email to record your progress and scores.Get certified, free

Testing landers and keywords: finding what works without fooling yourself

How to run fair A/B tests on articles, layouts and search terms, which metric to judge them on, and which tests policy does not allow.

Every arbitrage campaign is a stack of guesses: this article, this headline, these search terms, this page layout. Testing replaces guesses with evidence. Done badly, it replaces guesses with confident mistakes. This lesson covers how to test so that the result means something, and where the line is between optimising a page and manipulating a visitor.

A baker who wants to know whether a new recipe sells better does not bake it on Saturday and compare it with last Tuesday's sales. She puts both loaves on the counter on the same days and counts. An A/B test is the same idea: two versions, the same kind of customers, the same period, one difference.

What can be tested

LayerExamplesMeasured first by
AdImage, headline, ad angleCTR and cost per visitor (creative testing)
ArticleTopic, headline, length, how well it matches the adLander CTR
Search termsWhich related search terms are suggested and their wordingLander CTR and RPC per keyword
LayoutWhere the related search unit sits, within the styles your account is allowedLander CTR
FeedThe same traffic on two providers or two feed typesRPV after finalisation

Judge on revenue per visit, not on any single rate

A page change can raise one number and lower the total. The honest scoreboard is RPV: revenue divided by all visitors sent to that version. It captures every step of the funnel at once.

Illustrative numbers. B wins on lander CTR and loses on revenue, because its terms attract lower-value searches.
VisitorsLander CTRSearchesAd CTRPaid clicksRPCRevenueRPV
Version A2,00040%80030%240$0.60$144$0.072
Version B2,00046%92030%276$0.50$138$0.069

Version B persuaded more people to click a search term, which looks like a win. But the terms it showed drew lower-paying advertisers, and revenue per visitor fell from 7.2 cents to 6.9 cents. Had you judged on lander CTR alone you would have rolled out the worse page.

Running a fair test

  1. Change one thingIf the new version has a different article and different terms, you will not know which caused the result.
  2. Split the same traffic at the same timeLet the tracker send visitors from the same ads randomly to each version. Do not compare this week with last week: demand shifts by day, and so does the ad auction.
  3. Decide the finish line firstFix in advance how many visitors or paid clicks each version must receive. Stopping the moment one version edges ahead is the commonest way to find winners that are not real.
  4. Wait for the money to settleJudge on revenue that has had time to firm up, and for big decisions on finalised revenue. A version that earns more estimated revenue but draws more deductions is not better.
  5. Check it holdsRun the winner against the old version once more, or on a second country. Real improvements repeat.

Tests that policy does not allow

Testing happens inside the rules of the feed and the traffic source. Some things that would lift a number are simply not available to test.

Legitimate to test

  • Different articles on the same topic
  • Clearer headlines that match the ad
  • Which relevant search terms to suggest
  • Unit position and styling within what your account may use
  • Page speed and mobile layout

Not a test, a violation

  • Arrows, wording or design that push people to click (incentivised click)
  • Making the search unit look like navigation, a download or a video player (implied functionality)
  • Terms chosen to trigger particular ads regardless of the content
  • Thin pages that exist only to hold the unit (Made-for-arbitrage)
  • Showing reviewers a different page from users (cloaking)

Two practical limits follow from Google's rules. First, Google's help pages say a related search implementation must be shown to your account manager as a mock-up before it goes live, so layout variants are not something to improvise. Second, under the Restricted Access Features introduced in August 2025, only qualified accounts may show more than five suggested terms, place more than one unit on a page, supply their own terms or use certain style options. Many operators reach Google through a feed provider, whose templates already reflect what is allowed; in that case your testing room is the article, the ad and the choice of relevant terms.

Testing keywords

  • Start from the article. List the searches a reader of that page would plausibly make next.
  • Test wording, not just topics: 'compare broadband deals' and 'broadband deals near me' can attract different advertisers.
  • Judge each term on revenue per search with enough clicks to matter, using keyword-level reporting.
  • Retest by country. A term that pays in the United States may have few advertisers elsewhere.
  • Retire terms that visitors click but advertisers do not value, and ask whether the ad is attracting the wrong audience.

Key takeaways

  • An A/B test compares two versions on the same traffic at the same time with one difference.
  • Judge page tests on revenue per visit after revenue has settled, not on lander CTR alone.
  • Fix the sample size in advance; small gaps need very large samples to be real.
  • Design tricks that push or mislead visitors into clicking are policy violations, not optimisations.
  • On Google feeds, layout and term options depend on account standing and approved mock-ups.

Questions people ask

How do I A/B test a landing page for search arbitrage?

Create two versions that differ in one thing, have your tracker split the same ad traffic between them at random over the same days, and decide beforehand how many visitors each needs. Compare revenue per visitor once the feed's numbers have settled. Rerun the winner once to confirm. Keep both versions within the feed's and the traffic source's policies throughout.

What is a good lander CTR for RSOC pages?

There is no universal benchmark, and chasing one is risky. It varies by traffic source, topic, device and country. What matters is revenue per visit and whether clicks convert for advertisers. A very high click rate achieved through confusing or pushy design tends to bring lower revenue per click and policy trouble. Compare your pages against each other on the same traffic instead.

How long should a keyword test run?

Long enough to collect meaningful paid clicks on each term and to cover different days of the week, since advertiser demand changes through the week. For most campaigns that means several days at least, and longer for low-traffic terms. Then wait for revenue to settle before deciding. Ending a test early because one option is ahead is a common source of false winners.

Can I test different numbers of related search terms?

Only within what your account is permitted. Under Google's Restricted Access Features, showing more than five suggested terms or more than one related search unit per page is reserved for qualified AdSense for Search accounts. If you work through a feed provider, its templates set the options. Test relevance and wording of terms first; that is where most of the gain is.

Previous: Automation rules and bid strategies: letting machines mind the shopGet certified by ClearTrust