The SEO conversation runs on claims. A tactic worked on one site, a correlation held across a dataset, a conference talk turned it into a rule, and by the time it reaches your roadmap it reads like settled physics. Some of these claims are true, many are true only under conditions nobody mentions, and a few were never true at all. The cost of acting on a bad one is not abstract: it is a quarter of production time spent moving pages in a direction the evidence never supported. The defence is not scepticism as a mood. It is method as a habit.
How do you phrase an SEO claim so it can be checked?
A checkable claim names a change, a metric and a horizon. “Improve the titles” is not a claim; “rewriting the titles of these forty pages will raise their click-through rate within six weeks” is. The transformation is the entire trick: once the claim has a metric and a clock, it can fail, and a claim that can fail is a claim that can be tested. The claims that resist this phrasing, the ones that stay vague no matter how you push them, are telling you something useful: they were never operational, and no test will rescue them.
It helps to separate the claim from its origin. “Google rewards helpful content” is a direction, not a hypothesis; the testable version is the specific change you would ship and the signal you would watch. The primary documentation, Google’s own page on creating helpful, reliable content, is written as questions to ask of a page, which is already closer to a checklist than most of the commentary built on it.
When is a small pilot test enough?

A pilot is enough when the change is reversible and the slice is representative: ten or twenty pages drawn from the same template family, changed together, watched against an unchanged sibling group. It is not enough when the claim is about scale itself, since a site-wide effect cannot be simulated on a shelf of twenty pages, or when the metric needs a season rather than a fortnight. The research tradition has a mature vocabulary for exactly this step: how a small trial reduces risk before a result is generalised, how the research question is framed so the trial can answer it. A Serbian desk devoted to verifiable research keeps pilot study guides on precisely that habit, and the discipline transfers cleanly to a site migration.
The failure mode to refuse is the pilot that cannot lose: ambiguous metric, no comparison group, horizon quietly extended until something moves. A pilot that ends in “hard to say” is not a failed experiment; it is the experiment reporting that the claim, as phrased, was not supported at that scale. That is a result, and acting on it means not shipping the change.
What counts as evidence for a visibility claim?
Evidence is a difference you can locate. The cleanest form is the pilot above: pages changed, siblings unchanged, a divergence that survives the obvious confounders of seasonality and news. Weaker but still useful is the natural experiment, the change you shipped for other reasons whose timing lets you read the before and after. Weakest is the anecdote, including the confident industry anecdote, which belongs at the bottom of the pile no matter how many times it is retweeted. The honesty tests are simple: could this result have happened anyway, and did I decide what would count before I looked?
None of this requires a statistics degree; it requires the posture the research guide already asks for with demand data, and the same suspicion of single numbers we apply to zero-click headlines. A claim earns its way onto the roadmap by surviving a test you designed to be fair to its failure. The companion note on citable guides makes the mirror argument from the other side: the pages that get cited are the ones whose claims can be traced, and the claims worth tracing are the ones somebody bothered to check.