AI

How Does AI Search Pick Between Two Pages That Say the Same Thing?

Written by
Pravin Kumar
Published on
Sep 12, 2026

How does AI search choose between two pages that say the same thing?

It mostly does not choose between them. It chooses the one that is not interchangeable. Google's own guidance for generative AI features tells publishers to provide a unique point of view, because its AI systems look at a variety of sources and a viewpoint that stands out is what gets surfaced. Sameness is the disqualifier.

This frustrates people, and I understand why. You wrote a clear, accurate, well structured page, a competitor wrote a clear, accurate, well structured page, and only one of you shows up in the answer. It feels arbitrary. It is not arbitrary, but the deciding factor is usually upstream of anything you can fix with formatting.

What follows comes from Google's guide to optimizing for generative AI features on Google Search, which is the closest thing to a rulebook any of these systems publish. I am going to quote what it actually says, then talk about which parts you can act on and which parts are just the cost of entry.

What does Google actually require before a page can appear at all?

Three things, and they are gates rather than tiebreakers. Google states that to be eligible for generative AI features, a page must be indexed, must be eligible to be shown in Google Search with a snippet, and must fulfil the Search technical requirements. It adds that the site must be included in Search generative AI features in Search Console.

Notice the snippet condition. If you have suppressed snippets on a page, whether deliberately through a robots meta tag or accidentally through an inherited template setting, you have removed that page from consideration. I find this on audits more often than you would expect, usually on the exact pages a team most wants cited, because somebody once read that suppressing snippets protects content.

Google is also blunt about what eligibility buys you, which is not much. In its words, just because a page meets all requirements, best practices, and complies with the policies, does not mean Google will crawl, index, or serve its content, and indexing and serving are not guaranteed. That sentence should be printed above every AEO dashboard. Eligibility is permission to compete, not a result.

What is commodity content and why does it lose?

Google names it directly and gives examples. It describes commodity content as something like an article titled seven tips for first time homebuyers, which is based on common knowledge that could originate from anyone and typically adds little unique insight. It contrasts that with non commodity content that provides expert or experienced takes beyond common knowledge.

The example Google uses for the good version is worth sitting with, because it is unusually specific for a documentation page. It offers a hypothetical titled around why the author waived the inspection and saved money, with a look inside the sewer line. That is not a better written version of the tips article. It is a fundamentally different artefact. Someone did a thing and reported what happened.

Here is what I take from that distinction, and it is the single most useful reframe I have found for this work. If your page could have been written by anyone with a search bar, it will be treated as one of many. If it could only have been written by you, because you did the work or made the decision or saw the outcome, it is not competing with the others at all. Most content strategy is an elaborate way of avoiding that conclusion.

Does matching the exact wording of a query still help?

Less than it used to. Google states that its AI systems have advanced further and improved on their ability to understand the relevance of pages even when there is no exact match between the query and the page's primary content. Keyword matching as a tiebreaker is losing force, and semantic understanding is doing more of the work.

The practical consequence is that writing twelve near identical pages to catch twelve phrasings of one question is a worse bet each year. The system is explicitly built to recognise that your one good page answers all twelve. You are not gaining coverage, you are diluting a single asset across a dozen weaker ones and giving yourself a maintenance problem.

Where wording still matters is clarity rather than matching. Use the term your readers and your industry actually use, once, consistently, and define it where a newcomer would need it. That is not keyword optimisation, it is just writing that can be understood without guessing. I covered the mechanics of this in more detail in a piece on writing answer blocks that get cited by AI.

Should you build a page for every fan-out query?

No, and Google explicitly warns against it. The guide says that creating separate content for every possible variation of how people might search, including by focusing on fan-out queries, primarily to manipulate rankings or generative AI responses, violates Google's scaled content abuse spam policy. It also calls the approach ineffective long term.

This matters because fan-out has become a popular thing to sell. The observation underneath it is real, that these systems break a question into subtopics and search several of them at once. The leap from that observation to build a landing page for each subtopic is where it goes wrong, and Google has now named that leap as a policy violation rather than merely a bad idea.

Google's reasoning is worth quoting because it applies far beyond this tactic. A high quantity of pages does not make a website higher quality or more relevant to users. I have audited sites with nine hundred thin pages that would trade all of them for thirty good ones, and I have never audited the reverse.

What tiebreakers can you actually control?

Structure, crawlability, and page experience. Google tells publishers to organise content with paragraphs, sections, and headings that give clear structure, to follow crawling best practices so content is accessible, and to provide a good page experience including displaying well across devices and reducing latency. None of these win on their own.

The honest framing is that these are hygiene factors. They cannot make an interchangeable page distinctive, but their absence can disqualify a distinctive one. A page with a genuinely original finding, rendered by JavaScript that blocks the crawler, contributes nothing. Google notes it can process content within JavaScript as long as it is not blocked, which puts the responsibility squarely on the implementation.

Duplicate content sits in the same bucket. Google advises reducing it, noting it can be a bad user experience and that search engines might waste crawling resources on URLs you do not care about. On a large CMS driven site, duplicate and near duplicate URLs are usually accidental, generated by filters, pagination, or tag archives that nobody decided to create. Those are worth an afternoon.

How do you tell whether your page is the commodity one?

Run a substitution test. Take your page, remove your brand name, and ask whether a competent competitor could publish it verbatim tomorrow without anyone noticing. If the answer is yes, you have written the seven tips article, whatever it is titled, and no amount of schema or heading structure changes that.

The second test is harder and more useful. Go through your page and mark every sentence that contains something you learned rather than something you looked up. On most business blog posts that count is zero, and once you see the zero it becomes very difficult to unsee. That is the actual problem, and it is a sourcing problem rather than a writing problem.

What fixes it is unglamorous. Do the thing, record what happened, and report it including the parts that did not work. A post about an approach that failed and why is non commodity by construction, because nobody else has your failure. This is also why teams that only publish wins struggle here, and why I wrote about why AI answer engines cite competitors instead of you.

What does this change about how you plan content?

It moves the constraint from production to sourcing. If distinctiveness is the tiebreaker, then the limiting factor on your content programme is how much genuine experience you can access, not how fast you can write. Most teams are organised the other way around, with a calendar and no pipeline of things actually worth reporting.

I would rather publish twelve pieces a year that each contain something only we could know than fifty that summarise the field. That is an unpopular recommendation because the twelve are much harder to commission. You cannot brief someone into having done something. You have to go and do it, or go and talk to the people who did.

The practical version inside a company is to attach content to work that is already happening. Every project produces decisions, tradeoffs, surprises, and numbers. Almost none of that reaches a blog, because nobody's job is to capture it. Fixing that is a process change, not a marketing campaign, and it is the highest leverage thing most teams could do here.

What should you do next?

Pick your three most important pages and check two things. First, confirm each is indexed and eligible to be shown with a snippet, because that is a hard gate. Second, run the substitution test on each one and be honest about the result. Those two checks take under an hour together.

If a page fails the substitution test, do not rewrite it yet. Go and find the missing ingredient first, whether that is a real project, a real number, a real decision, or a real conversation with someone who has done the thing. Rewriting commodity content produces better written commodity content, which loses the same tiebreak more elegantly.

If you want a sharper read on why a specific page is not getting picked up, or you are dealing with pages being quoted badly rather than not at all, I wrote about that in stopping AI search from misquoting your pages. And if you would rather have someone go through this with you against your actual site, reach out and let's chat.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.