Startup Blogs

215,128 pages were written so an AI would have something to cite. Your pricing page gets cited first 12% of the time.

Three linked sites published 215,128 machine-generated buying guides, and Perplexity cites them. Of its 7,534 citations, 59.8% point at domains outside the 100,000 most-visited sites. Your own pricing page gets cited first 12% of the time.

Makersclaw — Startup Blogs — 215,128 Pages

Three linked websites published 215,128 machine-generated buying guides between them, and Perplexity cites them.

Trellner Research asked two web-grounded Perplexity models for the best products in 380 buyer-intent categories and kept every URL they retrieved. That produced 7,534 citations across 2,055 distinct domains. Only Perplexity was measured, and the report is explicit that nothing in it should be read as a claim about any other engine.

Here is where those citations went. 59.8% point at domains outside the 100,000 most-visited sites on the web. Inside that number, 23.4% point at domains that are not in the Tranco top million at all. The median rank of the citations that land on a ranked domain is 71,611.

Three of those 2,055 domains are the ones this post is about, and between them they drew 181 of the 7,534: wifitalents.com 71, worldmetrics.org 60, gitnux.org 50. Wikipedia drew three. None of the three existed before December 2023.

Concentration is not the story, and the report says so plainly: the ten most-cited domains take only 17.3% of citations. The story is what fills the other four-fifths. Of the 2,055 cited domains, 751 of them — 36.5% — do not appear in the top million.

Trellner traced them, and the trail is not subtle. All three registered through NameCheap between December 2023 and May 2024, delegate DNS to the same pair of Cloudflare nameservers, and run one page template with identical navigation.

Two of them, worldmetrics.org and gitnux.org, return an HTML title of the form <Brand> — Facts & Grounding Page. Grounding is the retrieval step these models perform. The sites are named after the job they are there to do.

Each site also keeps a blog of exactly six posts. All eighteen posts are about the other brands in the set. A fourth site, zipdo.co, sits on the same nameserver pair.

Shared nameservers are circumstantial. The output is not.

One template, three verdicts

Fetch the same category page, "project estimation software", from all three. Each states its ranking in JSON-LD, so there is nothing to interpret.

  • worldmetrics.org — Float · Scoro · Teamwork.com
  • wifitalents.com — Float · Scoro · Teamwork.com
  • gitnux.org — Saviom · Mosaic · Buildertrend

It is one operation using one template. Yet the third ranking disagrees with the other two. There is no methodology here. There is a page shaped like a methodology.

The operation is also selling its work. worldmetrics.org advertises custom market research "from €5,000", ready-made reports "from €499", and vendor selection "from €2,500". Those offers sit directly above the generated taxonomy that the models retrieve. The citations are the demo.

The free takeaway, before anything else: go read your own category page on those three domains. If your product is in a software category with a "best X" list, you are already ranked somewhere in that 215,128. Nobody ranked it, there were no criteria, and an answer engine may be reading it today. The check takes four minutes. It is the single most useful thing in this post.

The part that should worry you more

A content farm winning citations is annoying. It is unsurprising, and SEO has worked this way for twenty years.

Here is what is new. Kyle Poyar and Nikolas Laskaris tested 7,600 AI responses to pricing questions about the Cloud 100 across six engines between 29 July and 10 August 2026. Each engine received the same flat question: evaluate this company on pricing.

One disclosure before the numbers, because it changes how you should hold them. The dataset came from Profound, an AI marketing platform that is also a commercial partner of the publication that ran the study, and Profound appears later in this post as a success story. The method is disclosed: six engines, 7,600 responses, a stated collection window. The raw data is not. Trust the ordering; treat any single percentage as single-sourced.

A company's own pricing page appeared somewhere in 46% of answers. It was cited first in 12%.

Not one company in the Cloud 100 had its own pricing page cited first in a majority of runs. Not one.

The engines behave differently, which matters when deciding where to spend effort. ChatGPT defaults to owned pricing pages and cites them first 38% of the time. Google AI Mode, Google AI Overviews and Gemini bury them.

The sources filling the gap make sense once you see the numbers: Vendr at 18.7% of runs, Reddit at 18.6%, G2 at 15.9%. Two of those three are places you cannot edit.

Twenty of seventy-seven pages were not fully readable

The obvious explanation is that engines prefer third-party sources because those sources look neutral. That explains some of the result. It does not carry it.

Of the Cloud 100, 77 have a public pricing page. Fifty-seven of those 77 were fully readable to bots. Ten hid at least 40% of their body content.

Four causes account for nearly all of it. Each began as a decision made for a human reason:

  • Client-side rendering. The pricing table populates after JavaScript executes. Most crawlers execute little or none, so they receive an empty shell.
  • Interactive elements. Tabs, accordions, "calculate your price" sliders. If the content only enters the DOM after a click, it does not exist for a crawler.
  • robots.txt blocks. Sometimes deliberate. Often inherited from a template and never revisited.
  • iFrames. The pricing widget is served from a second domain that is not itself crawlable.

Nobody shipped any of those to hide their prices. They wanted a nicer page. It still looks correct to every human who visits, which is why the problem can survive for years without anyone noticing.

So the picture is not "engines snub owned pages." It is closer to this: a quarter of these companies handed the engine a partial page, and in ten cases at least 40% of it was missing. Something else was standing there to fill the rest.

What happens in the vacuum

It is not neutral. Poyar and Laskaris found that the answers frequently read as indictments.

For one company they decline to name, negative pricing language appeared in 72 of 76 responses. Google AI Mode described it as "notoriously opaque, premium-priced, and scales aggressively."

For a vertical AI unicorn with no public pricing page at all, negative language appeared in 67 of 76. The engines assembled a number from third-party threads: pricing "ranges from roughly $100 to $2,400 per user per month depending on platform and scenario." That is a 24x spread for the same product, quoted with confidence and sourced largely from Reddit.

Think about the next sales call. Your prospect arrives with a number. That number is wrong in whichever direction is worse for you.

The volume is not marginal either. Prompts containing cost or pricing keywords are 1–2% of all LLM queries. The share of ChatGPT queries carrying commercial intent went from 13.9% to 19.2% over the past year.

What actually wins is the boring document

The best performer in the study was Plaid, whose pricing page was cited first 42% of the time. Then came Fireworks AI, Fal AI, ElevenLabs.

But look at what the engines pulled from Plaid. The ranking is the less interesting part:

  • Billing docs — 70% of answers
  • Pricing page — 64%
  • FAQs — 50%

The documentation out-cites the pricing page. The reference material wins over the marketing surface. Plaid's docs define one-time, subscription, per-request and flexible billing at the product and endpoint level. They also cover the expensive edge cases and segment by plan type. Density beats simplification because a model wants something specific enough to quote.

Then comes a move that sounds like a mistake. Plaid's docs state explicitly which information is not public. That gives the engine an authoritative reason for a missing dollar figure. It is strictly better than leaving the engine to infer one from a Reddit thread, as the vertical unicorn above discovered.

Fireworks uses a second version of the same idea. It publishes scheduled price changes as separate dated columns, "up to Aug 31" and "from Sep 1". It also exposes an llms.txt documentation index with copy-for-LLM controls on the page.

The cheapest fix anyone has published

Profound, whose dataset this study runs on, built a tool to check whether AI could crawl a given page. Sensibly, they ran it against themselves first.

Their pricing rendered client-side. Visitors saw a working page. Crawlers saw no prices. An external post about Profound, published around the same time, independently found the same problem. That detail makes the example useful instead of merely ironic: outsiders could see the failure before the company could.

They moved pricing to server-side rendering so the numbers appear in the initial HTML response. Deployed 25 June 2026.

Citation bot traffic rose 13% week over week and kept climbing. The pricing page became their second most-cited page across the entire site.

One rendering change. The caveat is worth stating plainly: this is a vendor measuring itself, and the 13% is not independently verified. But the mechanism costs nothing to check on your own site. That is the half that matters.

Two things to do today

One. curl your own /pricing, or load it with JavaScript disabled, and read what comes back. If the prices are not in the initial HTML, no crawler has ever seen them. You have been optimising a page that does not exist for the reader you are worried about.

Two. If exact pricing genuinely is not public, write that on the page. Say what is not published and why. The vertical AI unicorn in this study did not withhold a number; it delegated the number to strangers, and they invented a 24x range.

Here is the uncomfortable synthesis. Three sites generated 215,128 pages so that an answer engine would have something to retrieve. The supply-side bet is paying. Somebody was always going to meet the demand.

The question is not whether you can outrank a content farm. It is whether the thing you actually wrote is legible enough to be quoted instead.

Shreyans BhansaliPublished 3 Sept 2026 · updated 8 Sept 2026

More in Startup Blogs

Get Makersfuel in your inbox

Makersfuel is the Makersclaw newsletter: a five-minute briefing for founders building with AI, with the tools, resources and reads worth saving, and what actually happened. Five mornings a week, Tuesday to Saturday.

Double opt-in. One click in the confirmation email, then Tuesday to Saturday. Unsubscribe from any issue.