GEO content formats:
the 52% rule
AI engines do not cite content types evenly. Three formats — listicles, articles and product pages — collect half of every citation across the major engines. The rest of your library still matters, but it is doing a different job.
What the citation data shows
Across roughly 75,000 AI answers and more than a million citations analysed on ChatGPT, Google AI Mode and Perplexity, the pattern is concentration: listicles take 21.9% of citations, articles 16.7%, product pages 13.7%. Three formats, 52% of everything quoted.
This is not a quirk of one dataset. Two independent 2026 studies — Wix Studio's AI Search Lab and HubSpot's State of AEO — arrive at the same shortlist. The citation economy is not a meritocracy of effort; it is a market with a few dominant containers.
The biggest block of the 52% is not yours to publish. Listicle citations go overwhelmingly to third-party round-ups — the "best X" pages that feature you, not the ones on your own blog. Of the 52%, only articles and product pages — about 30 points — are inventory you own outright. The other 22 has to be earned through placement, and that is a PR and partnerships budget, not a content-calendar line.
Intent picks the format
The strongest predictor of which content gets cited is not your industry, and not which AI model you optimise for. It is the intent of the query — and the pattern holds from SaaS to health.
Articles win
"What is X" and "how does X work" queries route overwhelmingly to articles — around 45% of citations, 2.7× more than any other format.
Listicles win
"Best X" and "top X" queries hand roughly 40% of citations to listicles — nearly double any other type. Mostly third-party ones: the round-ups that feature you, not the ones you publish.
Product & category pages win
When the query is ready to buy or looking for a specific brand, around 40% of surfaced pages are product or category pages. Your catalogue is citation inventory.
Comparison content tops ChatGPT at a 95% citation rate — the highest of any format on any engine. Yet standalone comparison and alternatives pages pull under 3% of total citation share. High rate, low volume: when a comparison page is relevant it almost always gets quoted, but those moments are rare. Build them for the buying queries that matter, not as a content programme.
The matrix is a citation heat map
Take the standard 30-content-types matrix. Its two axes were drawn for human marketing — but they turn out to predict machine behaviour.
The horizontal axis — awareness to purchase — is the intent axis. It decides which format an engine reaches for, exactly as the data in section 02 describes. The vertical axis — emotional to rational — is the citation divide. Citations pool in the rational half: blogs and editorial content are the most-cited source types across AI engines, while forums and social platforms rank lowest. The emotional half builds demand humans feel; machines rarely quote it.
Two emotional-half types still earn citations: community forums (#10) and user-generated content (#11). Perplexity in particular leans on community discussion. If your category gets debated on Reddit and forums, that conversation is quietly part of your citation footprint — and it is the one part of your footprint you cannot write yourself.
The citation shortlist
Of the thirty types, eight map directly onto the formats and elements the citation research rewards. These are the ones to brief for citability first.
Articles
The informational workhorse; largest single winner on "what is" queries.
Ratings & Reviews
Third-party proof that engines treat as a trust signal.
Product Features
Feeds the product-page citations that own transactional queries.
Checklists
Step structure is among the easiest things for an engine to lift cleanly.
Case Studies
Named outcomes and numbers — the claims engines quote.
Trend Reports
Original statistics are the single biggest visibility lever available.
Whitepapers
The dense, data-rich guides some engines favour most.
Comparison Guides
The 95%-rate format — build for buying decisions, not for volume.
The other twenty-two are not wasted spend — they are the demand layer. The emotional half makes people ask AI about you by name; the rational half decides whether the answer quotes you. Cut either side and the other underperforms. What changes is the brief: everything in the rational half must now be built to be quoted, not just read.
All thirty types compete for humans.
Only half compete for citations.
The emotional half of the matrix creates the demand that makes people ask AI about you at all. The rational half decides whether the answer quotes you. In the GEO era the funnel map and the visibility map are the same drawing — read it left to right for intent, top to bottom for citability.
What AI checks inside the format
Format gets you considered. Structure gets you quoted. Cited pages share the same interior fittings, whatever the container.
Statistics in the body
Adding statistics is the largest single visibility gain identified in the foundational Princeton / Georgia Tech GEO research — roughly a 40% lift. Numbers under twelve months old earn multiples more citations than stale ones. Verifiable claims are safer for an engine to repeat, so they get repeated.
A visible last-updated date
The three-signal rule from GEO Insight No. 01 applies inside every format: dateModified in the schema, a visible date on the page, and genuinely new substance behind both. A perfect container with a stale timestamp still loses the slot.
Answer-first structure
Intent-matched titles ("What is X", "X vs Y", "Best X"), the answer in the first two sentences of each section, short paragraphs, FAQ blocks with schema, and a named author with a bio. Engines quote what they can cleanly extract and attribute.
Engines have preferences. ChatGPT spreads citations fairly evenly across guides, blogs and listicles. Perplexity leans on blog content and community threads. Claude gives its highest rates to comprehensive, data-dense guides. Know which engine your buyers actually ask, and weight the shortlist accordingly.
The format-mapping workflow
Three steps, applied to the queries that carry commercial weight — not to the whole keyword list.
Classify the money queries
Take the questions that lead to revenue and sort them by intent: informational, commercial, transactional, navigational. The intent label is the routing decision — it tells you which container the engine will reach for before you write anything.
Match the container from the matrix
Informational → articles and guides. Commercial → comparison guides, plus placement in third-party round-ups. Transactional → product and category pages. Fill format gaps before adding volume: another article cannot win a query that wants a comparison page.
Load the citation signals
Every rational-half asset ships with statistics, a live last-updated date, FAQ schema and a named author — as standard fittings, not enhancements. Then keep it inside the freshness ceiling from GEO Insight No. 01.
In the GEO era, your content's citability is not decided by how well it is written — but by which container it ships in, and by whether that container matches the intent of the question.