Ranking first and being cited are two different games now, and most sites are still only playing the old one.
You can hold position one for a query and never appear when someone asks the same question to ChatGPT. The model reads a handful of sources, writes a single answer, and names two or three of them. If you are not among those, the click never existed — and nothing in your rank tracker would have told you.
This guide is about the mechanics: how an answer gets assembled, why some pages get quoted and others do not, and the specific changes that make yours quotable. It is deliberately technical rather than a tool list.
What "being cited" actually means
Traditional search gives the user ten options and lets them choose. An AI answer gives one synthesised response and attributes a few sources.
The consequences are worth sitting with:
- There is no page two. Being the eleventh-best source is identical to not existing.
- Position matters less than clarity. A model quotes whichever source states the answer most unambiguously, not necessarily the one that ranks highest.
- You compete on being quotable, not on being comprehensive. A 4,000-word guide that buries the answer loses to a 600-word page that states it in the first paragraph.
That last point reverses a decade of SEO instinct, and it is the single most useful thing to internalise here.
How an AI answer is actually assembled
Simplified, but accurate enough to act on. Three stages:
1. Retrieval. The system turns the user's question into one or more searches and pulls a set of candidate pages. If a crawler cannot reach your page, you are eliminated here — before quality is ever considered.
2. Synthesis. The model reads those candidates and composes an answer. It is looking for statements it can lift with confidence: definitions, numbers, comparisons, direct answers. Ambiguous or heavily hedged text is hard to use, so it gets skipped in favour of something cleaner.
3. Attribution. It names the sources that contributed. Pages that supplied a specific, checkable claim are far more likely to be named than pages that supplied general context.
Every recommendation below maps to one of those three stages. Fix retrieval first, because nothing downstream matters if you fail it.
Step 1 — Let the crawlers in
This is the most common silent failure, and it takes ten minutes to check.
AI systems use their own crawlers, separate from Googlebot. If your robots.txt blocks them — or your CDN's bot protection does — you are invisible regardless of how good your content is. The main ones to know:
- GPTBot — OpenAI
- ClaudeBot — Anthropic
- PerplexityBot — Perplexity
- Google-Extended — Google's control for AI training and grounding, separate from Googlebot
Two things routinely go wrong. Some sites copied a robots.txt that disallows these agents, often from a template written when blocking AI crawlers was fashionable. Others never touched robots.txt but sit behind aggressive bot protection that challenges anything unfamiliar — so the crawler gets a CAPTCHA instead of your page.
Check both. Read your robots.txt yourself, then confirm your CDN or firewall is not blocking these agents at the edge. The second one is invisible in robots.txt and catches people out constantly.
The llms.txt question
A newer convention: a file at your root that tells AI systems what your site contains and which pages matter, in plain markdown. Think of it as a sitemap written for a reader rather than a parser.
Adoption is still uneven and no engine formally guarantees it changes anything. But it costs an hour to write, does no harm, and forces you to articulate what your site is actually about — which is a useful exercise regardless. Treat it as cheap insurance, not a growth lever.
What to actually put in it
The format is plain markdown at /llms.txt. Keep it short and factual:
- One line saying what the site is — the same description you use everywhere else, word for word
- A short list of your most important pages, each with a sentence explaining what question it answers
- Any key facts you want stated correctly: what you sell, who it is for, where pricing lives
- Links to documentation if you have it, because docs are cited more than marketing pages
Resist the temptation to dump your whole sitemap in. The value is in curation — you are saying "these fifteen pages are the ones that matter", which is a signal a full sitemap cannot send.
Step 2 — Make your pages quotable
This is where most of the work is, and almost none of it is technical.
Answer the question in the first paragraph
Not after the anecdote. Not after the positioning. The first 50 words should contain the direct answer to the question the page targets.
The classic SEO pattern — build tension, establish authority, deliver the answer in the middle — is actively harmful here. A model scanning for a quotable statement finds nothing in your opening and moves to a source that leads with the answer.
You can still write the full argument afterwards. Lead with the conclusion, then earn it.
One claim per paragraph
Paragraphs carrying three loosely related ideas are hard to quote, because extracting one means dragging in the other two. A paragraph making a single clean point can be lifted whole.
This is why the more useful writing advice for AI visibility is simply "be clearer", not "add keywords".
Publish the specifics you are tempted to hide
Concrete, checkable facts are what get quoted. Vague claims are not.
- Prices. "Contact us for pricing" is unquotable. "$29 a month for 1,000 credits" is exactly what a model reaches for.
- Numbers and limits. Capacities, thresholds, timeframes, supported formats.
- Direct comparisons. Naming alternatives and stating the honest difference.
- Explicit limitations. Saying what your product does not do makes everything else you say more credible — to readers and models alike.
There is a business objection here: publishing prices helps competitors. That has been true forever, and it is now weighed against not appearing in the answer at all. For most businesses the arithmetic has changed.
Use question-shaped headings
Structure the page as the questions people actually ask, with each heading followed immediately by its answer. This maps directly onto how retrieval works, and it makes your page easier to read for humans too — a rare case where both audiences want the same thing.
Keep one H1, and make it say what the page is
Basic, and frequently broken. One H1 that states the page's subject plainly. Clever headlines that hide the topic cost you here.
The same page, before and after
Abstract advice about "being quotable" is hard to act on, so here is one paragraph rewritten.
Before — a typical opening for a page targeting "how much does invoice automation cost":
In today's fast-moving business landscape, managing invoices efficiently has become more critical than ever. Companies of all sizes are discovering that manual data entry is holding them back. Our team has spent years helping businesses transform their financial operations, and we understand the challenges you face. In this guide, we will explore everything you need to know about invoice automation pricing.
Sixty-eight words, and not one checkable fact. There is nothing here a model can lift. It will read this, find no answer, and move to the next candidate.
After:
Invoice automation costs between $0 and $30 a month for most small businesses processing under 200 documents. Free tiers typically cover 30 to 100 documents monthly. Paid plans start around $24 a month and are billed per document parsed, usually one to three credits per page depending on whether the layout is fixed or needs AI extraction. Enterprise volumes above 5,000 documents a month move to custom pricing.
Sixty-six words — almost identical length — and it contains six extractable facts: a range, a free-tier size, an entry price, a billing unit, a cost driver, and the threshold where pricing changes shape.
Ask yourself which of those two paragraphs a model would quote when someone asks what invoice automation costs. Then check your own top pages against the same test. The rewrite is rarely about adding words; it is about replacing throat-clearing with specifics you already know.
The comparison page nobody wants to write
Comparison pages are disproportionately cited, and most businesses write them badly because they are trying to win rather than inform.
A comparison page that only says you are best is not useful to a model, because it is indistinguishable from every vendor's page about themselves. One that genuinely differentiates gets quoted because it answers the question the user asked.
The structure that works:
- State the honest one-line summary first. "X is better for teams under ten; Y is better if you need SSO." That single sentence is what gets lifted.
- Give a real table. Price, the limits that actually bite, the one feature each side wins on.
- Name the cases where you lose. This is the part that feels wrong and does the most work. A page that says "if you need X, use the competitor" reads as trustworthy to a model and to a human.
- Keep it current and date it. Pricing changes; a comparison with stale numbers gets quoted with stale numbers, which is worse than not being quoted.
If writing that feels uncomfortable, notice that your competitors feel the same way — which is exactly why the slot is usually open.
Step 3 — Be a defined entity
Models need to know what you are before they can recommend you. That means structured data and consistency.
Schema.org JSON-LD is the practical mechanism. At minimum: an Organization entity describing who you are, and appropriate types for your content — Article, Product, SoftwareApplication, whatever fits. FAQ markup where you genuinely have questions and answers.
A word of caution learned the expensive way: mark up what is actually on the page. Schema describing things a visitor cannot see is a manual-action risk, and fabricated review markup in particular has cost sites their rich results.
Consistency across the web matters as much as your own markup. If your company name, description and category are stated differently on your site, your LinkedIn, your directory listings and your docs, you are a fuzzy entity. Models resolve fuzzy entities poorly. Pick one description and use it everywhere, word for word.
Step 4 — Get cited where models actually look
Here is the uncomfortable finding when you examine real AI answers: the citation is often not the vendor's own homepage.
It is usually one of these:
- A comparison or listicle page that names several options and states differences
- A directory or review listing with structured, consistent data
- Documentation that answers a specific question directly
- A community discussion where the answer was argued out
The pattern is that these sources answer the question in the form it was asked. Homepages sell; they rarely answer.
The practical implication: getting cited is partly an off-site problem. Being present and accurately described in the places that already get quoted matters, and publishing your own honest comparison page — one that genuinely acknowledges when a competitor is the better choice — is one of the highest-return pages you can write for this.
Step 5 — Check your work
Do the manual version first, because it teaches you more than any dashboard.
Write the ten questions a buyer would type before choosing something like you — category questions, not your brand name. Ask each one in ChatGPT and Perplexity. Note who gets named and, crucially, what the answer cited. That is your competitive brief.
Then automate the check.
AISEO USA AI is the fastest place to start, because it is completely free with no signup and it checks precisely the things covered above. It scores a URL live across five dimensions: crawler access for GPTBot, ClaudeBot, PerplexityBot and Google-Extended plus whether you publish an llms.txt; structured data including Schema.org JSON-LD, an Organization entity and FAQ markup; answer-readiness covering title, meta, a single H1, question-format headings and depth; and entity and E-E-A-T signals. It returns the fixes in priority order. Submitting an email is optional and returns a fuller report.
For ongoing monitoring, the options split by budget. AnswerScout tracks your appearance across ChatGPT, Claude, Gemini, Grok and Perplexity with a daily 0-100 score, competitor intelligence and citation tracking, from $39.99 a month. Morningscore bundles AI visibility with conventional SEO from $69, including a GEO Score and AI Overviews tracking. SE Ranking does the same at a larger scale from $103.20. And if you want the work executed rather than reported, Corank covers seven engines and runs the optimisation for you, starting at $2,500 a month.
For the traffic that does arrive, Page Pulse is free and tracks visitors, conversions and channels with automatic click tracking — useful for spotting AI referral traffic before it is large enough to appear anywhere else.
The 90-minute audit you can run today
Everything above, compressed into one sitting. No purchases required.
Minutes 0-10 — the gate. Open yoursite.com/robots.txt and read it. Look for any Disallow affecting GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Then check your CDN or firewall settings for bot rules. If you find a block you did not intend, stop here and fix it — nothing else on this list matters until you do.
Minutes 10-25 — the manual reality check. Write your ten category questions and ask them in ChatGPT and Perplexity. Record three things for each: were you named, who was, and what was cited. Fifteen minutes of this tells you more about your position than any dashboard will in a month.
Minutes 25-40 — the first-paragraph test. Open your five most important pages. For each, read only the first fifty words and ask whether they contain the answer the page promises. Most will not. Mark the ones that fail; those are your rewrite queue, in order of business importance.
Minutes 40-55 — the specifics sweep. On those same pages, count the checkable facts — prices, limits, numbers, named alternatives. If a page has fewer than three, it is unlikely to be quoted for anything. Note what you know but have not published, which is usually more than you expect.
Minutes 55-70 — entity consistency. Open your homepage, your LinkedIn, and two directory listings side by side. Compare how each describes what you do. If the four descriptions differ meaningfully, write one canonical sentence now and diary the updates.
Minutes 70-85 — the automated check. Run your homepage and your two most important pages through a free visibility checker. It will confirm what you found manually and catch the technical items you cannot see — schema gaps, missing Organization entity, heading structure problems.
Minutes 85-90 — decide the first three fixes. Not ten. Three. In almost every audit the same three surface: unblock a crawler, rewrite one opening paragraph, publish a price. Do those, then wait a fortnight before looking again.
The discipline that matters here is stopping at three. Audits generate long lists, long lists generate paralysis, and the sites that improve are the ones that shipped a small change this week rather than planning a large one for next quarter.
What does not work
Time and money get wasted on these, so they are worth naming.
Keyword stuffing, in any modern form. The model is reading for meaning. Density does nothing.
Publishing volume for its own sake. Forty thin AI-generated pages do not increase your odds. They dilute your entity and create a quality problem you will have to clean up.
Schema on content that is not there. Covered above, and worth repeating because it carries real downside rather than merely wasting effort.
Chasing every engine at once. Find where your buyers actually ask. For most B2B that is ChatGPT and Perplexity; for consumer research it is increasingly Google's AI Overviews. Optimising for all seven simultaneously is how small teams achieve nothing anywhere.
Waiting for a definitive playbook. There is not one yet, and the fundamentals — be reachable, be clear, be specific, be consistent — have not changed in twenty years of search and are unlikely to now.
A realistic timeline
Expectations are where most of these projects die.
Days. Crawler access fixes take effect quickly. If you were blocked, unblocking is the fastest win available.
Two to six weeks. Page-level rewrites — leading with the answer, publishing specifics, question headings — start showing up as your pages are recrawled.
Two to six months. Entity consistency and off-site citations. This is slow because it depends on other people's publishing schedules, not yours.
Never, if the answer is not on your site. No amount of technical work makes a model cite you for a question you have not answered. Sometimes the honest fix is to write the page.
Common questions
Does traditional SEO still matter?
Yes, considerably. Retrieval leans on the same signals — crawlability, authority, relevance. AI visibility is a layer on top of SEO, not a replacement for it. A site that ranks nowhere rarely gets cited.
Should I block AI crawlers to protect my content?
That is a genuine strategic choice, not an obvious one. Blocking protects your text from being used, and guarantees you are never cited. Publishers with paywalled archives reasonably choose to block; businesses that want to be recommended reasonably do not. What you should not do is block by accident, which is what usually happens.
How do I know if AI search is sending me traffic?
Look for referrals from chatgpt.com, perplexity.ai and similar domains in your analytics. Volumes are small but the visitors tend to convert well, because they arrived after a recommendation rather than a search result.
Is llms.txt worth doing?
It costs an hour and the downside is zero. Just do not treat it as the strategy — crawler access and quotable content matter far more.
What if a model says something wrong about my business?
Usually it is reading outdated or inconsistent information somewhere. Fix the source: update your own pages, correct directory listings, and make sure your description is identical everywhere. Monitoring tools that flag incorrect statements exist for exactly this.
Do I need to pay for a tool to start?
No. The manual pass costs nothing, and the free checker covers the technical audit. Buy monitoring once you have made changes worth measuring.
The bottom line
Getting cited is not a new discipline so much as an old one with the padding removed.
Let the crawlers in. Answer the question in the first paragraph. Publish the specifics you were tempted to keep vague — especially your prices. Describe yourself identically everywhere. Write the comparison page your competitors are too proud to write.
Then check it with a free tool, and give it two months before deciding whether it worked.
