Query Fan-Out and AI Crawlers: A Deep Dive for Nonprofits
2026 edition. Covers the United States, the United Kingdom, German-speaking Europe, the EU and smaller markets.
This is the deep dive behind the pillar guide, Getting Cited in AI Search. The pillar gives you the full sequence: readable pages, fan-out coverage, authority, Search Console reporting, measurement, a 30-day plan and FAQ. This article goes further into two of its parts: how query fan-out decides which organizations get named, and which AI crawlers you are actually making decisions about. If you have not read the pillar yet, start there and come back.
Your nonprofit can rank on page one of Google and still be functionally invisible inside ChatGPT, Google AI Overviews, AI Mode, Claude, Perplexity and Copilot.
That is not a contradiction. It is a symptom of two different systems asking two different questions.
Google's classic index asks: which page best matches this query?
An AI answer engine asks: which organizations are real, relevant and verifiable enough that I can safely name them?
Those are not the same test, and most nonprofit websites are built to pass only the first one.
Key takeaways
- AI systems do not grade your page against the question the person typed. They split it into sub-queries and grade you against all of them. That is query fan-out.
- No single page answers every sub-query. Some are answered by your site, others only by third-party sources that confirm who you are.
- AI crawlers come in three kinds (training, search, user-triggered). Blocking the wrong one either removes you from answers or does nothing at all.
- The default stance is the same as in the pillar: allow the crawlers. Block training bots only with a concrete reason, and leave
Google-Extendedalone.- Your firewall is more likely to lock crawlers out than your robots.txt is.
- A wrong sub-answer is worse than no answer, because it can send a real person to a closed service.
Table of contents
| # | Section | What you get |
|---|---|---|
| 1 | The click is not coming back | Why the funnel changed |
| 2 | Page visibility versus entity visibility | Why fan-out rewards organizations, not just URLs |
| 3 | How query fan-out decides who gets named | The pipeline, a worked sub-query example, a coverage check |
| 4 | Who is fetching your site: the AI crawler fleet | Crawler types, a robots.txt that matches the pillar, firewall, rendering |
| 5 | Answering every sub-query with evidence | Dated numbers, where to find them, audience vocabulary |
| 6 | Schema that resolves the sub-queries | Working JSON-LD and a page-type map |
| 7 | Entity consistency across the web | Register and seal bodies by jurisdiction |
| 8 | When a sub-answer is wrong | Failure modes and a correction protocol |
| 9 | What we still do not know | Documented vs. inferred vs. unproven |
| 10 | Where the rest lives | Search Console, measurement, 30-day plan and FAQ, all in the pillar |
| 11 | Where to start |
1. The click is not coming back
Nonprofit search strategy has run on the same five-step model for twenty years:
Search → ranking → click → website → action.
A prospective donor searched "best homelessness charities in Boston," opened four tabs, compared them, and gave to one.
AI search collapses that sequence:
Question → synthesis → three to six organizations named → one or two citations → action.
The user can learn your mission, location, target population and rough size without ever loading your homepage.
The scale of the shift is no longer speculative. Google reported in June 2026 that AI Overviews had passed 2.5 billion monthly active users and AI Mode had surpassed one billion monthly users. ChatGPT's weekly active user count reached roughly 900 million by February 2026.
The research on click behaviour is not flattering. Pew Research Center's browsing-data analysis found that when an AI summary appears, clicks to websites roughly halve, and links inside the summary are clicked about one percent of the time. Semrush data put the share of AI Mode sessions ending without a click to an external site at 92 to 94 percent.
Google's public position is that AI features still drive substantial traffic to the web. Both things can be partly true at once: total volume can grow while your individual click-through rate falls. What is not in dispute is that the composition of your visibility has changed.
Your website now does two jobs at the same time:
| Role | Audience | Optimized for |
|---|---|---|
| Destination | Humans | Persuasion, trust, conversion |
| Evidence repository | Machines | Extraction, verification, attribution |
An AI system may build one answer about your organization out of a sentence from your impact report, a line from a regulator's register, a funder's grant listing and your programme page. You influenced the answer. You did not get the session.
The organization that wins AI discovery is often not the one with the biggest homepage. It is the one whose identity, programmes, geography and results are easiest to verify.
2. Page visibility versus entity visibility
Several acronyms are circulating. They are less different than the people selling them suggest.
| Term | Full name | What it optimizes for | Unit of competition |
|---|---|---|---|
| SEO | Search Engine Optimization | Ranking a page in results | A URL |
| AEO | Answer Engine Optimization | Making information easy to extract as a direct answer | A passage |
| GEO | Generative Engine Optimization | Raising the odds of being referenced in a generated response | A claim |
| AI visibility | (umbrella term) | Being understood, retrieved, named and cited | The organization |
Google has publicly treated AEO and GEO as vocabulary layered on top of ordinary SEO rather than a separate discipline. That framing is broadly right, with one important exception for nonprofits.
Traditional SEO optimizes a URL. AI systems often need to resolve an entity.
An entity is the organization itself, independent of any one page. Take a fictional example, Hope Children Foundation. It exists across:
- its own website
- a Google Business Profile
- a national charity register
- a donor-rating or seal-of-approval body
- funder and grant databases
- news coverage
- partner organizations' websites
- annual reports and audited accounts
A machine has to decide whether all of those references describe one organization or several. When it cannot, it either skips you or blends you with someone else.
So the real optimization target is not:
example.org/youth-programme
It is a fact that holds up everywhere:
Hope Children Foundation is a registered nonprofit based in Chicago providing after-school education for children aged 6 to 14. In 2025 it served 1,247 students across 14 schools.
The more consistently and verifiably that sentence appears, the easier your organization is to name. Fan-out is what makes this matter, as the next section shows.
3. How query fan-out decides who gets named
Generative search does not rank ten results. A reasonable simplified model of the pipeline:
Query → query expansion → retrieval → evidence selection → synthesis → citation.
Google has described its AI Search systems as using query fan-out, decomposing one complex question into multiple related searches before assembling an answer. The pillar guide explains the eight types of sub-query and how to build your own list of them, so I will not repeat that here. What follows is the part it leaves out: who has to answer each sub-query, you or someone else.
Suppose someone asks:
"Which nonprofits provide job training for refugees in New York?"
The system may effectively be running several searches at once:
| Implicit subquery | What satisfies it | Who answers it |
|---|---|---|
| Refugee-serving nonprofits in New York | Programme pages, directories | Your site and directories |
| Employment and vocational training programmes | Service descriptions | Your site |
| Eligibility and referral criteria | Explicit eligibility sections | Your site |
| Registration and legitimacy | Statutory register entries | A register, not you |
| Programme outcomes | Dated impact statistics | Your site, ideally confirmed by a funder or partner |
| Service geography | Address and areaServed data | Your site and structured data |
One beautifully optimized landing page cannot answer all six. That is the structural reason single-page SEO underperforms here. And notice that at least one row cannot be answered by your own website at all.
Why third-party registers punch above their weight
Structured, externally maintained data is disproportionately useful to a retrieval system because it is attributable and hard to fake.
Compare two signals.
Your website says:
"We transform thousands of lives every year."
A regulator's register says:
Organization: Hope Children Foundation Registration number: 12-3456789 Activity: Youth development Income: 2.3M Location: Chicago, Illinois
The second block is machine-readable, dated and independently maintained. It resolves an entity. The first block resolves nothing.
Your own site is still the anchor. It is simply much stronger when independent sources agree with it. Section 7 maps those sources across jurisdictions.
Run a coverage check on one programme
Pick your most important programme and write out its sub-queries using the table above as a template. For each one, fill in a row:
| Sub-query | URL on your site that answers it | External source that confirms it | Gap? |
|---|---|---|---|
| Who is eligible? | |||
| Where is it available? | |||
| Is the organization registered? | |||
| What were the results, and when? | |||
| How do I apply or refer someone? | |||
| How is it funded? |
Every empty cell in either column is a slot a competitor can fill. Most gaps turn out to be in the second column, and that is where sections 5 to 7 come in.
If your organization works in more than one language, remember that sub-queries may run in other languages too. The same facts, with the same numbers and dates, need to be there in every language you serve. The pillar covers the translation type of sub-query in more detail.
4. Who is fetching your site: the AI crawler fleet
Optimizing content on a site that blocks retrieval bots is decorating a locked room. Before any writing, know who is at the door.
4.1 Know which crawler does what
The most expensive mistake in this field is treating all AI crawlers as one category. Each major vendor runs a small fleet split by job: a training bot that collects content for future models, a search bot that indexes pages for AI answers, and a user bot that fetches a page the moment someone asks about it.
| Vendor | Training crawler | AI search indexing | User-triggered fetch |
|---|---|---|---|
| OpenAI | GPTBot |
OAI-SearchBot |
ChatGPT-User |
| Anthropic | ClaudeBot |
Claude-SearchBot |
Claude-User |
| Perplexity | not declared separately | PerplexityBot |
Perplexity-User |
Google-Extended (a control token, see 4.4) |
Googlebot |
not applicable | |
| Apple | Applebot-Extended |
Applebot |
not applicable |
| Common Crawl | CCBot |
not applicable | not applicable |
Bot names and their jobs change. Check each provider's crawler documentation before you copy anything into your own file.
Two consequences follow.
First, blocking GPTBot does nothing for ChatGPT search visibility, and blocking OAI-SearchBot does everything. OpenAI's documentation tells publishers that sites blocking OAI-SearchBot will not appear in ChatGPT search answers, though navigational links may still appear.
Second, the training and search bots are separate for OpenAI and Anthropic, so those two decisions can be made separately. Google is different, and 4.4 explains why.
4.2 The default stance
This matches the pillar guide. For most nonprofits, allow all of them. Your mission benefits when an assistant can explain what you do and point people to you, and being part of what models already know about your cause has value of its own.
Block training bots only if you have a concrete reason, such as paid training materials or a curriculum you license to others. Blocking search crawlers removes you from the answers people already use to choose where to volunteer, refer and donate.
Treat it as a governance decision, not a technical one. Decide the two questions separately (may models train on us, and may assistants cite us), and write the answer down at board level.
4.3 A robots.txt that matches the pillar
For a nonprofit with no objection to training use, the simplest robots.txt says nothing about AI bots at all, because anything not blocked is allowed. If you prefer to document the decision explicitly, this is the equivalent:
# Search, answers and user-triggered fetches: allowed
User-agent: Googlebot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
Sitemap: https://www.example.org/sitemap.xml
If your board has a concrete reason to block model training, add this and leave everything above untouched:
# Model training: blocked only if you have a reason.
# Google-Extended is deliberately NOT listed (see 4.4).
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
Applebot-Extended is listed in the table for completeness. It is not in the sample block above, so check Apple's own documentation before you add it.
Two honest caveats. Compliance is opt-in: a directive only works if the bot reads and honors it, some crawlers have historically ignored robots.txt, and a user-agent string can be spoofed. And when a person asks an assistant to summarize a specific URL, providers often treat that as user-directed access rather than crawling, so your crawl rules may not apply the way you expect. robots.txt is a norm, not a lock.
It is also not a place to protect anything sensitive. Beneficiary case files, internal PDFs, staff lists and safeguarding documents belong behind a login or off the public server entirely, not behind a Disallow line.
4.4 Google-Extended is not the Google control
Google-Extended is the one entry in the training column that behaves differently, and it is the easiest to get wrong.
- It is not a Google Search control. Google states that it is not a Search ranking signal and does not affect inclusion in Google Search.
- Blocking it does not remove you from AI Overviews or AI Mode. Those are part of Google Search, fed by the same
Googlebotthat ranks you in normal results. - Google also says it governs grounding in the Gemini app, so blocking it can keep your pages out of the answers that app builds from search. For most nonprofits that is the opposite of what they want.
This is why the sample block above leaves it out, and why "do not train on us, but do cite us" cannot be cleanly expressed for Google through Google-Extended the way it can for OpenAI and Anthropic.
The control that actually governs Google's AI features is a setting in Search Console, and today it is all-or-nothing at property level. The UK's Competition and Markets Authority has set a deadline of March 2027 for Google to introduce page-level controls for generative AI features. The pillar explains how to find that setting and why to check who has access to it, so I will not repeat the steps here.
4.5 The firewall is probably the real problem
This surprises nonprofit teams most often. Your robots.txt can say Allow while Cloudflare, Sucuri, Akamai or your host's bot protection returns a 403 to the same crawler.
Check your logs and edge rules for:
| Symptom | Where to look |
|---|---|
| 403 or 429 to AI user agents | Server access logs, CDN analytics |
| Managed challenge or CAPTCHA on non-browser traffic | WAF bot-management rules |
| Geo-blocking left over from an old spam incident | Firewall country rules |
| Aggressive rate limiting on key paths | Rules applied to /programmes/, /impact/ |
Small nonprofits are especially exposed, because bot protection is usually switched on by a volunteer or agency during an incident and never revisited.
4.6 Render the things that matter
Mission, programmes, eligibility criteria, impact figures, locations, donation information and contact details should not depend on client-side JavaScript.
Google can render JavaScript, but most other AI crawlers read the HTML your server sends and stop there, so anything your scripts add afterwards is invisible to them. Real-time search crawlers also have lower tolerance for slow pages and redirect chains than training crawlers do. One extra hop can be enough for a page to be dropped from a generated answer.
The rule: if the information matters for discovery, serve it in crawlable HTML, at a stable URL, in one hop.
4.7 A crawler access check
Fifteen minutes, no special tools:
| Check | How |
|---|---|
| Read your robots.txt | curl -s https://yoursite.org/robots.txt |
| Test for a firewall block | curl -A "OAI-SearchBot" -I https://yoursite.org/ and expect 200, not 403 |
| Repeat for other retrieval bots | Same command with PerplexityBot, Claude-SearchBot |
| Confirm server-rendered content | View page source (not the browser inspector) on a programme page and search for a sentence you care about |
| Confirm the sitemap | https://yoursite.org/sitemap.xml returns and is current |
A 200 to curl shows your rules do not block that user-agent string. Your server logs and CDN analytics show what real crawlers actually receive, so look there as well.
5. Answering every sub-query with evidence
Fan-out rewards pages whose facts survive being lifted out of context. The general craft of writing answer-first, self-contained sections is covered in the pillar and in How to Write Content That AI Models Actually Cite?. Three points matter especially for the sub-queries in section 3.
Date every number, and name its source
| Do not write | Write instead |
|---|---|
| We serve more than 10,000 people. | In the 2025 financial year we provided food assistance to 10,482 people. |
| Most of our funding goes to programmes. | According to our 2025 audited financial statements, 82 percent of expenditure supported programme delivery. |
| We work across the region. | We deliver services in Cook County, Illinois, from six fixed locations and two mobile units. |
An undated statistic ages into an unreliable one. A dated statistic stays citable forever, because it describes a period rather than a claim about now.
Where to find your facts when your M&E is weak
This is where nonprofits actually get stuck. You are not writing "thousands of lives" because you love vague prose. You are writing it because nobody can tell you the number.
The numbers almost always exist already, just not in publishable form.
| Source | What you will find there |
|---|---|
| Donor and grant reports | Counted outputs, because funders demand them |
| Logframe indicator tables | Indicators already defined, baselined and measured |
| Attendance and registration records | Sign-in sheets, enrolment lists, case files |
| Annual accounts | Expenditure by programme, beneficiary counts, staff and volunteer numbers |
| Procurement and distribution logs | Meals served, kits delivered, sessions run |
| Partner reporting | Schools, clinics and municipalities counting you in their own reports |
Pull three numbers, verify each with the person who owns the record, add the period, publish. Three verified dated facts on your Impact page will do more for the "programme outcomes" sub-query than ten new blog posts.
Use your audience's vocabulary alongside your own
Your programme may be "Community Resilience Initiative III" internally. People search for "free food parcels for families in Leeds."
Use both, in that order: plain language first, formal name second. Semantic retrieval bridges concepts well. It cannot bridge a gap you never wrote down.
6. Schema that resolves the sub-queries
Structured data does not buy you a citation. It removes ambiguity, which is a quieter and more durable advantage. In fan-out terms, it is how your site answers the "registration", "geography" and "eligibility" sub-queries in a form a machine does not have to guess at.
Schema.org provides the NGO type and the nonprofitStatus property. A workable homepage implementation:
{
"@context": "https://schema.org",
"@type": ["Organization", "NGO"],
"@id": "https://www.example.org/#organization",
"name": "Hope Children Foundation",
"legalName": "Hope Children Foundation Inc.",
"url": "https://www.example.org/",
"logo": "https://www.example.org/logo.png",
"description": "A nonprofit providing after-school education and family support in Chicago.",
"foundingDate": "2008",
"nonprofitStatus": "https://schema.org/Nonprofit501c3",
"identifier": "12-3456789",
"address": {
"@type": "PostalAddress",
"streetAddress": "120 W Madison St",
"addressLocality": "Chicago",
"addressRegion": "IL",
"postalCode": "60602",
"addressCountry": "US"
},
"areaServed": {
"@type": "AdministrativeArea",
"name": "Cook County, Illinois"
},
"sameAs": [
"https://www.linkedin.com/company/example",
"https://projects.propublica.org/nonprofits/organizations/123456789",
"https://www.charitynavigator.org/ein/123456789"
]
}
Non-US organizations use the same structure, swap nonprofitStatus for the closest applicable value or omit it, and rely on identifier plus sameAs pointing at the national register entry. A UK charity should carry its Charity Commission register page in sameAs and its charity number in identifier.
Schema architecture across the site
| Page type | Schema type | Must include |
|---|---|---|
| Homepage | Organization + NGO |
legalName, identifier, address, areaServed, sameAs |
| Programme pages | Service |
provider, areaServed, audience, eligibility text |
| Articles and news | Article |
datePublished, author, publisher |
| Events | Event |
startDate, location, organizer |
| Team | Person |
jobTitle, worksFor |
| Locations | Place |
address, openingHours |
| Donation | DonateAction |
recipient, target |
On FAQPage markup
Use it if you genuinely have FAQ content. Do not use it expecting a Google FAQ rich result: Google restricted those to authoritative government and health sites in 2023 and has not reversed course.
The general principle applies to all schema. Mark up what is true because it is true. Any claim that a specific markup type "increases AI citations" is, as of today, an assertion without public evidence behind it.
7. Entity consistency across the web
This is the most neglected and highest-leverage work in the article, and it costs nothing but attention. It is also the answer to the sub-queries your website cannot answer by itself.
Search your own organization and you may find:
| Where | Name as published |
|---|---|
| Website | Hope for Children Foundation |
| Statutory register | HOPE CHILDREN FOUNDATION, INC. |
| Hope Foundation USA | |
| Hope4Children | |
| Funder database | Hope Children Fdn |
A human reconciles those instantly. A machine has to decide, on evidence, whether they are one organization or five.
A 30-second test for your own pages
Open your homepage and About page. Can a stranger find, inside 30 seconds:
- legal name
- working name
- registration number
- country and city
- mission in one sentence
- who you serve
- where you serve them
- how to contact you
If any of those requires two clicks or opening a PDF, that is a gap.
Build an Entity Source of Truth
One page, one owner, reviewed annually.
| Field | Canonical value |
|---|---|
| Public name | Hope Children Foundation |
| Legal name | Hope Children Foundation Inc. |
| Abbreviation | HCF |
| Registration number | 12-3456789 |
| Website | https://example.org |
| Headquarters | Chicago, Illinois, United States |
| Founded | 2008 |
| Mission, one sentence | After-school education for children aged 6 to 14 |
| Population served | Children aged 6 to 14 |
| Service geography | Cook County, Illinois |
| Executive Director | Jane Smith |
| Only official donation URL | https://example.org/donate |
Audit every external profile against it. You do not need identical sentences. You need identical facts. A founding year of 2008 on your site, 2011 in a register and 2009 on LinkedIn is not a cosmetic inconsistency. It is an instruction to a machine to treat your entity as unreliable.
Where your verification actually lives, by jurisdiction
Most guides on this topic assume a US organization and stop at Form 990. Here is the broader map.
| Region | Primary statutory register | Seal, rating or transparency body | Other strong sources |
|---|---|---|---|
| United States | IRS Form 990 filings | Charity Navigator, Candid | ProPublica Nonprofit Explorer, state AG registries |
| England and Wales | Charity Commission register | Fundraising Regulator | 360Giving GrantNav and UKGrantmaking (free, open), Companies House |
| Scotland | OSCR | Fundraising Regulator | 360Giving, Companies House |
| Ireland | Charities Regulator | Triple Lock (Charities Institute Ireland) | Revenue CHY listing |
| Switzerland | Cantonal Handelsregister; Stiftungsverzeichnis for foundations | Zewo quality seal, awarded to NPOs meeting its 21 standards | Cantonal tax-exemption listings, SwissFoundations |
| Germany | Vereinsregister; Freistellungsbescheid | DZI Spenden-Siegel; Deutscher Spendenrat transparency certificate | Initiative Transparente Zivilgesellschaft |
| Austria | Vereinsregister; Spendenbegünstigungsliste | Österreichisches Spendengütesiegel | Firmenbuch for gGmbH structures |
| Netherlands | KVK; ANBI status register | CBF Erkenning | Belastingdienst ANBI publication duty |
| France | Journal Officiel des Associations | Comité de la Charte / Don en Confiance | RNA, SIRENE |
| EU level | EU Transparency Register, recording who represents which interests at Union level and with what resources | not applicable | CORDIS and the Funding and Tenders Portal for EU-funded projects |
| Western Balkans | National NGO registers, for example APR in Serbia | rarely available | Donor project databases, public procurement portals, municipal partner listings |
Prioritize in this order:
Statutory register → seal or rating body → major funder listings → LinkedIn → partner and coalition sites → media → social profiles.
If you have received EU, UN or large institutional funding, your CORDIS or donor project page is one of the strongest third-party entity signals you own, and most organizations never link to it.
Fix what you control. For what you do not control, most registers and rating bodies have a correction process. It is slow and worth it.
8. When a sub-answer is wrong
Fan-out has a downside that few people write about. Each sub-query is answered from whatever source the system found, and if that source is stale, the generated answer is stale. For nonprofits this is the highest-stakes part of the subject.
Visibility without accuracy is not a win.
| Failure mode | What it looks like | Who it harms |
|---|---|---|
| Stale programme | A service closed in 2023 described as current | Someone in crisis arrives at an empty building |
| Wrong eligibility | Incorrect age range, income threshold or catchment | People self-exclude, or are turned away |
| Entity blending | Merged with a similarly named organization | Your reputation, their reputation, both |
| Wrong donation route | An old campaign page or retired platform surfaced | Donors, and your income |
| Impersonation | A fraudulent donation site named in your place | Donors, directly |
| Invented specifics | A plausible number or quote you never published | Your credibility with funders and press |
A correction protocol
1. Detect. Record not just whether you were mentioned but whether the description was correct. The prompt panel in the pillar includes an accuracy check for exactly this.
2. Diagnose the source. Ask the system to cite its source, then open it. Nine times out of ten the error is real and lives on your own site, an outdated PDF or an old press release. Fabrications with no source are the minority.
3. Fix the evidence, not the model. You cannot edit the answer. You can correct the source, update the register entry, publish a dated correction, and make the current fact unambiguous and easy to retrieve.
4. Publish an explicit status statement. For discontinued programmes, keep the URL live and replace the content:
"The Riverside Drop-in Centre closed in June 2024. Services previously delivered there are now available at [location]. Referrals should go to [contact]."
This gives retrieval systems a current, dated fact to prefer over a cached one. Deleting the page and returning a 404 is worse, because it leaves stale third-party copies unchallenged.
5. Escalate impersonation. A fraudulent donation site appearing in AI answers about you is a trademark and platform-abuse issue, not an SEO issue. Report it to the platform, your payment provider and your national fraud body, and publish a clear statement naming your only official donation URL.
Accuracy drift is silent, and by the time a beneficiary tells you, the answer has been wrong for months. The quarterly cadence for checking it is in the pillar.
9. What we still do not know
This field is full of confident claims and thin evidence. An honest guide should say where the line is.
| Status | Claim |
|---|---|
| Documented | Which crawlers exist, what each is for, and what blocking each one does |
| Documented | Google-Extended does not affect Google Search inclusion or AI Overviews, and it also covers grounding in the Gemini app |
| Documented | Google's control over its AI search features is set at property level today |
| Reasonable inference | Structured data improves accurate entity resolution. Plausible and cheap, but no vendor publishes a causal claim |
| Reasonable inference | Third-party corroboration increases citation likelihood. Consistent with how retrieval is described, not independently measured |
| Reasonable inference | The sub-queries in section 3 are a model of the behaviour. Google documents that fan-out is used, but does not publish the sub-queries a given question produces |
| Unproven | llms.txt. No major provider has publicly confirmed using it for retrieval or ranking. Nearly free to publish; not a strategy |
| Unproven | Any "AI citation score" or guaranteed placement in AI answers. The systems are not deterministic and vendors do not expose mechanics |
| Unproven | That a specific word count, heading structure or "GEO format" causes citations |
| Actively changing | Crawler economics. Some CDNs now offer pay-per-crawl and bot monetization |
| Actively changing | Page-level AI controls in Google, expected under regulatory pressure by March 2027 |
If someone sells you AI visibility services without distinguishing between these categories, that is your answer about their rigour.
10. Where the rest lives
Several topics that used to sit in this article now live in the pillar guide, Getting Cited in AI Search, so that there is one version of each to keep current.
| If you need | Go to the pillar section |
|---|---|
| The Search Console AI performance report and the generative AI control | "Google's opt-out now lives in Search Console" and "How to tell whether it's working" |
| Measurement: baselines, a monthly prompt panel, server logs | "How to tell whether it's working" |
| A step-by-step plan for a small comms team | "A 30-day plan for a small comms team" |
| Common questions, including whether to block GPTBot | "Frequently asked questions" |
| The eight types of sub-query and where to find yours | "Layer 2: Does your page survive the fan-out?" |
11. Where to start
The nonprofits that will be visible in AI search are not the ones producing the most content. They are the ones easiest to discover, understand, verify, retrieve, quote and trust, across every sub-query a system might generate about them.
Think of it as an evidence network:
| Asset | What it establishes |
|---|---|
| Your website | What the organization is |
| Your programme pages | What it does |
| Your dated impact data | What actually happened |
| Your structured data | How it all relates |
| Your external profiles | That someone else agrees |
Start with the crawler access check in section 4.7 and the coverage check in section 3. Find out what machines can reach, and which sub-queries currently have someone else's answer instead of yours.
Then fix it in this order:
Access → Entity → Evidence → Structure → Corroboration.
Anything done out of that order is decoration. For how to measure the result and how to sequence the work over a month, continue with the pillar: Getting Cited in AI Search.