AI search glossary · 46 terms
Every AI search term,
defined in plain English.
AEO, GEO, AI Overviews, query fan-out, llms.txt, grounding, sameAs. The vocabulary of AI search in one page, written for business owners, with the source behind each claim. Bookmark it; we update it as the engines change their documentation.
Definitions describe how each term is used in 2026. Where a company's documentation is the source, the wording follows the documentation. Statistics are from the studies listed at the end.
- AEO (answer engine optimization)
- The practice of making a business easy for AI answer engines to find, understand and recommend: crawler access, structured data that matches the page, answer-shaped content, consistent business details, and reviews and profiles a machine can read. See the complete guide.
- AggregateRating
- A schema.org property that states a rating value and a review count for a business or product. Legitimate only when the same rating and reviews are visible on the page; marking up a rating the page does not show violates Google's structured data policies.
- AI Mode
- Google's conversational search experience, launched in 2025 and passing one billion monthly users by May 2026 according to Google. It answers with a synthesized response and links, drawing on the regular Google Search index through query fan-out.
- AI Overviews
- The AI-generated summary Google shows at the top of many search results, reported by Google's CEO at 2.5 billion monthly users in May 2026. It is Google's product, not a discipline; businesses appear in it through normal Google indexing.
- AI search optimization
- A broad label for the same work as AEO and GEO: being findable and citable by AI-driven search products. Used mostly in advertising; not a distinct method.
- AI SEO
- The plain-English label DiamondBack uses for the service that installs the AEO checklist on a business website. Google's view is that all of it is still SEO; the name tells a business owner what it is for.
- AI visibility
- The outcome AEO is trying to produce: whether, and how accurately, AI systems can read, describe and recommend a business. Distinct from citation frequency, which is how often they actually do.
- AI Visibility Score
- DiamondBack's 0 to 100 grade from a 20-point technical scan of a website: structured data, crawler access, rendering, answerability, entity consistency, reviews visibility, sitemap and page health. It measures readability by machines, not how often the business is cited. Bands: At risk, Weak, Developing, Strong.
- Answer engine
- Any system that responds to a question with a written answer rather than a list of links: ChatGPT, Perplexity, Google AI Overviews and AI Mode, Microsoft Copilot, Claude. Most combine a language model with live retrieval from the web.
- Bingbot
- Microsoft's search crawler. The Bing index it builds grounds Microsoft Copilot's answers, so blocking it removes a site from both.
- ChatGPT-User
- OpenAI's user-triggered fetcher: it loads a page when a ChatGPT user's request needs it. OpenAI says it is not an automatic crawler and that robots.txt rules may not apply to it.
- Citation
- A source an answer engine attaches to a statement in its response, usually as a link or a footnote. Being cited is the AEO equivalent of ranking.
- ClaudeBot, Claude-User, Claude-SearchBot
- Anthropic's three agents: ClaudeBot collects content that may contribute to training, Claude-User fetches pages when a Claude user asks, Claude-SearchBot indexes content for Claude's search results. All three honour robots.txt, and blocking one does not block the others.
- Client-side rendering
- A page whose text is assembled by JavaScript in the visitor's browser rather than delivered in the HTML. Human visitors see the page; many crawlers and fetchers see an empty shell. The most common silent cause of AI invisibility.
- Crawler vs fetcher
- A crawler visits pages on its own schedule to build an index or a training set. A fetcher loads one page because a person just asked a question that needs it. Companies document them separately and they obey different rules.
- E-E-A-T
- Experience, Expertise, Authoritativeness and Trustworthiness: the qualities Google's search quality rater guidelines describe for judging content. Not a ranking factor as such, but a useful checklist for what to make visible: named people, credentials, first-hand detail, real reviews.
- Entity
- A specific thing a machine can identify unambiguously: this business, at this address, with this phone number, offering these services. Entity clarity is what structured data and consistent business details produce. Ambiguous entities do not get recommended.
- FAQPage schema
- Schema.org markup that pairs each visible question with its answer. Google limited FAQ rich results to government and health sites in August 2023, so it no longer produces dropdowns in results; it still gives a machine a clean, quotable question and answer with the business attached.
- Generative engine
- The research term, from the 2023 Princeton paper, for a search system that uses a language model to write a synthesized, cited answer from retrieved sources. Perplexity, ChatGPT search and Google AI Mode all fit the definition.
- GEO (generative engine optimization)
- Increasing how often and how prominently a source is cited in generative engine responses. Coined by Aggarwal et al. (Princeton, Nov 2023). Their research found statistics, quotations, cited sources and clear writing raised visibility by up to 40 percent while keyword stuffing lowered it. See what GEO is.
- Google Business Profile
- Google's free listing for a local business (formerly Google My Business). Google's own AI optimization guide points local businesses to it as the place its AI features learn local details. Completeness, categories, hours, photos, posts and reviews all live here.
- Google-Extended
- A robots.txt token, not a separate crawler, that tells Google whether content Googlebot already fetched may be used for Gemini training and grounding. Google states it does not affect inclusion or ranking in Google Search, and it does not control AI Overviews.
- Googlebot
- Google's search crawler. AI Overviews and AI Mode draw on the index it builds, so a site blocked from Googlebot is absent from Google's AI features as well as from the blue links.
- GPTBot
- OpenAI's training crawler. Blocking it excludes a site's content from future OpenAI model training; it does not remove the site from ChatGPT search, which uses OAI-SearchBot.
- Grounding
- Giving a language model retrieved, current documents to base its answer on, instead of relying on what it memorized in training. Every major answer engine grounds its answers in a live index; that is why a crawlable website matters more than a model's training data.
- Hallucination
- A confident statement by a language model that is not supported by its sources. Consistent, machine-readable business details are the practical defence: an engine with clear facts about your hours and address has less to invent.
- JSON-LD
- The format Google recommends for structured data: a small block of JSON inside a script tag in the page's HTML. Invisible to visitors, readable by every crawler that reads HTML.
- Knowledge graph
- A structured database of entities and the relationships between them. Google's Knowledge Graph is the best known. Businesses enter it through consistent signals: Business Profile, structured data, citations and third-party mentions that all agree.
- Large language model (LLM)
- The AI model that reads text and writes text: GPT (OpenAI), Gemini (Google), Claude (Anthropic) and others. In an answer engine the LLM writes the answer; a separate retrieval system finds the sources.
- llms.txt
- A proposed plain-text file at a site's root describing the site for language models, published as a proposal at llmstxt.org in September 2024. Google states it does not use it. Ahrefs found in May 2026 that 97 percent of llms.txt files on 137,210 domains received zero requests. Cheap and harmless; not a ranking lever.
- LocalBusiness schema
- The schema.org type (and its more specific subtypes such as Dentist, Plumber, Attorney, AutoRepair, MedicalBusiness) that identifies a business, its contact details, address or service area, hours and profiles. The single most valuable structured data block on a local website, and absent from 47 percent of the Triad businesses we scanned.
- NAP consistency
- Name, address and phone number matching exactly across the website, Google Business Profile, directories and structured data. Multiple numbers or spellings read to a machine as multiple businesses or a stale one.
- nosnippet, max-snippet, data-nosnippet
- Google's preview controls. Because AI Overviews and AI Mode use the same indexing as Search, these tags limit how much of a page can be shown in AI features too. A business that wants to be recommended leaves them off.
- OAI-SearchBot
- OpenAI's search crawler. It builds the index used to surface websites in ChatGPT's search features and honours robots.txt. This is the OpenAI agent that decides whether ChatGPT search can find you.
- PerplexityBot and Perplexity-User
- Perplexity's search crawler (obeys robots.txt, not used for training) and its user-triggered fetcher (which Perplexity says generally ignores robots.txt because a user asked).
- Query fan-out
- Google's term for how AI Overviews and AI Mode retrieve: issuing multiple related searches across subtopics and data sources for one question, which surfaces a wider set of pages than a single search would.
- Rendering
- Turning a page's code into the text and layout a reader sees. Search engines render; many AI fetchers read the raw HTML only. Content that depends on rendering to appear is at risk.
- Retrieval-augmented generation (RAG)
- The architecture behind answer engines: retrieve relevant documents, then generate an answer from them. AEO is, in effect, optimization for the retrieval step.
- robots.txt
- The plain-text file at a site's root that tells crawlers what they may fetch. Each AI company publishes the user-agent tokens it honours. See the AI crawler guide.
- sameAs
- A schema.org property listing the other web addresses that represent the same entity: the Google Business Profile, Facebook, LinkedIn, Yelp, Wikipedia. It is how a machine confirms that the business on the site is the business on the map.
- Schema.org
- The shared vocabulary for structured data, maintained by Google, Microsoft, Yahoo and Yandex since 2011. Types like LocalBusiness, Service, FAQPage and Review come from here. Only types that exist in the vocabulary count; invented types are ignored.
- Service schema
- A schema.org Service node on a service page naming what is offered, who provides it and where. One per service page, tied to the business's LocalBusiness node by @id.
- Structured data
- Machine-readable labels added to a page, usually as JSON-LD, describing what the page is about. Google says it is not required for its AI features and that no special markup exists for them; it remains the clearest way to remove ambiguity about what a business is.
- Training data vs search index
- Two different ways a model can know about you. Training data is what the model read while being built, frozen at a cutoff. The search index is what the engine looks up at answer time. For a business, the index matters far more: it is current, and it is what gets cited.
- User agent
- The name a crawler or browser sends when it requests a page. robots.txt rules are written against user-agent tokens such as GPTBot or PerplexityBot.
- Zero-click search
- A search that ends without a click on any result, because the answer was on the results page. Pew found users clicked a traditional result on 8 percent of visits when an AI summary was present versus 15 percent without one, and ended the session outright on 26 percent of AI-summary pages versus 16 percent.
A
B
C
E
F
G
H
J
K
L
N
O
P
Q
R
S
T
U
Z
The next step
Know the words.Now check the score.
The free scan grades the twenty checks behind these terms on your actual site, out of 100, in fifteen minutes.
Get your free AI Visibility ScanBook a callSources
- Google Search Central, "Google's guide to optimizing for generative AI features" (May 2026, updated Jul 2026)
- Google Search Central, "AI features and your website" (updated Dec 2025)
- Google Search Central blog, changes to HowTo and FAQ rich results (Aug 2023)
- Google Search Central, overview of Google crawlers and fetchers (Googlebot, Google-Extended)
- OpenAI, "Overview of OpenAI crawlers"
- Anthropic, "Does Anthropic crawl data from the web?" (Claude Help Center)
- Perplexity, "Perplexity crawlers" (developer docs)
- CNBC, Sundar Pichai: AI Overviews now has over 2.5 billion monthly users (May 19, 2026)
- Google, Search at I/O 2026 (May 2026)
- Pew Research Center, "Google users are less likely to click on links when an AI summary appears" (Jul 2025)
- Ahrefs, "We analyzed 137K sites: 97% of llms.txt files never get read" (Jun 2026)
- llmstxt.org, the llms.txt proposal (Sep 2024)
- Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024 (arXiv 2311.09735)
- Schema.org vocabulary
- DiamondBack Advertising, "The State of AI Search in the Piedmont Triad" (79 local businesses, Aug 2026)
DiamondBack scan scores measure what AI systems can read on a website (a 20-point technical check), not how often a business is cited. Third-party figures are quoted from the linked studies as published.