If you run a law firm, someone has probably offered to get you "ranked in ChatGPT". The offer usually comes with a screenshot and a confident timeline.
The difficulty is not that the pitch is always wrong. It is that most firms have no way to judge it.
My view is this. Answer engine optimisation — the work of improving how accurately and prominently a firm is represented in AI-generated answers — is commercially important. It is also immature. The systems change often, the providers behave differently, and a lot of what is sold as knowledge is inference or guesswork.
So rather than another list of tactics, this piece grades the claims. There are four buckets:
- What we know — documented by the providers themselves or by credible research.
- What seems sensible — reasonable practice with indirect support, but no proven causal link to AI recommendations.
- What we are testing — open research questions.
- What is currently overclaimed — claims I would not accept from anyone.
If you want the foundations first, start with AEO for Law Firms. This piece is about how much weight each idea can bear.
The two commercial risks
There are two ways to get this wrong.
The first is paying for outcomes nobody can verify. If a vendor cannot show you a repeatable measurement, you are buying a story.
The second is ignoring the channel entirely. AI systems can put a small number of firms into someone's consideration set before that person visits a website. If that is happening in your practice areas, you want to know.
Both risks have the same remedy: measure before you optimise.
Part 1: What we know
Confidence: high. Sourced to primary documentation or published research, with limits stated.
Google says there is no secret AI Overviews technique. Google's Search Central guidance on AI features says there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary", and "no special schema.org structured data that you need to add". A page must be indexed and eligible to show in Search with a snippet. (Google Search Central, AI features and your website, last updated 10 December 2025.)
That is the plainest statement from the largest search provider. Anyone selling a special Google AI technique should be asked how it squares with it.
Traditional search and AI answers are linked — at least for search-grounded systems. Google describes AI Overviews and AI Mode as using "query fan-out" — running multiple related searches across subtopics to build a response. Microsoft's Bing team says content must first be indexed before it can be cited in AI answers (Microsoft, Optimizing Your Content for Inclusion in AI Search Answers, 8 October 2025). What neither provider publishes is how much ordinary ranking position determines what gets cited. We don't know that yet.
AI crawler controls are separate, and they differ by provider. This is where firms most often make a technical mistake.
- OpenAI runs OAI-SearchBot for ChatGPT search and GPTBot for model training. They are independent: a site can allow OAI-SearchBot to appear in search results while disallowing GPTBot. ChatGPT-User acts on user requests and is not used for automatic crawling. (OpenAI, Overview of OpenAI Crawlers)
- Anthropic runs ClaudeBot (training), Claude-User (fetching pages when a user asks) and Claude-SearchBot (search indexing). Anthropic says blocking Claude-SearchBot or Claude-User may reduce a site's visibility in user search results. (Anthropic, Claude Help Center)
- Perplexity says PerplexityBot surfaces sites in search results and is not used to crawl content for foundation models, while Perplexity-User "generally ignores robots.txt rules" because a user initiated the request. (Perplexity, Perplexity Crawlers)
The practical point: "block AI" is not one decision. A firm that blocks everything to stop model training may also be removing itself from AI search. A firm that allows everything may be making a training decision it never considered. Both should be deliberate.
Mass-produced content is a documented risk. Google's spam policies list "using generative AI tools or other similar tools to generate many pages without adding value for users" as an example of scaled content abuse (Google Search Central, Spam policies). The policy targets purpose and value, not AI as such. But it cuts directly against the idea that publishing hundreds of AI-written pages is an AEO strategy.
AI recommendations are highly variable between runs. In a study published in January 2026, SparkToro and Gumshoe had 600 volunteers run 12 recommendation prompts through ChatGPT, Claude and Google's AI 2,961 times. They reported less than a 1 in 100 chance of getting the same list of brands in any two responses, and roughly 1 in 1,000 of getting the same list in the same order (SparkToro, January 2026).
Two caveats matter. The authors say the work was not peer-reviewed, and the prompts covered consumer categories, not law firms. It should not be read as a measurement of legal recommendations. But it is strong evidence against treating one answer as a ranking. Notably, the authors concluded that visibility measured across many prompts run many times was a reasonable metric, while a single "ranking position" was not.
Content changes can affect generative visibility — in controlled settings. The best-known academic paper, GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024), reported that certain content changes could boost visibility in generative engine responses by up to 40% on the authors' benchmark (arXiv 2311.09735). The authors also found that effectiveness varied across domains. That is useful evidence that content matters. It is not evidence that a given law firm's changes will produce a given result in ChatGPT.
Part 2: What seems sensible
Confidence: moderate. Good practice with indirect support. None of it has a proven causal link to being recommended.
This is where most sensible AEO work sits. I would do all of it. I would not promise any of it produces a recommendation.
Entity clarity. The firm's name, practice areas, locations and lawyers should be described consistently across its website, directories, profiles and media. If a person has to work out whether two listings are the same firm, a machine may struggle too. Basis: reasoning from how retrieval and synthesis work; no published causal study for law firms.
Expertise architecture. A firm should make it obvious what it is genuinely expert in — not a list of 40 practice areas, but clear pages that match how clients describe their problems, connected to the lawyers who do the work. Basis: consistent with Google's and Microsoft's emphasis on clear, helpful, well-structured content.
Lawyer expertise, visibly evidenced. Lawyer profiles that show jurisdiction, admission, experience, publications and matters (where permitted) give any system — and any client — something concrete to work with. Basis: opinion informed by provider guidance on clear, trustworthy content.
Useful content. Pages that answer real client questions plainly. Microsoft's guidance favours clear headings and self-contained answers that can be lifted into a response. That is also just good writing for clients.
Technical accessibility. Pages that load, render without heavy scripts, and are not accidentally blocked. Check your robots.txt against each provider's documentation above. Basis: documented — a page that cannot be crawled cannot be retrieved.
Third-party authority and citations. What credible external sources — law society listings, legal directories, courts' published judgments where appropriate, reputable media — say about the firm. AI answers draw on more than the firm's own site. Basis: inference. Which sources carry weight, and how much, is not published by any provider.
Digital reputation and reviews. Reviews and reputation signals are part of what a system may retrieve about a firm. Basis: inference. Note that review practices are regulated differently across jurisdictions; this is not advice on soliciting reviews.
Structured data as hygiene. Schema markup helps machines interpret a page. Google describes structured data as making content eligible for enhanced display, not guaranteed (Google, Intro to structured data), and says no special schema is needed for AI features. Microsoft says schema helps machines interpret content but that there is "no secret sauce that guarantees selection in AI answers". I treat schema as hygiene, not strategy.
Strong traditional search visibility. Given how search-grounded AI systems describe themselves, I think neglecting SEO in favour of "AEO" is a mistake. They are overlapping work, not rivals.
Part 3: What we are testing
Confidence: unknown. These are research questions, not findings.
These are the questions I think matter most for law firms, and the kind of questions we are designing FirmRanker to examine. I am not reporting results here.
- Recommendation stability. When the same legal question is asked repeatedly, how stable is the set of firms recommended — and does that vary by practice area or city?
- Mentions versus recommendations. How often is a firm merely mentioned, suggested as one option, or explicitly recommended? These are different outcomes. (See What Actually Counts as an AI Recommendation?)
- Cross-provider agreement. Do different AI systems converge on the same firms, or does each have its own view? Providers are not equivalent, and results from one should not be assumed for another.
- Sources retrieved, cited and influential. A source can be retrieved without being cited, and cited without being what actually shaped the answer. Can we tell these apart well enough to guide investment?
- Entity resolution. When an answer names "Smith Lawyers", which Smith Lawyers is it? Measurement is only as good as its ability to match mentions to the right firm.
- Change over time. If a firm improves its entity clarity or third-party presence, does its measured visibility change — and can that be separated from the systems themselves changing?
The methodological principles are straightforward even where the answers are not. One answer is not a ranking. Repeated observations matter. Raw observations should be preserved so conclusions can be checked. You can read how we approach this on the methodology page.
Part 4: What is currently overclaimed
Confidence: high that these claims are not supported by current evidence.
"We guarantee you'll rank in ChatGPT." No provider offers a mechanism that would allow this. Given how much answers vary between runs, a guarantee of position is not something anyone can honestly give. I set out why ranking may be the wrong model in Why "Ranking in ChatGPT" May Be the Wrong Mental Model.
"Once you're in, you stay in." Models, retrieval systems and sources change. There is no published basis for permanent AI positions.
"Add schema and AI will recommend you." Google says no special schema is needed for its AI features. Microsoft says nothing guarantees selection. Schema can help interpretation; it is not a cause of recommendation.
"Publish hundreds of AI-generated pages." At best this adds noise. At worst it fits Google's own description of scaled content abuse.
"One citation will get you recommended." A single directory listing or article might contribute. There is no evidence that one source reliably causes recommendations.
"Here's a screenshot — it worked." A screenshot shows one answer, to one prompt, on one system, at one moment. Given the variability documented above, it is a demonstration, not measurement.
The operating model: Measure → Diagnose → Improve → Re-measure → Learn
This is the model I use.
- Measure. Establish a baseline across realistic client questions, relevant practice areas and markets, more than one AI system, and repeated runs.
- Diagnose. Work out where the firm is absent, misdescribed or outnumbered — and which sources and competitors appear instead.
- Improve. Make specific, recorded changes: entity clarity, expertise pages, lawyer profiles, technical access, third-party presence.
- Re-measure. Repeat the same measurement under the same conditions.
- Learn. Treat the result as evidence, not proof. Keep what appears to work; drop what doesn't; note what you still can't explain.
Why measurement comes first. Without a baseline, there is no way to know whether any change made a difference — or whether the systems simply moved. Measurement also stops firms spending on problems they don't have. Some firms will find they are already well represented. Others will find they are invisible, or described incorrectly. Those need very different responses.
Questions to ask anyone selling you AEO
- How do you measure visibility — how many prompts, how many runs, on which systems?
- Do you distinguish a mention from a recommendation?
- What was our baseline before you started?
- How will you separate the effect of your work from changes in the AI systems themselves?
- Do you guarantee any position? (If yes, ask what that guarantee is based on.)
- Will we see the raw answers, or only a score?
What we don't know
We don't know how much each provider weights ordinary search rankings, third-party sources or reviews. We don't know how stable legal recommendations are in particular markets. And the systems are changing quickly enough that some of this piece will date. I will update it when the evidence changes.
Where this fits
FirmRanker measures how law firms appear across AI systems. I interpret what the evidence means here on Law By Dan. Practice Proof helps firms act on it. For the wider context, see AI & Law and Research.
Disclosure: FirmRanker was founded by Dan Toombs. Practice Proof, also founded by Dan, provides digital and AI visibility services to law firms.
About the author Dan Toombs is a lawyer, founder of FirmRanker and Practice Proof, and a researcher focused on how technology is changing the way people find and choose lawyers. His current work examines how AI systems discover, evaluate and recommend law firms, and what that means for legal marketing, digital authority and the future of legal discovery. About Dan →
This article is commentary on legal marketing and technology, not legal advice.
Sources
- Google Search Central — AI features and your website (last updated 10 Dec 2025; accessed 28 Sep 2026)
- Google Search Central — Spam policies for Google web search (accessed 28 Sep 2026)
- Google Search Central — Introduction to structured data markup (accessed 28 Sep 2026)
- OpenAI — Overview of OpenAI Crawlers (accessed 28 Sep 2026)
- Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler? (accessed 28 Sep 2026)
- Perplexity — Perplexity Crawlers (accessed 28 Sep 2026)
- Microsoft (Krishna Madhavan, Bing) — Optimizing Your Content for Inclusion in AI Search Answers (8 Oct 2025; accessed 28 Sep 2026)
- Microsoft Bing — Introducing AI Performance in Bing Webmaster Tools (Public Preview) (10 Feb 2026; accessed 28 Sep 2026)
- SparkToro (Rand Fishkin with Patrick O'Donnell, Gumshoe) — AIs are highly inconsistent when recommending brands or products (28 Jan 2026; accessed 28 Sep 2026)
- Aggarwal et al. — GEO: Generative Engine Optimization, KDD 2024 (accessed 28 Sep 2026)
Editor's note: source 8 (Bing AI Performance) is cited in the SEO pass below as a recommended addition to Part 1 — see §5.