Imagine a managing partner is forwarded a screenshot. Someone in the firm typed "best employment lawyer near me" into ChatGPT. The firm came third. Or didn't appear. By the afternoon, the screenshot is in a partners' email thread with the words "we rank third in ChatGPT".
That example is hypothetical, but the pattern will be familiar to anyone who works in legal marketing, in Australia, the UK or the US. And I think the word doing the damage in that sentence is "rank".
My view is that "ranking in ChatGPT" is a borrowed mental model. It feels natural because we have spent two decades talking about search rankings. But it may lead firms to ask the wrong question, trust the wrong evidence and spend money on the wrong things.
This piece sets out why, and what might replace it. I've tried to keep four things separate throughout:
- (a) What we can observe about generative AI systems generally.
- (b) What external evidence supports.
- (c) What FirmRanker is designed to test.
- (d) My hypothesis — which is exactly that, a hypothesis.
Why "ranking" made sense in search
A traditional search results page gave everyone a shared reference point: a query, an ordered list of links, a position. Results were always personalised and localised to a degree, and local results in particular move around. But "we're in position four for this query" was a meaningful, repeatable statement that a firm could track month to month.
The whole vocabulary of SEO reporting — rank, position, page one — rests on that assumption of a relatively stable ordered list.
Why an AI answer is a different kind of object
(a) Observed. An AI assistant doesn't return a list of documents. It returns a piece of generated text. Retrieval may feed that text — Google says its AI Overviews and AI Mode may use "query fan-out", issuing multiple related searches across subtopics to build a response (Google Search Central). But what the user reads is written fresh each time.
(b) External evidence. The providers themselves say their outputs aren't fully repeatable. Anthropic's API documentation notes that even at a temperature of 0.0, "the results will not be fully deterministic" (Anthropic). OpenAI offers a seed parameter for what it calls "mostly" deterministic outputs and states that determinism is not guaranteed (OpenAI Cookbook). Researchers testing five models in supposedly deterministic settings found outputs still varied across runs, with accuracy moving by up to 15% (Atil et al., 2024).
Two cautions. First, those statements concern developer APIs. The consumer ChatGPT, Gemini or Claude apps add further variables — location, account memory, whether web search is triggered, which underlying model answers. Second, Atil et al. measured task accuracy, not law-firm recommendations. The point is narrower: variation between identical requests is a documented property of these systems, not a glitch.
What happens to brand lists
The most directly relevant external study I've found is industry research from SparkToro and Gumshoe, published in January 2026 (SparkToro). Volunteers ran 12 recommendation prompts nearly 3,000 times across ChatGPT, Claude and Google's AI features in late 2025.
They reported less than a 1-in-100 chance that ChatGPT or Google's AI would return the same list of brands in two responses, and identical order was rarer still. But how often a given brand appeared across many runs was more consistent — which led the authors to suggest visibility percentage across many repeated prompts is a more reasonable metric than position.
That study deserves careful handling. It wasn't peer-reviewed; the authors themselves flag limited samples and non-specialist methods; Gumshoe sells AI-visibility tracking; and none of the 12 categories was legal services (the nearest were cancer-care hospitals and digital marketing consultants). It supports a caution, not a law-firm conclusion. But the caution is important: if the same question rarely produces the same list, a single list isn't a ranking.
Why "one screenshot = one ranking" is methodologically weak
If you strip it back, a screenshot is one observation of one prompt, on one product, one model version, one day, in one location, for one account. Treating it as a ranking quietly assumes:
- The answer would be the same if asked again. External evidence says often it wouldn't.
- The prompt is representative. Real clients ask "do I need a lawyer for a workplace investigation?" at least as often as "best employment lawyer". A firm can be absent from one and prominent in the other.
- All AI systems behave alike. ChatGPT, Gemini, Claude, Perplexity and Copilot use different models and retrieval arrangements. There's no reason to assume they agree.
- Being named means being recommended. "Firms in this area include X" is not the same as "I'd suggest X for your situation".
- The firm in the answer is your firm. Similar names, old trading names, individual lawyers and merged practices get conflated.
Each is a testable assumption. Most screenshots test none of them.
A better vocabulary: eight candidate measures
What follows is my proposed framework, not an industry standard. Some of these measures require judgement calls in how they're coded, and those rules need to be published for any number to mean anything.
| Measure | Question it answers |
|---|---|
| Appearance frequency | Across repeated runs, how often is the firm named at all? |
| Recommendation frequency | How often is it actually recommended, not merely mentioned? |
| Recommendation strength | When recommended, is it the lead suggestion, one of several, or a caveated afterthought? |
| Top-three share | How often does it appear among the first few firms presented? |
| Prompt coverage | Across the range of questions real clients ask, how many produce an appearance? |
| Model coverage | Across different AI systems, where does the firm appear and where is it absent? |
| Recommendation stability | Does the pattern hold across days, weeks and model updates? |
| Cross-model consensus | Do independent systems converge on the same firms for the same need? |
Notice what's missing: a single number called "rank". Top-three share keeps the useful part of the ranking idea — prominence matters — while acknowledging that position is a distribution, not a fixed point.
The academic literature already leans this way. The paper that coined "generative engine optimisation" measured visibility as a share of the generated response rather than a position, reporting improvements of up to 40% in its benchmark (Aggarwal et al., KDD 2024). That was a controlled benchmark, not live legal queries, but the measurement choice is telling.
Sources are a separate question
Firms increasingly ask "which sources does ChatGPT cite about us?" That's worth asking, but it's another area where the mental model matters.
(b) External evidence suggests citations aren't a reliable map of what's true, let alone of what influenced an answer. Stanford researchers found that only 51.5% of sentences produced by four generative search engines were fully supported by their citations, and 74.5% of citations supported the statement they were attached to (Liu, Zhang & Liang, 2023). Those were 2023 systems. More recently, the Tow Center at Columbia tested eight AI search tools on identifying the source of news excerpts and found they gave incorrect answers to more than 60% of queries (Columbia Journalism Review, 2025).
Neither study measures influence. The principle I work from is methodological: a source retrieved is not necessarily a source cited, and a source cited is not necessarily the source that shaped the answer. Those are three different things, and conflating them produces bad strategy.
What FirmRanker is designed to test
(c) To be explicit: this article reports no FirmRanker findings. What I can describe are the principles the system was built around, because building it forced the questions.
- One AI answer is an observation, not a ranking.
- Mentions, suggestions and recommendations are coded separately.
- Where outputs vary, repeated observations of the same prompt matter.
- Providers aren't treated as automatically equivalent.
- Raw outputs are preserved so any classification can be re-checked.
- Entity resolution — deciding which "Smith & Co" an answer refers to — is treated as a core problem, not a clean-up task.
- Retrieved, cited and influential sources are kept distinct.
The research questions follow. How stable are law-firm recommendations over repeated runs? Do different models agree? How much does wording change the answer? Does the pattern differ by practice area and market? We don't know yet. We're testing that, and the methodology will be published alongside any results.
My hypothesis
(d) Opinion, clearly labelled. I think AI visibility for law firms is better understood as a pattern of likelihoods than as a position. Some firms may appear consistently across prompts and models; many will appear intermittently; the interesting commercial question is the shape of that pattern and what moves it.
I'm deliberately not saying AI visibility is probabilistic as a finding. The external evidence on non-determinism and list variability is consistent with that view. Whether it holds for legal recommendations, and how much stability there is underneath the noise, is an empirical question. It's possible that for narrow, well-defined needs in smaller markets, recommendations prove quite stable. That would be a useful finding too.
What this means for law firms
Stop reporting screenshots as rankings. A screenshot is an anecdote. Useful for spotting a problem — an incorrect description, a wrong address, a competitor you didn't expect — not for measuring standing.
Ask any vendor five questions. How many times is each prompt run? Across which systems and model versions? How are mentions distinguished from recommendations? How do you resolve which firm is meant? Can I see the raw answers? If a product reports "your AI rank is 3" without answering these, I'm increasingly sceptical of the number.
Define the prompts that matter to your practice. Write down the real questions your clients ask, by practice area and location, including the ones that don't contain the word "lawyer". That list is more valuable than any single tool.
Look across systems, not just ChatGPT. Your clients don't all use the same assistant. Absence in one system and presence in another is information.
Do the low-regret work. Google states there are no special requirements or optimisations needed to appear in its AI features beyond sound search fundamentals (Google Search Central). Clear, consistent information about who your firm is, what it does and where — on your own site and in the places that describe you — is sensible regardless. I'd treat claims of a secret formula with caution; the evidence behind most of them is thin. (More on that in the forthcoming AEO for Law Firms: Evidence vs Hype.)
Don't react to one bad answer. Check whether it recurs before you change strategy.
Confidence summary
| Statement | Status |
|---|---|
| AI outputs can vary between identical requests | External evidence — provider documentation and academic research |
| Brand recommendation lists rarely repeat exactly | External evidence — one industry study, non-legal categories |
| Generative search citations are often inaccurate | External evidence — academic and journalism research, earlier product versions |
| A single screenshot is weak evidence of standing | Inference from the above |
| Law-firm AI visibility is best measured as frequency, coverage and stability | Dan's hypothesis — being tested through FirmRanker |
The better question
The interesting question isn't "where do we rank in ChatGPT?" It's "across the questions our clients actually ask, and the systems they actually use, how often are we recommended — and is that improving?"
That question is harder to answer. It's also the one worth paying for. The next piece in this series, What Actually Counts as an AI Recommendation?, takes up the hardest part of it: deciding what to count.
For the broader context, see AI & Law, AEO for Law Firms and Research. FirmRanker's work is at firmranker.com. Practice Proof, which I also founded, works with firms on the implementation side.
FirmRanker and Practice Proof were founded by Dan Toombs. Practice Proof provides digital and AI visibility services to law firms.
Sources
- Google Search Central, "AI features and your website". https://developers.google.com/search/docs/appearance/ai-features — accessed 2026-09-28.
- Anthropic, Messages API documentation (temperature parameter). https://platform.claude.com/docs/en/api/messages — accessed 2026-09-28.
- OpenAI Cookbook, "How to make your completions outputs consistent with the new seed parameter". https://developers.openai.com/cookbook/examples/reproducible_outputs_with_the_seed_parameter — accessed 2026-09-28.
- Atil, B. et al. (2024), "Non-Determinism of 'Deterministic' LLM Settings", arXiv:2408.04667. https://arxiv.org/abs/2408.04667 — accessed 2026-09-28.
- Fishkin, R., O'Donnell, P. et al. (28 Jan 2026), "AIs are highly inconsistent when recommending brands or products", SparkToro. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/ — accessed 2026-09-28.
- Aggarwal, P. et al. (2024), "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735. https://arxiv.org/abs/2311.09735 — accessed 2026-09-28.
- Liu, N. F., Zhang, T. & Liang, P. (2023), "Evaluating Verifiability in Generative Search Engines", arXiv:2304.09848. https://arxiv.org/abs/2304.09848 — accessed 2026-09-28.
- Jaźwińska, K. & Chandrasekar, A. (6 Mar 2025), "AI Search Has a Citation Problem", Columbia Journalism Review / Tow Center. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php — accessed 2026-09-28.