Cornerstone Explainer

What Actually Counts as an AI Recommendation?

Not every appearance in an AI answer is equal. Here's a working framework for telling a mention from a recommendation, and why the difference matters to a law firm's pipeline.

Published 9 min read
In this article 11 sectionsBack to top ↑

Someone will tell your firm, if they haven't already, that it "appeared in ChatGPT". It might be an agency report, a vendor dashboard or a partner who tried it on a Sunday night.

My view is that the sentence, as usually used, tells you almost nothing. It doesn't say what the system did with your firm. It might have been named in passing, put on a shortlist, recommended outright, relegated to a footnote link, or mentioned next to a warning. Those are commercially different events. Counting them all as "visibility" is like counting a newspaper correction and a front-page profile as the same kind of coverage.

This piece sets out the taxonomy I use to keep those events apart. It's a working framework, not a validated industry standard. I'm publishing it because it will underpin how I discuss AI discovery research here and through FirmRanker, and because I think firms need a sharper vocabulary before they spend money on this.

The short answer

An AI system recommends a law firm only when the response itself endorses that firm: it says, in substance, "this one", or vouches for the firm's suitability. Being listed, being cited as a source or being mentioned isn't a recommendation, even when the user asked for one.

Five ways a firm can appear in an AI answer

Category Definition Minimum test Illustrative (hypothetical) example
Mentioned The firm is named in the substantive answer, but not presented as an option for the user to act on. The name appears in the answer body. "Firm A was involved in a well-known class action on this issue in 2019."
Suggested / consideration set The firm is placed in a set of options the user is invited to consider or contact. The answer frames the firm as a candidate for the user's need. "Firms you could consider include Firm A, Firm B and Firm C."
Recommended The response explicitly endorses the firm, singles it out or vouches for its suitability. Endorsing or preferential language directed at this firm, not the whole list. "For your situation, Firm B is likely the strongest choice because of its specialist team."
Source-only The firm's domain or content appears in citations or supporting links, but the firm isn't substantively put forward as an option. Present in citations or links, absent (or merely mentioned) in the answer body. The answer explains limitation periods and cites an article on Firm C's website.
Negative / cautionary The firm appears in a materially adverse or cautionary context. The answer's treatment would reasonably deter a prospective client. "Some reviewers have reported communication problems with Firm D."

All firm names above are invented. The examples illustrate the categories and aren't drawn from any observed answer.

These categories aren't strictly exclusive. A firm can be recommended and cited, or suggested and flagged with a caution. In practice I code the treatment in the answer body as the primary category and record citation presence and adverse context as separate flags. Otherwise, a recommended firm that also happens to be cited gets counted twice, or counted as the wrong thing.

Prompt intent is not response treatment

This is the distinction I think the industry blurs most often.

Suppose a user asks: "Which personal injury firm in Brisbane would you recommend?" That's a direct-recommendation prompt. Now suppose the system replies:

"I can't recommend a specific firm, but some personal injury firms in Brisbane include Firm A, Firm B and Firm C. Check their experience, fee arrangements and reviews before deciding." (hypothetical)

A lot of reporting would count Firm A as "recommended by ChatGPT" because the question asked for a recommendation. I think that's wrong. The system explicitly declined to recommend and produced a consideration set. The correct code is Suggested for all three firms.

So the rule I apply is: classify the response, not the prompt. Prompt intent is still worth recording, because it's a dimension of the research design. A system that recommends when asked "who's best?" but only lists firms when asked "who does family law near me?" is telling you something. Intent and treatment belong in separate columns, though.

Decision rules for the hard cases

Hedging and disclaimers. AI answers about legal services often come wrapped in caveats: "this isn't legal advice", "consult a qualified lawyer", "do your own research". My rule is that a generic disclaimer doesn't downgrade a firm-specific endorsement, and it doesn't upgrade a list. If the answer says "Firm B is a strong choice for your matter (but this isn't legal advice)", I still code it Recommended. If it says "here are some firms, I can't recommend one", I code it Suggested.

Ordering and position. Being first on a list isn't, by itself, a recommendation. I record position, but I don't treat it as endorsement unless the text says so ("the top choice is…"). There's a practical reason as well as a principled one. SparkToro's January 2026 study ran 12 prompts nearly 3,000 times across ChatGPT, Claude and Google's AI. The authors report less than a 1-in-100 chance of the same brand list appearing twice, and roughly 1-in-1,000 for the same order. They also found the pool of frequently mentioned brands relatively stable. That study isn't legal-specific, and the authors note it isn't academically peer-reviewed. I don't know yet whether legal queries behave the same way. It is a good reason not to build a metric on list position.

Multiple firms. If the answer endorses a set ("any of these three firms would be a good choice"), I code each firm in that set as Recommended but flag it as a shared recommendation. Being one of three endorsed firms is different from being the only one.

Qualified endorsement. "Firm A is well regarded, though it may be expensive" is still a recommendation, with a flag attached. A caution only becomes Negative / cautionary when it would reasonably deter a prospective client: reported misconduct, serious complaints, or an explicit "you may want to avoid".

Conditional recommendations. "If your matter involves a trust, Firm C would be worth talking to" is a recommendation conditioned on facts. Whether the user meets the condition matters. I code it Recommended and record the condition.

Entity resolution. The same firm can appear as "Smith & Co", "Smith and Co Lawyers", a former name or a named partner. If those aren't resolved to one entity, one firm can look like three weak appearances, or a competitor's appearance can be credited to you. It's unglamorous work, and it changes results.

Why "source-only" deserves its own category

It's tempting to count a citation to your website as a win. Sometimes it is: your content informed the answer. But being cited as a source for an explanation of limitation periods isn't the same as being put forward as the firm to hire.

There are three things people tend to conflate:

  1. Sources retrieved: what the system fetched.
  2. Sources cited: what the answer displays or links.
  3. Sources that causally shaped the answer: what actually influenced the text.

These are different sets. Anthropic's web search tool documentation, for example, returns search results and in-text citations as separate structures in the API response. What was retrieved and what was cited are recorded separately. Google says its AI Overviews and AI Mode use "query fan-out", issuing multiple related searches across subtopics to develop a response, so the pages an answer draws on can extend well beyond the classic results.

Nor does a citation reliably mean the source supports the claim next to it. In a human evaluation of four generative search engines, Liu, Zhang and Liang (Stanford, Findings of EMNLP 2023) found that only 51.5% of generated sentences were fully supported by their citations, and 74.5% of citations supported their associated sentence. Those were 2023 systems and today's products may perform differently. The broader point still stands: a link isn't proof of influence, and it certainly isn't an endorsement.

Lawyers have thought carefully about what a "recommendation" is, for different reasons. Comment [2] to Rule 7.2 of the North Carolina Rules of Professional Conduct, which tracks the ABA Model Rule language, says a communication contains a recommendation if it "endorses or vouches for" a lawyer's credentials, abilities, competence, character or other professional qualities. It adds that directory listings by practice area, "without more", aren't recommendations.

I'm not suggesting that rule applies to AI outputs. I'm not aware of any authority that says it does, and other jurisdictions (England and Wales, the Australian states, Canadian provinces) have their own frameworks. This isn't legal advice. But as a definitional test, it's close to the line I draw between Suggested (a listing "without more") and Recommended (endorsing or vouching).

Why this matters commercially

My view, which is opinion and not measured evidence, is that the categories likely differ in commercial value:

  • Recommended is the closest analogue to a referral. It's plausibly the most valuable treatment, and plausibly the rarest.
  • Suggested may matter more than people assume. Many clients choose from shortlists, and a consideration set is where many enquiries begin.
  • Mentioned is brand presence, not a lead.
  • Source-only is evidence that your content is being used. It might build authority over time, but it doesn't put you in front of the client as a choice.
  • Negative / cautionary is a reputational risk that aggregate "visibility" scores can hide completely. A firm can appear constantly and be hurt by it.

If a report adds these together into a single visibility number, it can go up while the firm's commercial position gets worse. That's the core methodological problem.

Why "we appeared in ChatGPT" is ambiguous

When you hear the claim, it can hide at least six unknowns:

  1. Which category? Mentioned, suggested, recommended, source-only or cautionary.
  2. Which prompt, and what intent? A branded query ("tell me about Firm A") guarantees a mention and proves little.
  3. How many runs? One AI answer isn't a ranking. Given the variability documented above, a single screenshot is an anecdote.
  4. Which provider and surface? ChatGPT, Claude, Gemini, Perplexity and Google's AI features aren't automatically equivalent, and neither are their app, API and search modes.
  5. Which market? Location, jurisdiction and practice area change the question.
  6. When? Models and retrieval change. An observation has a date.

These are the principles I've found I can't avoid while building FirmRanker. Preserve the raw answers. Separate treatment from citation. Resolve entities before counting. Repeat observations before drawing conclusions. Don't assume one provider speaks for the others. I'm describing method here, not results. Whether these categories behave consistently across providers, practice areas and markets is exactly what we're testing through FirmRanker, and we don't know yet. You can read more about the approach on the methodology page.

Questions to ask whoever reports your AI visibility

  • Which of these categories does your number count? Can I see them separately?
  • Are results based on repeated runs? How many, over what period?
  • Which providers and surfaces, and are they reported separately?
  • Are the prompts branded or unbranded, and what intent do they represent?
  • Can I see the raw answers behind the summary?
  • How do you handle firm name variants?
  • Do you flag negative or cautionary appearances?

If the answers are vague, treat the number as vague.

What we don't know yet

  • Whether the recommendation rate in legal queries differs materially from other categories.
  • How often AI answers about legal services decline to recommend, and whether that varies by provider or jurisdiction.
  • Whether a source-only appearance increases the chance of later recommendation.
  • How much the categories differ in real enquiry value. That's a commercial question no AI output can answer by itself.

Where this leaves firms

The useful shift, I think, is from asking "are we in AI?" to asking "how are we treated in AI answers, how consistently, and where?" That's a harder question. It's also the one that connects to revenue.

I'll use these five categories consistently in future research notes on this site. If you think the lines are drawn in the wrong place, I'd genuinely like to hear it. A shared vocabulary only works if it survives argument.

For the wider context, see AI and law, AEO for law firms, and my earlier pieces on why "ranking in ChatGPT" may be the wrong mental model and AEO for law firms: evidence vs hype. Firms that want help acting on this can talk to Practice Proof. The research itself lives at FirmRanker and on /research/.

Disclosure: FirmRanker was founded by Dan Toombs. Practice Proof, also founded by Dan, provides digital and AI visibility services to law firms.

Sources

  1. Fishkin, R. (with O'Donnell, P., et al.), "NEW Research: AIs are highly inconsistent when recommending brands or products…", SparkToro, 28 January 2026. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/ (accessed 2026-09-28)
  2. Liu, N. F., Zhang, T., & Liang, P., "Evaluating Verifiability in Generative Search Engines", Findings of EMNLP 2023 (arXiv:2304.09848). https://arxiv.org/abs/2304.09848 (accessed 2026-09-28)
  3. Google Search Central, "AI features and your website" (last updated 2025-12-10). https://developers.google.com/search/docs/appearance/ai-features (accessed 2026-09-28)
  4. Anthropic, "Web search tool", Claude Platform documentation. https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool (accessed 2026-09-28)
  5. North Carolina State Bar, Rules of Professional Conduct, Rule 7.2 and Comment [2]. https://www.ncbar.gov/for-lawyers/ethics-and-governing-rules/rules-of-professional-conduct/71-76-information-about-legal-services/72-communications-concerning-a-lawyers-services-specific-rules/ (accessed 2026-09-28)

Final article word count: approximately 2,150 words (body, excluding Sources).


AI recommendationsrecommendation vs mentionAI citationsAI visibilitymethodologylegal marketingAEO for law firmsFirmRanker