How large language models decide which law firms to cite
By Mohammad Kashif, Chief Technology OfficerLast updated
A language model cites what it can retrieve, parse and corroborate. It retrieves from a search index, prefers passages that answer the question in place, and drops claims it cannot confirm elsewhere. Optimising for that is mostly retrieval and clarity, not prompt tricks.
This is why a firm can rank on page one and still never be quoted. Ranking gets a page into the retrieval set. Being quoted requires a passage inside it that stands on its own once lifted out of the page, and most law firm copy does not have one.
The three steps between your page and an answer
Almost every assistant that cites live sources runs the same rough pipeline. Knowing where you fall out of it tells you what to fix.
- Retrieval. The system issues one or more searches and pulls back a candidate set, usually the top handful of results. If you do not rank for the underlying query, nothing else on this list matters yet.
- Passage selection. It reads the candidates and picks the spans that answer the question. Self-contained paragraphs win; paragraphs that depend on the heading above them or the sentence before them usually lose.
- Corroboration and attribution. It checks the claim against other sources and attaches a citation. Specific, checkable facts survive this step. Unsupported superlatives get dropped, which is why "award-winning" never gets quoted and "licensed in Florida since 2009" does.
What this changes about how a page is written
The practical consequence is that a page optimised to be read by a person top to bottom is not automatically optimised to be quoted in fragments. Five habits make the difference, and none of them require new technology.
- Answer in the first forty words, before context, before credentials, before the story. If the answer is in paragraph six, paragraph six is competing against other sites' paragraph one.
- Write self-contained paragraphs. Assume any paragraph might be lifted out alone, and make sure it still makes sense and still names its subject.
- Use specific, checkable facts. Numbers, dates, jurisdictions, statute names. These are what survives corroboration.
- Put the question in the heading, in the words a client would use. Headings are part of how passages get matched to questions.
- Mark up what the page is with structured data, so a model is told rather than left to infer.
What does not work
The category has attracted a quantity of advice that does not survive contact with how these systems function.
- Hidden text addressed to the model. Instructions aimed at an assistant inside page copy are treated as spam by search crawlers and ignored by the models.
- Keyword density. Retrieval is embedding-based; repeating a phrase does not make a passage more likely to be selected, and it makes the passage worse to quote.
- Submitting your site to an assistant. There is no submission endpoint and no index you can join directly.
- Publishing volume without authority. More pages on a domain nothing links to produces more pages nothing retrieves.
Where law firms differ from other businesses
Legal queries carry a higher corroboration burden than most. Assistants are noticeably more conservative about naming a specific professional than about naming a restaurant, because the consequences of a bad recommendation are higher and because bar advertising rules make many claims unverifiable.
In practice that raises the weight of third-party confirmation: bar directory listings, court records, review platforms, and press. A firm whose details agree everywhere a machine can check them is materially easier to cite than one whose own site disagrees with its Business Profile, regardless of which has better copy.
Where firms drop out of the pipeline, and what fixes each stage
| Stage | Why a firm drops out | What fixes it | How fast it shows |
|---|---|---|---|
| Retrieval | The page does not rank for the underlying query | Conventional ranking work: authority, relevance, technical health | Quarters |
| Passage selection | No self-contained paragraph answers the question | Answer-first rewriting, question-shaped headings | Weeks, after recrawl |
| Corroboration | Claims cannot be confirmed anywhere else | Entity consistency, third-party profiles, checkable specifics | Weeks to months |
| Attribution | The model cannot tell what the page is about | LegalService, Person and FAQPage structured data | Days, after recrawl |
Common questions
- What is LLM SEO and is it different from regular SEO?
- LLM SEO describes optimising to be retrieved and quoted by a language model rather than clicked from a ranked list. The overlap with conventional SEO is large, because retrieval still runs through a search index and authority still decides what gets retrieved. The genuine differences are at the passage level: answer-first structure, self-contained paragraphs, and checkable specifics matter far more when your text is being lifted out of context.
- Do I need to rank on page one before an AI will cite my law firm?
- Usually yes for the live-search assistants, because retrieval pulls from the top results of an underlying query. It is not an absolute rule, since different assistants issue different queries and some pull from their own indexes, but a firm that ranks nowhere is rarely retrieved. Ranking is the entry ticket; structure decides whether you are quoted once you are in the set.
- Does adding more blog posts improve AI visibility?
- Only if the domain has enough authority for those posts to be retrieved. Publishing volume on a site with few referring domains produces more pages that nothing retrieves. For a newer site, the same effort spent earning citations from other credible sites moves the needle considerably further than adding posts.
- Can I put instructions for AI models in my website content?
- It does not work and it carries risk. Text addressed to a model inside page copy is ignored by the models and treated as cloaking or hidden text by search crawlers, which is a ranking liability. The supported way to communicate with AI crawlers is robots.txt directives and an llms.txt file, both of which describe your site rather than instructing a model what to say.
More on AI visibility
- How to get your law firm cited by ChatGPT and Google AI OverviewsThese two work differently. Google AI Overviews draws almost entirely from pages already ranking in Google, so conventio
- What structured data markup do law firms and lawyers need for AI search?Four types cover almost every law firm: LegalService for the firm, Person for each lawyer, FAQPage for question blocks,
- Attorney SEO for ChatGPT and PerplexityBoth read the live web and show their sources, so both reward pages that answer directly and can be verified. Perplexity
Search & AI Visibility
Rank in Google + cited in AI answers.
Learn moreExclusive Legal Leads
Pre-screened leads matched to your firm.
Learn moreFind out whether assistants name your firm
We run a fixed prompt set across the major assistants and report which sources they quoted instead of you.
Talk to us