All posts

Unknown Unknowns: What AI Legal Research Can’t Read

On July 16, 2026, I fetched the robots.txt file for 34 hostnames and read each one. The sample covers three classes of source a practicing lawyer relies on: the legal trade press (Law360, Bloomberg Law, ALM, JD Supra, and their peers), the general and financial press that transactional lawyers read alongside it (the Journal, the FT, Bloomberg, the private equity trade titles), and the primary sources both report on (congress.gov, govinfo, the Federal Register, SEC EDGAR, court sites, the free case-law databases). A robots.txt file is where a website tells crawlers what they may index, and since 2023 it is where publishers tell AI companies whether their content may be used for model training or AI-assisted search.

The general news media has been audited this way repeatedly: BuzzStream tracks which of the biggest UK and US news sites block AI bots, and Press Gazette has run the same exercise on the top 100 news websites. Nobody, as far as I can tell, had run it on the legal vertical, and the single-vertical version of the study turns out to be the less interesting one. Whether any of this endangers a lawyer’s research depends on a comparison: when the trade press blocks AI crawlers, can a research tool still reach the underlying primary sources? Usually it can, but the exceptions cluster in particular places: pending legislation, case law at scale, and nearly everything the profession pays to read.

Everything below comes from the files as they stood on July 16, 2026, except where I say otherwise. This data moves fast; a file can change tomorrow.

The legal trade press

Publisher AI crawlers named Effect
Law360 Google-Extended, Bytespider, Bytedance Full block of those three; every other AI bot unaddressed
Bloomberg Law (news) GPTBot, ChatGPT-User, ClaudeBot, Claude-Web, anthropic-ai, Google-Extended, CCBot, PerplexityBot, others Full block of named bots; newer search bots unaddressed
LexisNexis (www) Google-Extended, Bytespider, Bytedance, plus a longer list (ClaudeBot, GPTBot, PerplexityBot, CCBot, others) First three blocked site-wide; the longer list blocked from /en-gb/legal/ only
Thomson Reuters (both domains) None No AI rules of any kind
JD Supra GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, CCBot, Google-Extended, PerplexityBot, others Full block of named bots, including OpenAI’s search bot
Above the Law None No AI rules of any kind
Law.com / ALM (incl. americanlawyer.com) Roughly sixty bots, including every current OpenAI, Anthropic, Perplexity, Google, and Mistral agent Full block of all of them, except ALM’s online bookstore, which every blocked bot is allowed to visit
National Law Review None No AI rules of any kind
Reuters Allowlist model: ChatGPT-User and OAI-SearchBot admitted by name Everything not on the allowlist is blocked site-wide, sweeping in GPTBot, ClaudeBot, all Claude and Perplexity agents, CCBot, Google-Extended
ABA Journal None No AI rules; ten-second crawl delay for everyone
CourtListener Sixteen AI agents in three commented categories: training, search, user-initiated All three categories blocked from case data; a handful of informational pages allowed

The table looks chaotic, but three regularities run through it. First, the paywall is the access policy, inverted from the obvious guess: the publishers with the strongest subscription gates do the least robots.txt blocking. Law360 blocks only Google-Extended and ByteDance’s crawlers because its articles sit behind a hard paywall anyway; Thomson Reuters names no AI bot at all because Westlaw, the asset a crawler might want, lives behind authentication where robots.txt has no role. The heaviest blockers, JD Supra and ALM, are the outlets whose full text is free or lightly metered on the open web. The file is what you use when you have no better wall.

Second, the blocklists show their construction dates like tree rings. Bloomberg Law blocks anthropic-ai and Claude-Web, tokens from Anthropic’s earliest documentation that no current crawler sends, and files ClaudeBot, the real one, under a comment header reading “Analytics & Marketing.” JD Supra blocks a bare anthropic. Someone copied a blocklist when AI crawling first made news, and nobody has reconciled it against a live bot roster since; the newer retrieval bots slip through or get caught by accident of timing rather than by policy. Naming a bot can even loosen its leash: under the robots standard a crawler obeys only the most specific group that names it, so a bot on Lexis’s AI list answers to the single UK rule and escapes the general restrictions that bind anonymous crawlers across the rest of the site.

Third, the file understates the walls. Above the Law’s robots.txt has no AI rules at all, yet its CDN blocked my command-line fetches and I had to read the file through a real browser. Enforcement lives in layers this audit, and anyone’s audit, cannot see.

Law.com’s file also doubles as a quiet de-indexing ledger, with dozens of disallow lines for individual articles, mostly stories about attorney discipline, deaths, and settlements, presumably removed from search at someone’s request. Justia’s file has the same feature in miniature, disallowing four specific district-court case pages by citation.

The financial press is more blocked, and blocked differently

Lawyers do not read only legal news. At transactional firms, the daily diet runs heavily to financial and deal coverage: the Journal, the FT, Bloomberg, and for private equity practices the trade titles, PE Hub, Buyouts, Private Equity International, PitchBook. I pulled the same file for each.

Publisher Posture toward AI crawlers
Wall Street Journal Default block of every bot; an allowlist admits search engines plus exactly three AI agents: GPTBot, ChatGPT-User, OAI-SearchBot
Financial Times Blocks ClaudeBot, PerplexityBot, Perplexity-User, CCBot, Google-Extended, Meta’s agents, and more; names no OpenAI bot, so all of OpenAI’s crawl
New York Times Full block of the modern roster: OpenAI, Anthropic, Perplexity, Google, Meta agents alike
Bloomberg.com Full block of the modern roster, current to Claude-SearchBot and Claude-User
CNBC Full block of the modern roster
Axios Blocks dataset scrapers (CCBot, Bytespider, Diffbot, Amazonbot); GPTBot and ClaudeBot unnamed and therefore allowed; Google-Extended explicitly allowed
Fortune Blocks only CCBot and Google-Extended
PitchBook No AI rules; content behind subscription
PE Hub, Buyouts, Private Equity International (PEI Group) Identical template files; no AI rules; hard paywalls
The Deal No AI rules; paywall

The licensing deals show through the files like an X-ray. OpenAI has signed content agreements with News Corp, the Financial Times, and Axios, among many others; the New York Times is suing OpenAI instead (deal and lawsuit reporting per Press Gazette’s tracker, a secondary source). Now reread the table. The Journal blocks every AI bot on earth except OpenAI’s three. The FT blocks Anthropic, Perplexity, Google, and Meta by name and OpenAI not at all. Reuters, which licensed content to Meta in October 2024 per Axios, admits OpenAI’s and Meta’s bots and defaults everyone else out. The NYT, mid-litigation, blocks everything. Robots.txt in the financial press has stopped being a crawling policy and become a ledger of who has paid.

For a lawyer this cashes out as vendor asymmetry. As of July 16, 2026, an assistant built on OpenAI’s bots can retrieve from the Journal, the FT, and Reuters; an assistant built on Anthropic’s or Google’s can retrieve from none of the three. Neither product discloses the difference. It shows up only here, in the newspapers’ robots.txt files.

The private equity trade press did not need robots.txt at all. PEI Group’s three titles run identical eight-line files that never mention an AI bot, because the paywall does everything. Between the paywalled trade titles, the licensing-gated dailies, and Bloomberg’s full block, current private-markets intelligence is close to a dead zone for AI retrieval, for every vendor.

Testing the primary-source safety net

The strongest objection to worrying about any of this: legal trade press is secondary literature. If an AI assistant cannot read Law360’s writeup of an opinion but can read the opinion, the loss is convenience, and the argument that crawler blocking endangers legal research collapses. So I audited the primary sources too.

The safety net holds for enacted law. Govinfo.gov, the Federal Register, and supremecourt.gov have no AI rules; their content is open to any crawler. Justia allows everything (those four de-indexed cases aside), so a large free case-law corpus is reachable. SEC EDGAR goes further than neutrality: sec.gov’s file contains an explicit Allow: /Archives/edgar/data, an affirmative welcome to the filings archive, though the agency rate-limits automated fetching at the network layer (my first request bounced) and publishes fair-access rules instead.

Then there is congress.gov. The Library of Congress blocks roughly 130 named agents from the entire site, Disallow: /, and the list is the most current I have seen anywhere in this audit: GPTBot, ClaudeBot, both companies’ search and user bots, Gemini-Deep-Research, MistralAI-User, and beyond them the agent products themselves, Devin, Operator, NovaAct, Manus, ChatGPT Agent, Google-NotebookLM. The primary source for pending federal legislation is closed to AI tools as a category. The bill texts exist elsewhere, on GPO mirrors and through the official API, so a diligent tool can still get them, but whether your tool does is plumbing you cannot see from the answer.

Case law at scale is patchier than the doctrine-is-open story suggests. CourtListener, the largest free case-law corpus, blocks all sixteen AI agents it names from case data, steering them to its API and bulk downloads. Google Scholar disallows /scholar, its case-law search, to every crawler. PACER sits behind fees and a login. Individual opinions are findable on court sites and Justia, but the aggregate layers where research happens, dockets, citators, full-corpus search, are either dark to crawlers or were never on the open web to begin with.

For doctrine, then, the objection holds. The categories where it fails have something in common: no open primary source exists for them, and they are specific enough to list.

The topics that go dark

Market terms and deal intelligence. What sponsors are paying, how reps-and-warranties packages are trending, what a market rate for a management fee looks like this year: this knowledge lives in the PE trade press, PitchBook, and the deal coverage of the Journal, Bloomberg, and the FT. Every one of those is paywalled, licensing-gated, or blocked. No primary source backstops it, because private deal documents are private; SEC filings expose only the public sliver. A transactional lawyer asking a general-purpose assistant about current market practice is querying the category with the least crawler-reachable coverage in this entire audit.

Settlement values and verdicts. Most settlements are confidential; the ones that surface do so in trade reporting, ALM’s verdict products, Law360’s coverage, none of it crawlable. The judicial opinions that are open rarely state the number.

Attorney discipline and lateral diligence. The de-indexing ledgers cut here. The individual articles Law.com removes from crawlers are disproportionately discipline stories, misconduct reporting, and litigation involving lawyers; Justia removes specific case pages. Each removal is presumably justified on its own facts, but the aggregate effect is that the negative-information layer of the legal press is the most actively suppressed layer of it. An AI-assisted background check on opposing counsel or a lateral candidate draws on a corpus with the discipline stories selectively deleted.

Pending legislation. Congress.gov’s block means the tracking layer, bill status, amendments, cosponsors, is invisible to any agent that honors the file and lacks API plumbing.

Docket-level litigation intelligence. Who is suing whom this month, in front of which judge, represented by which firm: PACER is gated, CourtListener blocks AI agents from case data, and the analytics products built on both are subscription services. The trade press that summarizes them (Law360, Bloomberg Law) is paywalled and partly blocked.

The recent past, generally. Models trained before the blocking wave absorbed these publishers’ archives; the blocks took effect from 2023 onward. So a tool can discuss the 2021 deal market fluently and retrieve nothing about last quarter’s, with no change in the confidence of its prose.

The gaps follow the money

The list is not random, but you have to read the files from the publisher’s side to see why: a crawler posture is a pricing decision, and the three postures in this audit line up with three revenue models.

Ad-supported outlets block nothing, because traffic is the product and a block only shrinks it: Above the Law, the National Law Review, the ABA Journal. General news publishers sit in the middle: any individual story is substitutable, so the rational sale is the corpus, licensed wholesale to AI companies, and robots.txt is the negotiating lever. Block by default, open per deal, which is the WSJ-FT-Reuters ledger from the table above; the New York Times is litigating over the price of the same asset.

The professional intelligence publishers run a third model, and it explains the silence of their files. Law360, Bloomberg Law, PitchBook, and the PEI titles sell non-substitutable content to a small base of buyers who expense it. A wholesale deal with an AI company would cannibalize that base: if a general-purpose assistant can quote current sponsor deal terms, the five-figure subscription dies. So they neither block nor deal on the open web. The paywall already does the price discrimination, and the licensing happens downstream, into the profession’s own tools at the profession’s prices. Thomson Reuters builds CoCounsel on Westlaw content rather than licensing Westlaw to model builders. LexisNexis licensed its case law and Shepard’s into Harvey in June 2025, by the companies’ own announcement. PitchBook sells data to subscribers through APIs. The content reaches AI, but through metered pipes, per seat.

The consequence is that willingness to pay predicts crawler darkness. The topics an AI assistant cannot see are dark precisely because their publishers know law firms, and the third-party tools law firms buy, will pay for access; the openly crawlable web skews toward the information nobody could charge for. The gap persists for a consumer assistant and largely closes, at consumption prices, for a firm that buys the licensed tools. The unknown-unknowns burden falls hardest on whoever queries the general-purpose product and assumes it sees what the licensed one sees.

ROSS is a different question

Any piece on legal publishers and AI has to mention Thomson Reuters v. ROSS Intelligence, and has to keep it separate. ROSS is copyright litigation about training data: Judge Bibas held in February 2025 that ROSS’s use of Westlaw headnotes to train its legal research tool infringed and was no fair use, and the Third Circuit heard argument on June 11, 2026, with no decision as of this writing (status per the docket and reporting, not something I can verify from a file). None of that is a crawler block: ROSS bought its Westlaw content through a contractor after Thomson Reuters refused it a license, and robots.txt never entered the story. The files suggest the industry drew Thomson Reuters’ lesson rather than the news industry’s: control access through authentication and contracts, litigate misappropriation when it happens. The publishers doing the heavy robots.txt work are the ones with nothing behind a login and no deal to point to.

Unknown unknowns, revised

I began this audit expecting to write that AI legal research is dangerous because the legal trade press is invisible to it. The fuller dataset supports a sharper and stranger claim. The danger is not uniform, and it is not where the volume of blocking is. It is topic-shaped: doctrine is reachable while market intelligence, settlement data, discipline records, and pending legislation go dark, and the topic shape has a price logic, since the categories that go dark are the ones whose publishers can charge the profession for them. It is vendor-shaped: the same question put to assistants from two AI companies runs against different newspapers, under licensing deals disclosed in no product interface. And it is time-shaped: the training corpus remembers the open web of 2022, so the tool sounds equally informed about the periods it can and cannot see.

None of this appears anywhere a user can inspect. Westlaw and Lexis publish their coverage; a librarian can tell you what a database includes and when it starts. An AI assistant’s effective coverage is the residue of 34-plus robots.txt files, a handful of licensing deals, paywalls, and CDN rules, all changing without notice. The tool cannot flag the gap because the blocking happens upstream, before anything reaches it: a blocked source is missing from the answer the way a case outside a database’s coverage is missing from a search, silently. I have argued that AI research outputs are leads to verify rather than finished work; verification catches what the tool got wrong. Only knowing where the tool cannot look catches what it never saw.

The practical version, by topic. For doctrine, AI-assisted research fails soft; the primary sources are open, and verification against them works. For market terms, deal intelligence, settlement values, and anything else whose only home is the paywalled trade press, it fails hard and silently, and the subscription products, or the AI tools that license them, remain the only honest option. For diligence on people, assume the crawlable record has been curated. And date-stamp any claim about who blocks what, including every claim in this post, because the files change without notice and nobody announces the diff.


Method: I fetched robots.txt directly on July 16, 2026 from the legal set (law360.com, news.bloomberglaw.com, lexisnexis.com and www, thomsonreuters.com, legal.thomsonreuters.com, jdsupra.com, abovethelaw.com, law.com, americanlawyer.com, natlawreview.com, reuters.com, abajournal.com, courtlistener.com), the financial set (wsj.com, ft.com, bloomberg.com, nytimes.com, cnbc.com, axios.com, fortune.com, pitchbook.com, pehub.com, buyoutsinsider.com, privateequityinternational.com, thedeal.com), and the primary-source set (congress.gov, govinfo.gov, federalregister.gov, supremecourt.gov, sec.gov, pacer.uscourts.gov, law.justia.com, scholar.google.com), reading each file in full; americanlawyer.com redirects into law.com, whose file governs it, and abovethelaw.com and sec.gov had to be read through a browser because both block command-line fetching. Claims resting on secondary sources rather than the files: the OpenAI, Meta, and NYT deal-and-litigation postures (Press Gazette’s tracker, Axios) the appellate status of Thomson Reuters v. ROSS Intelligence (docket and reporting), and the downstream-licensing examples, which rest on the LexisNexis–Harvey announcement and the companies’ own product descriptions. Paywall characterizations reflect each site’s public subscription gates. Robots.txt is a request rather than an enforcement mechanism; compliance is voluntary, and access can also flow through licensing APIs this audit cannot see. Prior audits of the general news media: BuzzStream and Press Gazette. This post builds on earlier posts on reading the limitations section and the verification standard.