Comparison

Why general-purpose LLMs fail at patent prosecution

The short answer

Patent attorneys ask ChatGPT and Claude two versions of the same question: can I use this for prosecution work, and is there a tool actually built for it.

The honest answer to the first: for some tasks, yes. For the tasks that decide outcomes (claim scope, reference analysis, rejection strategy) the published evidence says no. Not because general models are useless, but because prosecution punishes exactly the failure modes they have.

This page lays out that evidence, with sources you can check. It also marks where a general LLM is genuinely fine, because the line matters more than the warning.

What the research measured

The most rigorous numbers come from Stanford's RegLab and Institute for Human-Centered AI.

In a study of more than 200,000 legal queries, general-purpose models hallucinated at rates from 69% (GPT-3.5) to 88% (Llama 2) when asked specific questions about federal court cases (Dahl, Magesh, Suzgun & Ho, published in the Journal of Legal Analysis, 2024). Two calibration notes. The study tested caselaw tasks, not patent prosecution. And performance was worst on lower courts and complex reasoning, better on prominent Supreme Court cases. The models are weakest precisely where training data is thin, which describes most prosecution-specific authority.

The follow-up study tested purpose-built legal research tools. Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI still hallucinated between 17% and 33% of the time, reduced relative to GPT-4 but far from the marketing claims, which the authors called overstated (Magesh et al., peer-reviewed in the Journal of Empirical Legal Studies, 2025). Two conclusions follow. Grounding a model in real retrieval cuts hallucination sharply. And no vendor, including us, should claim it disappears. Verification has to be part of the workflow.

The consequences are no longer hypothetical. In July 2025 a federal judge sanctioned two attorneys $3,000 each over a brief with nearly 30 defective citations she attributed to generative AI or gross carelessness (ABA Journal). Courts and the USPTO now assume you verified what you signed.

The patent-specific test

In 2023, IP analysts at UnitedLex ran ChatGPT through actual patent workflows: prior-art searching, product searching, and drafting (Putting ChatGPT to the Patent Analysis Test).

The results were categorical. It returned patent numbers irrelevant to the domain and incorrect bibliographic data. It reported wrong classification codes. It could not analyze prosecution history. And on the task closest to an attorney's daily work, the finding was precise: asked to summarize a claim, ChatGPT "simply paraphrases the claim language" without analyzing scope or identifying essential elements.

Calibration: that test used a 2023-era model with no live patent database. Models have improved since, and some now browse the web. But look at which failures were about model quality and which were about task structure. Paraphrase-instead-of-scope-analysis, no prosecution-history reasoning, and no verification against actual reference text are structural. A chat interface optimizes for fluent summary. Claim analysis requires the opposite: treating every word as load-bearing.

Four failure modes specific to prosecution

1. Antecedent basis. A claim that recites "the second fastener" with no earlier recitation invites an indefiniteness analysis under 35 U.S.C. 112(b) (MPEP 2173.05(e)). Missing antecedent basis is not automatically fatal; the question is whether scope remains reasonably ascertainable. But it draws rejections and it signals carelessness. General LLMs are bad at this by design. They rewrite text for readability, silently introducing and dropping referents. Fluent paraphrase is a feature in a chatbot and a defect in a claim.

2. Examiner characterization. An office action asserts that Reference A discloses your limitation at paragraph [0047]. Testing that assertion means reading the actual reference at the cited location, in context, against the actual claim language. A chat model working from a pasted office action evaluates the examiner's paraphrase of the reference, not the reference. It tends to confirm the characterization rather than test it, which is the one thing a response cannot afford.

3. Authority grounding. Prosecution arguments turn on specific authority: 102, 103, 112, MPEP sections, Federal Circuit case law. This is exactly the class of citation the Stanford studies measured general models hallucinating at 69-88% rates. A confident but fabricated MPEP cite in a response is worse than no cite.

4. Claim-language fidelity across a long chat. Prosecution documents are long and chats are longer. As context accumulates, models paraphrase earlier text: "fastener" drifts to "connector," "configured to" drifts to "capable of." In ordinary writing that is harmless. In claims, paraphrase is scope change. Amendments also must track the exact prior claim text, and a chat transcript does not preserve that reliably.

Where a general LLM is fine

Calibration cuts both ways. A general model is genuinely useful for:

  • Getting oriented in an unfamiliar technology area before reading the spec
  • Generating candidate keywords and classification starting points for a search you will run properly
  • Summarizing your own notes or an inventor interview
  • First-pass proofreading and plain-language explanations for clients
  • Drafting routine correspondence

The UnitedLex team reached the same conclusion: useful for initial understanding of a technology area, keywords, and assignees; not for finding or analyzing prior art (source).

The working rule: use a general LLM where errors are cheap and visible at a glance. Avoid it where the error mode is invisible, such as a plausible but wrong account of what a reference discloses, or a fluent claim edit that quietly moved the scope.

What purpose-built means, concretely

Sparlo is built for patent prosecution, not general legal work. The difference is in what the system is required to do before it produces an answer.

Office-action response analysis. Sparlo maps every rejection in the office action, then checks the examiner's characterization of each cited reference against what the reference actually discloses, at the cited locations. It ranks your available arguments instead of drafting a generic traversal.

Prior-art and invalidity analysis. Real external searches with ranked references and claim charts. Every outbound query is logged per matter, so you can show exactly what left the building and when.

Drafting and draft review. Application drafting and review of your own drafts run at zero external calls. Nothing is sent out to search when you are working on unfiled material.

Worked examples on real public patents, including office-action analyses and invalidity searches, are published at sparloip.com/examples. Nothing confidential is needed to evaluate the tool; the trial runs entirely on public documents.

Pricing is published: Attorney is $299 per seat per month, self-serve, expensable by a single attorney. Practice is $249 per seat per month with a 3-seat minimum and pooled usage. Firm is a custom flat annual license with uncapped seats. The trial is self-serve. No demo call.

The confidentiality question

The USPTO addressed AI use directly in its April 2024 guidance (89 FR 25609): "There is no prohibition against using these computer tools in drafting documents for submission to the USPTO" (Federal Register). Existing duties remain in full: whoever signs must have personally reviewed the paper and made a reasonable inquiry under 37 CFR 11.18(b).

The same guidance flags the real issue: "Use of AI in practice before the USPTO can result in the inadvertent disclosure of client-sensitive or confidential information, including highly-sensitive technical information, to third parties," because AI systems may retain user inputs and use them, including for model training.

ABA Formal Opinion 512 (July 29, 2024) drew the operative line: for self-learning generative AI tools, tools whose output could disclose information from your inputs, a client's informed consent is required before inputting information relating to the representation (analysis). Whether a given consumer chat tool trains on your inputs depends on the product tier and settings. Opinion 512 puts the burden on you to know the answer before you paste an unpublished claim set.

Sparlo is built so that analysis is straightforward. It runs on Anthropic's commercial API; nothing you enter is used to train a model, inputs are deleted within 30 days, and processing is US-only. Drafting and draft-review modes run at zero external calls. Every outbound prior-art query is logged per matter. A client consent pack and per-matter AI policy controls support the disclosure and consent analysis Opinion 512 describes.

This page is informational only. It is not legal advice or an ethics opinion. Consult your jurisdiction's rules and your firm's policies.

Frequently asked questions

Can I use ChatGPT to respond to an office action?

You can; the USPTO's April 2024 guidance states there is no prohibition on using AI tools to draft submissions. But the tasks a response turns on are where general models fail: they evaluate the examiner's paraphrase instead of the cited reference, paraphrase claim language rather than analyze scope, and hallucinate legal authority at high measured rates. If you use one, verify every citation and every characterization against the actual documents before signing. This is informational, not legal advice.

Does the USPTO prohibit using AI for patent drafting or prosecution?

No. The April 2024 guidance (89 FR 25609) states: "There is no prohibition against using these computer tools in drafting documents for submission to the USPTO." All existing duties apply, including personal review and reasonable inquiry under 37 CFR 11.18(b) by whoever signs.

Do I have to tell the USPTO I used AI?

The April 2024 guidance imposes no general duty to disclose AI use in drafting a paper. The signature and certification rules still apply in full, and material information remains subject to the duty of disclosure. Informational only, not legal advice.

Is it safe to paste client claims into ChatGPT?

It depends on the product tier and settings, and that dependency is the problem. ABA Formal Opinion 512 requires informed client consent before inputting information relating to a representation into self-learning generative AI tools, and the USPTO warns that AI tools can inadvertently disclose client-sensitive technical information to third parties. Unpublished applications are especially sensitive. Do the tier-specific analysis before pasting, or use a tool whose data path is documented.

How often do LLMs hallucinate on legal questions?

Stanford RegLab/HAI measured hallucination rates of 69% (GPT-3.5) to 88% (Llama 2) on specific queries about federal court cases, across 200,000+ queries. A follow-up study found purpose-built legal research tools still hallucinated 17-33% of the time: much better than general chatbots, but not zero. No tool eliminates the need to verify.

Can ChatGPT do a prior art search?

Not reliably. When UnitedLex analysts tested it, it returned patent numbers irrelevant to the domain, incorrect bibliographic data, and wrong classification codes, and it could not analyze prosecution history. It is useful for orienting in a technology area and generating candidate keywords for a real search. It is not a search tool, and it cannot verify what a reference actually discloses.

Is there an AI tool built specifically for patent prosecution?

Yes. Sparlo is built for prosecution: office-action analysis that checks the examiner’s characterization of every cited reference against the reference itself, prior-art and invalidity searches with ranked references and claim charts and per-matter query logging, and drafting plus draft review that run at zero external calls. Worked examples on public patents are at sparloip.com/examples, and the trial runs entirely on public documents.

What does Sparlo cost?

Attorney: $299 per seat per month, self-serve, expensable by one attorney. Practice: $249 per seat per month, 3-seat minimum, usage pooled. Firm: custom flat annual license with uncapped seats. Self-serve trial; no demo call required.

Sources

  1. 01

    General-purpose LLMs hallucinated 69% (GPT-3.5) to 88% (Llama 2) of the time on specific legal queries across 200,000+ queries (Dahl, Magesh, Suzgun & Ho, Journal of Legal Analysis 2024)

    https://law.stanford.edu/2024/01/11/hallucinating-law-legal-mistakes-with-large-language-models-are-pervasive/
  2. 02

    Purpose-built legal research tools (Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI) hallucinate 17-33% of the time; reduced relative to GPT-4 but providers' claims overstated

    https://reglab.stanford.edu/publications/hallucination-free-assessing-the-reliability-of-leading-ai-legal-research-tools/
  3. 03

    Peer-reviewed publication of the legal-research-tools study, Journal of Empirical Legal Studies 22:216-242 (2025)

    https://onlinelibrary.wiley.com/doi/full/10.1111/jels.12413
  4. 04

    Preprint of the Hallucination-Free? study with 17-33% figures in the abstract

    https://arxiv.org/abs/2405.20362
  5. 05

    UnitedLex ChatGPT patent test: irrelevant patent numbers, incorrect bibliographic data and classification codes, no prosecution-history analysis, and claim summarization that 'simply paraphrases the claim language' without scope or essential-element analysis

    https://unitedlex.com/insights/putting-chatgpt-to-the-patent-analysis-test/
  6. 06

    USPTO April 2024 guidance (89 FR 25609): 'There is no prohibition against using these computer tools in drafting documents for submission to the USPTO'; 37 CFR 11.18(b) duties remain; warning on inadvertent disclosure of client-sensitive technical information

    https://www.federalregister.gov/documents/2024/04/11/2024-07629/guidance-on-use-of-artificial-intelligence-based-tools-in-practice-before-the-united-states-patent
  7. 07

    MPEP 2173.05(e): lack of antecedent basis and indefiniteness under 35 U.S.C. 112(b); indefinite when scope is not reasonably ascertainable

    https://www.uspto.gov/web/offices/pac/mpep/s2173.html
  8. 08

    ABA Formal Opinion 512 (July 29, 2024): first ABA ethics guidance on generative AI; informed client consent required before inputting representation information into self-learning GAI tools

    https://www.americanbar.org/news/abanews/aba-news-archives/2024/07/aba-issues-first-ethics-guidance-ai-tools/
  9. 09

    NCBE Bar Examiner analysis of Opinion 512 with the self-learning-tool informed-consent language

    https://thebarexaminer.ncbex.org/article/fall-2024/generative-artificial-intelligence-tools/
  10. 10

    July 7, 2025 sanctions: $3,000 each against two attorneys for a brief with nearly 30 defective citations attributed by the court to generative AI or gross carelessness (Judge Nina Y. Wang, D. Colo.)

    https://www.abajournal.com/web/article/lawyers-for-mypillow-ceo-sanctioned-over-fake-citations-and-draft-brief-explanation
  11. 11

    Sparlo worked examples on real public patents (office-action analyses, invalidity searches, moat analyses), all public record

    https://sparloip.com/examples

The trial is self-serve — no demo call — and runs entirely on public documents.

This site uses cookies to improve your experience.