Source-linked AI summary
PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination
Qiyao Wang, Xinyi Chen, Longze Chen, Hongbo Wang, Hamid Alinejad-Rokny, Yuan Lin, Min Yang
TL;DR
Patent examination benchmarks have largely missed the interactive, iterative reasoning between examiners and applicants. PatRe models the full lifecycle through Office Action and rebuttal generation, and experiments show strong task- and ownership-dependent differences while current LLMs remain insufficient as independent examiners.
Problem
Prior benchmarks focus on discriminative classification or static extraction, overlooking multi-turn examination and fine-grained correctness of examination suggestions.
Method
PatRe benchmarks Office Action and applicant rebuttal generation across a multi-turn lifecycle, using oracle or retrieved prior-art settings and legal-semantic evaluation.
Results
Proprietary models generally outperform open-source models, with GPT-5-mini achieving the highest Decision Accuracy in OA-DP (51.4%) and OA-RO (50.0%).
Takeaways & Limitations
PatRe indicates that current LLMs are not sufficient as independent patent-examination systems and that open-source models may suit privacy-sensitive settings.
Abstract
from arXiv · showhide
Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising application volumes. Prior benchmarks predominantly view patent examination as discriminative classification or static extraction, failing to capture its inherently interactive and iterative nature, similar to the peer review and rebuttal process in academic publishing. In this paper, we introduce PatRe, the first benchmark that models the full patent examination lifecycle, including Office Action generation and applicant rebuttal. PatRe comprises 480 real-world cases and supports both oracle and retrieval-simulated evaluation settings. Our benchmark reframes patent examination as a dynamic, multi-turn process of justification and response. Extensive experiments across various LLMs reveal critical insights into model performance, including differences between proprietary and open-source models, as well as task asymmetries between examiner analysis and applicant-side rebuttal. These findings highlight both the potential and current limitations of LLMs in modeling complex, real-world legal reasoning and technical novelty judgment in patent examination. We release our code and dataset to facilitate future research on patent examination modeling.
1. Introduction
Patent examination requires both technical and legal judgment, but prior benchmarks largely model it as static classification or extraction. PatRe instead evaluates the full, interactive lifecycle of Office Actions and applicant rebuttals under realistic evidence conditions.
- Motivation: Patent examination demands judgments about novelty, non-obviousness, usefulness, and statutory compliance while application volumes increase examiner pressure.The paper cites 475,223 USPTO applications, 837,928 unexamined applications, and 20.5 months of first-action pendency in 2025.
- Limitations of Prior Work: Prior AI benchmarks mainly use discriminative classification and lack interpretable, detailed analysis of rejection or grant decisions.Existing tasks include acceptance prediction and fine-grained rejection-reason classification, but remain classification-oriented.
- Motivation: Patent examination is iterative: applicants rebut Office Actions and revise patent versions until a final decision, unlike one-shot review of an initial application.The paper compares this interaction with discussion and rebuttal in academic peer review.
- PatRe: PatRe is the first full-stage benchmark covering Office Action generation and applicant rebuttal across the patent examination lifecycle.It contains 480 recent examination records spanning diverse IPC fields and legal attributes.
- PatRe: Office Action generation is evaluated with oracle citations and retrieval-simulated prior art, while rebuttal generation tests legally and technically persuasive responses.Retrieval-simulated evaluation requires identifying and assessing relevant prior art before drafting the response.
- Experiments: Experiments compare proprietary and open-source LLMs and examine asymmetries between proactive examination and reactive rebuttal.The benchmark is designed to expose differences in legal reasoning and technical novelty judgment across these tasks.
2. Related Work
Earlier patent benchmarks emphasize classification, static legal annotations, or claim-version alignment rather than modeling the examiner-applicant dialogue that drives prosecution. PatRe is positioned against these narrower dataset designs.
- PatRe’s Positioning: PatRe distinguishes itself by combining explicit legal basis, claim evolution, and multi-turn adversarial interaction in one benchmark.The comparison table labels these dimensions as Statute, Evolution, and Adversarial.
- Classification and Static Extraction: Prior benchmarks treat patent examination primarily as post-hoc classification or static justification extraction.Examples include acceptance prediction, modern-LLM classification, and IRAC-aligned board-decision datasets.
- Claim Revision and Drafting: Claim-revision datasets align initial applications with granted versions but omit the explicit examiner-applicant discussion behind those revisions.These resources capture prosecution outcomes without representing the interaction that produces claim changes.
3. PatRe Benchmark
PatRe models examination as repeated examiner-applicant interaction over evolving claims and prior art, then evaluates generated Office Actions and rebuttals with deterministic and semantic-logic measures. Its dataset reconstructs complete prosecution histories from USPTO records with expert quality control.
- Task Taxonomy and Formalization: At each turn, the examiner evaluates current claims C_t against prior art R to produce an Office Action, after which the applicant submits rebuttal B_t and may amend claims to C_t+1.The process is formalized as a multi-turn strategic interaction between Examiner E and Applicant A.
- Office Action Generation: Office Action generation tests examiner decision-making from current claims and prior rebuttals under Direct Prompting, Reference-Oracle, and retrieval-simulated information settings.These settings vary the amount and source of prior-art guidance available to the model.
- Rebuttal Generation: Applicant rebuttal generation produces substantive arguments responding to an Office Action and associated prior art while aligning legal grounds, technical scope, and counterarguments.The task focuses on overcoming objections rather than reproducing procedural filing form.
- Evaluation Metric Design: Evaluation combines deterministic verification of decisions, statutes, and lexical overlap with LLM-as-a-Judge auditing of semantic and logical quality.The framework includes Decision Accuracy, Statute Precision, Rouge-L, and five 1–10 auditing dimensions.
- Evaluation Metric Design: Point-wise Coverage measures how thoroughly rebuttals respond to atomic rejection points in the Office Action.It is introduced specifically as a semantic measure of defense thoroughness.
- Dataset Construction: PatRe reconstructs longitudinal USPTO examination histories, including Office Actions, applicant responses, claim revisions, cited references, and legal metadata.A multi-stage quality-control protocol adds automated filtering, expert audits, and personally identifying information redaction.
- Dataset Statistics: The benchmark contains 480 recent patents across all eight IPC sections and reports distributions of fields, examination rounds, and document lengths.Additional rejection-type, Office Action-type, and cited-reference distributions appear in Appendix A.
4. Experiment
Experiments benchmark diverse proprietary and open-source LLMs using deterministic and LLM-as-a-judge metrics across Office Action and rebuttal generation. Results show task-dependent gaps, strong surface language but weaker legal reasoning, sensitivity to evidence quality, and systematic citation and rejection errors.
- Main Results: GPT-5-mini achieves the highest Decision Accuracy in OA-DP (51.4%) and OA-RO (50.0%), while reaching 52.7% in OA-RS.It also reports 90.5% Point-wise Coverage and an 8.71 Soundness score for rebuttal generation.
- Main Results: Proprietary models outperform open-source models more clearly in rebuttal generation, although the Office Action gap remains relatively narrow for structured decisional logic.The larger rebuttal gap concerns technical precision and global logical alignment in applicant-examiner discourse.
- Analysis: Models score well on Language Style and Clarity but lag in Soundness, Constructiveness, and Completeness, revealing a gap between professional form and legal content.When moving from Office Action generation to rebuttal, Soundness and Constructiveness increase more than twofold, while other dimensions also improve.
- Analysis: Oracle references improve Statute Precision but do not consistently improve Decision Accuracy, while top-tier models in OA-RS can filter noise and preserve decisional stability.External evidence strengthens formal legal alignment without inherently reinforcing logical consistency in patentability determinations.
- Analysis: LLaMA-3.3-70B-it reaches 54.7% Statute Precision but as little as 9.7% Decision Accuracy, reflecting hyper-critical false-rejection bias in allowance cases.The analysis reports that 37.47% of its errors come from prematurely classifying non-final cases as final rejections.
- Analysis: Citation accuracy depends strongly on external evidence, while lexical overlap poorly tracks substantive validity, with Rouge-L and Decision Accuracy showing Kendall’s τ=0.0258 correlation.Models may hallucinate citations when relying only on internal knowledge, and high textual overlap can coexist with legal inconsistency.
5. Conclusion
PatRe reframes patent examination as a dynamic, multi-turn process covering Office Action and rebuttal generation. Experiments show current LLMs are not yet sufficient as independent examiners, while proprietary and open-source models differ notably.
- PatRe is the first full-stage benchmark for patent Office Action and rebuttal generation, evaluating legal reasoning and technical novelty judgment.
- The benchmark models patent examination as a dynamic, multi-turn process of justification and response rather than binary classification or static extraction.
- Current LLMs remain insufficient as independent systems for patent examination.
- Experiments reveal a notable performance gap between proprietary and open-source models, with open-source models more suitable for privacy-sensitive settings.
A. More Details about Data Statics
The dataset consists of recent USPTO examination records characterized by rejection types, Office Action types, and cited-reference statistics. These distributions capture both legal outcomes and the evidence used in novelty assessment.
- The dataset contains patent examination records published after 2024 and collected from the public USPTO data website.
- Rejection categories include §103 Obviousness, §112 Written Description/Enablement, §102 Lack of Novelty, §101 Patent-Eligible Subject Matter, and double patenting.
- Multiple rejection types may appear within a single examination case.
- Office Action records cover Notice of Allowance, Non-Final Rejection, Final Rejection, and Ex Parte Quayle Action.
- Cited-reference distributions are analyzed to reflect novelty requirements in patent examination.
B. More Details about Evaluated Models
The evaluation compares models using reported capacity, context, output, and access characteristics, with separate API-based and locally deployed settings. Costs are also reported for proprietary-model access.
- Evaluated models are described by size, maximum context length, maximum output length, and access method.
- Proprietary models and DeepSeek-V3.2 are evaluated through official APIs for a fair and consistent comparison.
- Open-source models with sizes up to 70B are deployed with vLLM on 8 NVIDIA A800 GPUs using the same hyperparameters.
- The benchmark reports total costs incurred by proprietary models accessed through APIs.
C. More Results
Human evaluation uses trained IP-specialist PhD students to assess generated Office Actions and rebuttals under blind conditions. The reported results include five dimensions and strong inter-rater consistency.
- Three trained PhD students specializing in Intellectual Property evaluate generated Office Actions and rebuttals.
- The blind evaluation samples 100 generated Office Actions and 100 generated rebuttals across 7 models and all IPC sections.
- Human evaluation covers Soundness, Clarity, Constructiveness, Completeness, and Language Style on a 1–10 scale.
- Inter-rater agreement among the three human experts is reported across five dimensions and demonstrates strong consistency.
D. Detailed Case Study
The case study examines generated Office Actions across three task settings and a generated rebuttal, revealing recurring weaknesses in patent examination reasoning. Models often identify relevant material but struggle to rigorously connect claims, prior art, and technical distinctions.
- Case coverage: The study analyzes generated Office Actions in OA-RO, OA-RS, and OA-DP settings, alongside a generated rebuttal case.Examples are presented in Tables 13–16.
- Observed limitations: In OA-RO, models can identify relevant prior art but often fail to perform rigorous claim-to-art mapping.They instead reduce complex technical distinctions to superficial similarities.
- Observed limitations: Across settings, models show consistent limitations in patent examination performance.The supplied analysis highlights recurring problems in reasoning about prior art and technical distinctions.
E. Prompts
The paper supplies prompts for Office Action generation, rebuttal generation, and LLM-as-a-judge evaluation. It also provides example outputs for the three Office Action settings and rebuttal generation.
- Generation prompts: Detailed prompts are provided for Office Action generation across three task settings.These prompts are shown in Figure 8.
- Generation prompts: A detailed prompt is provided for the rebuttal generation task.The rebuttal-generation prompt appears in Figure 9.
- Evaluation prompts: LLM-as-a-judge prompts assess generated Office Actions and rebuttals.The evaluation prompts appear in Figures 10 and 11.
- Examples: Examples cover Office Action generation under OA-RO, OA-RS, and direct-prompting settings, plus a generated rebuttal.The examples are provided in Tables 13–16.