Source-linked AI summary

GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

Linh Le, Melanie Bui, My Chiffon Nguyen, Zachary Schlosser, David Williams-King

arXiv:2609.03553v1cs.AIcs.CY

TL;DR

LLM policy simulations need validation against observed actors, actions, and impacts rather than plausibility alone. GPS-Bench grounds policy states and personas in dated public records, then compares inference modes under a common schema. Weight-level fine-tuning gives the strongest actor-level impact prediction, while decomposition contributes an interpretable mechanism for coalition formation.

  • Problem

    LLM policy simulations can generate plausible behaviour without reproducing real actors, actions, or outcomes, leaving their validity difficult to establish.

  • Method

    GPS-Bench reconstructs dated policy states and actors from public evidence, scores passage, affected actors, and impacts, and compares inference modes over the same graph and schema.

  • Results

    Weight-level fine-tuning is strongest for actor-level impact prediction, reaching macro-F1 0.716 on the Gold-and-Silver union versus 0.604 untrained on the human-curated pool.

  • Takeaways & Limitations

    GPS-Bench provides a common empirical setting for testing when evidence, actor modelling, and multi-agent interaction improve policy-outcome prediction and interpretation.

  • Takeaways & Limitations

    The corpus is English-heavy with substantial U.S. coverage, and public records represent visible behaviour better than private negotiation.

Abstract

from arXiv · show

Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns "does multi-agent simulation help?" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.

Introduction

GPS-Bench addresses the validity gap in LLM policy simulation by grounding actors and policy states in dated records and evaluating the full actor-to-impact chain. It provides a controlled comparison of inference strategies over the same evidence-grounded state.

  • Motivation: LLM policy simulations can produce plausible interactions without reproducing real actors, actions, or outcomes.Existing evaluations often rely on qualitative plausibility, while roleplayed agents may converge, mistrack groups, or exhibit political bias.
  • Problem and approach: Governance simulation is reformulated as evidence-grounded structured prediction over actors, actions, and impacts.Actors are reconstructed from records, conditioned on the dated situation, and scored against predictor-independent targets.
  • Benchmark design: GPS separates state construction from inference, reconstructing a pre-decision graph and predicting legislative outcomes, affected actors, actions, and impacts.The observed graph contains the policy, antecedent events, and policy-specific actor states; the predicted continuation covers downstream components.
  • Evaluation targets: The benchmark prioritizes three linked targets: legislative passage, affected-actor identification, and actor-level impact direction.Mechanisms, world-state changes, and multi-step edges are emitted and structured but not scored because retrospective labels are unavailable.
  • Inference comparison: Because all inference modes use the same graph and output schema, GPS enables controlled comparison of joint, independent, collaborative, and communicating agents.The study also compares training-free inference with weight-level adaptation over identical evidence.

Related Work and Dataset Gap

Prior work either studies emergent interaction judged by plausibility, predicts isolated outcomes, or generates scenarios without testing correspondence to historical policy records. GPS-Bench fills this gap by reconstructing bill-specific actors from dated public evidence and scoring downstream consequences independently.

  • LLM-based simulation: LLM social and multi-agent simulations study political interaction, cooperation, escalation, law-making, and policy scenarios but do not test historical correspondence.Their emphasis is emergent interaction rather than whether simulated actors and consequences match a policy record.
  • Validity concerns: Roleplayed agents can over-converge, mistrack represented populations, and exhibit documented political bias.A review found that 15 of 35 generative agent-based studies relied on plausibility alone.
  • Adjacent prediction settings: Outcome-prediction work forecasts legislative labels or votes, treating actors as features and stopping at enactment.Scenario-generation work reconstructs consequences but evaluates perceived impact rather than historical outcomes.
  • Dataset gap: GPS-Bench reconstructs the bill-specific cast from public records, conditions on the dated situation, and scores passage and consequences against independently built targets.Its inputs include provenance-graded actors and antecedents, while scoring uses official status and implementation records.

Problem Formulation

GPS represents policy simulation as predicting the continuation of a partially observed typed temporal graph. Its scored spine runs from policy to affected actor to impact, while actions, mechanisms, and world-state changes remain inspectable.

  • Graph formulation: GPS models a policy episode as a partially observed typed temporal graph with provisions, dated antecedents, and actor states.The observed state G≤t(b) contains information available by the prediction time.
  • Graph formulation: The inference mechanism predicts a continuation containing passage, relevant actors, actions, mechanisms, world-state changes, and impacts.Edges encode direction, probability, lag, controlling actor, and provenance.
  • Temporal boundary: The temporal boundary prevents post-decision events and consequences from entering the held-out policy’s input.Pre-t information is input; later events are prediction targets or evaluation evidence.
  • Scored targets: The benchmark scores passage, actor relevance, and impact direction along the policy →affected actor →impact spine.Passage uses jurisdiction-stratified balanced accuracy, while actor relevance and impact direction use macro-F1.
  • Scored targets: Actor relevance and impact direction form the benchmark’s core, while actions, mechanisms, and world-state changes are emitted and auditable but not scored.This separates the scored actor layer from inspectable trajectory components.

The GPS-Bench Benchmark

GPS-Bench builds a provenance-tracked temporal benchmark from curated public records, with strict temporal splits and distinct supervision tiers. It reconstructs policy-specific actors and standardizes actions and impacts for retrospective evaluation.

  • Construction: GPS constructs each policy graph through public-record data collection and customization, requiring inspectable inputs and predictor-independent evaluation targets.Targets range from observed outcomes to human-verified and Silver relevance labels.
  • Corpus: Curators verify official sources and metadata before inclusion; automation supports retrieval and validation but does not determine admission.The corpus contains 1,233 instruments with full text.
  • Evaluation design: The benchmark uses a strict temporal split, training on policies before 2024 and testing on policies from 2024 onward.The purpose-built sets are distinct rather than nested.
  • Supervision: The primary Gold pool contains 679 human-annotated cases, supplemented by 4,235 evidence-grounded Silver records.Silver labels are produced by a separate annotation LLM from retrieved trusted evidence.
  • Actor representation: Actors are selected from 34 stakeholder roles, and responses are coded using a controlled 15-way action vocabulary.Per-bill casts are chosen by relevance against actors named in the record, with recall treated as the meaningful criterion.
  • Actor and impact representation: Record-grounded personas encode dated behavioural memories rather than stipulated objectives, while impacts are organized into typed mechanisms and world perspectives.The corpus contains 402 typed impact nodes and 214 graded causal maps across 1,233 instruments.

Governance Simulation Methods

GPS separates evidence-grounded policy-state construction from inference, allowing joint and actor-decomposed methods to be compared over identical inputs and outputs. Actor agents use dated, provenance-carrying states and interact through independent, collaborative, or multi-round communication regimes.

  • Inference design: All inference modes receive the same evidence-grounded graph and emit the same schema, isolating how each LLM call’s context affects predictions.Joint inference sees the full bill and cast; independent inference partitions the graph by actor; collaborative inference adds exchanged messages before revision.
  • Graph representation: GPS represents mechanisms and world-state changes separately, while scoring the welfare direction of each affected actor’s impact.Mechanisms describe how a bill acts; world dimensions describe what it changes. Both are structured in the predicted graph but are not scored.
  • Actor states: Each actor agent is initialized from dated public-record evidence covering roles, positions, prior actions, resources, relationships, and relevant indicators.The state is restricted to information available before the simulated decision point, and each field carries source provenance.
  • Historical conditioning: Agents are conditioned on labelled pre-2024 examples through few-shot prompts or weight-level LoRA, without exposing the ≥2024 test period.The few-shot setup approximates hosted-model fine-tuning, while LoRA provides genuine weight-level adaptation for an open 7B model.
  • Interaction regimes: Independent, collaborative, and communication regimes differ in whether agents work alone, revise after one shared exchange, or target partners across several rounds.Communication includes partner selection, messages, updates, and a final synthesis pass over the transcript.

Experimental Setup

The evaluation uses a strict temporal split and prevents models from seeing future test instruments or their evaluation labels. Forecasts are assessed on held-out human or observed Gold outcomes, while Silver supervision is restricted to training.

  • Temporal evaluation: All forecasting uses train <2024 and test ≥2024, ensuring test instruments are not seen during training.The primary predictor is separated from the test window by its model cutoff and is not fitted.
  • Label integrity: Headline evaluation uses held-out Human Gold, while Evidence-Silver is produced by a separate annotator and never serves as a predictor’s test label.Models receive outcome-scrubbed context, preventing evaluation against labels they generated themselves.

Results and Analysis

Results distinguish prediction from mechanism: grounded supervision is strongest for actor-level impact, joint context helps global actor selection, and interaction exposes coalitions and pathways. Communication can recover some context lost through decomposition, but coalition patterns depend partly on instruction and pairwise representations can overstate institutional arrangements.

  • Evaluation validity: Every test sample is held-out Gold, while Silver enters training only, so supervision measures transfer rather than agreement with its own labels.The evaluation varies supervision and inference settings across joint, actor-agent, graph-based, and weight-level methods.
  • Grounding: Documented precedent predicts better than narrative persona detail when actor history exists.Conditioning on dated prior actions helps where history is available, whereas rich narrative personas alone are less reliable.
  • Legislative passage: 0.80 balanced accuracy is achieved by a calibrated joint forecast for passage, rising to 0.89 with dated antecedents; a jurisdiction-only baseline reaches 0.86.The multi-agent score of 0.96 is retained only as a contamination-bounded upper bound.
  • Actor relevance: Relevance is a global selection problem: joint inference benefits from the full cast, while actor-decomposed environments lose cross-actor context; grounded LoRA reaches relevance F1 0.62.The benchmark therefore separates actor identification from actor-level impact rather than collapsing them into one score.
  • Actor impact: 0.77 macro-F1 is achieved by training-free joint impact inference, versus 0.65 for independent agents, 0.64 after one collaborative revision, and 0.70 after three communication rounds.Communication recovers part of the decomposition gap by restoring shared context.
  • Weight-level grounding: 0.716 impact-direction macro-F1 is reached by Gold+Silver LoRA on the human-curated pool, above 0.604 untrained, 0.673 Gold, and 0.630 Silver.On the separate role-level surface, the same model reaches 0.85 when trained on Gold and Silver together.
  • Multi-agent mechanisms: Interaction adds mechanistic resolution: targeted communication yields homophily 0.77 and exposes actor-to-actor responses, coalitions, and pathways that joint forecasts do not represent explicitly.Adversarial debate preserves impact accuracy better than targeted or broadcast communication, while additional rounds provide no monotonic gain.
  • Coalition formation: Fourteen of 74 proposals are reciprocated, forming 7 mutually agreed coalitions across 15 instruments; NVIDIA and Microsoft independently name each other for the NAIRR pilot.The record confirms their committed contributions at $30M and $20M, respectively.

Limitations

GPS-Bench is a first version with bounded geographic, observational, and evaluation scope. Its evidence primarily captures publicly visible behaviour, while mechanisms and world-state edges remain unscored.

  • The English-heavy, substantially US corpus makes legislative passage partly jurisdiction-confounded, requiring jurisdiction-aware baselines.
  • Public-record actor histories capture visible behaviour better than private negotiation, so the actor layer models what outside analysts could observe.
  • GPS-Bench scores passage, actor relevance, and impact direction, while emitted mechanisms and world-state edges await prospective evaluation.

Conclusion

GPS-Bench frames governance simulation as evidence-grounded prediction from policy to affected actors and actor-level impacts. Its corpus and label design emphasize inspectable provenance, while the benchmark’s current evaluation scope remains limited.

  • Conclusion: GPS-Bench turns governance simulation into evidence-grounded prediction of policy trajectories from affected actors to actor-level impacts.
  • Conclusion: The precursor system and appendix materials document corpus construction, experimental setup, evidence sources, results, figures, and prompts.
  • Conclusion: The corpus contains 1,233 instruments, with authoritative source text, attached renderings, provenance records, and independently published analyses.
  • Conclusion: Silver labels expand supervision but do not serve as headline evaluation targets, which use observed or human-validated labels.
  • Conclusion: The benchmark’s evaluation is organized around what kinds of ground truth claims can have and is asymmetric between its questions.

C.1 Datasets and splits

The forecasting experiments use temporally separated data, outcome-scrubbed inputs, multiple predictor families, and controlled interaction modes. Evaluation uses task-specific metrics and examines how communication depth affects impact prediction.

  • Datasets and splits: Training uses measures dated before 2024 and testing uses measures dated 2024 or later, rather than a random split.
  • Datasets and splits: Only three post-cutoff resolved negatives support contamination-controlled passage evaluation, so the primary results exclude a fully controlled passage arm.
  • Evaluation: Passage uses jurisdiction-stratified balanced accuracy, while actor relevance and impact direction use macro-F1 and action labels use micro/macro/family-F1.
  • Predictors: The study compares temporally separated and robustness models, including Qwen2.5-7B-Instruct LoRA and partially or fully leaky upper-bound models.
  • Inputs and modes: Inputs withhold outcome tells, expert analysis, gold labels, and evidence quotes, while prompts compare zero-shot, independent, collaborative, and few-shot modes.
  • Interaction analysis: Additional communication raises interaction without corresponding predictive improvement: DeepSeek balanced accuracy rises 0.59 →0.65, while macro-F1 peaks near 0.62 before declining to 0.61.
  • Interaction analysis: Across 20 rounds, no communication protocol exceeds the independent baseline of 0.735; convince/ally declines by 0.09 by round ten.

H.2 Analysis: mechanisms behind the results

The analysis explains how grounding and interaction shape GPS-Bench predictions and coalition structure. Actor decomposition exposes mechanisms and joint commitments, while grounded behavioural evidence stabilizes prediction and aggregate context remains important.

  • Coalition formation: 77 coalitions formed over six rounds, but none paired a U.S. actor with a Chinese one.The resulting commitments split into Western secure-supply and Chinese self-reliance build-outs.
  • Effect of actor grounding: Behaviour-conditioned actor representations yield consistently stronger impact-direction performance across models and inference modes, reaching up to 0.74.Rich interest profiles alone improve GPT-3.5 and Qwen but reduce DeepSeek, whereas observed behavioural conditioning removes much of this instability.
  • Context aggregation: Joint context improves impact prediction from 0.65 to 0.77 relative to self-focused actor decomposition, while communication recovers only part of the lost shared context.Impact direction is relative to the cast, so splitting actors can miss who benefits when a rule harms another group.
  • Context aggregation: Self-interested relevance judgments collapse precision from 0.39 to 0.25 by over-claiming affected actors against a true rate of approximately 5 of 34.The joint call calibrates affected-actor identification across the full roster.
  • Coalition formation: Forced rival-partner bargaining produced 37 of 72 cross-bloc coalitions, converging on managed competition without trading the advanced frontier.The coalitions proposed safe-harbor and legacy-node arrangements, retaliation truces, and compute-for-safety exchanges.
  • World-state roll-up: The world-state roll-up converts actor impacts into signed totals across ten policy axes, revealing effects that a single bill-level label would flatten.For the Chips and Science Act, the vector is economic/fiscal −3, labour +2, trade +1, governance +1, geopolitical −1.

K Worked example: one bill through each inference mode

The worked example runs one export-control bill through independent, collaborative, and joint inference using the same grounded graph. Independent and collaborative modes expose actor-level interaction, while joint inference pools the roster for the strongest aggregate prediction.

  • Step 0: Input: All three inference modes receive the same outcome-scrubbed bill context and evidence-grounded graph, differing only in what each LLM call sees.The required output gives every actor an action and welfare direction.
  • Step 1: Independent mode: Independent mode queries each actor alone and returns its response and self-impact in one pass.The example produces distinct positions for U.S. political actors, commercial actors, and China’s government.
  • Step 2: Collaborative mode: Collaborative convince/ally inference exchanged 270 pitches, but 0 of 270 changed a stance, leaving every final response and impact direction unchanged.The bill-specific interaction clustered within blocs, while targeted communication eroded macro-F1 relative to independent inference in aggregate.
  • Step 3: Joint mode: Joint mode presents the whole bill and actor roster in one call, calibrating relative welfare signs and achieving the strongest aggregate score of the three modes.Decomposed modes sacrifice shared context for an inspectable interaction record.
  • What each setting buys the reviewer: The three settings serve different review questions: per-actor prediction, interaction structure, or the most accurate aggregate world state.The paper presents them as complementary readings of one grounded graph rather than competing architectures.
  • Prompt variants: Independent relevance and impact prompts ask each actor whether it is materially affected and whether the bill benefits or harms it.The collaborative variant adds other actors’ first-read answers before revision.
  • Prompt variants: The leak-free title-only passage prompt predicts adoption from jurisdiction and bill title while explicitly excluding the constructed summary.The title-only form is reported per open model in Table A6.

M Data-to-Agent Validation Matrix

GPS-Bench validates governance-policy reasoning against public records and independent labels, while separating robust pathway identification from more contested fine-grained repair diagnosis. Its pilot shows cross-domain transfer and useful diagnostic checks, but several outputs remain exploratory or bounded by observable-record and jurisdiction limits.

  • Grounding and validation: GPS-Bench treats actors as falsifiable predictive models grounded in public records, with independent validation rather than stipulated role-play.The benchmark records each component’s data source, prediction, and validator; pilot rows are scored on real data, while target rows are released for held-out evaluation.
  • Diagnostic layers: The closure model is a conservative bottleneck diagnostic, not a probability model, and its diagnosed bottleneck is invariant across the tested aggregators.The workflow strengthens one closure link at a time; for the retraining levy, coverage and enforcement provide the largest improvements in diagnostic failure score.
  • Certification: 27/57 chains pass strict edge-backed certification, while 30 remain exploratory after blind cross-model re-certification.The blind re-certification agrees on endpoint→pathway at 0.81, pathway→bottleneck at 0.77, and all five edges at 36/57 (0.63 [0.50, 0.75]).
  • Transfer and baselines: 0.83 held-out pathway accuracy follows from adding non-AI historical exemplars, up from 0.75 without them.A blind zero-shot labeler reaches 0.56 pathway accuracy (κ 0.32), while the exemplar result supports recurrence across governance domains rather than AI-specific stories.
  • Transfer and baselines: 0.61 leave-one-out top-1 accuracy is achieved by a contrastive few-shot mapper, compared with 0.49 for kNN.The mapper uses one exemplar per candidate pathway together with its discriminator; top-2 accuracy is 0.83.
  • Transfer and baselines: 0.54 versus 0.51 top-1 pathway naming shows keyword classification can match embedding retrieval, but it passes only 3/7 adversarial controls and transfers at 0.37.Abstaining on low-margin cases raises accuracy on the confident 40% to 0.65; memo generation reaches actor-role F1 0.61 and pathway top-1 0.67 on twelve held-out events.

N.5 Detailed Evaluation Metrics

The evaluation tests robustness of the decision objective, adoption predictions, world understanding, stakeholder coverage, and policy consequences. Results combine sensitivity analyses, backtests, and staged measures of political feasibility and effectiveness.

  • Decision-objective sensitivity: 99.9% of top-1 choices remain unchanged under ±0.05 noise, falling to 89.1% under ±0.20 noise.The ranking is also reproduced for 15/15 winners with min-link aggregation and 12/15 with additive aggregation.
  • Decision-objective sensitivity: Only two decision sets are genuinely fragile at ±0.10 noise: election-deepfake governance and AI incident response.In both cases, two candidates score nearly equally, so the benchmark flags rather than presents a spurious winner.
  • Adoption robustness: Adoption accuracy is 0.78 ± 0.035 across neutral phrasings, with only 4/48 cases changing binary prediction.By phrasing, accuracy is 0.81, 0.79, and 0.73.
  • Perception and coverage: World understanding reaches recall 0.94, precision 0.59, and F1 0.73 for involved role-level actors, with over-inclusion more common than misses.The blind model records 117 hits, 8 misses, and 80 over-inclusions.
  • Perception and coverage: Stakeholder coverage recovers key affected parties at 0.63 (19/30), finding primary actors more reliably than secondary and tertiary ripples.The structured error pattern indicates that most-affected actors are easier to recover than downstream stakeholders.
  • Political feasibility: Every shared-interest measure was adopted (3/3), while expert-favored high-cost, low-salience measures mostly failed (4/5).Adopted measures had mean public salience 0.68 versus 0.47 for non-adoptions; the lone high-cost adoption reflects regional regulatory culture.

O Benchmark Protocol and Leakage Control

The benchmark fixes temporal boundaries, evidence sources, and pathway-assignment rules to prevent leakage and make policy forecasts auditable. It evaluates predictions against a hierarchy of observed outcomes, official records, research, experts, and only finally LLM judgments.

  • Temporal design and leakage control: The protocol develops components through 31 December 2024 and freezes proposals from 1 January 2025 before outcomes are inspected.Frozen cases store proposal text, actors, and cutoff world state before later scoring against observed developments.
  • Evidence hierarchy: Scoring prioritizes observed outcomes, official records, credible causal estimates, and expert judgment before using an LLM evaluator.The hierarchy makes experts decisive where real-world outcomes cannot yet settle a case.
  • Label provenance: Silver labels are generated independently from incident attributes, while extracted fields remain auditable against cited sources.Author-coded fields are treated as hypotheses validated downstream rather than as direct observations.
  • Pathway construction: A pathway is a recurring mechanism traced through dated real events toward a catastrophic endpoint, not a single causal event.Its bottleneck is the link whose closure would break the chain.
  • Pathway construction: Events enter a pathway only when they match its invariant, uniquely beat competing pathways, remain source-grounded, and survive adversarial review.Events failing unique discrimination are assigned to neither pathway and flagged for audit; failures of later robustness criteria remain exploratory.
  • Pathway taxonomy: The codebook distinguishes governance failures including capability leakage, open-weight proliferation, systemic dependency, rivalry acceleration, governance lag, foreign free-riding, and enforcement gaps.Each pathway is defined by a different failed control or coverage condition, such as excessive response delay or missing authority.

Q.2 External Grounding in Existing Risk Datasets and Expert Elicitations

External grounding tests whether the pathway codebook captures governance mechanisms rather than merely harm categories or asserted analogies. Evaluation combines taxonomy cross-walks, incident-corpus coverage, expert elicitation, concrete policy chains, and inquiry-based labels.

  • External taxonomy grounding: The pathway taxonomy is orthogonal to the MIT AI Risk Repository because governance mechanisms cut across multiple harm domains.A single harm domain can span several pathways, and a pathway can span several harm domains.
  • Incident-corpus coverage: Only 7/26 (0.27) external incident reports map to one of seven pathways, while 19/26 are rejected as product harms without failed governance closure links.The explicit “other” option is intended to prevent force-fitting incidents into the taxonomy.
  • Expert calibration: Expert elicitation reports a median 5% probability of extremely bad outcomes and a median 10% probability of outcomes from inability to control advanced AI.Among surveyed AI authors, 38–51% assign at least a 10% chance to these outcomes.
  • Policy-chain analysis: Concrete policy examples trace governance gaps through hazard movement, failed controls, endpoints, pathways, bottlenecks, and repairs.Examples include offshore compute evading domestic reporting, public weights becoming non-revocable, and provider concentration creating systemic dependency.
  • Official-inquiry validation: A blind LLM recovers official governance links in 8/11 inquiry-backed cases, while hand-coded labels match inquiries in 4/8 overlapping corpus cases.The comparison uses official inquiry or root-cause labels independently of the codebook.
  • Mechanism-level evaluation: Surface methods fail invariant checks and strict domain transfer, whereas a mechanism-aware model reaches 0.86 on the same transfer.Recent-event pathway matching is about 0.8 at top-2, while exact single-label and per-class balance remain difficult.
Loading 2609.03553v1…