Source-linked AI summary
The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure
Kaiqiao Han, Yizhou Sun
TL;DR
The paper asks whether reported model failures on the imperfective-paradox benchmark reflect genuine semantic biases or problems in benchmark construction and evaluation. It reexamines the benchmark with controlled minimal pairs and multi-step diagnostics, finding Sufficiency Bias, Decision Shift, and additional aspectual and surface-form errors. The results provide initial evidence that larger models can classify aspect context-sensitively and perform comparably to human annotators under the benchmark standard.
Problem
The paper examines limited evidence about whether models' apparent Teleological Bias and prompting-related Calibration Crisis reflect semantic reasoning failures rather than benchmark and evaluation mis-specifications.
Method
The paper identifies three benchmark mis-specifications, constructs Lexically Matched Minimal Pairs, and evaluates event-semantic NLI through intermediate and oracle-guided multi-step reasoning analyses.
Results
Models often accept telic simple-past hypotheses without affirming culmination, while intermediate analyses reveal aspectual misclassification and Surface-form Attraction; GPT-5.4 and Qwen2.5-72B perform comparably to human annotators.
Takeaways & Limitations
The findings support reexamining categorical benchmark conclusions and distinguishing genuine semantic reasoning improvements from label Decision Shifts.
Takeaways & Limitations
Human annotations use a single background-informed protocol, without comparison to alternative annotation settings.
Abstract
from arXiv · showhide
The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly rule out culmination. In our native-speaker annotation, 38% of Group A examples and 29% of the Group C examples were judged to permit an alternative interpretation. To control these issues and lexical variation, we construct Lexically Matched Minimal Pairs. At the evaluation level, we formulate event-semantic NLI as a Multi-step Reasoning Problem and assess both intermediate semantic decisions and final predictions. Our results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern we characterize as Sufficiency Bias. We further show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understanding and reasoning. Intermediate and oracle-guided analyses identify two additional failure modes: errors in compositional aspectual classification and Surface-form Attraction toward surface-associated answers. Our experiments on Qwen-7B with suitable prompts, GPT-5.4, and Qwen-72B provide initial evidence for the context sensitivity of aspectual classification and suggest that these models can achieve performance comparable to that of human annotators.
1 INTRODUCTION
The paper argues that the imperfective-paradox benchmark and its conclusions are affected by conceptual and evaluation mis-specifications. It replaces direct classification with controlled minimal pairs and multi-step diagnostics, identifying Sufficiency Bias and Decision Shift rather than simply Teleological Bias and Calibration Crisis.
- Background: The imperfective paradox tests whether progressive descriptions entail completed telic events, unlike atelic activities whose ongoing occurrence entails the corresponding simple past.For example, building a gazebo does not entail completion, whereas running in the park entails running.
- Benchmark concerns: Aspectual Reduction assigns aspectual classes to isolated verbs even though lexical aspect is compositional and depends on the complete predicate.The same verb can occur in telic and atelic predicates, such as run home versus run in the park.
- Benchmark concerns: 76% of Group A instances do not explicitly establish non-completion under a strict NLI standard, while native-speaker annotation found alternative interpretations in 38% of Group A and 29% of Group C examples.These findings motivate reexamining categorical benchmark labels.
- Findings: Models often accept a telic simple-past hypothesis without affirming culmination, which the paper calls Sufficiency Bias.Prompting can shift output-label preferences without reliably improving semantic understanding or reasoning, producing Decision Shift.
- Findings: Intermediate and oracle-guided evaluations identify Aspectual Misclassification and Surface-form Attraction as additional failure modes.Experiments on GPT-5.4 and Qwen2.5-72B provide initial evidence of context-sensitive aspectual classification and performance comparable to human annotators.
- Methods: Lexically Matched Minimal Pairs vary arguments or outcome evidence while preserving other lexical and syntactic context.This design is intended to isolate compositional aspect from lexical associations.
2 RELATED WORK
Related work frames the imperfective paradox as a test of aspectual reasoning and reports inconsistent model sensitivity to context. The benchmark crosses predicate telicity with outcome information, using four labeled groups to evaluate progressive-to-simple entailment.
- Aspectual theory: Aspectual class depends on the complete predicate rather than the isolated verb, as contrasts such as run and run a mile demonstrate.Telicity arises through interactions among verbs, arguments, and other sentential material.
- Human judgments: Human judgments for telic predicates are heterogeneous, while atelic predicates generally yield consistent progressive-to-simple entailment judgments.Responses vary with the predicate, interruption scenarios, and world knowledge.
- Prior model evaluations: Context does not consistently improve model performance in CoRE, suggesting limited sensitivity to aspectual distinctions.The evaluation compares ongoing and completed event descriptions with and without narrative context.
- Benchmark design: The benchmark uses a 2 × 2 design crossing predicate telicity with contextual outcome information.The four groups are interrupted accomplishments, interrupted activities, ambiguous accomplishments, and ambiguous activities.
- Benchmark labels: Group A labels interrupted accomplishments as False, whereas Group B labels interrupted activities as True.Group C labels ambiguous accomplishments as Unknown and Group D labels ambiguous activities as True.
3 CONCEPTUAL MIS-SPECIFICATIONS
The benchmark contains three conceptual mis-specifications: aspect is reduced to isolated verbs, completion is treated too rigidly, and interruption is incorrectly mapped to contradiction. Lexically matched minimal pairs and validity analyses expose these confounds and clarify why model errors may reflect benchmark design rather than semantic failure.
- Aspectual Reduction: Aspectual Reduction assigns aspectual classes to isolated verbs, although aspect is compositionally determined by the complete predicate.The same verb can occur in telic or atelic predicates depending on arguments, paths, goals, and other constituents.
- Semantic Mis-specification: Semantic Mis-specification treats completion as a rigid canonical-endpoint condition, despite context-sensitive and potentially ambiguous completion judgments.Contextually negligible remainders may permit completion, whereas substantial partial events ordinarily do not.
- Relation Misassignments: Relation Misassignments label interruption as contradiction even though interruption alone leaves eventual completion unresolved.Without information about what happens afterward, interrupted construction should be neutral rather than false.
- Empirical Consequences: Human disagreement and model performance patterns indicate that benchmark labels, especially in Groups A and C, can confound semantic reasoning with annotation validity.Group A remains low and unstable while Group C improves sharply with model size, consistent with relation-label conflict and improved culmination recognition.
- Controlling Semantic Confounds with Lexically Matched Minimal Pairs: Lexically matched minimal pairs hold subjects, verbs, tense, and scenarios constant while manipulating predicate telicity and outcome evidence.The 2 × 2 design uses complete predicates, such as walk to the station versus walk in the station, to test compositional interpretation.
4 EVALUATION MIS-SPECIFICATIONS
The paper reframes event-semantic NLI as a multi-step reasoning problem because direct labels conflate distinct semantic and logical decisions. This analysis challenges Teleological Bias as the sole explanation, identifying Sufficiency Bias and showing that prompting can shift labels without necessarily improving semantic reasoning.
- Multi-step Reasoning Framework: Multi-step Reasoning Framework decomposes event-semantic NLI into distinct semantic and logical decisions, enabling stage-specific error localization.Intermediate and oracle-guided evaluations measure individual reasoning stages and error propagation rather than treating every final mistake as one bias.
- Sufficiency Bias: Models often identify culmination as unresolved or absent while still labeling the corresponding telic simple-past hypothesis entailed.The final error can arise when unresolved culmination is incorrectly converted into entailment from mere event occurrence.
- Sufficiency Bias: Occurrence-to-completion experiments show the largest decision shift when event occurrence is established, not when culmination is established.Models frequently accept the simple-past statement once the event occurred, even when its endpoint remains unguaranteed.
- Sufficiency Bias: The resulting failure mode, Sufficiency Bias, treats event occurrence as sufficient evidence for a telic simple-past predicate.This provides an alternative explanation to the claim that models systematically hallucinate culmination.
- Prompting Effects: Prompting interventions can produce Decision Shift among output labels without reliably resolving the underlying semantic distinction.The paper distinguishes changing evidence-to-label mappings from improving semantic discrimination and reasoning.
- Prompting Effects: Incomplete prompts omit operational criteria for aspectual classification and do not explicitly require classifying the complete predicate as telic or atelic.The zero-shot, DAP, and chain-of-thought prompts therefore fail to fully specify the reasoning process.
5 RETHINK THE IMPERFECTIVE PARADOX
The paper decomposes event-semantic NLI into intermediate decisions and evaluates where models fail. Results implicate aspectual classification, unstable label mapping, and surface-form attraction rather than a single culmination-representation failure.
- Multi-step Reasoning: Models must explicitly predict telicity, process occurrence, culmination, non-culmination, and the final NLI label.This structured output tests consistency across the complete reasoning process.
- Aspectual Classification: Aspectual classification is a necessary but difficult prerequisite for resolving the Imperfective Paradox, especially for 7B–9B models.The paper therefore evaluates it as a separate task.
- Prompting: Prompt modifications primarily shift the decision boundary between predicate classes while producing only modest gains in distinguishing telic from atelic predicates.This pattern suggests calibration of decision behavior rather than substantial improvement in semantic understanding.
- Oracle-Guided Evaluation: Oracle-guided evaluation supplies gold intermediate states and a deterministic procedure to isolate instruction-following and reasoning failures.It compares semantic and opaque representations under different prompt and reasoning settings.
- Failure Modes: Surface-form Attraction describes predictions influenced by answer-label associations even when the underlying decision rule remains constant.Explicit reasoning generally helps, but gains depend strongly on how intermediate states and final answers are represented.
- Large Models: GPT-5.4 and Qwen2.5-72B generally reason correctly after predicate classification, but their predictions diverge mainly at aspectual classification.In Group C, models sometimes choose True despite not judging the event completed, and their average benchmark performance is below 90.
6 DISCUSSION ABOUT EVERYDAY ENGLISH AND LANGUAGES WITH WEAK CULMINATION INFERENCES
Culmination judgments vary across predicates, contexts, and languages. Mandarin and several other languages permit bounded descriptions without guaranteeing that an inherent endpoint was reached, unlike canonical English cases such as kill.
- Cross-Linguistic Contrast: Mandarin permits a perfective killing description followed by denial of death because -le marks occurrence or boundedness without necessarily entailing the expected result state.The example is contrasted with the usual English judgment.
- Everyday English: The English sentence “James killed Hannibal twice, but Hannibal did not die” is normally unacceptable because kill lexically entails death.The continuation contradicts the first clause’s culmination entailment.
- Weak Culmination Inferences: Mandarin Chinese, Hindi, Tamil, and several Salish languages permit non-culminating interpretations of accomplishment predicates, with availability and grammatical source varying across languages.A perfective or otherwise bounded description therefore does not always guarantee endpoint attainment.
- Caveat: The linguistic background and underlying explanation are more complex than the paper’s brief presentation, and the example may admit alternative analyses.The paper directs readers to cited work for fuller discussion.
- Implications for Models: Multilingual parameter sharing may make distinctions between event occurrence and endpoint attainment harder for models to learn reliably.This is presented as one possible contributor among several.
7 CONCLUSION
The paper reexamines the imperfective-paradox benchmark and replaces its central single-step framing with controlled data and multi-step evaluation. It characterizes the resulting model behavior through distinct benchmark and reasoning failure modes.
- Conclusion: The paper identifies aspectual reduction, semantic mis-specification, and relation misassignment as three benchmark mis-specifications.It also constructs lexically matched minimal pairs to control annotation quality and lexical variation.
- Conclusion: Multi-step evaluation reveals Sufficiency Bias: models may reject culmination while accepting the corresponding telic simple-past statement.Prompting and ablation distinguish semantic improvement from Decision Shift among output labels.
- Conclusion: Intermediate-output and oracle-guided evaluations reveal sensitivity to semantically meaningful versus opaque label representations.This separates reasoning behavior from the effects of label representation.
LIMITATIONS AND FUTURE WORK
The paper’s controlled examples improve diagnostic clarity but may not capture the full interpretive range of natural language. Its background-informed annotation also trades theoretical control for possible protocol-induced priming.
- Background-informed annotation provides theoretical control while potentially priming annotators through the introduced linguistic distinctions.The protocol is treated as a deliberate trade-off rather than a procedure that eliminates annotation bias.
- Aspectual interpretation depends compositionally on predicates, arguments, paths, quantities, result phrases, and contextual standards of completion.The same verb can appear in telic or atelic predicates, and maximalization depends on the relevant linguistic and contextual scale.
- Predicates may support multiple linguistically plausible aspectual construals, so disagreement with a single gold label need not indicate a semantic error.This ambiguity is especially relevant when culmination depends on contextual standards or partial completion.
B DATA
The study compares three ImperfectiveNLI versions, using the Original and Lexically Matched Minimal-Pair datasets as primary diagnostics and the selectively filtered VSS only as auxiliary sensitivity analysis.
- The experiments use the Original, VSS, and Lexically Matched Minimal-Pair versions of ImperfectiveNLI.The Original dataset contains 400 controlled examples distributed across four benchmark groups.
- The Original and Lexically Matched Minimal-Pair datasets serve as the primary controlled diagnostics.The examples are organized as minimal pairs that vary event-semantic information relevant to the target inference.
- The VSS is reported only as an auxiliary sensitivity analysis because selective filtering may introduce bias.Its construction therefore cannot serve as an independent source of evidence for the main conclusions.
- Lexically matched pairs use explicit path, quantity, or result expressions while holding the lexical verb and general scenario constant.Examples are retained only when the aspectual contrast and culmination status are explicitly supported by the complete predicate and context.
- The human annotation study measures support for original labels and the distribution of alternative interpretations across benchmark conditions.It does not replace the original labels with a new categorical annotation scheme.
C.1 ANNOTATION PROTOCOL
Three native English-speaking annotators evaluated all 400 premise–hypothesis pairs using a background-informed, multi-label protocol that separates culmination uncertainty from aspectual ambiguity.
- Three native English-speaking annotators independently annotated all 400 premise–hypothesis pairs, with 100 examples in each benchmark group.The original labels assign FALSE to interrupted accomplishments, TRUE to interrupted activities, UNKNOWN to outcome-underspecified accomplishments, and TRUE to outcome-underspecified activities.
- Annotators could mark every linguistically acceptable NLI relation rather than selecting exactly one label.They also indicated a preferred or tendency interpretation, preserving plausible alternatives without forcing categorical agreement.
- UNKNOWN denotes uncertainty about whether an inherent endpoint was reached under a fixed telic interpretation.This differs from uncertainty about the predicate’s aspectual interpretation itself.
- The protocol introduces linguistic background before annotation to support distinctions relevant to the task.The study acknowledges that this preparation may influence how annotators construe ambiguous examples.
- Table 10 reports alternative-label support and preferred-label agreement as separate annotation outcomes.The first measure allows multiple acceptable relations, while the second tests reproduction of the original gold label by majority preference.
C.2 ALTERNATIVE INTERPRETATIONS ARE CONCENTRATED IN ACCOMPLISHMENTS
Alternative interpretations are concentrated in accomplishment conditions rather than distributed uniformly across the benchmark. This concentration weakens the security of categorical gold labels for the benchmark’s central semantic contrast.
- 17.5% of benchmark examples received alternative-label support from at least two of three annotators.Across the 400 examples, this uncertainty was not uniformly distributed across conditions.
- 38% of Group A and 29% of Group C examples received support for an alternative interpretation.By contrast, only 2% of Group B and 1% of Group D examples received such support; accomplishment conditions pooled at 33.5% versus 1.5% for activities.
- Majority voting reproduced the original gold label for 71% of Groups A and C, compared with 98% and 99% for Groups B and D.The asymmetry remains when annotators are reduced to a single preferred judgment.
- The findings indicate an annotation-validity gap concentrated in accomplishment conditions, not uniform label noise.Groups B and D show near-complete agreement with intended labels, while disagreement clusters in Groups A and C.
- Model–gold disagreement is harder to interpret for accomplishment examples because multiple plausible relations may be accepted.Accuracy on these items can conflate model errors with differences in predicate or culmination interpretations.
- The annotations identify where categorical evaluation is supported without establishing a replacement single gold label for every disputed example.The paper distinguishes strict evaluation under original labels from analyses that preserve relevant alternative interpretations.
D PROMPT
The paper argues that existing prompts underspecify the event-semantic reasoning required for three-way NLI. It introduces DAPCoT, which explicitly decomposes predicate classification, event-status analysis, truth-condition application, and final label mapping.
- D.1 BASELINE PROMPTS: Existing prompts define NLI labels but provide incomplete guidance for determining predicate type, process occurrence, and culmination.Zero-shot prompting supplies only final-label definitions, while DAP invokes aspectual terms without operational criteria.
- D.1 BASELINE PROMPTS: Lexical aspect is compositional: the complete predicate, including objects and complements, determines whether an event is telic or atelic.For example, run is atelic, whereas run home is a telic accomplishment because reaching home supplies an inherent endpoint.
- D.1 BASELINE PROMPTS: For atelic predicates, process occurrence normally suffices for the simple-past hypothesis, whereas telic predicates require established culmination.The distinction concerns truth conditions, not merely whether progressive descriptions imply completion.
- D.1 BASELINE PROMPTS: Chain-of-thought prompting asks about temporal status and endpoints but does not explicitly require classifying the complete predicate before reasoning about completion.The prompt leaves terms such as temporal status and defined endpoint underspecified and does not enforce the theoretically relevant reasoning order.
- D.2 REVISED EVENT-SEMANTIC PROMPT: DAPCoT separates process occurrence, endpoint identification, culmination status, aspectual truth conditions, and NLI-label alignment into explicit reasoning steps.The prompt extracts and classifies the complete predicate, determines whether its process occurred, applies the relevant rule, and maps the semantic judgment to TRUE, FALSE, or UNKNOWN.
- D.2 REVISED EVENT-SEMANTIC PROMPT: DAPCoT defines accomplishments as predicates with inherent endpoints and activities as predicates without inherent required endpoints.It instructs models to judge the complete predicate and avoid inferring endpoints from common sense, purposes, stopping, or contextual reactions.
- D.2 REVISED EVENT-SEMANTIC PROMPT: The revised rules state that progressive accomplishments do not imply completion, while activity occurrence supports the corresponding simple-past hypothesis without requiring completion.These rules distinguish inherent endpoint requirements from mere process occurrence.
- D.4 DIAGNOSTIC PROMPTS: The endpoint probe directly tests whether the premise guarantees the hypothesis’s required culmination, treating absent completion evidence as UNKNOWN rather than non-completion.It uses TRUE, FALSE, and UNKNOWN to assess endpoint status independently of the final NLI decision.
E IMPLEMENTATION DETAILS
The implementation evaluates multiple instruction-tuned models and prompting conditions while separating endpoint judgments, event-status classifications, intermediate semantic decisions, and final NLI labels. These diagnostics are designed to identify whether errors arise from semantic analysis or label mapping.
- Models and inference: Four instruction-tuned models at the 7B–9B scale are evaluated with zero temperature and a maximum output length of 2,048 tokens.The models are Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, GLM-4-9B-0414, and DeepSeek-R1-Distill-Qwen-7B.
- Prompting baselines: The prompting baselines compare zero-shot, definition-aware, chain-of-thought, and counterfactual prompts within the same TRUE/FALSE/UNKNOWN label space.Zero-shot supplies label definitions, while the other prompts add progressively different reasoning guidance.
- Completion and event-status diagnostics: Completion diagnostics compare direct-answer and reasoning prompts on whether the hypothesis endpoint is established, reporting UNKNOWN accuracy and the non-TRUE rate.The direct completion probe evaluates endpoint evidence rather than only the final sentence-pair label.
- Completion and event-status diagnostics: The group-wise event-status diagnostic samples twenty examples from each of Groups A–D and distinguishes unresolved occurrence, interruption, and completion.Groups A and B use interruption as the gold state, while Groups C and D use unresolved status.
- Occurrence-to-completion intervention: The occurrence-to-completion intervention varies information from intention through occurrence, interruption, explicit non-completion, and completion while holding the hypothesis fixed.This design attributes decision changes to added event information across five premise variants.
- Revised prompts and ablations: Revised-prompt ablations remove occurrence rules, aspectual definitions, or the distinction between activity and accomplishment truth conditions.All ablations retain the same examples, decoding parameters, and output-label normalization as the main experiments.
- Intermediate-output evaluation: Structured evaluation scores predicate telicity, process occurrence, completion, interruption, predicted relation, final label, and the percentage of responses with every field correct.This separates intermediate semantic discrimination from errors mapping semantic judgments to NLI labels.