Source-linked AI summary

Contextual Embedding Evidence for Main--Light Verb Distinctions in Urdu

Farah Adeeba, Miriam Butt

arXiv:2608.23645v1cs.CLcs.LG

TL;DR

It remains unclear whether contextual language models encode Urdu main–light verb distinctions in ways predicted by linguistic accounts. This study analyzes contextual embeddings across seven verbs and three models, finding systematic main–light separation alongside lexical relatedness and recoverable verb-specific structure in masked contexts.

  • Problem

    It remains unclear whether contextual language models encode Urdu main–light distinctions and the representational predictions of linguistic accounts.

  • Method

    The study analyzes contextual embeddings from three language models across naturally occurring main- and light-verb sentences, including masked light-verb identity prediction.

  • Results

    Main and light uses separate systematically across all model–verb comparisons, while same-lemma proximity persists and UrduBERT reaches 0.866 masked accuracy and 0.782 under preceding-form-disjoint evaluation.

  • Takeaways & Limitations

    The findings provide computational evidence consistent with Urdu light verbs having distinct yet lexically related and verb-specific contextual representations.

  • Takeaways & Limitations

    The study cannot isolate semantic information from constructional or distributional cues, and overt target forms may partly explain same-lemma proximity.

Abstract

from arXiv · show

Urdu light verbs contribute schematic event-structural meaning while remaining lexically related to corresponding main verbs. This study tests representational predictions derived from Butt's analysis using contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT across 1,126 naturally occurring sentences containing seven Urdu verbs. Main and light uses show significant representational separation in all 21 verb--model comparisons. At the same time, same-lemma main and light centroids are consistently closer than mismatched main--light lemma pairs, supporting continued lexical relatedness. In a seven-way prediction task restricted to light uses, verb identity remains recoverable after the target is masked, with UrduBERT achieving 0.866 accuracy and 0.852 macro-F1. UrduBERT also retains 0.782 accuracy under a preceding-form-disjoint evaluation, indicating generalization beyond repeated local verb combinations. These findings provide computational evidence consistent with Butt's account that Urdu light verbs differ systematically from their main uses while retaining lemma-specific and verb-specific representational structure.

1 Introduction

This study tests whether contextual language models encode Urdu main–light verb distinctions in ways predicted by Butt’s analysis and lexical-relatedness account. Across seven verbs and three models, it combines geometric, probing, clustering, and masking analyses to assess separation, structure, and recoverability.

  • Research motivation: The study asks whether contextual language models distinguish main and light uses of Urdu verbs and whether their embedding geometry reflects Butt’s theoretical predictions.The analysis is motivated by uncertainty about whether language models encode this linguistically established contrast.
  • Study design: The dataset contains 1,126 naturally occurring Urdu sentences with seven canonical light verbs, evaluated using UrduBERT, DunbaaBERT, and multilingual BERT.Representative main and light uses for all seven verbs are provided in Appendix E.
  • Analytical approach: Complementary analyses measure main–light separation and characterize its structure through centroid cosine distance, silhouette scores, logistic probing, KMeans clustering, and PCA.Embedding distance is interpreted as representational differentiation rather than direct semantic distance.
  • Lexical relatedness: Same-lemma main and light centroids remain closer than mismatched cross-lemma pairs across all three models, supporting continued lexical relatedness.This comparison tests a prediction associated with the lexical-relatedness account developed by Butt and Lahiri.
  • Light-verb identity: Individual light verbs retain distinguishable contextual profiles after target masking, including under a preceding-form-disjoint split.The masking analysis tests whether light-verb identity remains recoverable beyond repeated local verb combinations.

2 Background and Related Work

Prior work characterizes Urdu light verbs as lexically related to main verbs but contributors of reduced, schematic event-structural meaning. Corpus and contextual-model research motivates studying the construction through naturally occurring data and representational comparisons.

  • Theoretical foundation: Butt’s account treats Urdu light verbs as modifying event structure with weak but non-vacuous schematic content rather than full main-verb meaning.Despite semantic reduction, light verbs remain in the lexical domain and retain verb-specific contributions.
  • Corpus evidence: A relatively small inventory of verbs accounts for a substantial proportion of naturally occurring Urdu light-verb constructions, motivating the study’s selection of seven verbs.The seven verbs are drawn from an inventory supported by theoretical and empirical work.
  • Computational approaches: Contextual language models can test whether different uses of the same surface form receive systematically different representations.English studies reported separable patterns for light-verb and full-verb uses and distinguished verb-particle constructions through contextual representations.

3 Dataset

The dataset comprises held-out Urdu sentences drawn from a deduplicated 1.6 GB corpus and centered on seven verbs with documented main and light uses. Labels were assigned through a GPT-5.1 and expert-human agreement protocol grounded in Butt’s criteria, excluding complex predicates and ambiguous cases.

  • Approximately 1.6 GB of deduplicated Urdu text was compiled from OSCAR 2019, GitHub, and Kaggle, with evaluation sentences held out from UrduBERT pre-training.
  • Seven high-frequency Urdu verbs were selected because linguistic literature identifies them as exhibiting both main and light verb behavior.The verbs are āyā, uṭhā, baiṭhā, paṛā, diyā, gayā, and liyā.
  • Sentences were extracted with the target verb in sentence-final position, consistent with Urdu’s SOV word order.
  • A strict three-way annotation protocol combined GPT-5.1 classification with independent review by two expert Urdu linguists.The labels were main, light, or skip; skip excluded noun-verb complex predicates and structurally ambiguous cases, following Butt’s 1995 criteria.

4 Methodology

The methodology tests three representational predictions about main–light differentiation, lexical relatedness, and verb-specific light-verb structure using three Urdu-to-multilingual encoder models. It combines contextual-embedding extraction, quantitative separation and relatedness analyses, masked-target classification, grouped generalization, and an exploratory tense-auxiliary control.

  • The study evaluates main–light differentiation, continued lexical relatedness, and verb-specific structure, while treating tense auxiliaries only as an exploratory grammatical reference condition.
  • Three encoder models span Urdu-specific to multilingual pre-training: UrduBERT, DunbaaBERT, and multilingual BERT.UrduBERT was trained on 5.8 GB of Urdu data, DunbaaBERT on 17 GB, and mBERT on Wikipedia text from 104 languages without Urdu-specific optimization.
  • Target-verb contextual embeddings are extracted from the final hidden layer, with attention-weighted pooling for targets spanning multiple subword tokens.Token positions are identified through tokenizer offset mappings and the target’s final occurrence in each sentence.
  • Main–light separation is assessed with cosine centroid distance, silhouette scores, permutation tests, logistic-regression probing, and annotation-independent k = 2 KMeans clustering.PCA is used only for qualitative visualization; quantitative separation claims rely on full-dimensional analyses.
  • Masked-target light-verb identity is evaluated with seven-way class-weighted logistic regression using accuracy and macro-F1, majority and 1/7 chance baselines, and label permutations.A preceding-form-disjoint StratifiedGroupKFold evaluation uses 153 groups to prevent shared immediately preceding surface forms across training and test data.

5 Results

Main and light uses are significantly separable across all verb–model comparisons, while same-lemma uses remain closer than mismatched lemmas. Masked-context prediction further shows that light-verb identity remains recoverable, especially with UrduBERT, though clustering and auxiliary proximity vary by model and verb.

  • Main–light separation: All 21 main–light centroid distances are significant at p < 0.001, with UrduBERT showing the largest mean distance (0.177) and mean silhouette score (0.272).mBERT and DunbaaBERT have mean distances of 0.045 and 0.034, respectively.
  • Main–light separation: UrduBERT achieves the highest mean F1 (0.997) in cross-validated probes, while all models exceed the majority-class baseline (0.675).DunbaaBERT reaches 0.986 mean F1 and mBERT 0.952.
  • Clustering: UrduBERT’s mean ARI is 0.778, compared with 0.017 for DunbaaBERT and 0.141 for mBERT, showing that linear recoverability does not guarantee dominant spherical clusters.Unsupervised two-cluster alignment varies substantially by model and verb.
  • Lexical relatedness: For every verb and model, same-lemma main–light centroids are closer than cross-lemma pairs, with top-1 accuracy and mean reciprocal rank both equal to 1.0.The exact one-sided sign-flip test gives p = .016 for each model; lexical identity may contribute because matching targets share the same orthographic form.
  • Light-verb identity prediction: After masking the target, UrduBERT predicts seven-way light-verb identity at 0.866 accuracy and 0.852 macro-F1, above the majority baseline (0.252) and uniform chance (0.143).DunbaaBERT reaches 0.616/0.585 and mBERT 0.633/0.599 for accuracy/macro-F1.
  • Light-verb identity prediction: Under preceding-form-disjoint evaluation, performance decreases from 0.616 to 0.520 for DunbaaBERT and from 0.633 to 0.571 for mBERT.The evaluation excludes preceding-form overlap across folds.
  • Auxiliary proximity: UrduBERT’s main and light centroids remain far from the auxiliary (dMA = 0.637; dLA = 0.621) but close to each other (dML = 0.177), with mean rv = 0.967.Four verbs have ratios below 1.0, indicating light embeddings displaced toward the auxiliary relative to main uses, although direction varies.

6 Discussion

The discussion finds that Urdu light verbs are systematically distinguishable from their main uses while retaining same-lemma relatedness and verb-specific contextual profiles. These patterns are compatible with Butt’s account, but the analyses do not by themselves establish a specific lexical architecture or semantic scale.

  • Representational separation: Across all seven verbs and three models, main and light uses occupy measurably different embedding regions, with the distinction linearly recoverable.Centroid permutation tests make the separation unlikely under random labels, while logistic probes establish linear recoverability.
  • Model geometry: DunbaaBERT combines high probe F1 with small centroid distances and near-zero mean ARI, showing that accessible distinctions need not form widely separated spherical clusters.UrduBERT instead combines large centroid distances, higher silhouette scores, high probe accuracy, and comparatively strong KMeans alignment.
  • Lexical relatedness: For every model and verb, matching main–light centroids are closer than mismatched cross-lemma pairs, supporting continued lexical relatedness.Each main centroid retrieves its corresponding light centroid at rank one, and bootstrap intervals show stable cross-minus-same differences.
  • Interpretive limits: The proximity analysis is compatible with lexical relatedness but does not prove a shared formal lexical entry because models preserve surface-token identity alongside contextual information.The conclusion should therefore be read as a representational correlate rather than proof of a particular lexical architecture.
  • Masked-target prediction: UrduBERT achieves 0.866 masked accuracy and retains 0.782 accuracy under the preceding-form-disjoint split, indicating generalization beyond repeated immediate V1–V2 pairs.All models predict light-verb identity above majority and uniform-chance baselines; DunbaaBERT also remains above baseline under grouped evaluation, while mBERT declines more.
  • Interpretive limits: The masked classifier may exploit lexical, syntactic, genre, or constructional cues, so the result establishes distinct and partially generalizable contextual profiles rather than isolated semantic content.These profiles remain compatible with verb-specific schematic event-structural content.
  • Verb-specific variation: Paṛā shows the largest representational distances, but embedding distance alone cannot establish graded semantic weakening or a semantic-distance scale.Independent human judgments or controlled semantic-similarity experiments would be required to test that interpretation.
  • Auxiliary comparison: Four of seven UrduBERT light-verb centroids have relative auxiliary-proximity ratios below 1.0, a directional pattern consistent with partial semantic reduction.Light-verb centroids nevertheless remain substantially distant from the tense-auxiliary centroid across all models.

7 Conclusion

Across 1,126 sentences and three contextual encoders, Urdu main and light uses are systematically distinguished while retaining same-lemma lexical relatedness. Light-verb representations also preserve verb-specific structure, supporting consequences of Butt’s account without directly proving lexical-entry structure, diachrony, or numerical semantic distance.

  • Representational differentiation: Main and light uses of the same seven Urdu verb forms are systematically differentiated in embedding space across three contextual encoders.The magnitude and geometry of this distinction vary across verbs and models.
  • Lexical relatedness: Same-lemma main and light centroids are consistently closer than mismatched cross-lemma pairs across all models, indicating continued lexical relatedness.This pattern is compatible with the lexical relationship predicted between corresponding main and light verbs.
  • Verb-specific structure: Individual light verbs remain recoverable from masked contexts, and UrduBERT retains high performance when immediate preceding surface forms are disjoint between training and test data.These analyses indicate preservation of verb-specific structure beyond repeated local combinations.
  • Theoretical interpretation: Taken together, the findings provide computational evidence consistent with Butt’s account of differentiated uses, continued lexical relatedness, and verb-specific schematic structure among light verbs.The evidence concerns specific consequences of the account rather than a direct demonstration of its full internal analysis.
  • Limitations: The study does not directly prove the internal structure of lexical entries, establish a diachronic pathway, or provide a numerical measure of semantic distance.These are explicit limits on the conclusions drawn from the representational analyses.

Limitations … D.3 Model Architecture

The study’s conclusions are limited by its narrow, naturally occurring dataset and by analyses that cannot fully disentangle lexical, semantic, constructional, or distributional cues. The paper also documents ethical safeguards, computational requirements, UrduBERT’s pre-training resources and tokenizer, and its BERT-base architecture.

  • Limitations: The study covers only seven canonical Urdu light verbs and naturally occurring text, so lexical, syntactic, genre, and source cues may influence the embedding geometry.Its conclusions should not be generalized to the full light-verb inventory or to other languages.
  • Limitations: The same-lemma proximity analysis retains overt target forms, while masked classification cannot isolate semantic information from constructional or distributional cues.Shared orthographic and lexical identity may partly explain matching-pair proximity, so the analysis supports relatedness rather than a shared formal lexical entry.
  • Limitations: The grouped masked evaluation blocks exact preceding-form overlap but may retain inflectional and complex-chain inconsistencies requiring manual normalization.Future work should use manually verified V1 lemmas and construction types.
  • Limitations: The tense-auxiliary comparison is exploratory and cannot establish a diachronic pathway or quantify semantic reduction.Auxiliaries and light verbs differ in lexical identity, morphology, syntax, and grammatical function.
  • A Ethical Consideration: All GPT-generated labels were independently reviewed by two Urdu linguistics experts, and only unanimously agreed instances were retained.The study used publicly available Urdu text and involved no human participants or private data.
  • B Computational Cost: UrduBERT pre-training used a single NVIDIA H100 GPU for approximately three days, while reported experiments took approximately 30 minutes and used roughly half the available memory.These figures describe the study’s computational cost.
  • D.1 Pre-training Corpus; D.2 Tokenizer: The pre-training corpus combined OSCAR 2019, GitHub, and Kaggle sources, was deduplicated to approximately 5.8 GB, and used five training files plus one validation file of approximately 1.16 GB each at an approximately 5:1 ratio.A WordPiece tokenizer trained on this corpus used 64,000 tokens and a maximum sequence length of 512.
  • D.3 Model Architecture: UrduBERT follows the standard BERT-base architecture, with its configuration summarized in Table 8.The supplied passage does not provide the individual configuration values.

D.4 Pre-training Procedure

UrduBERT was pre-trained with masked language modeling under a 0.15 masking probability using the HuggingFace Transformers Trainer API. Training ran on one NVIDIA H100 GPU for approximately 3–4 days, with checkpointing and evaluation at specified intervals.

  • Pre-training objective: 0.15 masking probability was used for masked language modeling pre-training through the HuggingFace Transformers Trainer API.The supplied passage indicates that the remaining training hyperparameters were specified separately.
  • Training execution: Approximately 3–4 days of training were performed on a single NVIDIA H100 GPU.The get_last_checkpoint utility enabled resumption from intermediate checkpoints.
  • Checkpointing and evaluation: Checkpoints were saved every 2,000 steps and evaluated every 4,000 steps.Checkpoint resumption used the get_last_checkpoint utility.

D.5 Downstream Evaluation · E Main and Light-Verb Examples

UrduBERT was evaluated on named entity recognition and sentiment analysis, outperforming multilingual BERT in both tasks. Table 10 provides one representative main and light use for each target verb.

  • D.5 Downstream Evaluation: UrduBERT outperformed multilingual BERT (mBERT) on both named entity recognition and sentiment analysis.These downstream evaluations were used to validate the quality of the pre-trained model.
  • D.5 Downstream Evaluation: The two downstream tasks were Named Entity Recognition (NER) and Sentiment Analysis.The evaluation covered both tasks explicitly.
  • D.5 Downstream Evaluation: Urdu-specific pre-training yielded representations better suited to Urdu language understanding than multilingual BERT.This conclusion follows from UrduBERT’s stronger performance in both downstream tasks.
  • D.5 Downstream Evaluation: Detailed downstream evaluation results will be released with the accompanying repository.The passage does not provide task-specific scores.
  • E Main and Light-Verb Examples: Table 10 presents one representative main use for each target verb.The examples are part of the paper’s main and light-verb examples section.
  • E Main and Light-Verb Examples: Table 10 presents one representative light use for each target verb.The passage identifies the table as covering paired main and light uses.

F Visualizations

Figures 3–5 visualize per-verb relationships among main, light, and AUXT uses across UrduBERT, DunbaaBERT, and mBERT. The visualizations encode cosine distances and display the ratio rv relative to 1.0 for each subplot.

  • Three-way distance visualizations: Figures 3–5 present per-verb triangle visualizations for all three models, with vertices representing main, light, and AUXT uses.Main vertices are blue, light vertices red, and AUXT vertices green.
  • Visualization encoding: Edge labels report cosine distance, while the ratio rv above each subplot is colored green when rv < 1.0 and red when rv > 1.0.The color coding distinguishes ratios below and above the 1.0 reference value.
  • UrduBERT: Figure 3 shows the three-way distance visualization for UrduBERT.The figure uses the shared triangle format described for Figures 3–5.
  • DunbaaBERT: Figure 4 shows the three-way distance visualization for DunbaaBERT.The figure uses the shared triangle format described for Figures 3–5.
  • mBERT: Figure 5 shows the three-way distance visualization for mBERT.The figure uses the shared triangle format described for Figures 3–5.
Loading 2608.23645v1…