Source-linked AI summary

A French Corpus Annotated for Multiword Expressions with Adverbial Function

Eric Laporte, Takuya Nakamura, Stavroula Voyatzi

arXiv:2606.04828v1cs.CL

TL;DR

Research on French multiword adverbs needs broader annotated resources because these expressions are difficult to distinguish from other prepositional phrases and existing corpora are limited. This paper defines an annotation target, constructs a corpus using lexical and grammatical resources with manual revision, and represents morphosyntactic, semantic, and discourse features. The resulting corpus is intended for information retrieval, lexical acquisition, and shallow and deep syntactic parsing.

  • Problem

    French multiword-adverb corpora are limited, while distinguishing these expressions from arguments or noun modifiers is difficult and relevant to syntactic analysis.

  • Method

    The authors annotate French multiword expressions with adverbial function using a syntactic-semantic lexicon, local grammars for named entities, automatic tagging, and manual revision.

  • Results

    The paper produces a French corpus whose annotations encode morphosyntactic structure, named-entity semantic types, and conjunctive discourse function.

  • Takeaways & Limitations

    The corpus can be used with the French Treebank for information retrieval and extraction, automatic lexical acquisition, and shallow and deep syntactic parsing.

  • Takeaways & Limitations

    The compositionality criterion can be empirically checked only after the lexicon and grammar for the same language are complete and compatible.

Abstract

from arXiv · show

This paper presents a French corpus annotated for multiword expressions (MWEs) with adverbial function. This corpus is designed for investigation on information retrieval and extraction, as well as on deep and shallow syntactic parsing. We delimit which kind of MWEs we annotated, we describe the resources and methods we used for the annotation, and we briefly comment the results. The annotated corpus is available at http://infolingu.univ-mlv.fr/ under the LGPLLR license.

1. Introduction

The paper presents a French corpus annotated for multiword adverbs, motivated by their potential value for information retrieval, extraction, and syntactic parsing.

  • Recognising multiword adverbs may improve information retrieval and extraction because these adverbials convey information.
  • Recognition may also help resolve prepositional attachment because many multiword adverbs superficially resemble prepositional phrases.Identifying them can rule out analyses treating them as arguments or noun modifiers.
  • The authors create and analyse a French corpus annotated with multiword adverbs using specified resources and methods.The corpus is intended for research and is to be freely available under the LGPLLR license.

2. Related work

Existing corpora provide limited support for studying French multiword adverbs, because such units are rarely annotated and often receive only coarse treatment.

  • French and other-language corpora annotated with multiword units are rare and small.

1 Several reasons explain this lack of interest. Firstly, adverbials

Researchers have shown limited interest in multiword adverbs because they seem less useful than nouns for retrieval and are difficult to distinguish from other prepositional phrases.

  • Adverbials are often considered less useful than nouns for information retrieval and extraction.
  • Distinguishing multiword adverbs from arguments or noun modifiers is difficult because the distinction lacks strong textual markers and involves complex linguistic notions.
  • The distinction is nevertheless essential to identifying a sentence’s semantic core, and larger annotated corpora may clarify the associated problems.
  • The French Treebank contains function annotations for 350 000 words, and the authors report no other available French corpora annotated with multiword adverbs.

3. Target of annotation

The annotation targets French multiword expressions with adverbial function, defining them through lexical freezing and circumstantial-complement criteria, and representing their structural and discourse properties.

  • Target of annotation: The annotation target is the intersection of multiword expressions and adverbial function.
  • Multiword expression criterion: A phrase is treated as a multiword expression when some or all elements are frozen and their combination does not follow productive syntactic and semantic compositionality.
  • Multiword expression criterion: The compositionality criterion is empirically checkable only once a language’s lexicon and grammar are complete and compatible.
  • Adverbial function: Adverbial expressions are circumstantial complements rather than predicate objects, identified through optionality, broad predicate compatibility, and some specific pronominalization patterns.
  • Features: Each occurrence receives one morphosyntactic structure or semantic type among 19, based on the number, category, and position of frozen and free components.
  • Annotation scope: The annotation includes frozen-plus-free expressions and named entities of date and duration, whose specific grammatical rules support treating them as multiword expressions.
  • Features: A binary feature records whether the adverbial connects its clause to the preceding clause through a conjunctive discourse function.

4. Methodology

The corpus annotation combines lexicon-based and local-grammar tagging with manual revision. It records morphosyntactic and syntactic-semantic information about French multiword adverbs.

  • Resources and annotation workflow: The annotation targeted adverbial occurrences identified through a syntactic-semantic lexicon, local grammars for temporal named entities, and manual revision.The local grammars covered date, duration, time, and frequency named entities.
  • Resources and annotation workflow: The shared lexicon contains 6 800 entries and is available under the LGPLLR license for research and business.It was constructed from dictionaries, grammars, corpora, and introspection using the Lexicon-Grammar methodology.
  • Lexicon representation: Lexicon-Grammar tables represent lexical items by morphosyntactic elements and binary syntactic-semantic features.Examples include conjunctive discourse function and restrictions requiring occurrence in a negative clause.
  • Lexicon representation: The resources include 15 tables, one for each morphosyntactic structure, with examples helping readers find sentences containing the adverbial.The features supplied by the lexicon were used for annotation.
  • Automatic tagging: Unitex tagging handled fixed expressions and variants involving agreement, permutations, omissions, and noun-phrase or clause complements.Examples distinguish complements tagged with NP and S structures.
  • Manual revision: Three experts manually reviewed the automatic annotation, deleting embedded tags, adding coordinated cases, and searching for lexicon-absent adverbs.Meetings during annotation were used to maintain consistency.

5. Results

The supplied passage identifies Table 4 as presenting annotated occurrences of multiword expressions with adverbial function.

  • Table 4 concerns annotated occurrences of multiword expressions with adverbial function.
  • The supplied table reference does not state the occurrence counts or other results.
  • No comparison or quantitative finding is specified in the supplied passage.

6. Conclusion

The paper presents a French corpus annotated for multiword expressions with adverbial function and includes several annotation feature types. The corpus supports research in information retrieval, lexical acquisition, and syntactic parsing.

  • The corpus annotates multiword expressions with adverbial function in French.
  • Annotations include morphosyntactic structure, special discourse functions, and semantic types of named time entities.
  • The corpus can be used with the French Treebank for information retrieval and extraction, automatic lexical acquisition, and deep or shallow syntactic parsing.
Loading 2606.04828v1…