A French Corpus for Event Detection on Twitter - Archive ouverte HAL Access content directly
Proceedings Year : 2020

A French Corpus for Event Detection on Twitter

Béatrice Mazoyer
  • Function : Author
  • PersonId : 1217618
Julia Cagé
Nicolas Hervé
  • Function : Author
  • PersonId : 1217619
Céline Hudelot

Abstract

We present Event2018, a corpus annotated for event detection tasks, consisting of 38 million tweets in French (retweets excluded) including more than 130,000 tweets manually annotated by three annotators as related or unrelated to a given event. The 257 events were selected both from press articles and from subjects trending on Twitter during the annotation period (July to August 2018). In total, more than 95,000 tweets were annotated as related to one of the selected events. We also provide the titles and URLs of 15,500 news articles automatically detected as related to these events. In addition to this corpus, we detail the results of our event detection experiments on both this dataset and another publicly available dataset of tweets in English. We ran extensive tests with different types of text embeddings and a standard Topic Detection and Tracking algorithm, and detail our evaluation method. We show that tf-idf vectors allow the best performance for this task on both corpora. These results are intended to serve as a baseline for researchers wishing to test their own event detection systems on our corpus.
Fichier principal
Vignette du fichier
2020_mazoyer_cage_herve_hudelot_a_french_corpus_for_event_detection_on_twitter.pdf (1.12 Mo) Télécharger le fichier
Origin : Publisher files allowed on an open archive

Dates and versions

hal-03947820 , version 1 (19-01-2023)

Licence

Attribution - NonCommercial - CC BY 4.0

Identifiers

  • HAL Id : hal-03947820 , version 1

Cite

Béatrice Mazoyer, Julia Cagé, Nicolas Hervé, Céline Hudelot. A French Corpus for Event Detection on Twitter. European Language Resources Association (ELRA), pp.6220-6227, 2020. ⟨hal-03947820⟩
9 View
1 Download

Share

Gmail Facebook Twitter LinkedIn More