---
_id: '10284'
abstract:
- lang: eng
  text: We study text reuse related to Wikipedia at scale by compiling the first corpus
    of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia
    text in a sample of the Common Crawl). To discover reuse beyond verbatim copy
    and paste, we employ state-of-the-art text reuse detection technology, scaling
    it for the first time to process the entire Wikipedia as part of a distributed
    retrieval pipeline. We further report on a pilot analysis of the 100 million reuse
    cases inside, and the 1.6 million reuse cases outside Wikipedia that we discovered.
    Text reuse inside Wikipedia gives rise to new tasks such as article template induction,
    fixing quality flaws, or complementing Wikipedia's ontology. Text reuse outside
    Wikipedia yields a tangible metric for the emerging field of quantifying Wikipedia's
    influence on the web. To foster future research into these tasks, and for reproducibility's
    sake, the Wikipedia text reuse corpus and the retrieval pipeline are made freely
    available.
author:
- first_name: Milad
  full_name: Alshomary, Milad
  id: '73059'
  last_name: Alshomary
- first_name: Michael
  full_name: Völske, Michael
  last_name: Völske
- first_name: Tristan
  full_name: Licht, Tristan
  last_name: Licht
- first_name: Henning
  full_name: Wachsmuth, Henning
  id: '3900'
  last_name: Wachsmuth
- first_name: Benno
  full_name: Stein, Benno
  last_name: Stein
- first_name: Matthias
  full_name: Hagen, Matthias
  last_name: Hagen
- first_name: Martin
  full_name: Potthast, Martin
  last_name: Potthast
citation:
  ama: 'Alshomary M, Völske M, Licht T, et al. Wikipedia Text Reuse: Within and Without.
    In: Azzopardi L, Stein B, Fuhr N, Mayr P, Hauff C, Hiemstra D, eds. <i>Advances
    in Information Retrieval</i>. Cham: Springer International Publishing; 2019:747-754.'
  apa: 'Alshomary, M., Völske, M., Licht, T., Wachsmuth, H., Stein, B., Hagen, M.,
    &#38; Potthast, M. (2019). Wikipedia Text Reuse: Within and Without. In L. Azzopardi,
    B. Stein, N. Fuhr, P. Mayr, C. Hauff, &#38; D. Hiemstra (Eds.), <i>Advances in
    Information Retrieval</i> (pp. 747–754). Cham: Springer International Publishing.'
  bibtex: '@inproceedings{Alshomary_Völske_Licht_Wachsmuth_Stein_Hagen_Potthast_2019,
    place={Cham}, title={Wikipedia Text Reuse: Within and Without}, booktitle={Advances
    in Information Retrieval}, publisher={Springer International Publishing}, author={Alshomary,
    Milad and Völske, Michael and Licht, Tristan and Wachsmuth, Henning and Stein,
    Benno and Hagen, Matthias and Potthast, Martin}, editor={Azzopardi, Leif and Stein,
    Benno and Fuhr, Norbert and Mayr, Philipp and Hauff, Claudia and Hiemstra, DjoerdEditors},
    year={2019}, pages={747–754} }'
  chicago: 'Alshomary, Milad, Michael Völske, Tristan Licht, Henning Wachsmuth, Benno
    Stein, Matthias Hagen, and Martin Potthast. “Wikipedia Text Reuse: Within and
    Without.” In <i>Advances in Information Retrieval</i>, edited by Leif Azzopardi,
    Benno Stein, Norbert Fuhr, Philipp Mayr, Claudia Hauff, and Djoerd Hiemstra, 747–54.
    Cham: Springer International Publishing, 2019.'
  ieee: 'M. Alshomary <i>et al.</i>, “Wikipedia Text Reuse: Within and Without,” in
    <i>Advances in Information Retrieval</i>, 2019, pp. 747–754.'
  mla: 'Alshomary, Milad, et al. “Wikipedia Text Reuse: Within and Without.” <i>Advances
    in Information Retrieval</i>, edited by Leif Azzopardi et al., Springer International
    Publishing, 2019, pp. 747–54.'
  short: 'M. Alshomary, M. Völske, T. Licht, H. Wachsmuth, B. Stein, M. Hagen, M.
    Potthast, in: L. Azzopardi, B. Stein, N. Fuhr, P. Mayr, C. Hauff, D. Hiemstra
    (Eds.), Advances in Information Retrieval, Springer International Publishing,
    Cham, 2019, pp. 747–754.'
date_created: 2019-06-21T08:59:32Z
date_updated: 2022-01-06T06:50:34Z
department:
- _id: '600'
editor:
- first_name: Leif
  full_name: Azzopardi, Leif
  last_name: Azzopardi
- first_name: Benno
  full_name: Stein, Benno
  last_name: Stein
- first_name: Norbert
  full_name: Fuhr, Norbert
  last_name: Fuhr
- first_name: Philipp
  full_name: Mayr, Philipp
  last_name: Mayr
- first_name: Claudia
  full_name: Hauff, Claudia
  last_name: Hauff
- first_name: Djoerd
  full_name: Hiemstra, Djoerd
  last_name: Hiemstra
language:
- iso: eng
main_file_link:
- url: https://webis.de/downloads/publications/papers/stein_2019c.pdf
page: 747-754
place: Cham
publication: Advances in Information Retrieval
publication_identifier:
  isbn:
  - 978-3-030-15712-8
publisher: Springer International Publishing
status: public
title: 'Wikipedia Text Reuse: Within and Without'
type: conference
user_id: '82920'
year: '2019'
...
