--- _id: '10284' abstract: - lang: eng text: We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy and paste, we employ state-of-the-art text reuse detection technology, scaling it for the first time to process the entire Wikipedia as part of a distributed retrieval pipeline. We further report on a pilot analysis of the 100 million reuse cases inside, and the 1.6 million reuse cases outside Wikipedia that we discovered. Text reuse inside Wikipedia gives rise to new tasks such as article template induction, fixing quality flaws, or complementing Wikipedia's ontology. Text reuse outside Wikipedia yields a tangible metric for the emerging field of quantifying Wikipedia's influence on the web. To foster future research into these tasks, and for reproducibility's sake, the Wikipedia text reuse corpus and the retrieval pipeline are made freely available. author: - first_name: Milad full_name: Alshomary, Milad id: '73059' last_name: Alshomary - first_name: Michael full_name: Völske, Michael last_name: Völske - first_name: Tristan full_name: Licht, Tristan last_name: Licht - first_name: Henning full_name: Wachsmuth, Henning id: '3900' last_name: Wachsmuth - first_name: Benno full_name: Stein, Benno last_name: Stein - first_name: Matthias full_name: Hagen, Matthias last_name: Hagen - first_name: Martin full_name: Potthast, Martin last_name: Potthast citation: ama: 'Alshomary M, Völske M, Licht T, et al. Wikipedia Text Reuse: Within and Without. In: Azzopardi L, Stein B, Fuhr N, Mayr P, Hauff C, Hiemstra D, eds. Advances in Information Retrieval. Cham: Springer International Publishing; 2019:747-754.' apa: 'Alshomary, M., Völske, M., Licht, T., Wachsmuth, H., Stein, B., Hagen, M., & Potthast, M. (2019). Wikipedia Text Reuse: Within and Without. In L. Azzopardi, B. Stein, N. Fuhr, P. Mayr, C. Hauff, & D. Hiemstra (Eds.), Advances in Information Retrieval (pp. 747–754). Cham: Springer International Publishing.' bibtex: '@inproceedings{Alshomary_Völske_Licht_Wachsmuth_Stein_Hagen_Potthast_2019, place={Cham}, title={Wikipedia Text Reuse: Within and Without}, booktitle={Advances in Information Retrieval}, publisher={Springer International Publishing}, author={Alshomary, Milad and Völske, Michael and Licht, Tristan and Wachsmuth, Henning and Stein, Benno and Hagen, Matthias and Potthast, Martin}, editor={Azzopardi, Leif and Stein, Benno and Fuhr, Norbert and Mayr, Philipp and Hauff, Claudia and Hiemstra, DjoerdEditors}, year={2019}, pages={747–754} }' chicago: 'Alshomary, Milad, Michael Völske, Tristan Licht, Henning Wachsmuth, Benno Stein, Matthias Hagen, and Martin Potthast. “Wikipedia Text Reuse: Within and Without.” In Advances in Information Retrieval, edited by Leif Azzopardi, Benno Stein, Norbert Fuhr, Philipp Mayr, Claudia Hauff, and Djoerd Hiemstra, 747–54. Cham: Springer International Publishing, 2019.' ieee: 'M. Alshomary et al., “Wikipedia Text Reuse: Within and Without,” in Advances in Information Retrieval, 2019, pp. 747–754.' mla: 'Alshomary, Milad, et al. “Wikipedia Text Reuse: Within and Without.” Advances in Information Retrieval, edited by Leif Azzopardi et al., Springer International Publishing, 2019, pp. 747–54.' short: 'M. Alshomary, M. Völske, T. Licht, H. Wachsmuth, B. Stein, M. Hagen, M. Potthast, in: L. Azzopardi, B. Stein, N. Fuhr, P. Mayr, C. Hauff, D. Hiemstra (Eds.), Advances in Information Retrieval, Springer International Publishing, Cham, 2019, pp. 747–754.' date_created: 2019-06-21T08:59:32Z date_updated: 2022-01-06T06:50:34Z department: - _id: '600' editor: - first_name: Leif full_name: Azzopardi, Leif last_name: Azzopardi - first_name: Benno full_name: Stein, Benno last_name: Stein - first_name: Norbert full_name: Fuhr, Norbert last_name: Fuhr - first_name: Philipp full_name: Mayr, Philipp last_name: Mayr - first_name: Claudia full_name: Hauff, Claudia last_name: Hauff - first_name: Djoerd full_name: Hiemstra, Djoerd last_name: Hiemstra language: - iso: eng main_file_link: - url: https://webis.de/downloads/publications/papers/stein_2019c.pdf page: 747-754 place: Cham publication: Advances in Information Retrieval publication_identifier: isbn: - 978-3-030-15712-8 publisher: Springer International Publishing status: public title: 'Wikipedia Text Reuse: Within and Without' type: conference user_id: '82920' year: '2019' ...