A universal literary canon based on multilingual encyclopedic data: Proposal of a method for the ranking of literary works using quantitative data obtained from Wikidata and Wikipedia
DOI:
https://doi.org/10.3989/redc.2023.3.2013Keywords:
Literary canon, literary works, Wikidata, Wikipedia, Wiki3Drank, RankingAbstract
The research described in this article aims to verify the use of Wikidata and Wikipedia as a source to identify a universal literary canon. Both Wikimedia Foundation projects are placed in the context of data on literary works. The methodology used is based on the construction of a dataset from specific data on literary works retrieved from Wikidata and Wikipedia editions in all languages. The depth of description of the items of literary works in Wikidata and their presence and level of elaboration of the corresponding articles in Wikipedia are analyzed. The authors use K-means to define three clusters of literary works that allow the identification of a set of works that can be used to create a universal literary canon. Wiki3DRank is proposed as a metric that allows the literary works analyzed to be selected and ranked. The study deals with the analysis of the language of literary works and their presence in Wikipedia, their temporal distribution. The article includes a discussion section with reflections on the results obtained and concludes with the proposal to use Wikidata and Wikipedia as an alternative source for the elaboration of both global and language-specific literary canons.
Downloads
References
Algee-Heweitt, M., Allison, S., Gemma, M., Heuser, R., y Moretti, F. (2018). Canon/archivo: dinámicas de largo alcance y campo literario. En F. Moretti (Ed.), Literatura en el laboratorio: canon, archivo y crítica literaria en la era digital, 131-181. Gedisa.
Arthur, D., y Vassilvitskii, S. (2007). K-means++: the advantages of careful seeding. Proceedings of the eighteenth annual ACM-SIAM symposium on Discrete algorithms, 1027-1035.
Bianchini, C., y Sardo, L. (2022). Wikidata : a new perspective towards universal bibliographic control. JLIS, 13(1).
Bourdieu, P. (1995). The Rules of Art: Genesis and Structure of the Literary Field. Stanford University Press. https://doi.org/10.1515/9781503615861
Boxall, P., y Mainer, J. C. (2016). 1001 libros que hay que leer antes de morir: relatos e historias de todos los tiempos (7a ed.). Grijalbo.
Claes, F., y Tramullas, J. (2021). Estudios sobre la credibilidad de Wikipedia: una revisión. Área Abierta, 21(2), 187-204. https://doi.org/10.5209/arab.74050
Damrosch, D. (2009). How to read world literature. Wiley-Blackwell. https://doi.org/10.1002/9781444304596
Ding, C., y He, X. (2004). K-means clustering via principal component analysis. Twenty-First International Conference on Machine Learning - ICML '04, 29. https://doi.org/10.1145/1015330.1015408
Haider, J., y Sundin, O. (2019). Invisible Search and Online Search Engines: The Ubiquity of Search in Everyday Life (1.a ed.). Routledge. https://doi.org/10.4324/9780429448546
Hartigan, J. A., y Wong, M. A. (1979). Algorithm AS 136: A K-Means Clustering Algorithm. Journal of the Royal Statistical Society, 28(1), 100. https://doi.org/10.2307/2346830
Hill, B., y Shaw, A. (2020). The Most Important Laboratory for Social Scientific and Computing Research in History. En J. Reagle y J. Koerner (eds.), Wikipedia @ 20: Stories of an Incomplete Revolution. The MIT Press. https://doi.org/10.7551/mitpress/12366.003.0015 PMCid:PMC6980988
Hube, C., Fischer, F., Jäschke, R., Lauer, G., y Thomsen, M. R. (2017). World Literature According to Wikipedia: Introduction to a DBpedia-Based Framework. arXiv. Disponible en: http://arxiv.org/abs/1701.00991.
Jemielniak, D., y Wilamowski, M. (2017). Cultural diversity of quality of information on Wikipedias. Journal of the Association for Information Science and Technology, 68(10), 2460-2470. https://doi.org/10.1002/asi.23901
Lemus-Rojas, M., y Pintscher, L. (2018). Wikidata and Libraries: Facilitating Open Knowledge. En M. Proffitt (ed.), Leveraging Wikipedia: Connecting Communities of Knowledge, 143-158. IL: ALA Editions. Disponible en: https://scholarworks.iupui.edu/handle/1805/16690.
Lewoniewski, W., Węcel, K., y Abramowicz, W. (2019). Multilingual Ranking of Wikipedia Articles with Quality and Popularity Assessment in Different Topics. Computers, 8(3), 60. https://doi.org/10.3390/computers8030060
Minguillón, J., Lerga, M., Aibar, E., Lladós-Masllorens, J., y Meseguer-Artola, A. (2017). Semi-automatic generation of a corpus of Wikipedia articles on science and technology. El Profesional de la Información, 26(5), 995-1004. https://doi.org/10.3145/epi.2017.sep.20
Miquel-Ribé, M. (2019). The Sum of Human Knowledge? Not in One Wikipedia Language Edition. Wikipedia@20. Disponible en: https://wikipedia20.mitpress.mit.edu/pub/26ke5md7/release/15.
Miquel-Ribé, M., y Laniado, D. (2018). Wikipedia Culture Gap: Quantifying Content Imbalances Across 40 Language Editions. Frontiers in Physics, 6, Article 54. https://doi.org/10.3389/fphy.2018.00054
Miquel-Ribé, M., y Laniado, D. (2021). The Wikipedia Diversity Observatory: helping communities to bridge content gaps through interactive interfaces. Journal of Internet Services and Applications, 12(1), 10. https://doi.org/10.1186/s13174-021-00141-y
Moretti, F. (2013). Distant reading. Verso.
Muñoz Rico, M., García Rodríguez, A., y Cordón García, J. A. (2020). Hacia una teoría del bestseller canónico: la constitución de un modelo estructural. Revista General de Información y Documentación, 30(1), 149-165. https://doi.org/10.5209/rgid.69673
Nielsen, F. Å. (2019). Wikipedia research and tools: Review and comments. Disponible en: http://www2.imm.dtu.dk/pubdb/edoc/imm6012.pdf.
Piscopo, A., y Simperl, E. (2018). Who Models the World? Collaborative Ontology Creation and User Roles in Wikidata. Proceedings of the ACM on Human-Computer Interaction, 2, 1-18. https://doi.org/10.1145/3274410
Reagle, J., y Koerner, J. (eds.). (2020). Wikipedia @ 20: Stories of an Incomplete Revolution. The MIT Press. https://doi.org/10.7551/mitpress/12366.001.0001
Reznik, I., y Shatalov, V. (2016). Hidden revolution of human priorities: An analysis of biographical data from Wikipedia. Journal of Informetrics, 10(1), 124-131. https://doi.org/10.1016/j.joi.2015.12.002
Rousseeuw, P. J. (1987). Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20, 53-65. https://doi.org/10.1016/0377-0427(87)90125-7
Shatnawi, R. (2015). Deriving metrics thresholds using log transformation. Journal of Software: Evolution and Process, 27(2), 95-113. https://doi.org/10.1002/smr.1702
Shenoy, K., Ilievski, F., Garijo, D., Schwabe, D., y Szekely, P. (2022). A study of the quality of Wikidata. Journal of Web Semantics, 72, 100679. https://doi.org/10.1016/j.websem.2021.100679
Skiena, S. S., y Ward, C. (2014). Who's bigger? where historical figures really rank. Cambridge University Press. https://doi.org/10.1017/CBO9781139649605 PMid:23690603 PMCid:PMC3677463
Venuti, L. (2008). Translation, interpretation, canon formation. En A. Lianeri y V. Zajko (eds.), Translation and the Classic: Identity as Change in the History of Culture, 27-51. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780199288076.003.0002
Zschirnt, C. (2011). Libros: todo lo que hay que saber (1a ed). Taurus
Published
How to Cite
Issue
Section
License
Copyright (c) 2023 Consejo Superior de Investigaciones Científicas (CSIC)

This work is licensed under a Creative Commons Attribution 4.0 International License.
© CSIC. Manuscripts published in both the print and online versions of this journal are the property of the Consejo Superior de Investigaciones Científicas, and quoting this source is a requirement for any partial or full reproduction.
All contents of this electronic edition, except where otherwise noted, are distributed under a Creative Commons Attribution 4.0 International (CC BY 4.0) licence. You may read the basic information and the legal text of the licence. The indication of the CC BY 4.0 licence must be expressly stated in this way when necessary.
Self-archiving in repositories, personal webpages or similar, of any version other than the final version of the work produced by the publisher, is not allowed.








