<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD with OASIS Tables with MathML3 v1.1 20151215//EN" "JATS-journalpublishing-oasis-article1-mathml3.dtd">
<article article-type="research-article" dtd-version="1.1" xml:lang="en" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
	<front>
		<journal-meta>
			<journal-id journal-id-type="publisher-id">REDC</journal-id>
			<journal-title-group>
				<journal-title>Revista Espa&#xf1;ola de Documentaci&#xf3;n Cient&#xed;fica</journal-title>
				<abbrev-journal-title abbrev-type="publisher">Rev. esp. doc. cient.</abbrev-journal-title>
			</journal-title-group>
			<issn publication-format="print">0210-0614</issn>
			<issn publication-format="electronic">1988-4621</issn>
			<issn-l>0210-0614</issn-l>
			<publisher>
				<publisher-name>Consejo Superior de Investigaciones Cient&#xed;ficas</publisher-name>
			</publisher>
		</journal-meta>
		<article-meta>
			<article-id pub-id-type="publisher-id">redc.2022.4.1917</article-id>
			<article-id pub-id-type="doi">10.3989/redc.2022.4.1917</article-id>
			<article-categories>
				<subj-group subj-group-type="heading">
					<subject>ESTUDIOS / RESEARCH STUDIES</subject>
				</subj-group>
			</article-categories>
			<title-group>
				<article-title>Automatic indexing of scientific articles on Library and Information Science with SISA, KEA and MAUI</article-title>
				<trans-title-group xml:lang="es">
					<trans-title>Indizaci&#xf3;n autom&#xe1;tica de art&#xed;culos cient&#xed;ficos sobre Biblioteconom&#xed;a y Documentaci&#xf3;n con SISA, KEA y MAUI</trans-title>
				</trans-title-group>
			</title-group>
			<contrib-group>
				<contrib contrib-type="author" corresp="yes">
					<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-7175-3099</contrib-id>
					<name>
						<surname>Gil-Leiva</surname>
						<given-names>Isidoro</given-names>
					</name>
					<email xlink:href="isgil@um.es">isgil@um.es</email>
					<aff id="aff1"><institution>Universidad de Murcia</institution></aff>
				</contrib>
				<contrib contrib-type="author">
					<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-2975-766X</contrib-id>
					<name>
						<surname>D&#xed;az Ortu&#xf1;o</surname>
						<given-names>Pedro</given-names>
					</name>
					<email xlink:href="diazor@um.es">diazor@um.es</email>
					<aff id="aff2"><institution>Universidad de Murcia</institution></aff>
				</contrib>
				<contrib contrib-type="author">
					<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-9880-8678</contrib-id>
					<name>
						<surname>Fernandes Corr&#xea;a</surname>
						<given-names>Renato</given-names>
					</name>
					<email xlink:href="renato.correa@ufpe.br">renato.correa@ufpe.br</email>
					<aff id="aff3"><institution>Universidade Federal de Pernambuco</institution></aff>
				</contrib>
			</contrib-group>
			<pub-date pub-type="epub">
				<day>03</day>
				<month>10</month>
				<year>2022</year>
			</pub-date>
			<pub-date pub-type="collection">
				<month>12</month>
				<year>2022</year>
			</pub-date>
			<volume>45</volume>
			<issue>4</issue>
			<elocation-id>e338</elocation-id>
			<history>
				<date date-type="received">
					<day>02</day>
					<month>09</month>
					<year>2021</year>
				</date>
				<date date-type="rev-recd">
					<day>26</day>
					<month>10</month>
					<year>2021</year>
				</date>
				<date date-type="accepted">
					<day>02</day>
					<month>11</month>
					<year>2021</year>
				</date>
				<date date-type="pub">
					<day>18</day>
					<month>10</month>
					<year>2022</year>
				</date>
			</history>
			<permissions>
				<copyright-statement>&#xa9;2022 CSIC</copyright-statement>
				<copyright-year>2022</copyright-year>
				<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
					<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) License.</license-p>
				</license>
			</permissions>
			<self-uri xlink:href="http://redc.revistas.csic.es/index.php/redc/article/view/XXXX/XXXX"/>
			<abstract>
				<title>Abstract</title>
				<p>This article evaluates the SISA (Automatic Indexing System), KEA (Keyphrase Extraction Algorithm) and MAUI (Multi-Purpose Automatic Topic Indexing) automatic indexing systems to find out how they perform in relation to human indexing. SISA&#x2019;s algorithm is based on rules about the position of terms in the different structural components of the document, while the algorithms for KEA and MAUI are based on machine learning and the statistical features of terms. For evaluation purposes, a document collection of 230 scientific articles from the <italic>Revista Espa&#xf1;ola de Documentaci&#xf3;n Cient&#xed;fica</italic> published by the Consejo Superior de Investigaciones Cient&#xed;ficas (CSIC) was used, of which 30 were used for training tasks and were not part of the evaluation test set. The articles were written in Spanish and indexed by human indexers using a controlled vocabulary in the InDICES database, also belonging to the CSIC. The human indexing of these documents constitutes the baseline or golden indexing, against which to evaluate the output of the automatic indexing systems by comparing terms sets using the evaluation metrics of precision, recall, F-measure and consistency. The results show that the SISA system performs best, followed by KEA and MAUI.</p>
			</abstract>
			<trans-abstract xml:lang="es">
				<title>Resumen</title>
				<p>Este art&#xed;culo eval&#xfa;a los sistemas de indizaci&#xf3;n autom&#xe1;tica SISA (Automatic Indexing System), KEA (Keyphrase Extraction Algorithm) y MAUI (Multi-Purpose Automatic Topic Indexing) para averiguar c&#xf3;mo funcionan en relaci&#xf3;n con la indizaci&#xf3;n realzada por especialistas. El algoritmo de SISA se basa en reglas sobre la posici&#xf3;n de los t&#xe9;rminos en los diferentes componentes estructurales del documento, mientras que los algoritmos de KEA y MAUI se basan en el aprendizaje autom&#xe1;tico y las frecuencia estad&#xed;stica de los t&#xe9;rminos. Para la evaluaci&#xf3;n se utiliz&#xf3; una colecci&#xf3;n documental de 230 art&#xed;culos cient&#xed;ficos de la <italic>Revista Espa&#xf1;ola de Documentaci&#xf3;n Cient&#xed;fica,</italic> publicada por el Consejo Superior de Investigaciones Cient&#xed;ficas (CSIC), de los cuales 30 se utilizaron para tareas formativas y no formaban parte del conjunto de pruebas de evaluaci&#xf3;n. Los art&#xed;culos fueron escritos en espa&#xf1;ol e indizados por indizadores humanos utilizando un vocabulario controlado en la base de datos InDICES, tambi&#xe9;n perteneciente al CSIC. La indizaci&#xf3;n humana de estos documentos constituye la referencia contra la cual se eval&#xfa;a el resultado de los sistemas de indizaci&#xf3;n autom&#xe1;ticos, comparando conjuntos de t&#xe9;rminos usando m&#xe9;tricas de evaluaci&#xf3;n de precisi&#xf3;n, recuperaci&#xf3;n, medida F y consistencia. Los resultados muestran que el sistema SISA funciona mejor, seguido de KEA y MAUI. </p>
			</trans-abstract>
			<kwd-group>
				<kwd>automatic indexing</kwd>
				<kwd>automatic indexing systems</kwd>
				<kwd>SISA</kwd>
				<kwd>KEA</kwd>
				<kwd>MAUI</kwd>
				<kwd>indexing assessment</kwd>
			</kwd-group>
			<kwd-group xml:lang="es">
				<kwd>indizaci&#xf3;n autom&#xe1;tica</kwd>
				<kwd>sistemas de indizaci&#xf3;n autom&#xe1;tica</kwd>
				<kwd>SISA</kwd>
				<kwd>KEA</kwd>
				<kwd>MAUI</kwd>
				<kwd>evaluaci&#xf3;n de indizaci&#xf3;n</kwd>
			</kwd-group>
			<counts>
				<fig-count count="4"/>
				<table-count count="8"/>
				<equation-count count="0"/>
				<ref-count count="59"/>
				<page-count count="18"/>
			</counts>
		</article-meta>
	</front>
	<body>
		<sec id="sec1" sec-type="intro">
			<label>1.</label>
			<title>Introduction</title>
			<p>Document production has grown exponentially since the 1950s. Not only has publication increased significantly but increasing numbers of documents are processed and disseminated, giving rise to a need for more efficient and faster information processing systems. For instance, in large bibliographic databases such as Scopus, some three million documents are incorporated each year. Furthermore, libraries are integrating large quantities of electronic books, papers, theses and dissertations, which they are unable to process adequately in order to make them accessible through a catalogue or institutional repository. In the documentary management of the aforementioned information systems, the indexing of content to facilitate access plays a fundamental role. </p>
			<p>
				<xref ref-type="bibr" rid="B30">ISO standard 5963-1985</xref> defines indexing as &#x201c;The act of describing or identifying a document in terms of its subject content&#x201d;. To this it may be added that, on occasion, concepts are normalized and controlled by controlled vocabulary, as otherwise it would be natural language indexing, and likewise that indexing is carried out - be it consciously or unconsciously - according to the users&#x2019; information needs in order to convert these (in natural or controlled language) into a search query. Hence, indexing constitutes an essential process for storing documents and may also be so for retrieving information if the result of the indexing (keywords, descriptors, subjects indexing) is used later for retrieval. Indexing is therefore the cornerstone of document management systems as it is essential to represent the contents of documents and to facilitate their subsequent retrieval.</p>
			<p>Programs to increase workflow performance were first created in the late 1950s and the terminology used in the literature to refer to the process of making indexing automatic is varied, although the most used term is &#x201c;automatic indexing&#x201d;. The definition of automatic indexing can derive from three perspectives: a) computer aided indexing during storage; b) semi-automatic indexing; and c) automatic indexing (<xref ref-type="bibr" rid="B17">Gil-Leiva, 2008: 320</xref>). From the 1970s to the present, several automatic indexing programs have been developed. Without intending to be exhaustive, some of them are:</p>
			<list list-type="bullet">
				<list-item>
					<p>MAI (<xref ref-type="bibr" rid="B34">Klingbiel, 1973</xref>; <xref ref-type="bibr" rid="B54">Silvester et al., 1994</xref>). Machine-aided indexing (MAI) was developed at the US Defense Documentation Center.</p>
				</list-item>
				<list-item>
					<p>The SMART project by <xref ref-type="bibr" rid="B46">Salton (1989</xref>, <xref ref-type="bibr" rid="B47">1991)</xref>. One of the first to incorporate advances produced by the automatic processing of natural language, it brings in tools to extract word roots, thesaurus, morphological or syntactic analyzers.</p>
				</list-item>
				<list-item>
					<p>CLARIT (<xref ref-type="bibr" rid="B11">Evans, 1990</xref>; <xref ref-type="bibr" rid="B12">Evans et al., 1991a</xref>; <xref ref-type="bibr" rid="B13">Evans et al., 1991b</xref>). CLARIT (Computational-Linguistic Approaches to Indexing and Retrieval of Text) used a lexicon for general English which consists of approximately 100,000 root forms and hyphenated phrases, tagged for syntactic category and irregular morphological variation; a morphological analyzer; a lexical disambiguator; a noun phrase grammar; and various indexing algorithms, such as the ranking indexing terms.</p>
				</list-item>
				<list-item>
					<p>SIMPR (<xref ref-type="bibr" rid="B31">Karetnyk, Karlsson &amp; Smart, 1991</xref>). In the SIMPR (Structured Information Management: Processing and Retrieval) project, output from a morphological analyzer was used to provide the input to indexing software, from which some indexing terms are finally obtained, and are validated manually.</p>
				</list-item>
				<list-item>
					<p>SAPHIRE (<xref ref-type="bibr" rid="B22">Hersh &amp; Greenes, 1990</xref>; <xref ref-type="bibr" rid="B23">Hersh et al., 1991</xref>). SAPHIRE (Semantic and Probabilistic Heuristic Information Retrieval Environment) is a project in the field of biomedicine and the Medline database.</p>
				</list-item>
				<list-item>
					<p>Indexing Initiative/Text Indexer-MTI (<xref ref-type="bibr" rid="B25">Humphrey &amp; Miller, 1987</xref>; <xref ref-type="bibr" rid="B26">Humphrey, 1999</xref>; <xref ref-type="bibr" rid="B27">Humphrey et al., 2006</xref>; <xref ref-type="bibr" rid="B1">Aronson et al., 2000</xref>; <xref ref-type="bibr" rid="B40">Mork et al., 2017</xref>). Projects whose mission among library teams is to explore indexing methodologies to ensure the quality and currency of the document collections of the US National Library of Medicine (NLM). The NLM Medical Text Indexer (MTI) is the core product of this project and has been providing automated indexing recommendations since 2002 to the present day.</p>
				</list-item>
				<list-item>
					<p>CAIT, Luxid, AgNIC (<xref ref-type="bibr" rid="B29">Irving, 1997</xref>; <xref ref-type="bibr" rid="B45">Salisbury &amp; Smith, 2014</xref>). At the United States National Library of Agriculture, various projects have been implemented, including CAIT (Computer-Assisted Indexing Tutor), a program which sought to enhance the quality of indexing and the training of new indexers (<xref ref-type="bibr" rid="B29">Irving, 1997</xref>). This was followed by the acquisition of the Luxid indexing software from the TEMIS company in 2011 to considerably increase the production capacity of the six indexers from 75,000 articles a year to about 300,000. In addition, more recently there was AgNIC (Agriculture Network Information Collaborative), which bases automatic indexing on the Thesaurus, as Luxid does.</p>
				</list-item>
				<list-item>
					<p>CISMeF (<xref ref-type="bibr" rid="B7">Chebil et al., 2012</xref>). The <italic>Catalogue et Index des Sites M&#xe9;dicaux de langue Fran&#xe7;aise</italic>, a system implemented for the automatic indexing of medical information resources.</p>
				</list-item>
				<list-item>
					<p>Annif (<xref ref-type="bibr" rid="B57">Suominen, 2019</xref>). Developed by the National Library of Finland, Annif is an open source tool and microservice for automated subject indexing based on a combination of existing natural language processing and machine learning tools, combining multiple approaches and existing open source algorithms.</p>
				</list-item>
			</list>
			<p>The scientific literature on automatic indexing is extensive and mainly devoted to describing and evaluating prototypes. There is a long tradition of evaluation and competition of systems and prototypes, starting in the 1960s with the Cranfield experiments. Subsequently, from 1992 onwards, the Text Retrieval Evaluation Conference (TREC) initiative began to gain momentum, becoming the largest annual competition centring on information retrieval. The CLEF Initiative (Conference and Labs of the Evaluation Forum, formerly known as the Cross-Language Evaluation Forum) also emerged from this event and also promotes innovation and the development of information access systems in Europe on an annual basis. A third important evaluation-competition initiative is SemEval (Semantic Evaluation), which has focused on the meaning of language since 1998. Each of these three initiatives has featured spaces dedicated to indexing and automatic indexing. In a way, the work presented here is imbued with the spirit of the above initiatives in the sense of an evaluation-competition based on a test set, golden indexing and metrics.</p>
			<p>Furthermore, since the early 2000s, scientific publications have been impregnated with the Semantic Web (which uses XML [eXtensible Markup Language] , RDF [Resource Description Framework] and OWL [Ontology Web Language] ), with the addition in 2012 of the Journal Article Tag Suite (JATS) standard for describing the formal, textual and graphic content of scientific articles. According to the <xref ref-type="bibr" rid="B48">2020 Scholastica</xref> report, 35% of scientific journal publishers are already using the XML or XML JATS format. The <italic>Revista Espa&#xf1;ola de Documentaci&#xf3;n Cient&#xed;fica</italic> itself in 2013 extended its publication format from PDF to HTML and XML. The Redalyc platform has been using XML JATS extensively since 2016. Thus, this combination of Open Science and Semantic Web has given rise to the so-called semantic publishing that facilitates dissemination, exchange, reuse of information, retrieval and automatic information processing.</p>
			<p>The work presented here uses the web version of SISA, which represents a significant advance over the previous version in Windows. In some way, SISA imitates the work of a human indexer during the reading phase of the documents. Besides, this approach, based on the position occupied by terms in the documents and aligned with the increasing number of scientific articles published using XML JATS, makes SISA a different system to other automatic indexing prototypes. KEA and MAUI, for example, use a machine learning model that starts from a training set with documents and their indexing terms to then predict the indexing terms for new documents using a Na&#xef;ve Bayes classifier and Decision Tree classifier respectively. The machine learning model requires input consisting of statistical features of terms in a document, and its output comprises weighting of terms. A trained indexing model performs the term-weighting process, based on calculating characteristics for each term, such as occurrence, position, size, the probability of being a keyword and semantic relations. KEA uses the two first feature types, while MAUI can use all of them.</p>
			<p>Therefore, the main objective of this article is to find out how the SISA algorithm (based on the position that the terms occupy in the different parts of the document) performs compared to the KEA and MAUI algorithms, which are machine-learning based. This work may also enable further improvement of SISA through the feedback obtained from studying and analysing the human indexing of the InDICES database, which for the purposes of this study has been taken as a benchmark, as well as from analysing the outputs of KEA and MAUI. Thus, in addition to highlighting the importance of the methodology on which SISA is based compared to other indexing systems, this research seeks to answer specific questions such as: what is the average time taken to index a scientific article in SISA, KEA and MAUI compared to a human indexer?; what is the average number of terms assigned by these systems compared to human indexing?; and what maximum and average F-measure and consistency indices can these automatic indexing systems achieve compared to human indexing?</p>
		</sec>
		<sec id="sec2" sec-type="materials|methods">
			<label>2.</label>
			<title>Materials and method</title>
			<p>Our study consists of six main parts: a) selection of the prototypes to carry out the comparison-evaluation with SISA; b) collection of the sources and data to carry out the evaluation (document collection, controlled vocabulary and human-assigned golden indexing); c) selection of the evaluation metrics; d) configuration and indexing of the document collection using the three prototypes; e) application of the evaluation metrics to obtain the precision, recall, F-measure and consistency indices; f) analysis and discussion of the results. We provide more details of these below.</p>
			<sec id="sec2.1">
				<label>2.1</label>
				<title>Automatic indexing systems</title>
				<p>SISA, KEA and MAUI are three automatic information processing tools specifically for automatic document indexing. The reasons that led to KEA and MAUI being selected include the following: all three have their origins in academic settings; they are free, open-source software tools; they all have in common the option of performing automatic indexing by assignment, which means terms being assigned from a given controlled vocabulary; and, finally, KEA and MAUI use an automatic indexing model based on a machine learning algorithm and statistical features of terms.</p>
				<p>As has been pointed out above, the evaluation-comparison of these three indexing tools will make it possible to examine in depth two different approaches to the automation of indexing, and therefore to find out which of them is closer to human indexing. One approach is based on the position that terms occupy in the different parts of documents (keeping in mind here that the ISO standard for document indexing itself indicates to which parts of documents indexers should pay attention when extracting and assigning terms) while the other is a machine learning approach (based on a model which has been trained through a small collection of documents and the golden indexing of those documents to predict the indexing terms for a new set of documents). The three automatic indexing systems are presented below.</p>
				<sec id="sec2.1.1">
					<title>SISA</title>
					<p>The conceptual development of SISA began in the mid-1990s. SISA is an automatic indexing system developed in JAVA to extract information from documents. It performs automatic indexing of scientific articles, legislation (laws, decrees) and judicial sentences, although to date the document typology it works with most is scientific articles. It processes documents in TXT, HTML or XML formats. It also uses a controlled vocabulary in TXT or SKOS format.</p>
					<p>SISA is available on the Internet to users through a password (<ext-link ext-link-type="uri" xlink:href="http://fcd1.inf.um.es:8080/portal/">http://fcd1.inf.um.es:8080/portal/</ext-link>). SISA processes documents written in Spanish, Portuguese and English. It processes texts with tags that indicate each of the components of the documents depending on document typology: for example, for articles it processes title, abstract, keywords, headings, first paragraph, table title, chart title, conclusions and references. Currently, it directly processes XML articles from the <italic>Revista Espa&#xf1;ola de Documentaci&#xf3;n Cient&#xed;fica</italic> and HTML articles derived from the JATS standard of the Redalyc platform. If the documents do not contain the necessary marks, the software has a wizard to assist with labelling.</p>
					<p>These baseline fundamentals of SISA appear to align with the current development of the JATS standard for scientific articles. Marks and a set of rules based on heuristic (positioning - titles, abstracts, author keywords, headings, first paragraphs of headings, table titles, graph titles, conclusions and references) and statistical methods (frequency, TF-IDF) are the hallmark of this software.</p>
					<p>Ever since its first version in 2002, SISA has offered the option of processing using automatic indexing or semi-automatic indexing. With semi-automatic indexing, users can edit the proposed indexing for each document through a wizard that shows each indexing term or phrase in its context. The successive tasks for automatic indexing of a document are the following: labelling the documents or uploading pre-marked documents; processing (applying stemming, calculating the TF-IDF and IDF and recording the place where terms and phrases appear); and then indexing the documents, taking into account the activated rules. It also includes an evaluation module that measures recall, precision and F-measure in information retrieval (<xref ref-type="bibr" rid="B17">Gil Leiva, 2008</xref>, <xref ref-type="bibr" rid="B18">2017a</xref>).</p>
					<p>SISA has been used in several experiments. <xref ref-type="bibr" rid="B36">Lima and Boccato (2009)</xref> used SISA to evaluate the performance of descriptors in manual, automatic and semi-automatic indexing processes. In the study by <xref ref-type="bibr" rid="B56">Souza-Rocha and Gil-Leiva (2016)</xref>, a comparison was performed of the indexing of the same test set by SISA and by the PyPLN platform developed at the Getulio Vargas Foundation&#x2019;s School of Applied Mathematics in Rio de Janeiro, Brazil. Available functions of the platform include part-of-speech tagging, word and sentence level statistics, n-grams extraction and word matching. <xref ref-type="bibr" rid="B18">Gil-Leiva (2017a)</xref> investigated the capabilities of SISA by comparing its indexing of a test set of scientific articles on Agriculture with their indexing by the Agricola, WOS and SCOPUS databases. Later work by <xref ref-type="bibr" rid="B19">Gil-Leiva (2017b)</xref> aimed to determine which rules (rules of position of the terms in the document or TF-IDF rules) provide the best indexing terms, using SISA to obtain the automatic indexing of 200 scientific articles on fruit growing written in Portuguese. In addition, more recently the indexing of the desktop version of SISA has been compared with MAUI (<xref ref-type="bibr" rid="B52">Silva &amp; Correa, 2020</xref>; <xref ref-type="bibr" rid="B53">Silva et al., 2020</xref>).</p>
					<p>The work presented here uses the web version of SISA, which represents a significant advance over the Windows version in terms of processes, functionalities, document input formats and the incorporation of statistical rules, but mainly due to the greater number of heuristic rules based on the position of the terms in the documents. At the current time, SISA can process tagged PDF, TXT and XML documents published by the Revista Espa&#xf1;ola de Documentaci&#xf3;n Cient&#xed;fica and XML JATS articles produced and published by the editors of Redalyc.</p>
				</sec>
				<sec id="sec2.1.2">
					<title>KEA</title>
					<p>The Keyphrase Extraction Algorithm (KEA) is a project developed by the &#x201c;Digital Library&#x201d; and &#x201c;Machine Learning&#x201d; research groups at the University of Waikato. New Zealand Digital Library Project members have developed a range of practical software packages in the course of their research. The home page is available at <ext-link ext-link-type="uri" xlink:href="http://community.nzdl.org/kea/">http://community.nzdl.org/kea/</ext-link>.</p>
					<p>KEA is an algorithm for automatically extracting keyphrases from text documents, assuming that keyphrases provide semantic metadata that summarize and characterize documents. Implemented in Java, it is platform-independent and an open-source software distributed under the GNU General Public License. The system is simple, robust and publicly available. It can be used either for free indexing or for indexing with a controlled vocabulary (Thesaurus, Subject Headings). The KEA website provides examples of its use in different domains with a thesaurus: the AGROVOC thesaurus, Medical Subject Headings (MESH) thesaurus and High Energy Physics thesaurus (HEP).</p>
					<p>KEA includes a machine learning component. In experiments using KEA, it is utilized to split the document collection into a training and a test set. The training set is applied to train a machine learning model to classify candidate keyphrases. The test set makes it possible to evaluate the effectiveness of KEA in terms of how many author-assigned keyphrases are correctly identified.</p>
					<p>KEA performs the following processing steps: it identifies candidate keyphrases using lexical methods; it calculates the feature values for each candidate; and it uses a machine learning algorithm to predict which candidates are good keyphrases. The machine learning scheme first builds a prediction model using training documents with known keyphrases, and then it uses the model to find keyphrases in new documents (<xref ref-type="bibr" rid="B59">Witten et al., 1999</xref>). The machine learning model is constructed automatically from these labelled training examples using the WEKA machine learning workbench. KEA (<xref ref-type="bibr" rid="B15">Frank et al., 1999</xref>) uses the Na&#xef;ve Bayes classifier, which implicitly assumes that the features are independent of each other. </p>
					<p>
						<xref ref-type="bibr" rid="B38">Medelyan (2005)</xref> extended KEA into a new version called KEA++. The indexing process is extended through analyzing semantic information about a document&#x2019;s terms, i.e. the relationship between the terms and their relationship to other terms in the thesaurus. It combines two approaches into a single process whereby terms and phrases are extracted from documents (keyphrase extraction, phrases analyzed to select the most representative ones) but need to be part of a controlled vocabulary (assignment of keyphrases, documents are classified into a pre-defined number of categories corresponding to descriptors). </p>
					<p>As KEA++ is a supervised learning approach, it involves two phases: training and testing. It builds a learned model by applying a Naive Bayes algorithm using training data labelled with thesaurus terms. The extraction phase uses the learned model to identify, from a thesaurus, the most significant keyphrases based on certain properties (features) and assign them to the test documents. After computing the feature values for the training set, the model built is used to extract keyphrases from new documents. Each candidate keyphrase is marked as a positive or negative example, depending on whether users have assigned it as a keyphrase or tag to the corresponding document (<xref ref-type="bibr" rid="B39">Medelyan, 2009</xref>).</p>
					<p>KEA has been used in a number of different experiments. <xref ref-type="bibr" rid="B10">El-Haj et al. (2013)</xref> experimented with KEA in a different domain to the examples available on the project website, when they used KEA to examine the quality of the automated indexing process based on a controlled vocabulary called the Humanities and Social Sciences Electronic Thesaurus (HASSET). <xref ref-type="bibr" rid="B32">Khan et al. (2011)</xref> and <xref ref-type="bibr" rid="B28">Irfan et al. (2014)</xref> set out to improve the functioning of KEA++ based on more efficient exploitation of the existing hierarchical relationships in the domain vocabularies used by the system. <xref ref-type="bibr" rid="B58">Wang et al., (2015)</xref> designed DIKEA with the idea of improving the performance of KEA++ in two ways: by extracting keywords from documents and by not depending on a specific vocabulary. Other experiments in which KEA has been used include those by <xref ref-type="bibr" rid="B9">Duwairi and Hedaya (2016)</xref> for documents with Arabic news, those by <xref ref-type="bibr" rid="B2">Akhtar et al. (2017)</xref> to create a hierarchy of keywords extracted from documents and the research by <xref ref-type="bibr" rid="B20">Gopan et al. (2020)</xref>, who compared the keyword extraction algorithms of KEA, TextRank and PositionRank.</p>
				</sec>
				<sec id="sec2.1.3">
					<title>MAUI</title>
					<p>Olena Medelyan developed the MAUI (Multi-purpose Automatic Topic Indexing) as part of her doctoral project, under the supervision of Ian H. Witten and Eibe Frank, at the Department of Computer Science at the University of Waikato, New Zealand, in 2009.</p>
					<p>To represent the content of the documents, the software uses statistical and linguistic methods for extracting and weighting terms. The software extracts terms, or generates candidates, by identifying n-grams not containing punctuation marks and not beginning or ending with stopwords, after normalization/conflation by stemming. A trained supervised indexing model performs the term-weighting process, based on the calculation of characteristics for each term, such as occurrence, position, size, the probability of being a keyword and semantic relations.</p>
					<p>The software allows users to perform the following tasks (<xref ref-type="bibr" rid="B39">Medelyan, 2009</xref>): assigning terms with a controlled vocabulary or thesaurus; subject indexing; topic indexing with Wikipedia terms; keyword extraction; terminology extraction; automatic markup, terminology extraction and semi-automatic topic indexing; and keyword extraction from the text using a controlled vocabulary as a source of terms.</p>
					<p>MAUI uses a machine learning algorithm to generate a model for selecting index terms based on the intellectual indexing of a set of documents. For this reason, it requires input consisting of a training set, made up of documents and respective terms for human indexing.</p>
					<p>To train the indexing model, the data entry requirements for MAUI are the following:</p>
					<list list-type="bullet">
						<list-item>
							<p>Thesaurus in SKOS format: a text file containing the authorized terms and their non-preferred terms, as well as the semantic relations between authorized terms;</p>
						</list-item>
						<list-item>
							<p>Stemmer of words for the language used in the text of the documents: a software component that reduces the words to an approximation of their stem;</p>
						</list-item>
						<list-item>
							<p>Pre-established list of stopwords in the document text language: a text file containing words without thematic meaning for elimination, such as connectors and articles;</p>
						</list-item>
						<list-item>
							<p>A training set: a set of documents and respective indexing terms for training a machine learning model;</p>
						</list-item>
						<list-item>
							<p>All input files, including the texts for indexing, must be in text format with UTF-8 without BOM charset.</p>
						</list-item>
					</list>
					<p>After the model has been trained, the data entry requirements for automatic indexing of a text document using MAUI processing are: the full text of the document to be indexed; the trained model; a list of stopwords; a stemmer; and a thesaurus or controlled vocabulary.</p>
					<p>According to <xref ref-type="bibr" rid="B39">Medelyan (2009)</xref>, MAUI performs indexing through the following processing steps:</p>
					<list list-type="order">
						<list-item>
							<p>Generation of candidate topics - extraction of candidate terms for indexing;</p>
						</list-item>
						<list-item>
							<p>Calculation of characteristics - calculation of characteristics for candidate terms;</p>
						</list-item>
						<list-item>
							<p>Construction of the indexing model - training of an indexing model considering the terms assigned by the indexers to each document in the training set;</p>
						</list-item>
						<list-item>
							<p>Application of the learned model to select topics for other documents - application of the trained indexing model to propose indexing terms to other documents.</p>
						</list-item>
					</list>
					<p>The quality of the training set provided for the machine learning algorithm is the key to better quality in automatic indexing, whether the indexing process involves automatically extracting keywords or assigning controlled vocabulary descriptors.</p>
					<p>In addition to the experiments by <xref ref-type="bibr" rid="B39">Medelyan (2009)</xref>, several other studies have applied MAUI. For instance, MAUI was tested to extract keyphrases from scientific texts in English at the SEMEVAL 2010 conference (<xref ref-type="bibr" rid="B33">Kim et al., 2013</xref>), while <xref ref-type="bibr" rid="B41">Mynarz and &#x160;kuta (2010)</xref> used MAUI together with other applications to implement an automatic indexing system for Czech grey literature and <xref ref-type="bibr" rid="B55">Sinkkil&#xe4;, Suominen and Hyv&#xf6;nen (2011)</xref> tested MAUI and three different stemming and derivation tools on Finnish texts to see which obtained the best indexing terms. The report by <xref ref-type="bibr" rid="B50">Shams and Mercer (2012a)</xref> looked at the indexing performance of MAUI when paired with a text extraction procedure called text denoising, and MAUI was also applied to extract keyphrases from scientific texts written in Spanish (<xref ref-type="bibr" rid="B5">Aquino &amp; Lanzarini, 2015</xref>). <xref ref-type="bibr" rid="B53">Silva, Correa and Gil-Leiva (2020)</xref> implemented MAUI to assign thesaurus terms to scientific texts written in Portuguese, in addition to comparing MAUI indexing with indexing from a desktop version of SISA. In the research by <xref ref-type="bibr" rid="B20">Gopan et al. (2020)</xref>, MAUI&#x2019;s keyword extraction algorithm was compared with KEA, TextRank and PositionRank.</p>
				</sec>
			</sec>
			<sec id="sec2.2">
				<label>2.2</label>
				<title>Document collection</title>
				<p>A document collection was created, made up of 230 scientific articles published in the <italic>Revista Espa&#xf1;ola de Documentaci&#xf3;nCient&#xed;fica</italic> of the Consejo Superior de Investigaciones Cient&#xed;ficas (CSIC- the Spanish National Research Council). This long-running journal began to publish its articles in PDF, HTML and XML formats in 2013, using an XML format with DTD NLM evolved in the current JATS standard. The criteria for selecting the 230 documents that made up the document collection were simple: articles published in the &#x201c;Studies&#x201d; section from 2013 to 2020 that were written in Spanish. To complete the collection, the &#x201c;Note and experiences&#x201d; section, which publishes scientific articles of similar relevance and formal structure, was also included. </p>
				<p>The training set for KEA and MAUI software comprised 30 randomly selected articles, since SISA does not require an initial machine learning phase for the system. The remaining 200 articles constituted the test set that was applied to evaluate the automatic indexing performed by SISA, KEA and MAUI and later, to carry out the comparative studies. SISA used the 200 articles in XML format while KEA and MAUI worked with the extracted text in TXT format.</p>
			</sec>
			<sec id="sec2.3">
				<label>2.3</label>
				<title>Controlled vocabulary</title>
				<p>A controlled vocabulary composed of 10,981 terms was used, of which 8,214 were preferred terms and 2,767 non-preferred terms. The controlled vocabulary was provided by those responsible for the InDICEs database of the Consejo Superior de Investigaciones Cient&#xed;ficas. In this database, the controlled vocabulary is utilized for human indexing of scientific journals published in Spain, as well as conference proceedings on Library and Information Science. Therefore, the 230 articles that make up the test collection were indexed using this controlled vocabulary from 2013 to 2020. </p>
			</sec>
			<sec id="sec2.4">
				<label>2.4</label>
				<title>Gold indexing</title>
				<p>As golden indexing, we used the indexing assigned by the CSIC&#x2019;s indexers to each of the 200 documents that made up the test set. For this purpose, a query was made in the InDICEs database to retrieve all the documents published in the <italic>Revista Espa&#xf1;ola de Documentaci&#xf3;n Cient&#xed;fica</italic> from 2013 to 2020. Subsequently, the indexing assigned to each of the 200 documents was compiled and was used for the evaluation processes. It was also ascertained whether all the indexing terms assigned to the 200 documents in the test set were present in the controlled vocabulary. Various terms were identified which had been used as indexing terms but did not appear in the controlled vocabulary, and we proceeded to include them.</p>
			</sec>
			<sec id="sec2.5">
				<label>2.5</label>
				<title>Evaluation metrics</title>
				<p>The metrics used to compare indexing were of two types: on the one hand, precision, recall and F-measure, and on the other hand, the indexing consistency measure.</p>
				<p>Precision, recall and F-measure are used extensively in the field of information retrieval (<xref ref-type="bibr" rid="B21">Gupta et al.,2015</xref>; <xref ref-type="bibr" rid="B18">Gil-Leiva, 2017a</xref>; <xref ref-type="bibr" rid="B4">Al-Zoghby, 2018</xref>; <xref ref-type="bibr" rid="B49">Seiler, H&#xfc;bner &amp; Paech, 2019</xref>; and <xref ref-type="bibr" rid="B37">Lin et al., 2020</xref>, to cite some recent examples). These metrics have been adapted for other purposes such as evaluating information extraction, summarizing and comparing automatic indexing with golden indexing.</p>
				<p>In order to evaluate the systems&#x2019; automatic indexing output, the output keywords (the generated keywords) were compared with the human-assigned keywords for each document. We considered an output keyword to be &#x201c;relevant&#x201d; if it was an exact match with a human-assigned keyword. Precision, recall and F1 scores were calculated at document level, and then aggregated over the document collection.</p>
				<p>Examples of use of these metrics to compare indexing pairs include their use by <xref ref-type="bibr" rid="B35">Krapivin et al. (2008)</xref>, <xref ref-type="bibr" rid="B51">Shams and Mercer (2012b)</xref> and <xref ref-type="bibr" rid="B10">El-Haj et al. (2013)</xref> to compare KEA results; by <xref ref-type="bibr" rid="B6">Bandim and Correa (2019)</xref> and <xref ref-type="bibr" rid="B53">Silva, Correa and Gil-Leiva (2020)</xref> to compare SISA; by <xref ref-type="bibr" rid="B42">N&#xe9;v&#xe9;ol et al. (2005)</xref> and <xref ref-type="bibr" rid="B7">Chebil et al. (2012)</xref> to compare CISMEF automatic indexing; and by <xref ref-type="bibr" rid="B43">Rae et al. (2021)</xref> to evaluate the Medical Text Indexer (MTI) system.</p>
				<p>The indexing consistency measures applied are those proposed by <xref ref-type="bibr" rid="B24">Hooper (1965)</xref> and <xref ref-type="bibr" rid="B44">Rolling (1981)</xref>, specifically designed to compare indexing. Experiments using Hooper&#x2019;s measure include those by <xref ref-type="bibr" rid="B40">Mork, et al. (2017)</xref> in their evaluation of Medical Text Indexer system and by <xref ref-type="bibr" rid="B53">Silva, et al. (2020)</xref> to evaluate SISA; however, <xref ref-type="bibr" rid="B55">Sinkkil&#xe4; et al. (2011)</xref>, for example, used Rolling&#x2019;s measure to evaluate MAUI.</p>
				<p>In our experiment, we applied Hooper&#x2019;s measures to compare the indexing performed automatically by SISA, KEA and MAUI with human indexing (golden indexing).</p>
			</sec>
			<sec id="sec2.6">
				<label>2.6</label>
				<title>Experiments</title>
				<p>To apply KEA and MAUI, it was first necessary to convert the XML and PDF documents into text files (.txt files), as well as extract the descriptors of the human indexing and fill the key files (.key files).</p>
				<p>Having compiled the golden indexing for the test set and following the training phase required by KEA and MAUI using 30 documents, the automatic indexing of the test set of 200 articles in XML and TXT format by the three tools was begun. The use of the same test set and controlled vocabulary for the automatic indexing of articles allows comparisons to be made between the automatic indexing of the three software programs. </p>
				<p>We set a threshold of 10 as the maximum number of keywords that KEA and MAUI could generate. The threshold value was chosen to reflect the number of keywords that we estimated it would be reasonable for KEA and MAUI to extract from the documents in the test set, given the number of keywords assigned by the golden indexing. </p>
				<sec id="sec2.6.1">
					<title>SISA System configuration</title>
					<p>There were nine rules used in this experiment. Seven were position rules, that is, the position of a term in the text is considered. As they are scientific articles, the following parts were used: title, abstract, keywords, section headings, first paragraph of each section, the other paragraphs of that section, table titles, graph titles, conclusions and references. The position rules utilize a single part of the text (for example, Rule 1 proposes terms present in the title and Rule 2 proposes terms present in the author&#x2019;s keywords) or are activated when a term appears in several parts at the same time (for example, Rule 3 proposes the term if it appears in all these parts: abstract, section headings, first paragraph, other paragraphs, conclusions and references). In addition, the experiment used two statistical rules: one on the TF-IDF (for instance, Rule 8 proposes terms with a TF-IDF of at least 0.018) and the other on the total frequency of occurrence of a term in the document (for instance, Rule 9 proposes terms with a minimum frequency of 40 occurrences in the text).</p>
					<p>All the terms (or their synonyms) selected by the rules must be present both in the document and in the controlled vocabulary. Several rules can propose the same term. At present, in the set of terms proposed to index a document, there is no weighting of the primary descriptor or secondary descriptor type to try to give more or less importance to one descriptor over another.</p>
					<p>SISA does not currently allow users to set or limit the number of indexing terms for each document.</p>
				</sec>
				<sec id="sec2.6.2">
					<title>KEA System configuration</title>
					<p>The source code of KEA version 5.0 (KEA++) can be downloaded from <ext-link ext-link-type="uri" xlink:href="https://code.google.com/archive/p/kea-algorithm/">https://code.google.com/archive/p/kea-algorithm/</ext-link> and <ext-link ext-link-type="uri" xlink:href="https://github.com/EUMSSI/KEA">https://github.com/EUMSSI/KEA</ext-link>. For the experiments, KEA was downloaded from the latter URL for importation in Eclipse IDE.</p>
					<p>The following KEA input parameters were specified via changes in source code: training set directory, test set directory, vocabulary controlled in SKOS format, Spanish language selection, selection of the Spanish Snowball Stemmer for stemming words of the Spanish language, selection of stopwords list for Spanish, UTF-8 encoding for files, minimum frequency of occurrence for candidate terms equals 1 (training) and proposed indexing terms for each document set to 10.</p>
					<p>KEA was applied to the corpora in stemming mode using Spanish Snowball Stemmer.</p>
				</sec>
				<sec id="sec2.6.3">
					<title>MAUI System configuration</title>
					<p>The source code of MAUI used in the experiments was downloaded from <ext-link ext-link-type="uri" xlink:href="https://github.com/zelandiya/maui-standalone">https://github.com/zelandiya/maui-standalone</ext-link>.</p>
					<p>The following MAUI configuration was set via changes in source code: selection of the stemmer for the Spanish language provided by Apache Lucene; selection of stopwords list for Spanish coded in the MAUI software; selection of the types of characteristics for the candidate terms based on frequency, position and size.</p>
					<p>The following parameter values were provided for MAUI in the training phase: training set directory, vocabulary controlled in SKOS format, Spanish language selection, UTF-8 encoding for files and minimum frequency of occurrence for candidate terms equals 2.</p>
					<p>The following parameter values were provided for MAUI in the test phase: the indexing model, test set directory, vocabulary controlled in SKOS format, Spanish language selection, UTF-8 encoding for files, the number of terms equals 10 and probability threshold equals 0.05.</p>
				</sec>
			</sec>
		</sec>
		<sec id="sec3" sec-type="results|discussion">
			<label>3.</label>
			<title>Results and discussion</title>
			<sec id="sec3.1">
				<label>3.1</label>
				<title>Processing time</title>
				<p>
					<xref ref-type="bibr" rid="B16">Garcia Gutierrez (1984: 115)</xref> cited an experiment carried out in the early 1970s to learn about the reality of indexing in Great Britain, in which about twelve minutes were spent to obtain eleven to twenty keywords; regarding the findings of Garcia Gutierrez, the present study considers that the time taken to convert keywords in natural language into descriptors of a thesaurus or controlled vocabulary should be added. However, <xref ref-type="bibr" rid="B14">Farrow (1994: 158)</xref>, citing <xref ref-type="bibr" rid="B8">Cleverdon (1962)</xref>, indicated that the optimum time for indexing technical reports could be four minutes plus 60%, depending on the working conditions. On the other hand, <xref ref-type="bibr" rid="B3">Amat (1989: 176)</xref> stated that an average time of twenty minutes would be required to obtain about ten terms. And finally, in the AusLit repository they recommend using the following benchmarks: novel/drama, 30 minutes; biography/autobiography, 30 minutes; short story, 10-20 minutes; verse, 5 minutes; and critical article, 20-30 minutes. We considered that the indexing of a scientific article could take between ten and fifteen minutes, to try to establish some correlation between human and automatic indexing (<xref ref-type="table" rid="t1">Table I</xref>).</p>
				<table-wrap id="t1">
					<label>Table I</label>
					<caption>
						<title>Estimated processing times for items.</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="left"> </th>
								<th align="left">one article</th>
								<th align="left">200 articles</th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="justify">Manual indexing</td>
								<td align="justify">10-15 minutes</td>
								<td align="justify">33-50 horas</td>
							</tr>
							<tr>
								<td align="justify">Automatic indexing by SISA</td>
								<td align="justify">8.1 seconds</td>
								<td align="justify">1620 seconds (27 minutes)</td>
							</tr>
							<tr>
								<td align="justify">Automatic indexing by KEA</td>
								<td align="justify">0.45 seconds</td>
								<td align="justify">90 seconds</td>
							</tr>
							<tr>
								<td align="justify">Automatic indexing by MAUI</td>
								<td align="justify">0.1 second</td>
								<td align="justify">10 seconds</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
				<p>Among the automatic indexing systems, SISA takes by far the longest time to process documents. However, for all the systems, processing times are much shorter than the processing time required for human indexing.</p>
			</sec>
			<sec id="sec3.2">
				<label>3.2</label>
				<title>Mean number of indexing terms assigned</title>
				<p>
					<xref ref-type="fig" rid="ch1">Graph 1</xref> shows the mean number of indexing terms assigned in human indexing (golden indexing). In addition, for each automatic indexing system, it shows the average number of indexing terms assigned as well as the average number of terms in common with the terms assigned during golden indexing. Here we can see that KEA and MAUI have similar average values. SISA has an average number of assigned terms per document that is closer to the average number of golden indexing terms and an average number of terms shared with golden indexing per document almost one unit higher than KEA and MAUI.</p>
				<fig id="ch1">
					<label>Graph 1</label>
					<caption>
						<title>Assigned terms</title>
					</caption>
					<graphic id="gra-1" xlink:href="REDC-45-04-e338-gch1.png"/>
				</fig>
			</sec>
			<sec id="sec3.3">
				<label>3.3</label>
				<title>Precision, recall and F-measure values</title>
				<p>The classic formulas used to calculate the measurements in the experiment are shown in <xref ref-type="table" rid="t2">Table II</xref>, as well as the results obtained for article number 4 by the three automatic indexing systems.</p>
				<table-wrap id="t2">
					<label>Table II</label>
					<caption>
						<title>Measurements used and data for article number 4 in the three systems</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="left">System</th>
								<th align="left"> </th>
								<th align="left"> </th>
								<th align="left"> </th>
								<th align="left"> </th>
								<th align="center">Precision (%)</th>
								<th align="center">Recall (%)</th>
								<th align="center">F-measure (%)</th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="left"> </td>
								<td align="center">D</td>
								<td align="center">C</td>
								<td align="center">A</td>
								<td align="center">GI</td>
								<td align="center">P = (C/A)</td>
								<td align="center">R= ( C / GI)</td>
								<td align="center">F1=2*(P*R)/(P+R)</td>
							</tr>
							<tr>
								<td align="justify">SISA</td>
								<td align="center">A4</td>
								<td align="center">3</td>
								<td align="center">11</td>
								<td align="center">4</td>
								<td align="center">27.27</td>
								<td align="center">75.00</td>
								<td align="center">40.00</td>
							</tr>
							<tr>
								<td align="justify">KEA</td>
								<td align="center">A4</td>
								<td align="center">2</td>
								<td align="center">10</td>
								<td align="center">4</td>
								<td align="center">20.00</td>
								<td align="center">50.00</td>
								<td align="center">28.57</td>
							</tr>
							<tr>
								<td align="justify">MAUI</td>
								<td align="center">A4</td>
								<td align="center">3</td>
								<td align="center">10</td>
								<td align="center">4</td>
								<td align="center">30.00</td>
								<td align="center">75.00</td>
								<td align="center">42.86</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
				<p>In <xref ref-type="table" rid="t2">Table II</xref>, the D column value is the identifier of a document in the test set; the C column value is the number of correct terms assigned by each system in comparison to golden indexing for a document; the A column value is the total number of terms assigned by each system for a document; the GI column value is the number of golden indexing terms for a given document; the P column value is the precision; the R column value is the recall; and the F1 column value is the F-measure obtained by each system for a document.</p>
				<p>
					<xref ref-type="table" rid="t2">Table II</xref> shows that, overall, the values of the metrics obtained by each system for article number 4 are different, with the exception of the similarity observed in the results obtained by SISA and MAUI.</p>
				<p>
					<xref ref-type="fig" rid="ch2">Graph 2</xref> shows the values of the metrics for the three documents where each automatic indexing system has performed best.</p>
				<fig id="ch2">
					<label>Graph 2</label>
					<caption>
						<title>The three best values for each system for precision, recall and F-measure</title>
					</caption>
					<graphic id="gra-2" xlink:href="REDC-45-04-e338-gch2.png"/>
				</fig>
				<p>It can be observed that the best performance for each system is based on different documents. In addition, SISA has obtained higher values for precision, recall and F-measure. The values obtained by KEA and MAUI for F-measure are similar to each other, but KEA has better values for precision while MAUI has better values for recall.</p>
				<p>On the other hand, the three worst non-zero results for the F-measure range between 5.9% and 11.8% for KEA, 7.4% and 10.5% for SISA, and 8.3% and 9.5% for MAUI. Thus, SISA, KEA and MAUI have similar values for precision, recall and F-measure in the worst cases.</p>
				<p>
					<xref ref-type="table" rid="t3">Table III</xref> shows the statistics for the metric values achieved by the automatic indexing systems based on the test set.</p>
				<table-wrap id="t3">
					<label>Table III</label>
					<caption>
						<title>Statistics for SISA, KEA and MAUI performance on test set</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="left"> </th>
								<th align="left"> </th>
								<th align="center">Assigned terms</th>
								<th align="center">Gold indexing</th>
								<th align="center">Mean of common terms</th>
								<th align="center">Precision</th>
								<th align="center">Recall</th>
								<th align="center">F-measure</th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="justify" rowspan="4">SISA</td>
								<td align="justify">Minimum</td>
								<td align="center">3</td>
								<td align="center">2</td>
								<td align="center">0</td>
								<td align="center">0%</td>
								<td align="center">0%</td>
								<td align="center">0%</td>
							</tr>
							<tr>
								<td align="justify">Mean</td>
								<td align="center">8.9</td>
								<td align="center">7.3</td>
								<td align="center">3.4</td>
								<td align="center">39.1%</td>
								<td align="center">52.1%</td>
								<td align="center">42.5%</td>
							</tr>
							<tr>
								<td align="justify">Deviation</td>
								<td align="center">2.7</td>
								<td align="center">3.4</td>
								<td align="center">1.7</td>
								<td align="center">17.9%</td>
								<td align="center">27.7%</td>
								<td align="center">18.9%</td>
							</tr>
							<tr>
								<td align="justify">Maximum</td>
								<td align="center">18</td>
								<td align="center">24</td>
								<td align="center">8</td>
								<td align="center">100%</td>
								<td align="center">100%</td>
								<td align="center">85.7%</td>
							</tr>
							<tr>
								<td align="justify" rowspan="4">KEA</td>
								<td align="justify">Minimum</td>
								<td align="center">10</td>
								<td align="center">2</td>
								<td align="center">0</td>
								<td align="center">0%</td>
								<td align="center">0%</td>
								<td align="center">0%</td>
							</tr>
							<tr>
								<td align="justify">Mean</td>
								<td align="center">10</td>
								<td align="center">7.3</td>
								<td align="center">2.7</td>
								<td align="center">27.3%</td>
								<td align="center">41.3%</td>
								<td align="center">31.7%</td>
							</tr>
							<tr>
								<td align="justify">Deviation</td>
								<td align="center">0</td>
								<td align="center">3.5</td>
								<td align="center">1.3</td>
								<td align="center">12.9%</td>
								<td align="center">22.1%</td>
								<td align="center">14.6%</td>
							</tr>
							<tr>
								<td align="justify">Maximum</td>
								<td align="center">10</td>
								<td align="center">24</td>
								<td align="center">6</td>
								<td align="center">60.0%</td>
								<td align="center">100%</td>
								<td align="center">66.7%</td>
							</tr>
							<tr>
								<td align="justify" rowspan="4">MAUI</td>
								<td align="justify">Minimum</td>
								<td align="center">6</td>
								<td align="center">2</td>
								<td align="center">0</td>
								<td align="center">0%</td>
								<td align="center">0%</td>
								<td align="center">0%</td>
							</tr>
							<tr>
								<td align="justify">Mean</td>
								<td align="center">9.8</td>
								<td align="center">7.3</td>
								<td align="center">2.5</td>
								<td align="center">25.7%</td>
								<td align="center">38.6%</td>
								<td align="center">29.6%</td>
							</tr>
							<tr>
								<td align="justify">Deviation</td>
								<td align="center">0.6</td>
								<td align="center">3.4</td>
								<td align="center">1.3</td>
								<td align="center">13.3%</td>
								<td align="center">22.4%</td>
								<td align="center">14.9%</td>
							</tr>
							<tr>
								<td align="justify">Maximum</td>
								<td align="center">10</td>
								<td align="center">24</td>
								<td align="center">6</td>
								<td align="center">60.0%</td>
								<td align="center">100%</td>
								<td align="center">70.6%</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
				<p>According to unpaired t test results with conventional criteria, the difference between the means of the values of the metrics for SISA and those of each of the other systems is considered extremely statistically significant. However, according to unpaired t test results with conventional criteria, the difference between the means of the values of the metrics for KEA and MAUI is not statistically significant.</p>
				<p>
					<xref ref-type="fig" rid="ch3">Graph 3</xref> illustrates the similar mean performance of KEA and MAUI, and the superior mean performance of SISA.</p>
				<fig id="ch3">
					<label>Graph 3</label>
					<caption>
						<title>Summary table with the resulting data</title>
					</caption>
					<graphic id="gra-3" xlink:href="REDC-45-04-e338-gch3.png"/>
				</fig>
				<p>
					<xref ref-type="table" rid="t4">Table IV</xref> shows the number of documents for which the performance of the automatic indexing systems falls in the intervals of F-measure values. </p>
				<table-wrap id="t4">
					<label>Table IV</label>
					<caption>
						<title>F-measure values of the two hundred indexed documents</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="left"> </th>
								<th align="center">F-measure 0%</th>
								<th align="center">F-measure 1-15%</th>
								<th align="center">F-measure 16-49%</th>
								<th align="center">F-measure &#x2265;50%</th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="justify">SISA</td>
								<td align="center">6</td>
								<td align="center">11</td>
								<td align="center">111</td>
								<td align="center">
									<bold>72</bold>
								</td>
							</tr>
							<tr>
								<td align="justify">KEA</td>
								<td align="center">8</td>
								<td align="center">26</td>
								<td align="center">
									<bold>145</bold>
								</td>
								<td align="center">21</td>
							</tr>
							<tr>
								<td align="justify">MAUI</td>
								<td align="center">
									<bold>10</bold>
								</td>
								<td align="center">
									<bold>34</bold>
								</td>
								<td align="center">135</td>
								<td align="center">21</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
				<p>We can see in <xref ref-type="table" rid="t4">Table IV</xref> that the systems have the majority of the F-measure values falling in the interval of 16 to 49% for F-measure. SISA has roughly three times the number of Fmeasure values equal to or higher than 50% compared to KEA and MAUI. KEA and MAUI have similar numbers of documents in each interval of F-measure values. In addition, the number of documents with an F-measure of zero is as follows: SISA produced six, KEA eight, and MAUI ten.</p>
				<p>According to <xref ref-type="bibr" rid="B39">Medelyan (2009)</xref>, the similarity between the values achieved by KEA and the data for MAUI is explained by the fact that MAUI was developed from the foundations and structure of KEA. However, we can observe that besides their similar performance, there is a performance variation for each document in these systems, so that they capture both common and different aspects of the statistical nature of indexing terms.</p>
				<p>SISA was conceived with the aim of mimicking a human indexer by directing the focus towards places in the text where it can identify significant terms. <xref ref-type="bibr" rid="B30">ISO standard 5963-1985</xref> on human indexing of documents states that &#x201c;important parts of the text need to be considered carefully, and particular attention should be paid to the following: a) the title, b) the abstracts, if provided; c) the list of contents; d) the introduction, the opening phrases of chapters and paragraphs, and the conclusion; e) illustration, diagrams, tables and their captions&#x201d;. This recommendation by the standard derives from the experience of human indexers, hence SISA&#x2019;s imitation of the behaviour of human indexers is perhaps the reason why SISA has achieved better results than KEA and MAUI, which do not base their processing on the structural positions of terms in documents. </p>
			</sec>
			<sec id="sec3.4">
				<label>3.4</label>
				<title>Consistency</title>
				<p>Indexing consistency seeks to determine the similarity between sets of indexing terms from analysis of the same document. Therefore, there are a number of possible combinations for comparison: between human indexing, between automatic indexing, and between human indexing and automatic indexing.</p>
				<p>Hooper&#x2019;s measure has been extensively used to calculate indexing consistency. In the formula, C is the number of terms assigned by SISA, KEA and MAUI that match those assigned by golden indexing, A is the number of terms assigned by SISA, KEA or MAUI, and GI is the number of golden indexing terms for a given document. <xref ref-type="table" rid="t5">Table V</xref> summarizes the data achieved in our experiment.</p>
				<table-wrap id="t5">
					<label>Table V</label>
					<caption>
						<title>Hooper&#x2019;s measure</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="left"> </th>
								<th align="left"> </th>
								<th align="center">Hooper H = C / (A+GI-C)</th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="justify" rowspan="4">SISA</td>
								<td align="justify">Minimum</td>
								<td align="center">0%</td>
							</tr>
							<tr>
								<td align="justify">Mean</td>
								<td align="center">
									<bold>28.8%</bold>
								</td>
							</tr>
							<tr>
								<td align="justify">Deviation</td>
								<td align="center">15.8%</td>
							</tr>
							<tr>
								<td align="justify">Maximum</td>
								<td align="center">75.0%</td>
							</tr>
							<tr>
								<td align="justify" rowspan="4">KEA</td>
								<td align="justify">Minimum</td>
								<td align="center">0%</td>
							</tr>
							<tr>
								<td align="justify">Mean</td>
								<td align="center">19.7%</td>
							</tr>
							<tr>
								<td align="justify">Deviation</td>
								<td align="center">10.6%</td>
							</tr>
							<tr>
								<td align="justify">Maximum</td>
								<td align="center">50.0%</td>
							</tr>
							<tr>
								<td align="justify" rowspan="4">MAUI</td>
								<td align="justify">Minimum</td>
								<td align="center">0%</td>
							</tr>
							<tr>
								<td align="justify">Mean</td>
								<td align="center">18.3%</td>
							</tr>
							<tr>
								<td align="justify">Deviation</td>
								<td align="center">10.7%</td>
							</tr>
							<tr>
								<td align="justify">Maximum</td>
								<td align="center">54.5%</td>
							</tr>
						</tbody>
					</table>
					<table-wrap-foot>
						<fn id="TFN1">
							<p>C= Number of commons terms assigned by SISA, KEA and MAUI in the relation to gold indexing; A= Number of terms assigned by SISA, KEA or MAUI; GI= Number of gold indexing terms for a given document.</p>
						</fn>
					</table-wrap-foot>
				</table-wrap>
				<p>
					<xref ref-type="fig" rid="ch4">Graph 4</xref> shows the three documents with the highest consistency. SISA&#x2019;s levels of consistency are higher than those of KEA and MAUI. KEA and MAUI have similar values of consistency for the best cases, but for different documents.</p>
				<fig id="ch4">
					<label>Graph 4</label>
					<caption>
						<title>Hooper&#x2019;s best three values</title>
					</caption>
					<graphic id="gra-4" xlink:href="REDC-45-04-e338-gch4.png"/>
				</fig>
				<p>The three worst results achieved a range from 3.8% to 5.8% for SISA, 3% to 6.3% for KEA and 4.3% to 5% for MAUI. Again, the systems generate very similar values for the three worst cases.</p>
			</sec>
			<sec id="sec3.5">
				<label>3.5</label>
				<title>Indexing analysis</title>
				<p>In this section we analyze the three best performance results and the three worst non-zero results achieved by the three systems.</p>
				<p>In <xref ref-type="table" rid="t6">Table VI</xref> we show the indexing that achieved the highest F-measure.</p>
				<table-wrap id="t6">
					<label>Table VI</label>
					<caption>
						<title>Indexing with the best F-measures</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="left" colspan="2">SISA (Article 166). F-measure = 85.7% </th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="left">
									<bold>Gold indexing</bold>
								</td>
								<td align="left">
									<bold>SISA indexing</bold>
								</td>
							</tr>
							<tr>
								<td align="left"> 1. <bold>Cadena de valor</bold> 2. Desintermediaci&#xf3;n <break/>3. <bold>Digitalizaci&#xf3;n</bold> 4. <bold>Espa&#xf1;a</bold> 5. <bold>Industria editorial</bold> 6. <bold>Libros electr&#xf3;nicos</bold>
								</td>
								<td align="left"> 1. <bold>Cadena de valor</bold> 2. Desinformaci&#xf3;n <break/>3. <bold>Digitalizaci&#xf3;n</bold> 4. <bold>Espa&#xf1;a</bold> 5. Estudio <break/>6. Impacto 7. <bold>Industria editorial</bold>
									<break/>8. <bold>Libros electr&#xf3;nicos</bold>
								</td>
							</tr>
							<tr>
								<td align="left" colspan="2">
									<bold>KEA</bold> (Article 7). F-measure = 66.6% </td>
							</tr>
							<tr>
								<td align="left">
									<bold>Gold indexing</bold>
								</td>
								<td align="left">
									<bold>KEA indexing</bold>
								</td>
							</tr>
							<tr>
								<td align="left"> 1. <bold>Accesibilidad web</bold> 2. Andaluc&#xed;a <break/>3. <bold>Discapacidad</bold> 4. <bold>Dise&#xf1;o web</bold> 5. Sitios web 6. <bold>Universidades</bold> 7. <bold>W3C</bold> 8. <bold>WCAG 2.0</bold>
								</td>
								<td align="left"> 1. Accesibilidad 2. <bold>Accesibilidad web</bold>
									<break/>3. <bold>Discapacidad</bold> 4. <bold>Dise&#xf1;o Web</bold>
									<break/>5. Personas con discapacidad 6. Portales <break/>7. <bold>Universidad</bold> 8. <bold>W3C</bold> 9. <bold>WCAG 2.0</bold>
									<break/>10. World Wide Web</td>
							</tr>
							<tr>
								<td align="left" colspan="2">
									<bold>MAUI</bold> (Article 141). F-measure = 70.5% </td>
							</tr>
							<tr>
								<td align="left">
									<bold>Gold indexing</bold>
								</td>
								<td align="left">
									<bold>MAUI indexing</bold>
								</td>
							</tr>
							<tr>
								<td align="left"> 1. <bold>Bibliotecas</bold> 2. <bold>Bibliotecas escolares</bold>
									<break/>3. <bold>Blogs</bold> 4. <bold>Educaci&#xf3;n infantil</bold> 5. Espa&#xf1;a <break/>6. <bold>Extremadura</bold> 7. <bold>Indicadores de calidad</bold>
								</td>
								<td align="left"> 1. <bold>Bibliotecas</bold> 2. <bold>Bibliotecas escolares</bold>
									<break/>3. <bold>Blogs</bold> 4. Centros de ense&#xf1;anza <break/>5. Centros educativos 6. <bold>Educaci&#xf3;n infantil</bold> 7. <bold>Extremadura</bold> 8. <bold>Indicadores de calidad</bold> 9. Modelo 10. Modelos</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
				<p>In <xref ref-type="table" rid="t6">Table VI</xref>, we can see in bold type the terms that golden indexing and the automatic indexing had in common with each other for the best cases. The number of terms they had in common is five or six terms. The terms proposed by automatic indexing which were not common to both the automatic and manual systems are however linked in some way to the golden indexing terms for each document. In addition, the number of terms in golden indexing and the number proposed by the systems is similar. In Annex, <xref ref-type="table" rid="ta1">Table A</xref> shows the three best results with the same patterns as those indicated above.</p>
				<p>Table B in Annex shows the worst non-zero results for SISA, KEA and MAUI. The number of terms they had in common is one or two terms. The terms proposed by automatic indexing which were not common to both the automatic and manual systems are linked in some way to the golden indexing terms for each document. In addition, a greater difference between the number of terms in golden indexing and the number of those proposed by the systems can be observed for some documents. In Annex, <xref ref-type="table" rid="ta1">Table A</xref> shows the three worst non-zero results with the same patterns as those indicated here.</p>
				<p>In <xref ref-type="app" rid="app1">Annex 1</xref>, <xref ref-type="table" rid="ta1">Table A</xref> shows the three best and three worst non-zero results for SISA, KEA and MAUI. In the case of SISA, there is a considerable disparity between the number of terms assigned by SISA and the number of terms assigned by the database&#x2019;s human indexers. In the case of Article 91, the human indexers assigned almost three times as many terms as SISA while for Article 94 they assigned twice as many. For Article 87, SISA assigned almost twice as many terms as those assigned by human indexing. This large disparity between SISA and human indexing also occurred for other articles that achieved a low F-measure of between 10-16% (Articles 1, 97, and 168, for example), although in other articles with a very similar number of assigned terms (Articles 10, 36 and 138) the F-measure was equally low.</p>
				<p>If the indexing by a given tool gives an F-measure that is similar to the examples shown in <xref ref-type="table" rid="t6">Table VI</xref>, we consider that use of some automatic indexing software or semi-automatic could be of great help in order to increase document processing capacity in real working environments, or even by enabling the automatic indexing of documents that are not manually indexed.</p>
			</sec>
		</sec>
		<sec id="sec4" sec-type="conclusions">
			<label>4.</label>
			<title>Conclusions and further work</title>
			<p>This study proposed an evaluation-comparison between three automatic indexing systems, analyzing, on the one hand, the latest web version of SISA (which uses a rule-based algorithm focusing on the position occupied by the terms in the documents) and, on the other hand, KEA and MAUI (two indexing systems that produce indexing terms by means of a machine learning model, after a training process with 30 documents) and their respective indexing. </p>
			<p>The data achieved by SISA with a total F-measure of ten points above KEA and MAUI, more than three times as many documents with an F-measure &#x2265;50% and an average number of terms assigned per document of 8.8, indicate that the automatic indexing by SISA is more similar to the human indexing by the InDICES database professionals than the automatic indexing by KEA and MAUI. Thus, an algorithm that focuses on different parts of the texts (titles, abstracts, author keywords, headings, first paragraphs of headings, table titles, graph titles, conclusions and references) and exploits the structure of documents in XML format has outperformed algorithms based on machine learning and statistical features of terms. </p>
			<p>By exploiting markup tags and being capable of handling XML documents and documents generated from JATS, SISA is in line with the current trend of creating publications that allow further automatic processing and reuse. It is therefore also worth mentioning that JATS has been the basis for creating two extensions: the Book Interchange Tag Suite (BITS), which is an XML model to describe the structural and semantic content of books published by scientific, technical and medical publishers, and the STS (Standards Tag Suite), ANSI/NISO Z39.102-2017 and the ISOSTS (ISO Standards Tag Set) systems which provide a common tagging format that developers, publishers and distributors of standards can use to publish and exchange contents. Therefore, SISA seems to be moving in the right direction if these forms of publication are extended.</p>
			<p>With regard to future work, it would be useful to replicate the experiments presented here with texts from other disciplines, other professional human indexing and other controlled vocabularies in order to verify whether the higher performance level achieved by the methodology implemented in SISA still outperforms the KEA and MAUI machine learning algorithms. It would perhaps also be interesting to work jointly with the indexers of the InDICES database to carry out a detailed analysis of the indexing of the 200 articles performed by SISA, KEA and MAUI so as to determine, for example, which types of articles obtained the greatest similarity with human indexing or to establish what types of errors are made by the automatic systems, among other aspects. Moreover, during the execution of this work, several improvements have been identified that could improve SISA&#x2019;s performance. One of them could be to limit the maximum number of terms that can be assigned per document, which would perhaps allow higher F-measure indexes to be achieved. In this regard, it should be noted that MAUI has limited the maximum number of indexing terms it can assign to 10 and KEA systematically establishes 10 terms for each document, whereas SISA currently has neither a minimum nor a maximum number, and thus this study has observed one document with 18 terms assigned by SISA, several with three and many others with 12, 13 and 14 terms. Another improvement that is already being implemented is a new module to generate indexing rules automatically. In SISA, indexing terms are derived by applying a set of rules established by the user according to their experience and knowledge of the system. With this improvement, the aim is to eliminate user intervention in favour of a data-driven automatic process that generates rules from a collection of training documents with their respective golden indexing, so that it is possible to know which rules produce F-measure values above certain thresholds. This data-driven automatic process would permit greater adaptability by SISA to the characteristics of the texts and subject areas of each test set. Finally, it would also be interesting to find out if the automatically generated rules manage to improve on the results of the manual rules established and used for SISA in the execution of the experiments presented here.</p>
		</sec>
	</body>
	<back>
		<ack>
			<label>5.</label>
			<title>Acknowledgments</title>
			<p>The authors would like to thank Teresa Abej&#xf3;n Pe&#xf1;a for providing us with the controlled vocabulary used in the InDICES database to carry out the experiments. We would also like to extend our thanks to the Consejo Superior de Investigaciones Cient&#xed;ficas, which is responsible for this database. Lastly, we would like to thank two anonymous reviewers for their comments, which have helped improve the understanding of the manuscript.</p>
		</ack>
		<ack xml:lang="es">
			<title>Agradecimientos</title>
			<p>Los autores agradecen a Teresa Abej&#xf3;n Pe&#xf1;a por facilitarnos el vocabulario controlado utilizado en la base de datos InDICES para la realizaci&#xf3;n de los experimentos. Tambi&#xe9;n extendemos el agradecimiento al Consejo Superior de Investigaciones Cient&#xed;ficas como responsable de esta base de datos.Por &#xfa;ltimo, agradecemos a los dos revisores an&#xf3;nimos sus comentarios, que han ayudado a mejorar la comprensi&#xf3;n del manuscrito.</p>
		</ack>
		<ref-list>
			<label>6.</label>
			<title>References</title>
			<ref id="B1">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Aronson</surname>
							<given-names>A.R.</given-names>
						</string-name>
						<string-name>
							<surname>Bodenreider</surname>
							<given-names>O.</given-names>
						</string-name>
						<string-name>
							<surname>Chang</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Florence</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Humphrey</surname>
							<given-names>S.M.</given-names>
						</string-name>
						<string-name>
							<surname>Mork</surname>
							<given-names>J. G.</given-names>
						</string-name>
						<string-name>
							<surname>Stuart</surname>
							<given-names>J.N.</given-names>
						</string-name>
						<string-name>
							<surname>Rindflesch</surname>
							<given-names>T. C.</given-names>
						</string-name>
						<string-name>
							<surname>Wilbur</surname>
							<given-names>W. J.</given-names>
						</string-name>
					</person-group>
					<year>2000</year>
					<source>The NLM Indexing Initiative</source>
					<person-group person-group-type="editor">
						<string-name>
							<surname>Marc Overhage</surname>
							<given-names>J.</given-names>
						</string-name>
					</person-group>
					<conf-name>Proceedings of the AMIA Annual Symposium</conf-name>
					<fpage>17</fpage>
					<lpage>21</lpage>
				</mixed-citation>
			</ref>
			<ref id="B2">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Akhtar</surname>
							<given-names>N.</given-names>
						</string-name>
						<string-name>
							<surname>Javed</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Ahmad</surname>
							<given-names>T.</given-names>
						</string-name>
					</person-group>
					<year>2017</year>
					<source>Searching related Scientific Articles Using Formal Concept Analysis</source>
					<conf-name>International Conference on Energy, Communication, Data Analytics and Soft Computing (ICECDS)</conf-name>
					<fpage>2158</fpage>
					<lpage>2163</lpage>
					<pub-id pub-id-type="doi">10.1109/ICECDS.2017.8389834</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B3">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>Amat</surname>
							<given-names>N.</given-names>
						</string-name>
					</person-group>
					<year>1989</year>
					<source>Documentaci&#xf3;n y nuevas tecnolog&#xed;as de la informaci&#xf3;n</source>
					<publisher-name>Pir&#xe1;mide</publisher-name>
				</mixed-citation>
			</ref>
			<ref id="B4">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>Al-Zoghby</surname>
							<given-names>A.</given-names>
						</string-name>
					</person-group>
					<year>2018</year>
					<chapter-title>A New Semantic Distance Measure for the VSM-Based Information Retrieval Systems</chapter-title>
					<source>Intelligent Natural Language Processing: Trends and Application</source>
					<volume>740</volume>
					<fpage>229</fpage>
					<lpage>250</lpage>
					<pub-id pub-id-type="doi">10.1007/978-3-319-67056-0_12</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B5">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Aquino</surname>
							<given-names>G.</given-names>
						</string-name>
						<string-name>
							<surname>Lanzarini</surname>
							<given-names>L.</given-names>
						</string-name>
					</person-group>
					<year>2015</year>
					<article-title>Keyword Identification in Spanish Documents using Neural Networks</article-title>
					<source>Journal of Computer Science and Technology</source>
					<volume>15</volume>
					<fpage>55</fpage>
					<lpage>60</lpage>
				</mixed-citation>
			</ref>
			<ref id="B6">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Bandim</surname>
							<given-names>M. A. S.</given-names>
						</string-name>
						<string-name>
							<surname>Corr&#xea;a</surname>
							<given-names>R. F.</given-names>
						</string-name>
					</person-group>
					<year>2019</year>
					<article-title>Indexa&#xe7;&#xe3;o autom&#xe1;tica por atribui&#xe7;&#xe3;o de artigos cient&#xed;ficos em portugu&#xea;s da &#xe1;rea de Ci&#xea;ncia da Informa&#xe7;&#xe3;o</article-title>
					<source>Transinforma&#xe7;&#xe3;o</source>
					<volume>31</volume>
					<fpage>1</fpage>
					<lpage>12</lpage>
					<pub-id pub-id-type="doi">10.1590/2318-0889201931e180004</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B7">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Chebil</surname>
							<given-names>Wiem</given-names>
						</string-name>
						<string-name>
							<surname>Soualmia</surname>
							<given-names>L.</given-names>
						</string-name>
						<string-name>
							<surname>Dahamna</surname>
							<given-names>B.</given-names>
						</string-name>
						<string-name>
							<surname>Srmoni</surname>
							<given-names>S.</given-names>
						</string-name>
					</person-group>
					<year>2012</year>
					<article-title>Indexation automatique de documents ensant&#xe9;: &#xe9;valuation et analyse de sources d&#x2019;erreurs</article-title>
					<source>IRBM</source>
					<volume>33</volume>
					<fpage>316</fpage>
					<lpage>329</lpage>
					<pub-id pub-id-type="doi">10.1016/j.irbm.2012.10.002</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B8">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>Cleverdon</surname>
							<given-names>C.W.</given-names>
						</string-name>
					</person-group>
					<year>1962</year>
					<source>Aslib Cranfield Research Project: report on the testing and analysis of an investigation into the comparative efficiency of indexing systems</source>
					<publisher-loc>Cranfield</publisher-loc>
				</mixed-citation>
			</ref>
			<ref id="B9">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Duwairi</surname>
							<given-names>R.</given-names>
						</string-name>
						<string-name>
							<surname>Hedaya</surname>
							<given-names>M.</given-names>
						</string-name>
					</person-group>
					<year>2016</year>
					<article-title>Automatic keyphrase extraction for Arabic news documents based on KEA system</article-title>
					<source>Journal of Intelligent and Fuzzy Systems</source>
					<volume>30</volume>
					<issue>4</issue>
					<fpage>2101</fpage>
					<lpage>2110</lpage>
				</mixed-citation>
			</ref>
			<ref id="B10">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>El-Haj</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>Balkan</surname>
							<given-names>L.</given-names>
						</string-name>
						<string-name>
							<surname>Barbalet</surname>
							<given-names>S.</given-names>
						</string-name>
						<string-name>
							<surname>Bell</surname>
							<given-names>L.</given-names>
						</string-name>
						<string-name>
							<surname>Shepherdson</surname>
							<given-names>J.</given-names>
						</string-name>
					</person-group>
					<year>2013</year>
					<source>An Experiment in Automatic Indexing Using the HASSET Thesaurus</source>
					<conf-name>5th Computer Science and Electronic Engineering Conference (CEEC)</conf-name>
					<fpage>13</fpage>
					<lpage>18</lpage>
					<pub-id pub-id-type="doi">10.1109/CEEC.2013.6659437</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B11">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Evans</surname>
							<given-names>D. A.</given-names>
						</string-name>
					</person-group>
					<year>1990</year>
					<source>Concept Management in Text via Natural-Language Processing: the CLARIT Approach</source>
					<conf-name>Working Notes of the 1990 AAAI Symposium on &#x201c;Text-Based Intelligent Systems&#x2019;9</conf-name>
					<conf-sponsor>Stanford University</conf-sponsor>
					<conf-date>March</conf-date>
					<page-range>27-29, 93-95</page-range>
				</mixed-citation>
			</ref>
			<ref id="B12">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Evans</surname>
							<given-names>D.A.</given-names>
						</string-name>
						<string-name>
							<surname>Hersh</surname>
							<given-names>W.R.</given-names>
						</string-name>
						<string-name>
							<surname>Monarch</surname>
							<given-names>I.</given-names>
						</string-name>
						<string-name>
							<surname>Lefferts</surname>
							<given-names>R. G.</given-names>
						</string-name>
						<string-name>
							<surname>Handerson</surname>
							<given-names>S. K.</given-names>
						</string-name>
					</person-group>
					<year>1991a</year>
					<article-title>Automatic Indexing of abstracts via Natural-Language Processing Using a Simple Thesaurus</article-title>
					<source>Medical Decision Making</source>
					<volume>11</volume>
					<issue>4</issue>
					<fpage>108</fpage>
					<lpage>115</lpage>
				</mixed-citation>
			</ref>
			<ref id="B13">
				<mixed-citation publication-type="report">
					<person-group person-group-type="author">
						<string-name>
							<surname>Evans</surname>
							<given-names>D.A.</given-names>
						</string-name>
						<string-name>
							<surname>Handerson</surname>
							<given-names>S. K.</given-names>
						</string-name>
						<string-name>
							<surname>Lefferts</surname>
							<given-names>R. G.</given-names>
						</string-name>
						<string-name>
							<surname>Monarch</surname>
							<given-names>I.</given-names>
						</string-name>
					</person-group>
					<year>1991b</year>
					<source>A Summary of the CLARIT Project</source>
					<month>November</month>
					<year>1991</year>
					<gov>Report No. CMU-LCL-91-2</gov>
					<pub-id pub-id-type="doi">10.1184/R1/6490799.v1</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B14">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>Farrow</surname>
							<given-names>J.</given-names>
						</string-name>
					</person-group>
					<year>1994</year>
					<chapter-title>Indexing as a cognitive process</chapter-title>
					<person-group person-group-type="editor">
						<string-name>
							<surname>Kent</surname>
							<given-names>A.</given-names>
						</string-name>
						<string-name>
							<surname>Lancour</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Daily</surname>
							<given-names>J.E.</given-names>
						</string-name>
					</person-group>
					<source>Encyclopedia of Library and Information Science</source>
					<volume>53</volume>
					<fpage>155</fpage>
					<lpage>171</lpage>
				</mixed-citation>
			</ref>
			<ref id="B15">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Frank</surname>
							<given-names>E.</given-names>
						</string-name>
						<string-name>
							<surname>Paynter</surname>
							<given-names>G. W.</given-names>
						</string-name>
						<string-name>
							<surname>Witten</surname>
							<given-names>I. H.</given-names>
						</string-name>
						<string-name>
							<surname>Gutwin</surname>
							<given-names>C.</given-names>
						</string-name>
						<string-name>
							<surname>Nevill-Manning</surname>
							<given-names>C. G.</given-names>
						</string-name>
					</person-group>
					<year>1999</year>
					<source>Domain-specific Keyphrase Extraction</source>
					<conf-name>Proceedings of the 16th International Joint Conference on Artificial Intelligence</conf-name>
					<conf-loc>Stockholm, Sweden</conf-loc>
					<conf-loc>San Francisco, Sweden</conf-loc>
					<conf-sponsor>Morgan Kaufmann Publishers</conf-sponsor>
					<fpage>668</fpage>
					<lpage>673</lpage>
				</mixed-citation>
			</ref>
			<ref id="B16">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>Garc&#xed;a Guti&#xe9;rrez</surname>
							<given-names>A.</given-names>
						</string-name>
					</person-group>
					<year>1984</year>
					<source>Ling&#xfc;&#xed;stica documental</source>
					<publisher-loc>Barcelona</publisher-loc>
					<publisher-name>Mitre</publisher-name>
				</mixed-citation>
			</ref>
			<ref id="B17">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>Gil-Leiva</surname>
							<given-names>I.</given-names>
						</string-name>
					</person-group>
					<year>2008</year>
					<source>Manual de indizaci&#xf3;n. Teor&#xed;a y pr&#xe1;ctica</source>
					<publisher-name>Trea</publisher-name>
				</mixed-citation>
			</ref>
			<ref id="B18">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Gil-Leiva</surname>
							<given-names>I.</given-names>
						</string-name>
					</person-group>
					<year>2017a</year>
					<article-title>SISA: Automatic Indexing System for Scientific Articles. Experiments with Location Heuristics Rules versus TF-IDF Rules</article-title>
					<source>Knowledge Organization</source>
					<volume>44</volume>
					<issue>3</issue>
					<fpage>139</fpage>
					<lpage>162</lpage>
				</mixed-citation>
			</ref>
			<ref id="B19">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Gil-Leiva</surname>
							<given-names>I.</given-names>
						</string-name>
					</person-group>
					<year>2017b</year>
					<source>La indizaci&#xf3;n de art&#xed;culos cient&#xed;ficos con el sistema de indizaci&#xf3;n autom&#xe1;tica SISA comparada con la indizaci&#xf3;n en las Bases de datos Agricola, WoS y SCOPUS</source>
					<conf-name>Third Spanish-Portuguese ISKO Conference, Portugal, Thirteenth ISKO Conference</conf-name>
					<conf-loc>Spain</conf-loc>
					<conf-sponsor>University of Coimbra</conf-sponsor>
					<conf-date>23 and 24 November</conf-date>
					<fpage>510</fpage>
					<lpage>524</lpage>
				</mixed-citation>
			</ref>
			<ref id="B20">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Gopan</surname>
							<given-names>E.</given-names>
						</string-name>
						<string-name>
							<surname>Rajesh</surname>
							<given-names>S.</given-names>
						</string-name>
						<string-name>
							<surname>Gr</surname>
							<given-names>V.</given-names>
						</string-name>
						<string-name>
							<surname>Akhil</surname>
							<given-names>R. R.</given-names>
						</string-name>
						<string-name>
							<surname>Thushara</surname>
							<given-names>M.</given-names>
						</string-name>
					</person-group>
					<year>2020</year>
					<source>Comparative Study on Different Approaches in Keyword Extraction</source>
					<conf-name>2020 Fourth International Conference on Computing Methodologies and Communication (ICCMC)</conf-name>
					<fpage>70</fpage>
					<lpage>74</lpage>
					<pub-id pub-id-type="doi">10.1109/iccmc48092.2020.iccmc-00013</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B21">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Gupta</surname>
							<given-names>Y.</given-names>
						</string-name>
						<string-name>
							<surname>Saini</surname>
							<given-names>A.</given-names>
						</string-name>
						<string-name>
							<surname>Saxena</surname>
							<given-names>A.</given-names>
						</string-name>
					</person-group>
					<year>2015</year>
					<article-title>A new fuzzy logic based ranking function for efficient Information Retrieval system</article-title>
					<source>Expert Systems with Applications</source>
					<volume>42</volume>
					<issue>3</issue>
					<fpage>1223</fpage>
					<lpage>1234</lpage>
				</mixed-citation>
			</ref>
			<ref id="B22">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Hersh</surname>
							<given-names>W. R.</given-names>
						</string-name>
						<string-name>
							<surname>Greenes</surname>
							<given-names>R.</given-names>
						</string-name>
					</person-group>
					<year>1990</year>
					<article-title>SAPHIRE: An information Retrieval Environment Featuring Conceptmatching, Automatic Indexing, and Probabilistic Retrieval</article-title>
					<source>Computers and Biomedical Research</source>
					<volume>123</volume>
					<fpage>410</fpage>
					<lpage>425</lpage>
				</mixed-citation>
			</ref>
			<ref id="B23">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Hersh</surname>
							<given-names>W. R.</given-names>
						</string-name>
						<string-name>
							<surname>Hickam</surname>
							<given-names>D. H.</given-names>
						</string-name>
						<string-name>
							<surname>Haynes</surname>
							<given-names>R. B.</given-names>
						</string-name>
						<string-name>
							<surname>McKibbon</surname>
							<given-names>K. A.</given-names>
						</string-name>
					</person-group>
					<year>1991</year>
					<source>Evaluation of SAPHIRE: an Automated Approach to Indexing and Retrieving Medical Literature</source>
					<conf-name>ProceedingsSymposium on Computer Applications in Medical Care</conf-name>
					<fpage>808</fpage>
					<lpage>812</lpage>
				</mixed-citation>
			</ref>
			<ref id="B24">
				<mixed-citation publication-type="report">
					<person-group person-group-type="author">
						<string-name>
							<surname>Hooper</surname>
							<given-names>R.S.</given-names>
						</string-name>
					</person-group>
					<year>1965</year>
					<source>Indexer consistency tests: origin, measurement, results, and utilization</source>
					<publisher-name>IBM Corporation</publisher-name>
					<gov>TR95-56</gov>
				</mixed-citation>
			</ref>
			<ref id="B25">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Humphrey</surname>
							<given-names>S. M.</given-names>
						</string-name>
						<string-name>
							<surname>Miller</surname>
							<given-names>N. E.</given-names>
						</string-name>
					</person-group>
					<year>1987</year>
					<article-title>Knowledge-Based Indexing of the Medical Literature: The Indexing Aid Project</article-title>
					<source>Journal of the American Society for Information Science</source>
					<volume>38</volume>
					<issue>3</issue>
					<fpage>84</fpage>
					<lpage>196</lpage>
				</mixed-citation>
			</ref>
			<ref id="B26">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Humphrey</surname>
							<given-names>S. M.</given-names>
						</string-name>
					</person-group>
					<year>1999</year>
					<article-title>Automatic Indexing of Documents from Journal Descriptors: A Preliminary Investigation</article-title>
					<source>Journal of the American Society for Information Science</source>
					<volume>50</volume>
					<issue>8</issue>
					<fpage>661</fpage>
					<lpage>674</lpage>
				</mixed-citation>
			</ref>
			<ref id="B27">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Humphrey</surname>
							<given-names>S. M.</given-names>
						</string-name>
						<string-name>
							<surname>Rogers</surname>
							<given-names>W. J.</given-names>
						</string-name>
						<string-name>
							<surname>Kilicoglu</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Demner-Fushman</surname>
							<given-names>D.</given-names>
						</string-name>
						<string-name>
							<surname>Rindflesch</surname>
							<given-names>T. C.</given-names>
						</string-name>
					</person-group>
					<year>2006</year>
					<article-title>Word Sense Disambiguation by Selecting the Best Semantic Type Based on Journal Descriptor Indexing: Preliminary Experiment</article-title>
					<source>Journal of the American Society for Information Science and Technology</source>
					<volume>57</volume>
					<issue>1</issue>
					<fpage>96</fpage>
					<lpage>113</lpage>
				</mixed-citation>
			</ref>
			<ref id="B28">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Irfan</surname>
							<given-names>R.</given-names>
						</string-name>
						<string-name>
							<surname>Khan</surname>
							<given-names>S.</given-names>
						</string-name>
						<string-name>
							<surname>Qamar</surname>
							<given-names>A. M.</given-names>
						</string-name>
						<string-name>
							<surname>Bloodsworth</surname>
							<given-names>P. C.</given-names>
						</string-name>
					</person-group>
					<year>2014</year>
					<article-title>Refining Kea++ Automatic Keyphrase Assignment</article-title>
					<source>Journal of Information Science</source>
					<volume>40</volume>
					<issue>4</issue>
					<fpage>446</fpage>
					<lpage>459</lpage>
					<pub-id pub-id-type="doi">10.1177/0165551514529054</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B29">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Irving</surname>
							<given-names>H. B.</given-names>
						</string-name>
					</person-group>
					<year>1997</year>
					<article-title>Computer-assisted Indexing Training and Electronic Text Conversion at NAL</article-title>
					<source>Knowledge Organization</source>
					<volume>24</volume>
					<issue>1</issue>
					<fpage>4</fpage>
					<lpage>7</lpage>
				</mixed-citation>
			</ref>
			<ref id="B30">
				<mixed-citation publication-type="standard">
					<std>
						<std-organization>ISO</std-organization>
						<pub-id>5963</pub-id>
						<year>1985</year>
						<source>Documentation -- Methods for Examining Documents, Determining their Subjects, and Selecting Indexing Terms</source>
					</std>
					<publisher-loc>Geneva</publisher-loc>
					<publisher-name>ISO</publisher-name>
				</mixed-citation>
			</ref>
			<ref id="B31">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Karetnyk</surname>
							<given-names>D.</given-names>
						</string-name>
						<string-name>
							<surname>Karlsson</surname>
							<given-names>F.</given-names>
						</string-name>
						<string-name>
							<surname>Smart</surname>
							<given-names>G.</given-names>
						</string-name>
					</person-group>
					<year>1991</year>
					<article-title>Knolewledge-based Indexing of Morpho-Syntactically Analysed Language</article-title>
					<source>Expert Systems for Information Management</source>
					<volume>4</volume>
					<issue>1</issue>
					<fpage>1</fpage>
					<lpage>29</lpage>
				</mixed-citation>
			</ref>
			<ref id="B32">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Khan</surname>
							<given-names/>
						</string-name>
						<etal/>
					</person-group>
					<year>2011</year>
					<article-title>A Refined Methodology for Automatic Keyphrase Assignment to Digital Documents</article-title>
					<source>Journal of Digital Information Management</source>
					<volume>9</volume>
					<issue>2</issue>
					<fpage>55</fpage>
					<lpage>63</lpage>
				</mixed-citation>
			</ref>
			<ref id="B33">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Kim</surname>
							<given-names>S. N.</given-names>
						</string-name>
						<string-name>
							<surname>Medelyan</surname>
							<given-names>O.</given-names>
						</string-name>
						<string-name>
							<surname>Kan</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>Baldwin</surname>
							<given-names>T.</given-names>
						</string-name>
					</person-group>
					<year>2013</year>
					<article-title>Automatic Keyphrase Extraction from Scientific Articles</article-title>
					<source>Language Resources and Evaluation</source>
					<volume>47</volume>
					<fpage>723</fpage>
					<lpage>742</lpage>
					<pub-id pub-id-type="doi">10.1007/s10579-012-9210-3</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B34">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Klingbiel</surname>
							<given-names>P. H.</given-names>
						</string-name>
					</person-group>
					<year>1973</year>
					<article-title>A Technique for Machine-Aided Indexing</article-title>
					<source>Information Storage and Retrieval</source>
					<volume>9</volume>
					<issue>9</issue>
					<fpage>477</fpage>
					<lpage>494</lpage>
					<pub-id pub-id-type="doi">10.1016/0020-0271(73)90034-X</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B35">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Krapivin</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>Marchese</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>Yadrantsau</surname>
							<given-names>A</given-names>
						</string-name>
						<string-name>
							<surname>Liang</surname>
							<given-names>Y.</given-names>
						</string-name>
					</person-group>
					<year>2008</year>
					<source>Unsupervised Key-Phrases Extraction from Scientific Papers using Domain and Linguistic Knowledge</source>
					<conf-name>International Conference on Digital Information Management</conf-name>
					<fpage>105</fpage>
					<lpage>112</lpage>
				</mixed-citation>
			</ref>
			<ref id="B36">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Lima</surname>
							<given-names>V. M. A.</given-names>
						</string-name>
						<string-name>
							<surname>Boccato</surname>
							<given-names>V. R. C.</given-names>
						</string-name>
					</person-group>
					<year>2009</year>
					<article-title>O desempenho terminol&#xf3;gico dos descritores em Ci&#xea;ncia da Informa&#xe7;&#xe3;o do Vocabul&#xe1;rio Controlado do SIBi/USP nos processos de indexa&#xe7;&#xe3;o manual, autom&#xe1;tica e semi-autom&#xe1;tica</article-title>
					<source>Perspectivas em Ci&#xea;ncia da Informa&#xe7;&#xe3;o</source>
					<issue>1</issue>
					<fpage>131</fpage>
					<lpage>151</lpage>
				</mixed-citation>
			</ref>
			<ref id="B37">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Lin</surname>
							<given-names>N.</given-names>
						</string-name>
						<string-name>
							<surname>Kudinov</surname>
							<given-names>V.A.</given-names>
						</string-name>
						<string-name>
							<surname>Zaw</surname>
							<given-names>H.M.</given-names>
						</string-name>
						<string-name>
							<surname>Naing</surname>
							<given-names>S.</given-names>
						</string-name>
					</person-group>
					<year>2020</year>
					<source>Query Expansion for Myanmar Information Retrieval Used by WordNet</source>
					<conf-date>2020</conf-date>
					<conf-name>IEEE Conference of Russian Young Researchers in Electrical and Electronic Engineering (EIConRus)</conf-name>
					<fpage>395</fpage>
					<lpage>399</lpage>
				</mixed-citation>
			</ref>
			<ref id="B38">
				<mixed-citation publication-type="thesis">
					<person-group person-group-type="author">
						<string-name>
							<surname>Medelyan</surname>
							<given-names>O.</given-names>
						</string-name>
					</person-group>
					<year>2005</year>
					<source>Automatic keyphrase indexing with a domain-specific thesaurus</source>
					<comment content-type="degree">Master&#x2019;s thesis</comment>
					<publisher-name>Albert-Ludwigs University</publisher-name>
				</mixed-citation>
			</ref>
			<ref id="B39">
				<mixed-citation publication-type="thesis">
					<person-group person-group-type="author">
						<string-name>
							<surname>Medelyan</surname>
							<given-names>O.</given-names>
						</string-name>
					</person-group>
					<year>2009</year>
					<source>Human-competitive automatic topic indexing</source>
					<comment content-type="degree">PhD Thesis</comment>
					<publisher-name>University of Waikato</publisher-name>
					<publisher-loc>New Zealand</publisher-loc>
					<ext-link ext-link-type="uri" xlink:href="https://cds.cern.ch/record/1198029/files/Thesis-2009-Medelyan.pdf">https://cds.cern.ch/record/1198029/files/Thesis-2009-Medelyan.pdf</ext-link>
					<date-in-citation content-type="access-date" iso-8601-date="2021-05-05">05/05/2021</date-in-citation>
				</mixed-citation>
			</ref>
			<ref id="B40">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Mork</surname>
							<given-names>J. G.</given-names>
						</string-name>
						<string-name>
							<surname>Aronson</surname>
							<given-names>A.</given-names>
						</string-name>
						<string-name>
							<surname>Demner-Fushman</surname>
							<given-names>D.</given-names>
						</string-name>
					</person-group>
					<year>2017</year>
					<article-title>12 Years on - Is the NLM Medical Text Indexer Still Useful and Relevant</article-title>
					<source>Journal of Biomedical Semantics</source>
					<volume>8</volume>
					<pub-id pub-id-type="doi">10.1186/s13326-017-0113-5</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B41">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Mynarz</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>&#x160;kuta</surname>
							<given-names>C.</given-names>
						</string-name>
					</person-group>
					<year>2010</year>
					<source>Integration of an Automatic Indexing System within the Document Flow of a Grey Literature Repository</source>
					<conf-name>Twelfth International Conference on Grey Literature</conf-name>
					<conf-loc>Prague</conf-loc>
					<conf-date>December</conf-date>
					<ext-link ext-link-type="uri" xlink:href="http://www.nusl.cz/ntk/nusl-42005">http://www.nusl.cz/ntk/nusl-42005</ext-link>
					<date-in-citation content-type="access-date" iso-8601-date="2021-03-24">24/03/2021</date-in-citation>
				</mixed-citation>
			</ref>
			<ref id="B42">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>N&#xe9;v&#xe9;ol</surname>
							<given-names>A.</given-names>
						</string-name>
						<string-name>
							<surname>Mary</surname>
							<given-names>V.</given-names>
						</string-name>
						<string-name>
							<surname>Gaudinat</surname>
							<given-names>A.</given-names>
						</string-name>
						<string-name>
							<surname>Boyer</surname>
							<given-names>C.</given-names>
						</string-name>
						<string-name>
							<surname>Rogozan</surname>
							<given-names>A.</given-names>
						</string-name>
						<string-name>
							<surname>Darmoni</surname>
							<given-names>S. J.</given-names>
						</string-name>
					</person-group>
					<year>2005</year>
					<source>A Benchmark Evaluation of the French MeSH Indexers</source>
					<series>Lecture Notes in Computer Science</series>
					<fpage>251</fpage>
					<lpage>255</lpage>
					<pub-id pub-id-type="doi">10.1007/11527770_37</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B43">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Rae</surname>
							<given-names>A.</given-names>
						</string-name>
						<string-name>
							<surname>Pritchard</surname>
							<given-names>D.</given-names>
						</string-name>
						<string-name>
							<surname>Mork</surname>
							<given-names>J. G.</given-names>
						</string-name>
						<string-name>
							<surname>Emner-Fushman</surname>
							<given-names>D.</given-names>
						</string-name>
					</person-group>
					<year>2021</year>
					<source>Automatic MeSH Indexing: Revisiting the Subheading Attachment Problem</source>
					<conf-name>Annual Symposium proceedings. AMIA Symposium</conf-name>
					<conf-date>2020</conf-date>
					<fpage>1031</fpage>
					<lpage>1040</lpage>
				</mixed-citation>
			</ref>
			<ref id="B44">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Rolling</surname>
							<given-names>L. N.</given-names>
						</string-name>
					</person-group>
					<year>1981</year>
					<article-title>Indexing Consistency, Quality snd Efficiency</article-title>
					<source>Information Processing and Management</source>
					<volume>17</volume>
					<fpage>69</fpage>
					<lpage>76</lpage>
				</mixed-citation>
			</ref>
			<ref id="B45">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Salisbury</surname>
							<given-names>L.</given-names>
						</string-name>
						<string-name>
							<surname>Smith</surname>
							<given-names>J. J.</given-names>
						</string-name>
					</person-group>
					<year>2014</year>
					<article-title>Building the AgNIC Resource Database Using Semi-Automatic Indexing of Material</article-title>
					<source>Journal of Agricultural &amp; Food Information</source>
					<volume>15</volume>
					<issue>3</issue>
					<fpage>159</fpage>
					<lpage>176</lpage>
					<pub-id pub-id-type="doi">10.1080/10496505.2014.919805</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B46">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>Salton</surname>
							<given-names>G.</given-names>
						</string-name>
					</person-group>
					<year>1989</year>
					<chapter-title>The SMART system 1961-1976: Experiments in Dynamic Document Processing</chapter-title>
					<source>Encyclopedia of Library and Information Science</source>
					<volume>28</volume>
					<fpage>1</fpage>
					<lpage>28</lpage>
				</mixed-citation>
			</ref>
			<ref id="B47">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Salton</surname>
							<given-names>G.</given-names>
						</string-name>
					</person-group>
					<year>1991</year>
					<source>The Smart Document Retrieval Project</source>
					<conf-name>Proceeding SIGIR &#x2018;91 Proceedings of the 14th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</conf-name>
					<fpage>356</fpage>
					<lpage>358</lpage>
				</mixed-citation>
			</ref>
			<ref id="B48">
				<mixed-citation publication-type="webpage">
					<source>Scholastica survey: The State Of Journal Production And Access 2020</source>
					<ext-link ext-link-type="uri" xlink:href="https://lp.scholasticahq.com/journal-production-access-survey/">https://lp.scholasticahq.com/journal-production-access-survey/</ext-link>
					<date-in-citation content-type="access-date" iso-8601-date="2021-10-08">8/10/2021</date-in-citation>
				</mixed-citation>
			</ref>
			<ref id="B49">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Seiler</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>H&#xfc;bner</surname>
							<given-names>P.</given-names>
						</string-name>
						<string-name>
							<surname>Paech</surname>
							<given-names>B.</given-names>
						</string-name>
					</person-group>
					<year>2019</year>
					<source>Comparing Traceability through Information Retrieval, Commits, Interaction Logs, and Tags</source>
					<conf-date>2019</conf-date>
					<conf-name>IEEE/ACM 10th International Symposium on Software and Systems Traceability (SST)</conf-name>
					<fpage>21</fpage>
					<lpage>28</lpage>
				</mixed-citation>
			</ref>
			<ref id="B50">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Shams</surname>
							<given-names>R.</given-names>
						</string-name>
						<string-name>
							<surname>Mercer</surname>
							<given-names>R. E.</given-names>
						</string-name>
					</person-group>
					<year>2012a</year>
					<source>Investigating Keyphrase Indexing with Text Denoising</source>
					<conf-name>Proceedings of the 12th ACM/IEEE-CS Joint Conference on Digital Libraries - JCDL &#x2019;12</conf-name>
					<pub-id pub-id-type="doi">10.1145/2232817.2232866</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B51">
				<mixed-citation publication-type="book">
					<person-group person-group-type="author">
						<string-name>
							<surname>Shams</surname>
							<given-names>R.</given-names>
						</string-name>
						<string-name>
							<surname>Mercer</surname>
							<given-names>R.E.</given-names>
						</string-name>
					</person-group>
					<year>2012b</year>
					<chapter-title>Improving Supervised Keyphrase Indexer Classification of Keyphrases with text Denoising</chapter-title>
					<source>Lecture Notes in Computer Science</source>
					<fpage>77</fpage>
					<lpage>86</lpage>
				</mixed-citation>
			</ref>
			<ref id="B52">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Silva</surname>
							<given-names>S. R. de B.</given-names>
						</string-name>
						<string-name>
							<surname>Corr&#xea;a</surname>
							<given-names>R. F.</given-names>
						</string-name>
					</person-group>
					<year>2020</year>
					<article-title>Sistemas de Indexa&#xe7;&#xe3;o autom&#xe1;tica por atribui&#xe7;&#xe3;o: uma an&#xe1;lise comparativa</article-title>
					<source>Encontros Bibli: Revista Eletr&#xf4;nica De Biblioteconomia E Ci&#xea;ncia Da Informa&#xe7;&#xe3;o</source>
					<volume>25</volume>
					<fpage>1</fpage>
					<lpage>25</lpage>
					<pub-id pub-id-type="doi">10.5007/1518-2924.2020.e70740</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B53">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Silva</surname>
							<given-names>S. R. de B.</given-names>
						</string-name>
						<string-name>
							<surname>Corr&#xea;a</surname>
							<given-names>R. F.</given-names>
						</string-name>
						<string-name>
							<surname>Gil-Leiva</surname>
							<given-names>I.</given-names>
						</string-name>
					</person-group>
					<year>2020</year>
					<article-title>Avalia&#xe7;&#xe3;o direta e conjunta de Sistemas de Indexa&#xe7;&#xe3;o Autom&#xe1;tica por Atribui&#xe7;&#xe3;o</article-title>
					<source>Informa&#xe7;&#xe3;o &amp; Sociedade-Estudos</source>
					<volume>30</volume>
					<fpage>1</fpage>
					<lpage>27</lpage>
					<pub-id pub-id-type="doi">10.22478/ufpb.1809-4783.2020v30n4.57259</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B54">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Silvester</surname>
							<given-names>J. P.</given-names>
						</string-name>
						<string-name>
							<surname>Genuardi</surname>
							<given-names>M. T.</given-names>
						</string-name>
						<string-name>
							<surname>Klingbiel</surname>
							<given-names>P. H.</given-names>
						</string-name>
					</person-group>
					<year>1994</year>
					<article-title>Machine-Aided Indexing at NASA</article-title>
					<source>Information Processing &amp; Management</source>
					<volume>30</volume>
					<issue>5</issue>
					<fpage>631</fpage>
					<lpage>645</lpage>
				</mixed-citation>
			</ref>
			<ref id="B55">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Sinkkil&#xe4;</surname>
							<given-names>R.</given-names>
						</string-name>
						<string-name>
							<surname>Suominen</surname>
							<given-names>O.</given-names>
						</string-name>
						<string-name>
							<surname>Hyv&#xf6;nen</surname>
							<given-names>E.</given-names>
						</string-name>
					</person-group>
					<year>2011</year>
					<source>Automatic Semantic Subject Indexing of Web Documents in Highly Inflected Languages</source>
					<conf-name>Proceedings The Semantic Web: Research and Applications : 8th Extended Semantic Web Conference, ESWC 2011</conf-name>
					<conf-loc>Heraklion, Crete, Greece</conf-loc>
					<conf-date>May 29-June 2</conf-date>
					<fpage>215</fpage>
					<lpage>229</lpage>
					<pub-id pub-id-type="doi">10.1007/978-3-642-21034-1_15</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B56">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Souza-Rocha</surname>
							<given-names>R.</given-names>
						</string-name>
						<string-name>
							<surname>Gil-Leiva</surname>
							<given-names>I.</given-names>
						</string-name>
					</person-group>
					<year>2016</year>
					<source>Automatic Indexing of Scientific Texts: A Methodological Comparison</source>
					<person-group person-group-type="editor">
						<string-name>
							<surname>Chaves Guimar&#xe3;es</surname>
							<given-names>J. A.</given-names>
						</string-name>
						<string-name>
							<surname>Oliveira Milani</surname>
							<given-names>S.</given-names>
						</string-name>
						<string-name>
							<surname>Dodebei</surname>
							<given-names>V.</given-names>
						</string-name>
					</person-group>
					<conf-name>Knowledge Organization for a Sustainable World: Challenges and Perspectives for Cultural, Scientific, and Technological Sharing in a Connected Society: Proceedings of the Fourteenth International ISKO Conference</conf-name>
					<conf-date>27-29 September 2016</conf-date>
					<conf-loc>W&#xfc;rzburg, Rio de Janeiro, Brazil</conf-loc>
					<conf-sponsor>Ergon Verlag</conf-sponsor>
					<fpage>243</fpage>
					<lpage>250</lpage>
				</mixed-citation>
			</ref>
			<ref id="B57">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Suominen</surname>
							<given-names>O.</given-names>
						</string-name>
					</person-group>
					<year>2019</year>
					<article-title>Annif: DIY Automated Subject Indexing using Multiple Algorithms</article-title>
					<source>LIBER Quarterly</source>
					<volume>29</volume>
					<issue>1</issue>
					<fpage>1</fpage>
					<lpage>25</lpage>
					<pub-id pub-id-type="doi">10.18352/lq.10285</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B58">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Wang</surname>
							<given-names>D.X.</given-names>
						</string-name>
						<string-name>
							<surname>Gao</surname>
							<given-names>X.</given-names>
						</string-name>
						<string-name>
							<surname>Andreae</surname>
							<given-names>P.</given-names>
						</string-name>
					</person-group>
					<year>2015</year>
					<article-title>DIKEA: Exploiting Wikipedia for keyphrase extraction</article-title>
					<source>Web Intelligence</source>
					<volume>13</volume>
					<issue>3</issue>
					<fpage>153</fpage>
					<lpage>165</lpage>
				</mixed-citation>
			</ref>
			<ref id="B59">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Witten</surname>
							<given-names>I. H.</given-names>
						</string-name>
						<string-name>
							<surname>Paynter</surname>
							<given-names>G. W.</given-names>
						</string-name>
						<string-name>
							<surname>Frank</surname>
							<given-names>E.</given-names>
						</string-name>
						<string-name>
							<surname>Gutwin</surname>
							<given-names>C.</given-names>
						</string-name>
						<string-name>
							<surname>Nevill-Manning</surname>
							<given-names>C. G.</given-names>
						</string-name>
					</person-group>
					<year>1999</year>
					<source>KEA: Practical Automatic Keyphrase Extraction</source>
					<conf-name>Proceedings of the fourth ACM conference on Digital libraries</conf-name>
					<page-range>254-255, 243-250</page-range>
					<pub-id pub-id-type="doi">10.1145/313238.313437</pub-id>
				</mixed-citation>
			</ref>
		</ref-list>
		<app-group>
			<app id="app1">
				<label>7.</label>
				<title>Annex</title>
				<table-wrap id="ta1">
					<label>Table A</label>
					<caption>
						<title>The three best results and three non-zero worst results of SISA, KEA and MAUI</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="left"> </th>
								<th align="center">Article</th>
								<th align="center">Assigned terms</th>
								<th align="center">Terms of Gold indexing</th>
								<th align="center">Terms commons</th>
								<th align="center">Precision</th>
								<th align="center">Recall</th>
								<th align="center">F-measure</th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="left"> </td>
								<td align="center">166</td>
								<td align="center">8</td>
								<td align="center">6</td>
								<td align="center">5</td>
								<td align="center">75%</td>
								<td align="center">100%</td>
								<td align="center">85%</td>
							</tr>
							<tr>
								<td align="center">S</td>
								<td align="center">193</td>
								<td align="center">10</td>
								<td align="center">9</td>
								<td align="center">8</td>
								<td align="center">80%</td>
								<td align="center">88%</td>
								<td align="center">84%</td>
							</tr>
							<tr>
								<td align="center">I</td>
								<td align="center">61</td>
								<td align="center">7</td>
								<td align="center">5</td>
								<td align="center">5</td>
								<td align="center">71%</td>
								<td align="center">100%</td>
								<td align="center">83%</td>
							</tr>
							<tr>
								<td align="center">S</td>
								<td align="center">91</td>
								<td align="center">7</td>
								<td align="center">20</td>
								<td align="center">2</td>
								<td align="center">14%</td>
								<td align="center">5%</td>
								<td align="center">7.4%</td>
							</tr>
							<tr>
								<td align="center">A</td>
								<td align="center">87</td>
								<td align="center">13</td>
								<td align="center">7</td>
								<td align="center">1</td>
								<td align="center">7%</td>
								<td align="center">14%</td>
								<td align="center">10.0%</td>
							</tr>
							<tr>
								<td align="left"> </td>
								<td align="center">94</td>
								<td align="center">6</td>
								<td align="center">12</td>
								<td align="center">1</td>
								<td align="center">16%</td>
								<td align="center">8%</td>
								<td align="center">11.1%</td>
							</tr>
							<tr>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
							</tr>
							<tr>
								<td align="left"> </td>
								<td align="center">7</td>
								<td align="center">10</td>
								<td align="center">8</td>
								<td align="center">6</td>
								<td align="center">60%</td>
								<td align="center">75%</td>
								<td align="center">66%</td>
							</tr>
							<tr>
								<td align="center">K</td>
								<td align="center">147</td>
								<td align="center">10</td>
								<td align="center">6</td>
								<td align="center">5</td>
								<td align="center">50%</td>
								<td align="center">83%</td>
								<td align="center">62%</td>
							</tr>
							<tr>
								<td align="center">E</td>
								<td align="center">86</td>
								<td align="center">10</td>
								<td align="center">9</td>
								<td align="center">6</td>
								<td align="center">60%</td>
								<td align="center">66%</td>
								<td align="center">63%</td>
							</tr>
							<tr>
								<td align="center">A</td>
								<td align="center">133</td>
								<td align="center">10</td>
								<td align="center">8</td>
								<td align="center">1</td>
								<td align="center">10%</td>
								<td align="center">12%</td>
								<td align="center">11.1%</td>
							</tr>
							<tr>
								<td align="left"> </td>
								<td align="center">17</td>
								<td align="center">10</td>
								<td align="center">7</td>
								<td align="center">1</td>
								<td align="center">10%</td>
								<td align="center">14%</td>
								<td align="center">11.7%</td>
							</tr>
							<tr>
								<td align="left"> </td>
								<td align="center">18</td>
								<td align="center">10</td>
								<td align="center">6</td>
								<td align="center">1</td>
								<td align="center">10%</td>
								<td align="center">16%</td>
								<td align="center">12.5%</td>
							</tr>
							<tr>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
								<td align="left"> </td>
							</tr>
							<tr>
								<td align="left"> </td>
								<td align="center">141</td>
								<td align="center">10</td>
								<td align="center">7</td>
								<td align="center">6</td>
								<td align="center">60%</td>
								<td align="center">85%</td>
								<td align="center">70%</td>
							</tr>
							<tr>
								<td align="center">M</td>
								<td align="center">195</td>
								<td align="center">10</td>
								<td align="center">5</td>
								<td align="center">5</td>
								<td align="center">50%</td>
								<td align="center">100%</td>
								<td align="center">66%</td>
							</tr>
							<tr>
								<td align="center">A</td>
								<td align="center">193</td>
								<td align="center">10</td>
								<td align="center">9</td>
								<td align="center">6</td>
								<td align="center">60%</td>
								<td align="center">66%</td>
								<td align="center">63.1%</td>
							</tr>
							<tr>
								<td align="center">U</td>
								<td align="center">131</td>
								<td align="center">10</td>
								<td align="center">14</td>
								<td align="center">1</td>
								<td align="center">10%</td>
								<td align="center">7%</td>
								<td align="center">8.3%</td>
							</tr>
							<tr>
								<td align="center">I</td>
								<td align="center">119</td>
								<td align="center">10</td>
								<td align="center">11</td>
								<td align="center">1</td>
								<td align="center">10%</td>
								<td align="center">9%</td>
								<td align="center">9.5%</td>
							</tr>
							<tr>
								<td align="left"> </td>
								<td align="center">16</td>
								<td align="center">10</td>
								<td align="center">9</td>
								<td align="center">1</td>
								<td align="center">10%</td>
								<td align="center">11%</td>
								<td align="center">10.5%</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
			
				<table-wrap id="tb2">
					<label>Table B</label>
					<caption>
						<title>Indexing terms of the non-zero worst results of SISA, KEA and MAUI</title>
					</caption>
					<table>
						<colgroup>
							<col span="2"/>
						</colgroup>
						<thead>
							<tr>
								<th align="center" colspan="2">SISA (Document 91). F-measure = 7.4% </th>
							</tr>
							<tr>
								<th align="justify">Gold indexing</th>
								<th align="justify">SISA indexing</th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="justify">1. Acceso a la informaci&#xf3;n 2. An&#xe1;lisis bibliogr&#xe1;fico 3. Archivos abiertos 					<break/>4. Autenticaci&#xf3;n 5. Ciudadanos <break/>6. Datos abiertos vinculados 7. Demanda de informaci&#xf3;n <break/>8. Documentaci&#xf3;n 9. Documentos <break/>10. Estudios de casos 11. Gesti&#xf3;n de la informaci&#xf3;n 12. Informaci&#xf3;n 13. <bold>Informaci&#xf3;n p&#xfa;blica</bold> 14. Instituciones p&#xfa;blicas 15. Internet <break/>16. Publicaciones oficiales 17. Reciclaje <break/>18. Reutilizaci&#xf3;n 19. <bold>Sector p&#xfa;blico</bold> 20. Tecnolog&#xed;as de la informaci&#xf3;n y la comunicaci&#xf3;n (TIC)</td>
								<td align="justify">1. Acceso 2. Datos 3. <bold>Informaci&#xf3;n p&#xfa;blica</bold>
									<break/>4. Reutilizaci&#xf3;n de informaci&#xf3;n 5. <bold>Sector p&#xfa;blico</bold> 6. Tic <break/>7. Usuarios</td>
							</tr>
							<tr>
								<td align="center" colspan="2">
									<bold>KEA (Document 133). F-measure = 11.1% </bold>
								</td>
							</tr>
							<tr>
								<td align="justify">
									<bold>Gold indexing</bold>
								</td>
								<td align="justify">
									<bold>KEA indexing</bold>
								</td>
							</tr>
							<tr>
								<td align="justify">1. Comunicaci&#xf3;n 2. Deontolog&#xed;a 3. Editores <break/>4. Educaci&#xf3;n 5. Espa&#xf1;a 6. &#xc9;tica 7. Psicolog&#xed;a <break/>8. <bold>Revistas cient&#xed;ficas</bold>
								</td>
								<td align="justify">1. Aspectos &#xe9;ticos 2. Ciencias Sociales <break/>3. Colaboraci&#xf3;n cient&#xed;fica 4. CSIC <break/>5. Humanidades 6. Psicolog&#xed;a de la Educaci&#xf3;n <break/>7. Revistas 8. <bold>Revistas cient&#xed;ficas</bold> 9. Trabajo de investigaci&#xf3;n 10. Universidad</td>
							</tr>
							<tr>
								<td align="center" colspan="2">
									<bold>MAUI (Document 131). F-measure = 8.3%</bold>
								</td>
							</tr>
							<tr>
								<td align="justify">
									<bold>Gold indexing</bold>
								</td>
								<td align="justify">
									<bold>MAUI indexing</bold>
								</td>
							</tr>
							<tr>
								<td align="justify">1. Carrera profesional 2. CERN 3. Cient&#xed;ficos <break/>4. Colaboraci&#xf3;n cient&#xed;fica 5. Estudios de casos <break/>6. Experimentaci&#xf3;n cient&#xed;fica 7. Ginebra <break/>8. Historia de la ciencia 9. Investigadores <break/>10. Organizaci&#xf3;n de la investigaci&#xf3;n <break/>11. Proyecto Atlas 12. Relaciones laborales <break/>13. Suiza 14. <bold>Transmisi&#xf3;n de conocimientos</bold>
								</td>
								<td align="justify">1. Atlas 2. Cooperaci&#xf3;n 3. Desarrollo profesional 4. Experimento 5. F&#xed;sica de part&#xed;culas 6. Miembros 7. Organizaci&#xf3;n social <break/>8. Toma de decisiones 9. Trabajadores <break/>10. <bold>Transmisi&#xf3;n de conocimientos</bold>
								</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
			</app>
		</app-group>
	</back>
</article>