Next Article in Journal
Root Resorption of Adjacent Teeth Associated with Maxillary Canine Impaction in the Saudi Arabian Population: A Cross-Sectional Cone-Beam Computed Tomography Study
Previous Article in Journal
Development of a Pre-Evaluation and Health Monitoring System for FAST Cable-Net Structure
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Iterative Named Entity Recognition with Conditional Random Fields

1
Zentrale Stelle fuer Informationstechnik im Sicherheitsbereich, 81677 Muenchen, Germany
2
Faculty Applied Computer Sciences and Biosciences, University of Applied Sciences Mittweida, 09648 Mittweida, Germany
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Appl. Sci. 2022, 12(1), 330; https://doi.org/10.3390/app12010330
Submission received: 18 November 2021 / Revised: 20 December 2021 / Accepted: 23 December 2021 / Published: 30 December 2021
(This article belongs to the Section Computing and Artificial Intelligence)

Abstract

Named entity recognition (NER) constitutes an important step in the processing of unstructured text content for the extraction of information as well as for the computer-supported analysis of large amounts of digital data via machine learning methods. However, NER often relies on domain-specific knowledge, being conducted manually in a time- and human-resource-intensive process. These can be reduced with statistical models performing NER automatically. The current work investigates whether Conditional Random Fields (CRF) can be efficiently trained for NER in German texts, by means of an iterative procedure combining self-learning with a manual annotation–active learning–component. The training dataset increases continuously with the iterative procedure. Whilst self-learning did not markedly improve the performance of the CRF for NER, the manual annotation of sentences with the lowest probability of correct prediction clearly improved the model F1-score and simultaneously reduced the amount of manual annotation required to train the model. A model with an F1-score of 0.885 was able to be trained in 11.4 h.
Keywords: active learning; self-learning; text; annotation; language active learning; self-learning; text; annotation; language

Share and Cite

MDPI and ACS Style

Alves-Pinto, A.; Demus, C.; Spranger, M.; Labudde, D.; Hobley, E. Iterative Named Entity Recognition with Conditional Random Fields. Appl. Sci. 2022, 12, 330. https://doi.org/10.3390/app12010330

AMA Style

Alves-Pinto A, Demus C, Spranger M, Labudde D, Hobley E. Iterative Named Entity Recognition with Conditional Random Fields. Applied Sciences. 2022; 12(1):330. https://doi.org/10.3390/app12010330

Chicago/Turabian Style

Alves-Pinto, Ana, Christoph Demus, Michael Spranger, Dirk Labudde, and Eleanor Hobley. 2022. "Iterative Named Entity Recognition with Conditional Random Fields" Applied Sciences 12, no. 1: 330. https://doi.org/10.3390/app12010330

APA Style

Alves-Pinto, A., Demus, C., Spranger, M., Labudde, D., & Hobley, E. (2022). Iterative Named Entity Recognition with Conditional Random Fields. Applied Sciences, 12(1), 330. https://doi.org/10.3390/app12010330

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop