Indonesian cross-linguistic named entity recognition

Closed

Danang Arbian Sulistyo, Aji Prasetya Wibawa, Didik Dwi Prasetya, Fadhli Almu'iini Ahda

2025 Research Methods in Applied Linguistics Vol. 4 Issue 3 Article Cited by 4 Quartile

Abstract

This study examines the potential of Named Entity Recognition (NER) in translating cross-biblical texts of Indonesian, Madurese, and Javanese. The goal is to enhance translation precision by incorporating entity categorization. The approach involves training an NER model using Conditional Random Fields (CRF) and evaluating its performance on the Book of Joshua. The annotated dataset includes features such as word identity, shape, part-of-speech identifiers, and semantic information. Tagging the data with labels such as Person, Location, and Organization reveals variations in effectiveness across languages. Indonesian yields the highest F1 score (78.69), reflecting consistent performance across all parameters. Although Madurese achieves a high recall for Location entities (82.16), its precision is lower (74.99). Javanese demonstrates strong precision in identifying locations (77.46), but a slightly lower recall score (77.21). The findings suggest the need to tailor the NER model to suit the specific characteristics of low-resource languages for improved translation quality. © 2025

Affiliations

Universitas Negeri Malang, Indonesia; Institut Teknologi dan Bisnis Asia Malang, Indonesia