A real time Named Entity Recognition system for Arabic text mining.

Arabic is the most widely spoken language in the Arab World. Most people of the Islamic World understand the Classic Arabic language because it is the language of the Qur'an. Despite the fact that in the last decade the number of Arabic Internet users (Middle East and North and East of Africa) has i...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 46; no. 4; pp. 543 - 564
Autores principales: Al-Jumaily, Harith, Martínez, Paloma, Martínez-Fernández, José, Goot, Erik
Formato: Artículo
Publicado: Springer Nature Dec2012
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Arabic is the most widely spoken language in the Arab World. Most people of the Islamic World understand the Classic Arabic language because it is the language of the Qur'an. Despite the fact that in the last decade the number of Arabic Internet users (Middle East and North and East of Africa) has increased considerably, systems to analyze Arabic digital resources automatically are not as easily available as they are for English. Therefore, in this work, an attempt is made to build a real time Named Entity Recognition system that can be used in web applications to detect the appearance of specific named entities and events in news written in Arabic. Arabic is a highly inflectional language, thus we will try to minimize the impact of Arabic affixes on the quality of the pattern recognition model applied to identify named entities. These patterns are built up by processing and integrating different gazetteers, from DBPedia (, ) to GATE (A general architecture for text engineering, ) and ANERGazet ().