A pragmatic guide to geoparsing evaluation: Toponyms, Named Entity Recognition and pragmatics.
Empirical methods in geoparsing have thus far lacked a standard evaluation framework describing the task, metrics and data used to compare state-of-the-art systems. Evaluation is further made inconsistent, even unrepresentative of real world usage by the lack of distinction between the different typ...
| Publicado en: | Language Resources & Evaluation Vol. 54; no. 3; pp. 683 - 713 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=144950770&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 144950770 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2020 vid: 54 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 144950770 10.1007/s10579-019-09475-3 ppf: 683 ppct: 30 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.3MB tig: atl: A pragmatic guide to geoparsing evaluation: Toponyms, Named Entity Recognition and pragmatics. aug: au: Gritta, Milan Pilehvar, Mohammad Taher Collier, Nigel affil: Language Technology Lab (LTL), Department of Theoretical and Applied Linguistics (DTAL), University of Cambridge, 9 West Road, CB3 9DP, Cambridge, UK su: Geographic names Corpora Linguistic analysis Named-entity recognition Machine learning Taxonomy sug: subj: Geographic names Corpora Linguistic analysis Named-entity recognition Machine learning Taxonomy keyword: Evaluation framework Geocoding Geonames Geoparsing Geotagging Named Entity Recognition Natural language understanding Pragmatics Toponym resolution Toponyms ab: Empirical methods in geoparsing have thus far lacked a standard evaluation framework describing the task, metrics and data used to compare state-of-the-art systems. Evaluation is further made inconsistent, even unrepresentative of real world usage by the lack of distinction between the different types of toponyms, which necessitates new guidelines, a consolidation of metrics and a detailed toponym taxonomy with implications for Named Entity Recognition (NER) and beyond. To address these deficiencies, our manuscript introduces a new framework in three parts. (Part 1) Task Definition: clarified via corpus linguistic analysis proposing a fine-grained Pragmatic Taxonomy of Toponyms. (Part 2) Metrics: discussed and reviewed for a rigorous evaluation including recommendations for NER/Geoparsing practitioners. (Part 3) Evaluation data: shared via a new dataset called GeoWebNews to provide test/train examples and enable immediate use of our contributions. In addition to fine-grained Geotagging and Toponym Resolution (Geocoding), this dataset is also suitable for prototyping and evaluating machine learning NLP models. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|