Standardizing Extracted Data Using Automated Application of Controlled Vocabularies.

BACKGROUND: Extraction of toxicological end points from primary sources is a central component of systematic reviews and human health risk assessments. To ensure optimal use of these data, consistent language should be used for end point descriptions. However, primary source language describing trea...

Descripción completa

Detalles Bibliográficos
Publicado en:Environmental Health Perspectives Vol. 132; no. 2; pp. 027006-1 - 27019
Autores principales: Foster, Caroline, Wignall, Jessica, Kovach, Samuel, Choksi, Neepa, Allen, Dave, Trgovcich, Joanne, Rochester, Johanna R., Ceger, Patricia, Daniel, Amber, Hamm, Jon, Truax, Jim, Blake, Bevin, McIntyre, Barry, Sutherland, Vicki, Stout, Matthew D., Kleinstreuer, Nicole
Formato: algorithm glossary research tables/charts Journal Article
Publicado: National Institute of Environmental Health Sciences Feb2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=175957076&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 175957076
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        00916765
        3B5
      jtl: Environmental Health Perspectives
      issn: 00916765
      maglogo: N
    pubinfo:
      dt: Feb2024
      vid: 132
      iid: 2
      pid: 56539
      pub: National Institute of Environmental Health Sciences
      place: Research Triangle Park, North Carolina
    artinfo:
      ui:
        175957076
        175957076
        175957076
        10.1289/EHP13215
        175957076
      ppf: 027006-1
      ppct: 13
      formats:
      tig:
        atl: Standardizing Extracted Data Using Automated Application of Controlled Vocabularies.
      aug:
        au:
          Foster, Caroline
          Wignall, Jessica
          Kovach, Samuel
          Choksi, Neepa
          Allen, Dave
          Trgovcich, Joanne
          Rochester, Johanna R.
          Ceger, Patricia
          Daniel, Amber
          Hamm, Jon
          Truax, Jim
          Blake, Bevin
          McIntyre, Barry
          Sutherland, Vicki
          Stout, Matthew D.
          Kleinstreuer, Nicole
        affil: ICF, Durham, North Carolina, USA
      sug:
        subj:
          Automation
          Artificial Intelligence Utilization
          Vocabulary Standards
          Vocabulary, Controlled
          Human
          Male
          Female
          Unified Medical Language System
          Dose-Response Relationship
          Drug Toxicity
          Fetal Development
          Toxicity Tests
          Data Curation
          Data Analysis Software
          Reproductive Health
          Metadata
          Quality Assessment
          Descriptive Statistics
          Male
          Female
      ab: BACKGROUND: Extraction of toxicological end points from primary sources is a central component of systematic reviews and human health risk assessments. To ensure optimal use of these data, consistent language should be used for end point descriptions. However, primary source language describing treatment-related end points can vary greatly, resulting in large labor efforts to manually standardize extractions before data are fit for use. OBJECTIVES: To minimize these labor efforts, we applied an augmented intelligence approach and developed automated tools to support standardization of extracted information via application of preexisting controlled vocabularies. METHODS: We created and applied a harmonized controlled vocabulary crosswalk, consisting of Unified Medical Language System (UMLS) codes, German Federal Institute for Risk Assessment (BfR) DevTox harmonized terms, and The Organization for Economic Co-operation and Development (OECD) end point vocabularies, to roughly 34,000 extractions from prenatal developmental toxicology studies conducted by the National Toxicology Program (NTP) and 6,400 extractions from European Chemicals Agency (ECHA) prenatal developmental toxicology studies, all recorded based on the original study report language. RESULTS: We automatically applied standardized controlled vocabulary terms to 75% of the NTP extracted end points and 57% of the ECHA extracted end points. Of all the standardized extracted end points, about half (51%) required manual review for potential extraneous matches or inaccuracies. Extracted end points that were not mapped to standardized terms tended to be too general or required human logic to find a good match. We estimate that this augmented intelligence approach saved >350 hours of manual effort and yielded valuable resources including a controlled vocabulary crosswalk, organized related terms lists, code for implementing an automated mapping workflow, and a computationally accessible dataset. DISCUSSION: Augmenting manual efforts with automation tools increased the efficiency of producing a findable, accessible, interoperable, and reusable (FAIR) dataset of regulatory guideline studies. This open-source approach can be readily applied to other legacy developmental toxicology datasets, and the code design is customizable for other study types.
      pubtype: Academic Journal
      doctype:
        algorithm
        glossary
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N