Capturing and Structuring Data Mined from the Web.

The article discusses data mining methodology, with a focus on methods of capturing and structuring data collected from the Web as of March 2014. Information is provided on Web crawler datasets, uniform resource locator (URL) frontiers, which are URL collections, and scheduling and tracking URL craw...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 57; no. 3; pp. 10 - 12
Autor principal: Matsudaira, Kate
Formato: Artículo
Publicado: Association for Computing Machinery Mar2014
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=94802520&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 94802520
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00010782
        ACM
      jtl: Communications of the ACM
      issn: 00010782
      maglogo: N
    pubinfo:
      dt: Mar2014
      vid: 57
      iid: 3
      pid: 68
      pub: Association for Computing Machinery
    artinfo:
      ui:
        94802520
        10.1145/2567664
      ppf: 10
      ppct: 2
      formats:
      tig:
        atl: Capturing and Structuring Data Mined from the Web.
      aug:
        au: Matsudaira, Kate
        affil: Vice president of Decide.com
      su:
        Data mining
        Data entry
        Data structures
        Data
        Uniform Resource Locators
        Data extraction
        Computer scheduling
        Big data
      sug:
        subj:
          Data mining
          Data entry
          Data structures
          Data
          Uniform Resource Locators
          Data extraction
          Computer scheduling
          Big data
      ab: The article discusses data mining methodology, with a focus on methods of capturing and structuring data collected from the Web as of March 2014. Information is provided on Web crawler datasets, uniform resource locator (URL) frontiers, which are URL collections, and scheduling and tracking URL crawls. Data extraction methods, parsing Web pages, post-processing URL links, and suggested open source crawlers are also discussed. A list of further data mining computer network resources is also provided.
      pubtype: Periodical
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2014
    holdings:
      @attributes:
        islocal: N