Capturing and Structuring Data Mined from the Web.

The article discusses data mining methodology, with a focus on methods of capturing and structuring data collected from the Web as of March 2014. Information is provided on Web crawler datasets, uniform resource locator (URL) frontiers, which are URL collections, and scheduling and tracking URL craw...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 57; no. 3; pp. 10 - 12
Autor principal: Matsudaira, Kate
Formato: Artículo
Publicado: Association for Computing Machinery Mar2014
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:The article discusses data mining methodology, with a focus on methods of capturing and structuring data collected from the Web as of March 2014. Information is provided on Web crawler datasets, uniform resource locator (URL) frontiers, which are URL collections, and scheduling and tracking URL crawls. Data extraction methods, parsing Web pages, post-processing URL links, and suggested open source crawlers are also discussed. A list of further data mining computer network resources is also provided.