Capturing and Structuring Data Mined from the Web.
The article discusses data mining methodology, with a focus on methods of capturing and structuring data collected from the Web as of March 2014. Information is provided on Web crawler datasets, uniform resource locator (URL) frontiers, which are URL collections, and scheduling and tracking URL craw...
| Publicado en: | Communications of the ACM Vol. 57; no. 3; pp. 10 - 12 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
Mar2014
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| Sumario: | The article discusses data mining methodology, with a focus on methods of capturing and structuring data collected from the Web as of March 2014. Information is provided on Web crawler datasets, uniform resource locator (URL) frontiers, which are URL collections, and scheduling and tracking URL crawls. Data extraction methods, parsing Web pages, post-processing URL links, and suggested open source crawlers are also discussed. A list of further data mining computer network resources is also provided. |
|---|