An Efficient Approach for Web Indexing of Big Data through Hyperlinks in Web Crawling.
Web Crawling has acquired tremendous significance in recent times and it is aptly associated with the substantial development of the World Wide Web. Web Search Engines face new challenges due to the availability of vast amounts of web documents, thus making the retrieved results less applicable to t...
| Publicado en: | Scientific World Journal Vol. 2015; pp. 739286 - 739287 |
|---|---|
| Autores principales: | , , |
| Formato: | Journal Article |
| Publicado: |
Wiley-Blackwell
1/1/2015
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=109593803&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 109593803 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 1537744X 1BX5 jtl: Scientific World Journal issn: 1537744X maglogo: N pubinfo: dt: 1/1/2015 vid: 2015 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 109593803 NLM26137592 2013090022 10.1155/2015/739286 NLM26137592 PMC4475586 109593803 ppf: 739286 ppct: 1 formats: tig: atl: An Efficient Approach for Web Indexing of Big Data through Hyperlinks in Web Crawling. aug: au: Devi, R Suganya Manjula, D Siddharth, R K sug: ab: Web Crawling has acquired tremendous significance in recent times and it is aptly associated with the substantial development of the World Wide Web. Web Search Engines face new challenges due to the availability of vast amounts of web documents, thus making the retrieved results less applicable to the analysers. However, recently, Web Crawling solely focuses on obtaining the links of the corresponding documents. Today, there exist various algorithms and software which are used to crawl links from the web which has to be further processed for future use, thereby increasing the overload of the analyser. This paper concentrates on crawling the links and retrieving all information associated with them to facilitate easy processing for other uses. In this paper, firstly the links are crawled from the specified uniform resource locator (URL) using a modified version of Depth First Search Algorithm which allows for complete hierarchical scanning of corresponding web links. The links are then accessed via the source code and its metadata such as title, keywords, and description are extracted. This content is very essential for any type of analyser work to be carried on the Big Data obtained as a result of Web Crawling. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|