Web Scraping Techniques and Applications: A Literature Review
Authors: Chaimaa Lotfi, Swetha Srinivasan, Myriam Ertz and Imen Latrous
Publishing Date: 25-04-2022
ISBN: 978-93-91842-08-6
Abstract
Big data analytics gives organizations a way to analyze huge data sets and gather new information. It helps answer basic questions about business operations and business performance. It also helps discover unknown patterns in vast datasets or combinations thereof. In the current data-driven world, it becomes increasingly essential that big data techniques are applied and analyzed for organizational growth. More specifically, with the large availability of data on the Web, whether from social media, websites, online portals, or platforms, to name but a few, it is important for organizations to know how to mine that data in order to extract useful knowledge. Web scraping represents a fundamental approach in this regard. Therefore, this paper aims to provide an updated literature review about the most advanced Web Scraping techniques to better equip scholars and managers with helpful knowledge on how to mine most effectively online data. The paper starts with presenting the basic design of a web scraper and the applications of web scraping in diverse sectors and areas. Next, the different Web scraping methods and Web scraping technologies are presented. Finally, a procedure to develop Web scraping with various tools is proposed before a conclusion wraps up the paper.
Keywords
Big data, web scraping, business performance, web crawling, web mining.
Cite as
Chaimaa Lotfi, Swetha Srinivasan, Myriam Ertz and Imen Latrous, "Web Scraping Techniques and Applications: A Literature Review", In: Raju Pal and Praveen Kumar Shukla (eds), SCRS Conference Proceedings on Intelligent Systems, SCRS, India, 2022, pp. 381-394. https://doi.org/10.52458/978-93-91842-08-6-38