We turn web content into useful data to help you power your business with the data it needs.

Lead Generation

Competitor Analysis

Market research

Custom Solutions

Fell Free To contact Us
To let us know about your data extraction projects requirements and get access to accurate on target information.

info@scrapingbox.com

Top

Market Research & Article Generation for French Media Portal

Market Research & Article Generation for French Media Portal

Challenge : The client wanted content to be extracted on a continuous basis from French news sites to power their news portal. The list of websites included popular blogs, news sites, forums and a few content bookmarking sites. The required data points were date of publishing, author name, title, main text content and tags.

Solution : Once we were provided with the list of source websites and data points, our team started working on the project. As this use case was for news data, the frequency of crawls had to be very high. This meant fresh data sets had to be provided every day. Since each site in the list had a different structure and design, site specific crawl and extraction was the solution used for this case. Once our team finished setting up the web crawlers, the data started flowing in. The data was then cleaned and formatted to be uploaded to the client’s Dropbox servers in XML format. The number of records being delivered per day was above 300,000.

Benefits :

  1. Our team handled every technical aspect of the crawling process
  2. The initial setup only took 48 hours to get completed and the supply of data was consistent
  3. Since monitoring was set up for each site, the quality and consistency of data was top notch
  4. Although some of the sites used dynamic coding practices, our tech stack could handle them well
  5. The client could launch their news portal with the data in a short notice
  6. The cost incurred was way less than what an in-house crawling set up would have costed them