1. Home
  2. Marketing Glossary

Web scraping

What is web scraping?

Web scraping, often referred to as “web page scraping”, consists of automatically browsing a website while extracting the data found, so that it can later be analyzed and manipulated based on certain parameters.

The application or software created to scrape is called a bot, spider or crawler. Many websites try to protect themselves from these applications to safeguard their data. Captchas are one example: they appear on many subscription forms and not only stop fake email accounts from getting into our subscriber database, but also prevent crawlers from accessing certain areas of a website.

1. What are the goals of web scraping?

The information obtained is very valuable, which is why “data scraping” is carried out for a wide range of purposes. You could say they are as endless as the possibilities of data mining, but some of the most common are:

  • Building email databases: perhaps one of the most obvious uses. Those addresses are then used to create databases for sending spam.
  • Getting to know your competitors: by scraping their website you obtain data that isn’t visible at first glance and that is very valuable for positioning yourself in the market.
  • Monitoring and comparing online offers: staying aware at all times of the offers available on other websites.
  • Generating alerts to monitor the aspects of a website we want to keep an eye on. Finding broken links in order to fix them and improve your SEO strategy.
  • Monitoring competitors’ prices and spotting trends, which lets you determine websites’ pricing strategies and react if necessary.
  • Keeping track of any changes to a website, so we know about every change made to our own website or to others.
  • Tracking online reputation and presence, which shows the position search engines give to a particular blog’s posts.
  • Collecting product pages: for ecommerce businesses, it is very useful to know how competitors’ product pages are built in order to improve their own.
  • Collecting data from several websites and comparing it to learn about the trends and techniques those websites use in various areas of interest.

2. Is web scraping legal?

This question is very common and the answer is that sometimes it is legal and sometimes it is not.

In other words, scrapers must always take into account the intellectual property rights of the website so that this cannot be considered illegal, and it is legal as long as the data obtained is freely available to third parties on the website itself.

Website owners often offer an API so that scraping isn’t necessary and the data can be obtained easily. Nobody, or almost nobody, minds Google’s crawler accessing their website to index its content and, in doing so, earn the top positions in the SERPs. To scrape legally, these aspects should be taken into account:

  • The collected data cannot be used for illegal or harmful purposes.
  • The website’s intellectual property and legal rights must always be respected.
  • If user registration or a usage contract is required, such data may not be collected by scraping.
  • Website owners are entitled to put technical barriers in place to prevent web scraping, and these must not be ignored.

3. How can we protect ourselves from web scraping?

Even if you explicitly state on your website that you don’t allow web scraping, there will always be people who will try to do so, so it is necessary that you implement a series of actions to protect yourself, such as:

  • Adjusting the .htaccess file according to the patterns of the IPs that try to scrape your site, that is, blocking them.
  • Controlling incoming requests: identifying IPs and filtering them in the firewall is a very effective way to try to prevent your website from being “scraped”.
  • Detecting and preventing hotlinking, so that our server’s resources can’t be used in unauthorized places.
  • Limiting requests per IP address, so an attacker can’t open multiple connections from the same IP.
  • Modifying the HTML structure: since crawlers focus on parsing the HTML, changing it fairly often makes it harder for an attacker to scrape your website easily.
  • Offering an API, so you can monitor and restrict the data that can be extracted from your site. This doesn’t prevent malicious web scraping, but it greatly reduces the number of times our website is exposed to data scraping.
  • Using honeypots or links to fake content, that is, specific content that isn’t visible to a normal visitor of our website. This lets you detect unwanted crawlers; you need to disallow those links in the robots.txt file for search engine bots.
  • Using cross-site request forgery (CSRF) tokens, so you prevent bot automations from making abusive requests.

Related entries

Find out how Mailrelay can help you with your email marketing strategy.