Can Your Business Scrape Data From Other Websites? Legal Risks To Know

Alex Solo
byAlex Solo8 min read

Looking at what other businesses are doing is part of running a business. You might compare prices, check what competitors are offering or look at what products and services are already in the market.

But there’s a difference between browsing information online and using software to collect that information automatically, particularly at scale.

Just because data is publicly available does not necessarily mean your business can scrape, copy or reuse it however it wants. What matters can depend on what you collect, how you access it and what you plan to do with it afterwards.

If your business uses web scraping tools, or relies on data collected from other websites, there are a few legal issues worth checking before you start.

Can Businesses Legally Scrape Websites In The US?

First, it helps to be clear about what web scraping actually means.

Looking through a competitor’s website, checking its prices or researching its services is not usually what people mean by scraping. Web scraping generally involves using software, scripts, bots or other automated tools to pull information from webpages, often repeatedly or at scale.

For example, an online retailer might use a tool to track prices across hundreds of competitor product pages. A startup might collect publicly displayed business information to build a directory. Another business might use publicly available web data to help power a comparison tool or other product.

Web scraping itself is not automatically illegal in the US. However, that does not mean every type of scraping is allowed.

The answer can depend on how you access the website, what information you collect and what you do with it afterwards. A tool collecting basic information from pages available to anyone can raise quite different issues from one getting around access restrictions or copying large amounts of someone else’s original content.

That is why it helps to look at each part of the process separately.

Check The Website’s Terms Of Service

A good place to start is the website’s Terms of Service.

Some websites specifically restrict automated scraping, crawling, copying or commercial use of their content. Others may set conditions around how their platform or data can be accessed.

So, even if you can open a page without a password, that does not necessarily mean the website allows businesses to automatically collect everything on it and reuse that information commercially.

Whether particular website terms apply or can be enforced will depend on the circumstances. However, ignoring them altogether can create unnecessary risk.

Before setting up a scraper, it is worth checking the terms that apply to the website and whether they say anything about automated access, data extraction or reuse.

For businesses operating their own websites or platforms, clear Website Terms of Service can also help set expectations around how other people are allowed to access and use the site.

Are You Accessing Public Pages Or Bypassing Restrictions?

The way your business accesses the information matters too.

The federal Computer Fraud and Abuse Act (CFAA) deals with certain forms of unauthorized computer access.

In Van Buren v United States, the US Supreme Court considered when someone who is allowed to access a computer nevertheless “exceeds authorized access” under the CFAA. The Court took a relatively narrow approach, focusing on whether the person accessed areas of a computer system that were actually off limits to them rather than simply using information they were allowed to access for an improper purpose.

For web scraping, hiQ Labs v LinkedIn provides a more direct example.

hiQ operated a business that collected information from public LinkedIn profiles using automated tools. The Ninth Circuit found that hiQ had raised serious questions about whether accessing information available to anyone on the internet amounted to access “without authorization” under the CFAA.

That does not mean businesses have a general right to scrape anything that appears publicly online.

It does, however, highlight an important distinction. Collecting information from a page that anyone can access without logging in is different from trying to get around passwords, authentication requirements or other technical barriers to reach information that is not openly available.

So, if a scraping tool is circumventing access controls rather than simply collecting information from public pages, the legal risk can look quite different.

Are You Scraping Facts Or Someone Else’s Content?

How you access a website is only part of the issue. You also need to look at what your business is actually copying.

Copyright generally does not protect facts themselves. It can, however, protect the original way those facts or ideas are expressed, including writing, photographs, artwork and other original website content.

Imagine you run an online store and want to keep track of competitor pricing.

Automatically collecting publicly displayed prices is different from copying the competitor’s full product descriptions, photographs and other content and publishing those materials on your own site.

The same issue can arise with directories or databases. Individual pieces of factual information may not themselves be protected by copyright, while original content or, in some circumstances, the particular selection or arrangement of material may raise separate copyright questions.

The important point is that being able to access information does not automatically give your business the right to reproduce everything around it.

Does The Data Include Personal Information?

There is another issue to consider if the information you are collecting relates to people rather than businesses, products or prices.

A scraping tool might collect names, email addresses, professional profiles or other information linked to individuals. Depending on where your business and the people involved are located, privacy laws may then become relevant.

California is a useful example.

Under the California Consumer Privacy Act (CCPA), certain information that qualifies as “publicly available” is excluded from the definition of personal information. However, “publicly available” has a specific meaning under the law. The California Privacy Protection Agency explains that it can include information a business reasonably believes has lawfully been made available to the general public by the consumer, through widely distributed media, or in certain other circumstances.

That is different from assuming that anything you happen to find somewhere online is automatically outside privacy law.

If web scraping forms part of a broader customer or user data operation, businesses may need to look more closely at their  data and privacy obligations rather than treating scraping as a standalone technical process.

What Will Your Business Do With The Data?

Once you have worked out what can be collected, there is still another question: what are you planning to do with it?

Say you collect publicly displayed product prices so your team can understand where your products sit in the market. That is quite different from copying another website’s content and republishing it as your own.

Likewise, gathering business information for internal research is different from creating a commercial database and selling access to it.

The same applies if scraped data becomes part of a product. A startup might collect information from multiple public sources and use it to power a search tool, directory or analytics product. At that point, the business may be relying on that data as a core commercial asset rather than simply using it for research.

That makes it important to think about both sides of the question:

Can we collect this information, and do we have the right to use it in the way our business intends?

Where one business is giving another the right to use a dataset, a Data License Agreement can help set out what data can be used, for what purpose and subject to what restrictions.

Using A Third-Party Scraping Tool

Your business might not be running its own scraper at all.

There are plenty of tools that collect web data, generate prospect lists, monitor competitors or provide ready-made datasets. Using one of those services can make the technical side much easier, but it does not automatically answer the legal questions.

For example, a provider might tell you its data comes from “public sources”. That is useful information, but it does not necessarily tell you how the data was collected, whether another website placed restrictions on its use or what rights your business actually receives.

Before becoming dependent on a dataset or scraping provider, it is worth understanding where the information comes from and checking the agreement you have with the provider.

Does it give your business the right to use the data commercially? Can you put it into your own product? Can you provide it to customers? What happens if a third party claims the data was collected or used improperly?

Those questions become particularly important where the data is central to the product or service you are selling.

What Should You Check Before Scraping A Website?

You do not necessarily need to avoid web scraping altogether. The important thing is understanding what your business is actually doing before you build a process around the data.

Before you start, check:

  • what information you actually need to collect
  • whether the website places restrictions on automated access or reuse
  • whether your tool is getting around any login or technical access controls
  • whether you are extracting facts or copying protected content
  • whether the data includes information about individuals
  • how your business plans to use the information afterwards
  • whether any third-party data provider gives you the rights you need

For a one-off research exercise, some of these issues may be relatively straightforward. If scraping is going to power a product, build a lead database or become an ongoing part of your business model, it is worth working through the legal position before you become dependent on the data.

Key Takeaways

Web scraping is not automatically illegal in the US, but the fact that information is publicly visible does not necessarily mean your business can collect and use it without restriction.

The key questions are usually how you are accessing the website, what you are collecting and what you want to do with the data afterwards.

Website terms, access restrictions, copyright and privacy laws can all become relevant depending on the circumstances. If you are using a third-party scraping or data provider, it is also worth checking what rights you actually receive rather than assuming “public data” means unrestricted data.

If web scraping or third-party data is going to form an important part of your product or business model, getting the legal position clear early can help avoid having to rethink how the data is collected or used later.

If you have any legal concerns about website scraping, you can reach us at (888) 449-8437 or team@sprintlaw.com for a free, no-obligations chat.

Alex Solo
Alex SoloCo-Founder

Alex is Sprintlaw's co-founder and a legal technology leader. He holds law and media degrees from the University of Sydney and has been recognized by Australasian Lawyer, Lawyers Weekly and the Sydney Young Entrepreneur Awards for his work building Sprintlaw and improving access to business legal support.

Need legal help?

Get in touch with our team

Tell us what you need and we'll come back with a fixed-fee quote - no obligation, no surprises.

Need support?

Need help with your business legals?

Speak with Sprintlaw to get practical legal support and fixed-fee options tailored to your business.