BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/You Can Read Their Price. Reading It a Million Tim…
ToolsSeptember 16, 2026
Read · 5 min
price monitoring · competitor analysis

You Can Read Their Price. Reading It a Million Times Is Not

Watching a rival's prices is ordinary retail. Automating it crosses three separate legal questions, and only one of them is about whether the page was public.

Key takeaways
  • Looking at a competitor's published price is not a legal problem. Collecting it automatically, at volume, every day, raises three separate questions and only one of them is about access.
  • The three are: did you break into anything, did you agree to terms that forbid this, and is what you took a protected database.
  • The most cited scraping case is usually summarised backwards. hiQ won on the computer misuse claim and then lost on breach of contract.
  • Robots.txt is a request, not a permission. The standard says so in its own text, which cuts both ways.
  • In Europe a price list can attract a database right if the maker invested substantially in obtaining, verifying or presenting it, and repeated systematic extraction of small parts still infringes.
  • For most small shops the honest answer is to monitor a dozen products by hand or with a light touch, because the legal exposure scales with volume and the value does not.

Every shopkeeper in history has walked past a competitor and looked at the price in the window. Nobody thinks that is wrong. The instinct carries over to the web, where the price is public, the page loads for anyone and reading it feels identical.

Then you automate it. A script checks forty competitors twice a day, stores the history and feeds a repricing rule, and you have crossed from looking into something that three different areas of law have opinions about. The awkward part is that none of them turn on whether the page was public.

What are the three questions, and why do they get merged?

They get merged because the popular question is whether scraping is legal, which has no answer. The useful questions are narrower and they can be answered separately for your specific situation.

The first is about access. Did you get into somewhere you were not permitted to be, in the sense that computer misuse law understands permission. The second is about contract. Did you agree to terms that forbid automated collection, and did you agree to them in a way that binds you. The third is about the data itself. Is the collection you took protected in its own right, independently of any copyright in the individual items.

A price monitoring setup can pass the first and fail the second, which is the common outcome and the one most write ups miss.

Diagram comparing the two legal questions kept separate in price scraping, getting in through public pages and logins under computer misuse law, and using the data under contract and database law

Does public mean permitted?

For the access question, largely yes. For everything else, no. The distinction was drawn sharply by the American courts and the result is more nuanced than the headlines suggested.

The starting point is Van Buren v. United States, decided 3 June 2021 by six votes to three. The Supreme Court read the Computer Fraud and Abuse Act's phrase about exceeding authorised access as a gates up or down inquiry: one either can or cannot access a computer system, and one either can or cannot access certain areas within it. Misusing information you were entitled to reach is not the offence. Reaching areas you were not entitled to reach is.

That reading fed directly into the scraping case everyone quotes. In hiQ Labs v. LinkedIn, the Ninth Circuit had allowed hiQ to keep collecting publicly available profile data even after a cease and desist letter. The Supreme Court vacated and remanded in light of Van Buren, and in April 2022 the Ninth Circuit reaffirmed its earlier decision.

Then the part almost nobody quotes. In November 2022 the district court found that hiQ had breached LinkedIn's user agreement, and the parties settled. The company won the argument about computer misuse and lost the argument about the contract it had accepted. If you take one thing from the case law, take that: the access question was the weaker weapon all along.

QuestionWhat triggers itWhat reduces your riskWhat does not help
Did you break inCredentials, paywalls, rate limit evasion, areas behind a loginOnly collecting pages that load for a logged out visitorThe page being interesting or the data being factual
Did you accept termsCreating an account, clicking to accept, continued use after noticeNever registering an account with the site you monitorArguing the terms were unreasonable
Is it a protected databaseSubstantial investment in obtaining, verifying or presenting contentsTaking small, non systematic samplesThe individual prices being unoriginal facts
Did you cause harmLoad on their servers, degraded service for real customersLow request rates, caching, respecting error responsesSaying your traffic was a small share of theirs
Did you misleadForged user agents, rotating addresses to avoid blocksIdentifying your crawler honestlyEveryone else doing the same thing

The pattern in the right hand column is that almost every mitigation is a decision about volume and honesty rather than about technology. That is worth internalising before evaluating any tool, because the vendor's selling points are frequently in the column marked what does not help.

Is robots.txt permission?

No, and it is not prohibition either. It is a request, and the specification says so about itself, which is unusual and useful.

RFC 9309, published in September 2022 on the standards track, defines the robots exclusion protocol as rules that crawlers are requested to honour when accessing URIs. It sets out how matching works, that the most specific match must be used and the most specific match is the one with the most octets, that matching is case sensitive, and that where no rule matches, the URI is allowed. It also tells crawlers they may cache the file but should not use a cached copy for more than 24 hours unless the file is unreachable.

Then it states the thing that matters for this article: these rules are not a form of access authorization. A permissive robots.txt does not give you a licence, and a restrictive one does not by itself make you a lawbreaker. What it does is establish that you knew, which is why ignoring an explicit disallow makes every other argument harder to run.

The practical reading for a merchant is that robots.txt is evidence rather than law. Honouring it is cheap, and it removes an accusation that would otherwise sit underneath every other question.

The mirror image is worth thinking about at the same time, since your own shop is being read by other people's automation constantly. The decision about who you let in and on what terms is a real one with revenue attached, which we set out in the piece on whether to allow, block or charge the AI crawlers on your shop.

Why does Europe treat a price list differently?

Because Europe protects the investment in assembling a collection, separately from any rights in its contents. A list of facts nobody could copyright can still be a protected database.

Article 7 of the Database Directive gives the maker of a database the right to prevent extraction and re-utilisation of all or a substantial part of the contents, where there has been qualitatively or quantitatively a substantial investment in obtaining, verifying or presenting those contents. Extraction is defined as the permanent or temporary transfer of contents to another medium by any means or in any form. Re-utilisation is any form of making the contents available to the public.

Two features of that text decide real cases. The first is that the right attaches to investment in obtaining, verifying or presenting, which is why a catalogue that somebody spent money to compile and check can qualify while a list somebody generated as a by product of their own business may not. The second is the clause that catches monitoring specifically: repeated and systematic extraction of insubstantial parts is caught where it conflicts with normal exploitation of the database or unreasonably prejudices the maker's legitimate interests.

Read that last sentence with a daily price crawler in mind. Each individual pull is trivially insubstantial. Three hundred and sixty five of them, systematically, reconstructing the whole thing over time, is exactly the behaviour the clause was written for. Volume is what converts a defensible activity into an indefensible one, and that is the opposite of how most people intuit the risk.

Note

None of this is legal advice and the position varies by country, by the terms you actually accepted and by what you do with the data afterwards. The value of knowing the three questions is that it turns a vague worry into three specific things you can check about your own setup in an afternoon.

So what should a small shop actually do?

Monitor much less than a tool will offer, and be boring about how. The value of competitor price data falls off a cliff after the first handful of products, while the exposure rises with every additional request.

Start by working out which prices actually change your decisions. For most small catalogues that is a short list: the items where you and a competitor are close substitutes, where the buyer compares before purchase, and where you have room to move. An item that is unique to you needs no monitoring. An item where you cannot change the price without going below cost needs no monitoring either, because the information cannot lead to an action.

Then set the collection to match the frequency at which prices genuinely move in your category. Daily is right for consumer electronics. Weekly is right for most homeware, food and craft. Hourly is right for almost nobody selling fewer than a thousand lines, and it is the setting that turns a modest activity into repeated and systematic extraction.

Be visible. Use a user agent that names your shop and gives a contact address, obey robots.txt, and respond to a 429 by backing off rather than by rotating addresses. This gets you blocked more often and sued less often, and for a business with no legal budget that is the correct trade. The alternative, evasion, is the single fact most likely to turn a civil argument into an aggravated one.

Card naming the three behaviours that raise legal risk fastest in competitor price monitoring, registering an account, hourly collection, and rotating addresses to evade blocks

Is it safer to buy a monitoring service?

It moves the collection risk and leaves the use risk with you. That is a real benefit and it is smaller than the sales page implies.

A vendor collecting once and serving many customers is doing something more defensible than forty merchants each running their own crawler, and they usually have licensing arrangements or feeds you could not get alone. What transfers less neatly is what you do with the output. If you republish a competitor's prices on your own site, or build a comparison page from them, you are making the data available to the public, which is re-utilisation in the language of the directive regardless of who collected it.

Ask a vendor two questions before buying. Where does the data come from for the specific competitors you care about, feed, partnership or crawl. And what does the contract say about your right to display it publicly rather than to use it internally. Most answers to the second question are more restrictive than buyers assume.

Does a marketplace count as a competitor's site?

It counts as the marketplace's site, which is a harder target and a different relationship. The prices belong to sellers, the page belongs to the platform, and the terms you are bound by are the platform's.

This matters because many small merchants already have an account on the marketplace they want to monitor, often as a seller. That account is the contract question answered against you before you start. Having clicked to accept seller terms, you cannot credibly argue you never agreed to anything, and marketplace terms are usually explicit about automated collection. A merchant who sells on a platform and also scrapes it is in the weakest position of anyone in this article.

The alternative that is actually available is the platform's own tooling. Most large marketplaces publish an API, a reporting interface or a competitive pricing feed for sellers, and those routes are licensed rather than tolerated. They give you less than a crawler would, they are rate limited in ways that feel mean, and they carry no risk at all. For a seller already inside the platform, that trade is close to obvious once written down.

The second thing to know is that a marketplace price is frequently not the price. Shipping thresholds, platform promotions funded by the platform rather than the seller, and subscription discounts all sit between the number on the page and what a buyer pays. A monitoring setup that compares your delivered price against a competitor's list price is comparing two different things and will push you to cut prices you did not need to cut. If you are going to measure anything, measure the total a buyer would pay at checkout, which usually means fewer products monitored more carefully.

What do you do with the number once you have it?

Less than you think, and slower than the tool wants. A competitor's price is one input into your pricing, not an instruction, and matching automatically hands your margin to whoever is most willing to lose money.

The specific failure is a race with a competitor who is also automated. Two repricers pointed at each other will find the floor within days, and neither business chose that outcome. A rule that says never go below a margin floor, and never move more than once a week, keeps the information useful and the decision yours. The boundaries around that whole decision, including the discount reference price rules in Europe, are set out in the piece on dynamic pricing for a small shop and the line not to cross.

The other thing to do with it is nothing, deliberately. If your position is that you are more expensive because you ship faster and answer the phone, then a competitor's lower price is confirmation rather than a problem, and the correct response is to make the difference visible on the product page rather than to close the gap.

What if somebody is scraping you?

Then the same three questions run in your favour, and your terms are the strongest of the three. This is the part most merchants never set up, and it costs an hour.

Publish a robots.txt that says what you mean, so that ignoring it is a documented choice by whoever does. Put a clear clause in your own terms of service stating whether automated collection is permitted and on what conditions, because a contract term is what hiQ actually lost on. Rate limit by address and by pattern rather than by user agent string alone, since honest crawlers identify themselves and the ones you worry about do not.

All three require control of your own storefront's configuration, which on a rented platform often means whatever the platform decided. Running a shop whose headers, robots file and terms you can actually edit is the unglamorous version of this advice, and it is one of the reasons we generate stores with code the merchant owns and can change rather than a template with settings.

"These rules are not a form of access authorization."RFC 9309, the Robots Exclusion Protocol

The short version

Reading a competitor's public price is fine. The risk comes from three things you control: whether you ever accepted their terms, how much you take and how often, and whether you republish it. Keep the first at never, the second at modest, and the third at no, and competitor price monitoring stays the ordinary retail activity it has always been.

The tools will encourage the opposite on all three counts, because volume is what they sell. A dozen products checked weekly will change your pricing decisions almost as much as ten thousand checked hourly, and it will do so without putting a letter from somebody's lawyer in your inbox.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building