The Free And Open Web Is Under Attack At The IETF
from the the-open-web-includes-the-ability-to-scrape dept
The ability to access publicly available information using automated tools is a central value and benefit of a free and open internet. Automated access—often called crawling or scraping—powers important, useful tools for locating, preserving, and analyzing online information. For example, crawling and scraping helps journalists, researchers, and watchdog organizations report the news, find security flaws, and investigate discrimination. Crawling the web allows non-profits like the Internet Archive to preserve historical copies of websites. Tools for automated comparison shopping allow consumers to find the best deals on items they want to buy. And so on.
Yet the open internet access is increasingly under threat from publishers and Big Tech companies alike. Fearing lost advertising and licensing revenues, website operators increasingly claim that they need to lock down their sites from bots that crawl public web content to train or operate AI models. Some companies are even trying to embed their business models into internet standards by changing Internet Engineering Task Force (IETF) technical standards that shape much of the internet.
Many of their economic anxieties are understandable. AI bots can strain websites’ infrastructure, in some cases, degrading site performance or taking them offline altogether. Upgrading systems costs money that some sites may not have. And AI is likely to disrupt the business models many publishers adopted in response to the rise of the internet, if users rely on AI overviews instead of visiting source websites.
However reasonable these fears may be, the answer is not to change the IETF standards from neutral protocols that encourage openness to restrictive requirements designed to monetize internet access.
The worst of these proposed standards would give websites far greater ability to automatically block legitimate, lawful scraping and crawling. For example, the AI Preferences working group is working on proposals to give publishers a way to express “preference signals” against crawling web data for AI-related purposes, including to train models, generate outputs, and help users search the web. These preference signals would be expressed through robots.txt and could potentially become legally binding in some jurisdictions.
Another working group, called Web Bot Auth, is pursuing efforts to protect sites from overly-aggressive bots that strain website resources—a positive goal that could meaningfully improve the internet in the AI era. But Web Bot Auth is simultaneously pursuing a much more dangerous path as well: standards changes that would enable sites to cryptographically identify bots so that they can more easily block anyone they wish—not just “bad” actors, but competitors, dissidents, or anyone who hasn’t paid for the right to access sites using automated tools. If sites restrict crawling to a preapproved list of cryptographically authenticated bots, they could require licensing payments from those wishing to crawl their sites. This would close off the open web to researchers, archivists, and startups without the ability to pay for automated access.
Websites may have legitimate reasons to worry about AI’s impacts on their traffic and advertising revenue, but those reasons must be weighed against the benefits of the open web. These proposals would effectively give website operators veto power over a wide range of important uses—from the investigations and archival works described above to accessibility tools for people with disabilities, to research efforts aimed at holding governments accountable.
That is why we are fighting back against these threats to open access. EFF and our allies in the open internet community have successfully resisted some of the most dangerous IETF proposals thus far—and won’t stop working to protect the open web from efforts to manipulate internet standards to undermine the right to freely access the internet in any legal way, including with automated tools.
Republished from the EFF’s Deeplinks blog.
Filed Under: ai, open access, open web, scraping
Companies: ietf


Comments on “The Free And Open Web Is Under Attack At The IETF”
This is all going to happen regardless. AI and all the AI fanatics made the choice to be giant disrespectful assholes to everyone else.
Websites spamming ads lead to everyone using ad blockers.
Robo calls lead to most ignoring unknown numbers.
Unsolicited door to door salesman lead increased laws on trespassing and solicitation.
The answer isn’t to whine about how people are reacting to bad actors. The answer is to punish the bad actors so badly that no one will act that way.
Provide an alternative solution, then
Small websites are getting absolutely crushed by automated tools–most of which are AI bots–and I don’t see you providing a real solution to that.
You don’t like robots.txt? Propose a better solution.
You don’t like bot auth? Propose a better solution.
You don’t have a better solution? Shut up.
Because you know what’s worse than websites blocking bots? Websites shuttering because they can’t justify the expense of Altman’s or Amodei’s robot pulling down the whole site 100 times a minute.
You know what’s worse for free speech than blocking bots? Websites shuttering.
You know what’s worse for research than blocking bots? Websites shuttering.
Websites shuttering is a WAY worse outcome on LITERALLY EVERY METRIC, isn’t it? If websites don’t exist, they can’t be crawled, they can’t be researched, they can’t be seen by anyone.
I really appreciate the EFF most of the time, but sometimes, my god. Provide a real alternative solution, or shut up about it.
Re:
I keep hearing this… and yet, we haven’t seen that here. And we don’t really take any steps to block AI scrapers (even though we could).
Every so often we get a bot that goes haywire, but it’s always been relatively easy to deal with.
I’m sure that some sites are overwhelmed with AI scraping traffic, but I find it odd that so many people insist it’s killing smaller sites… and we just haven’t seen it at all.
So… can someone explain why there’s a supposed flood of bot traffic overwhelming sites, but doesn’t seem to be hitting us? I’m not saying that it’s not happening. I just don’t understand why we don’t see it here.
Re: Re:
I suspect it’s overstated or misattributed too. “Altman’s or Amodei’s robot” respects robots.txt, so they are definitely not responsible for these issues. Probably the issue is just some data harvesters fucked up their bot config because they replaced programmers with chatGPT.
Re: Re:
I’m going to presume that you’re asking a serious and sincere question, although I have difficulty believing that anyone could be this naive.
You’re not seeing it here for the same reason that certain other sites aren’t seeing it: this site’s on an exemption list. This shouldn’t surprise you or anyone else: abusers have been doing this for decades, e.g. spammers started building exemption lists in the early 1990’s. Some of the entries on such lists are intended to avoid critics with platforms, some are intended to avoid allies, and some are intended to avoid irritating people in a position to respond effectively.
Techniques for doing this are well-known and mature: this isn’t speculative or experimental. And the marginal loss of utility, whether not spamming users at one domain or not scraping one web site, is insignificant when compared to the payoff…which is neatly encapsulated in your own remarks.
The rest of the world is not so lucky, which is why there are literally hundreds of public and private projects working on blocking AI/LLM scrapers. Just as we’ve had to spend a fortune in time and effort dealing with spammers, we’ve now been forced to spend a fortune dealing with the thugs at the AI/LLM companies. Understand: nobody wants to do this. We have better things to do. But we have no choice: it’s either do this or watch our sites burn. And sadly, the people most affected by this are the ones just trying to do some good in the world: libraries, archives, and the like. All they’re trying to do is preserve human knowledge and culture, and the AI/LLM companies are doing their best to destroy them.
Re: Re: Re:
Why would we be on an exemption list? We’ve made it clear we’re fine with AI scraping and are well represented in most AI tools I know of.
So I don’t understand this claim that they would exempt us. Why?
Re: Re:
Racism doesn’t exisn’t, I haven’t seen it.
Could you have had a dumber fucking response?
Re: Re: Re:
This is not even remotely equivalent.
I asked a legitimate question: I run a small site. We have done nothing to prevent AI scraping. If it’s true that AI scraping is overwhelming small sites, I expect we would have seen it. We have not.
So I’m legitimately confused as to why.
I said (directly, which you ignore — why?) that it’s entirely possible other sites are experiencing it. I’m just legitimately confused why we haven’t.
It honestly makes no sense to me. We should see this influx of bot traffic. And we haven’t. So I was hoping someone could explain why.
The only thing someone said is that we’ve been “exempted” but that doesn’t make sense to me. Who would exempt us? And why?
Re: Re:
TechDirt’s a site that serves almost pure text (and small amounts of it, at that) and doesn’t change frequently. Scraping it impinges far less than it might other sites.
Re:
It was hilarious recently when the operator of a wiki written entirely in spoonerisms announced it was down because it had exceeded its traffic allowance in a single day when anthropic’s bot sent it almost a million requests, and wasn’t even all that annoyed, because they’d essentially poisoned their own training data in the process.
Re: Re:
When he realises, he’ll be Am Saltman.
I am not convinced there is, or should be, a right for bots to access whatever sites they want for whatever reason they want.
You're fighting the wrong battle
None of this would even be up for discussion if the AI/LLM companies hadn’t done their best not only to grab every bit content with no regard for copyright, copyleft, licensing, terms-of-use, or anything else, but if they hadn’t also launched ongoing 24×7 denial-of-service attacks against millions of web sites.
I run several web sites that are among the first to ever be created. They’ve had their issues over all these years (decades) but it wasn’t until now that I had to restrict them — as a last resort — because all other technical measures failed. (And please: spare me “…but did you try?” Of course I did. I’m not new to the Internet.)
The AI/LLM companies are using public clouds, private clouds, hijacked systems, hijacked IOT devices, workload distribution, obfuscation, and every other trick in the rather large book to evade defenses. I’ve had to spend thousands of hours trying to solve problems that I should never have had, and that equates to a lot of time that wasn’t spent doing something productive, like improving sites or upgrading hardware.
So stop this foolish victim-blaming nonsense and target the assholes responsible for this mess: the AI/LLM companies. They’re the ones who bear 100% of the responsibility/culpability for this mess. Which — if you had done ANY research on this issue at all — you would already know.
Re:
Can you share the evidence that your traffic is from AI companies and which ones they are?
It’s crazy how the open web somehow held on for so many years and it’s now being completely dismantled in so short a time. So many people who should know better have decided that emerging technology constitutes casus belli against cyberspace.
What is the answer, then?
It’s funny how the thing that gets weighed against the benefits of the open web, is always the response to irresponsible AI scraping. Never the actual irresponsible AI scraping that caused it in the first place.
Unsustainable costs is just as much a threat to the open web as a change in protocol. It would be nice if EFF prioritized it as such.
Re:
Why blame the abuser when the victim is right there?
Re: Re: Blame the victim
That’s par for the course here in the US.
Re: Re: Re: Blame the victim
This isn’t a problem that’s exclusive to the US!
Yes, I realize that many of the big LLM/AI companies are legally domiciled in the US, but they hire people from all over the world to work for them. Also, we keep hearing about how China is going to surpass us in AI, so presumably there is Chinese AI/LLM website scraping too. (not that I’ve heard of it being a problem for any Western Hemisphere or UK website owners.)
More relevant: The AI/LLM behemoths scrape websites regardless of where the website owners live. It isn’t a form of US mercantilism. Maybe there’s preference for English language sites. If so, I suspect it is temporary, because Sam Altman and Anthropic et al are going to want to sell their chatbots and tokens to everyone.
It is incredibly disingenuous for Mike and the EFF to say that pillaging of entire websites’ content is a form of “open access”!!! Not to mention the examples given here in the replies, of the automated AI/LLM company crawlers that are so aggressive they interfere with even keeping one’s site online.
IETF Web Bot Auth WG not exactly restrictive
I visited IETF datatracker for the Web Bot Auth working group. The timestamps for most of the draft documents were 26 June 2026 i.e. yesterday. Then I glanced at the most recent meeting minutes https://datatracker.ietf.org/doc/minutes-interim-2026-webbotauth-02-202604131700/
AI/LLM bots that scrape and gulp training data don’t seem under serious threat (i.e. being cut off from the “open web” given this:
The first item on the agenda was whether alleviating volumetric abuse by AI bots. The WG vote was 7 in favor and 2 against. It could have been worse, but… sheesh
robots.txt is a way of withdrawing consent for crawling and scraping. It’s reasonable for there to be a way to withdraw consent, and robots.txt is the least harmful way of doing so. Site owners are perfectly within their rights to deny access to bots by other measures (loginwalls, paywalls, IP bans…). If giving robots.txt the force of law means site owners don’t have to do those things, that sounds fine to me.
What I’d propose is to take a leaf out of the GDPR book, and acknowledge that there are other bases, besides consent, on which it should be lawful for a bot to crawl or scrape a site. You’re worried about cases where there is a public interest or other legitimate interest in crawling and scraping, even when the site owner withdraws consent. So then it should be legal to ignore robots.txt when those legitimate interests exist. But the bot should have to identify itself with a URL in the user-agent string pointing at a page which says who operates the bot, what the lawful basis is, and how to contact them about it; just as anyone collecting personal data has to explain their lawful basis for doing so under the GDPR.