The most interesting thing in tech: will the Wayback Machine become a casualty of the battle between bot blockers and scrapers? It is an amazing resource, but lots of publishers---including us---have had to cut off the site to keep AI companies from scraping it for their content. This is a very hard problem to solve, and I worry it’s a sign of things to come for services that rely on data from publishers and content creators.
I recently worked on a project that simply would not have been possible without the Wayback Machine or similar service, as I needed to see a company's web site and product offerings over time, and it just wasn't a big enough company for the press to be bothered following its every move. I was impressed just how far back they go! This point in time is looking like it will be a weird spot in web history when we use tools like the Wayback Machine.
Really thoughtful concern about where this leads for data-driven services overall 💙⚡ It feels like the harder we try to protect content, the more careful we need to be about what we unintentionally block in the process!
Wayback Machine is like an archival book in a library, rarely used, but very important. Ideally, a technical solution is the fix. AI companies should take note and adjust it so it doesn't spread widely online.
Of course, the Internet Archive itself has been found to violate copyright on a mass scale... https://authorsguild.org/news/how-to-tell-internet-archive-to-remove-your-books/#:~:text=Update%2C%20September%206%2C%202024%3A,ruling%20against%20the%20digital%20library.
The Wayback Archive is an institution and it’s a real shame we can’t have nice things anymore because of the greedy actions of some.
But then these are the same people using AI. Make it equitable, and get on the same page. it’s like all of us consumers, in a way we are hypocrites , driving all the demand for computing power then complaining about data centers. Or people complaining about child labor but ordering from SHEIN from the comfort of their homes in the most wasteful country in the world. But I digress…
This is clearly a problem with a likely, relatively simple solution (assuming it's like charging tolls, which I think Reddit is doing?). The problem is probably in the business model for it. Is this what saves local journalism?
happened to me the other day. went to go see what a website looked like in the past and was met with a captcha that wouldn’t validate.
Love this topic, it highlights how fragile access to shared web memory can become when incentives clash 💙⚡ It’s interesting to think how tools built for preservation start getting affected by the same systems they try to archive!
One person’s externality is another’s disaster. This is one of the more frustrating aspects about AI companies’ co-opting the internet. Important and useful grassroots efforts have been abused so that very few can capture the gains from what should be open data sources. I’d love for a technical solution to let us all go back to the way things used to be, but I think this one is going to take some concerted regulation.