Slashdot Mirror


Should Archive.org Ignore Robots.txt Directives And Cache Everything? (archive.org)

Archive.org argues robots.txt files are geared toward search engines, and now plans instead to represent the web "as it really was, and is, from a user's perspective." We have also seen an upsurge of the use of robots.txt files to remove entire domains from search engines when they transition from a live web site into a parked domain, which has historically also removed the entire domain from view in the Wayback Machine... We receive inquiries and complaints on these "disappeared" sites almost daily."
In response, Slashdot reader Lauren Weinstein writes: We can stipulate at the outset that the venerable Internet Archive and its associated systems like Wayback Machine have done a lot of good for many years -- for example by providing chronological archives of websites who have chosen to participate in their efforts. But now, it appears that the Internet Archive has joined the dark side of the Internet, by announcing that they will no longer honor the access control requests of any websites.
He's wondering what will happen when "a flood of other players decide that they must emulate the Internet Archive's dismal reasoning to remain competitive," adding that if sys-admins start blocking spiders with web server configuration directives, other unrelated sites could become "collateral damage."

But BoingBoing is calling it "an excellent decision... a splendid reminder that nothing published on the web is ever meaningfully private, and will always go on your permanent record." So what do Slashdot's readers think? Should Archive.org ignore robots.txt directives and cache everything?

4 of 174 comments (clear)

  1. No brainer by fnj · · Score: 5, Insightful

    Duh. Naturally it should. The notion that robots.txt should operate RETROACTIVELY is asinine.

    1. Re:No brainer by thsths · · Score: 5, Insightful

      But that is not the question asked, is it?

      robots.txt should apply to the page at the time. I do not see any decent argument against that.

      But arguable robots.txt should not be a way to retroactively mark previously archived content as inaccessible.

    2. Re:No brainer by blackest_k · · Score: 5, Insightful

      One problem i run into is with owner manuals for old film camera's a lot of the time they disappear from the company website when they get taken over by another company. Sometimes archive.org can come to the rescue if I can find where they used to be. Fair enough the new company may only be interested in the digital models and has no interest in the historical product made by the company they acquired but when they make boneheaded choices like erasing the historical information the original company put out for their customers..

      Worst still is when a domain name is lapsed and bought by another company who had zero access to the content of the former site they bought a name not a right to control the history of the former site.

      The other thing which bugs me is the white washing of old news articles how often that trick gets pulled, I might personally remember an event but find the contemporary records are missing that happens a lot especially in Politics when a past stance becomes embarrassing and then you get told black was white...

      At the very least when a website changes hands the new owner should not be able to erase the history of the site under the previous owner.

  2. YES!! by Vadim+Makarov · · Score: 5, Insightful

    I applaud the direction internet archive takes. They should fully implement it.

    A year ago one of my domain names was stolen, through negligence of the registrar. The site was a non-profit resource that I maintained for the past 15 years. The squatter who now owns the name put deny all in robots.txt. As the result the website with some quantity of useful information has totally disappeared from existence and from the archive record.

    I do not see sufficiently important reasons to remove information that was once in public access. There are some reasons, however the public benefits of having access to all past public information outweigh all them.

    --
    17779 eligible voters in a district, 17779 'vote' as one. This is Russia.