Researcher Warns of "Digital Dark Age"
alphadogg writes "A assistant professor from the University of Illinois at Urbana-Champaign is sounding a warning that companies, the government and researchers need to come up with a plan for preserving our increasingly digitized data in light of shifting document management and other software platforms (think WordPerfect and floppy disks). Jerome P. McDonough, who teaches at the Graduate School of Library and Information Science at the University of Illinois at Urbana-Champaign, says there exists about 369 exabytes worth of data, and that includes some pretty hard to replace stuff, including tax files, email and photos. Open standards could play a key role in any preservation effort, he says. 'If we can't keep today's information alive for future generations, we will lose a lot of our culture,' McDonough said. Even over the course of 10 years, you can have a rapid enough evolution in the ways people store digital information and the programs they use to access it that file formats can fall out of date.'"
And who needs to store pictures and movies on their computers anyway? In fact, I think the world would be a better place without them!
Now if you excuse me, I'm going back to watching Iron Man on my wrist watch.
We can just store everything in the cloud! Problem solved!
I am the richest astronaut ever to win the superbowl.
There have been instances when the metallurgy of times past was remarkably superior in some respects to later arts. Think of Damascus steel or Chinese bell-casting. Though the general trend of technology is constant progress forward, in certain cases the ancients were able to teach us a thing or two.
"I often ask, 'Everyone in the audience who thinks they're going to be using the same word processor in ten years, raise your hand.' No hands go up. 'Everyone who has data around that's going to have value in ten years?' After a minute's thought, every hand goes up. The lesson is clear: information outlives technology."
- Tim Bray
Parity: What to do when the weekend comes.
The only motivation for a company to invent new ways to preserve data long term is to provide it as a service so they can profit from it. Other than that, a companies main goals are deleting everything it legally can. Anything that no longer exists can't result in a lawsuit.
Everything that is preserved is a potential liability. For items requiring indefinite retention because they are critical to the business... They will be stored, redundant, and backed up appropriately. As the systems that provide those qualities age, they will be replaced in regular maintenance and upgrade schedules as economics and timing come together in the right proportions. In that way, reliability and long-term survivability are maintained - nothing stays on ancient systems that are unmaintainable forever. When systems go out of support, everybody has already been looking to the next solution to migrate to.
So what's wrong with this approach? Its essentially what all "big" companies are currently doing. I don't believe in this proprietary format FUD either - if the proprietary format is no longer supported, you migrate. Potential of future cost to migrate is the only concern, not survivability.
Migration is todays solution to long term storage and I see no reason it should be ignored. Like security, data retention is an ongoing objective that requires maintenance - its not some end-state. Dreaming of a solution that will just last forever seems archaic, no?
Overclockers
OPEN file formats and OPEN hardware, well documented.
Even if no program exists anymore to read your data, as long as you have the specs you can rebuild it. And I mean hard- AND software. If you know how to build it, you can build it provided you have the means. And I'm pretty confident that our future cousins will be able to build a current computer with their future technology, as long as they know WHAT they should build.
We used to have a Bill of Rights. Now, with the rights gone, all we have left is the bill.
I'm reminded of this story from a few years ago, where a 500 year old Leonardo drawing inspired improvements in mitral valve heart surgery.
This comment is for entertainment purposes only. Any similarity to real insight or information is purely coincidental.
If archeologists find knives and trash to be important in a search, I'd say the average pictures that we are taking today might actually be very intereting to future generations for they represent normal life.
Most of the garbage that we have now just isn't worth keeping. The biggest problem is filtering out the junk we have so that we know what is really valuable. That would be things like great music; writing; the origins of software freedom; works of history and biography etc. Then we could store that, but the problem is we mostly store SOX inspired lies for compliance audits. This garbage takes away from any effort to store serious stuff long term. Who could we trust to do the filtering? The govt? (no please don't answer that :-)
=~ s,(.*),<sarcasm>$1</sarcasm>,g if any_point_you_wish();
Historically, things that have been very uninteresting at the time, have been hugely valuable to researchers later on. We may not care about the countless people talking "crap" on bebo right now, but in a few hundred years it might be a different story. When people can easily analyse all those posts for meaningful psychological profiles that aren't currently understood never mind modelled and easily detected, all of that could tell a lot about our society. Even rubbish tips from thousands of years ago are hugely valuable to paleontologists.
This goes more so, for important government records, etc. Peter Quinn did a great job of explaining that, with his Sovereignty talk.
The article talks about two very distinct and different problems--hardware and file formats. The author has a point about the hardware--if the media goes bad or if there is no way to read the data, then the data is lost. However, the author is completely off-base when it comes to file formats...
The author specifically mentions WordPerfect files. Bad example! The default file format in Wordperfect X4 (released in April, 2008) is the same as what was used in WordPerfect 6--which came out in 1993 (DOS and Windows). While I can't speak for OpenOffice or Google Docs, MS-Word can read those files (and WordPerfect 5.x files) with a simple File/Open. Excel opens Lotus 1-2-3 files as well. So, Word can open popular formats in use since 1988 (WP 5.0) and Excel can open some formats in use since 1983 (1-2-3 r1a). You can also buy programs like FileMerlin to convert old documents.
Frankly, when it comes to file formats, conversion apps will exist for a LONG time. For DOS apps, you could even go so far as to create a v/m or use Dosbox, load up your obsolete word processor (I miss "Leading Edge Word Processor"!) and copy/paste the text into Word or Notepad...
Image files, sounds, & videos are no exception... GIF has been around since 1987, JPEG has been around since the early '90s (opening those on a 10Mhz 8088 was slow!), and MPEG/WMV/AVI/Quicktime videos are easily openable...
Finally, the more people that are affected by obsolete files, the more interest there is in some way to convert the data... But don't forget that a LOT of the data is junk--do you really care about your 7th grade paper you wrote on Hong Kong in 1989?
Windows 3.1x calc: 3.11 - 3.10 = 0.00
This is one of those fairly bogus, highly overblown stories that keeps cropping up every so often. A similar one is the supposed shortage of scientists and engineers in the US, which has never existed, and is always supposed to be coming Real Soon Now; in fact, the data to support this claim are always either nonexistent or wrong. (E.g., they compare Indian college graduates with US college graduates, but the Indian degree they're comparing with a U.S. bachelor's is more equivalent to an AA degree in the U.S.)
First off, the concern about incompatibility of physical media was valid 30 years ago, but it's a false analogy to try to apply it to today's situation. Thirty years ago, I had data on a mixture of 8-inch floppies and 9-track tapes. I can't read an 8-inch floppy anymore, and although 9-track tapes still exist, most 9-tracks from that era are no longer readable due to physical deterioration of the media. But that was all in an era when hard disks were expensive, and the internet didn't exist. Today, I have all my data on hard disks of various computers, and I use file synchronization software to keep them all in sync. If one of my hard disks dies, I replace it, and I haven't lost any of my data. (I also have backups on optical media, but I basically never need those.)
There's also the concern about formats. People tend to bring up, for example, the image of rooms full of physically deteriorating 9-track tapes with data from old NASA space probe missions. The formats are often not documented. The thing is, most of our data isn't at all analogous to the raw data from Mariner or Voyager or Viking. Those were unique historical events, and the only way to get more data like the data they collected is by sending another space probe. (People also tend to vastly overestimate the value of scientific raw data. It's extremely uncommon for raw data to be of interest decades later.)
Most of the world's data isn't in some obscure NASA format, it's stored in formats that are used by tons of people, and are extremely well documented. Sorry, but I just don't believe that the knowledge of how to decode Adobe Acrobat format is going to be lost to future generations. Ditto for html, jpeg, and mp3.
Another thing to keep in mind is that nowadays you can emulate old computers with excellent performance. For instance, my first home computer was a TRS-80. I can still run my old TRS-80 games on my linux box, using an emulator. Sure, emulation isn't perfect, and some information may be lost. But the claimed threat of data loss is vastly overblown.
The biggest threat to the preservation of information isn't technological change, it's copyright. The most likely reason that I wouldn't be able to get back an old piece of digital data is that the people who tried to preserve it and put it on the web got sued by the people who own the copyright -- the same people who let it go out of print. The economic incentives are to hold on to your copyrights (because that doesn't cost you any money) and send out DMCA notices to anyone who puts it on the net (because that doesn't cost you any money either), all in the hope that your content will be worth eleven cents fifty years from now. This is exactly what we see happening, for instance, with ROMs for old video games, which you can play in MAME, except that you have to find an illegal source for the data, because the owners of the copyrights aren't willing to sell you a copy.
Find free books.
Garbage isn't the problem.. the problem is that we have millions of copies of the same data. Think of the 50gb of video games you may have installed.. 10 million people have the same games as you. Music? Unless you performed it yourself or it's sub-underground, chances are millions of people each have multiple copies of it. The anime you've torrented has 10,000 downloads. .
No, see.. actually I'm just keeping a back up for the RIAA in case they lose their copy. PLus I keep it all transcoded to the next generation formats at no charge. And on top of that it's forward deployed for easy re-distribution without bottlenecking their servers. I even paythe lectric bill on the disks and internet connection. So copies are a good thing.
Some drink at the fountain of knowledge. Others just gargle.
I'm more concerned about losing the culture from the 20th century.
Everyone born after 1975 hates the RIAA, doesn't pay any attention to whatever they say, and file-shares gigabytes without a thought to the music industry definition of 'piracy'. This is as it should be. It means that the music and movies of the (for now) young people is safe because it is widely circulated outside the control of those who have deluded themselves into believing that they own it.
It's all the stuff from the first 2/3rds of the 20th century that will disappear. Because the people who like it are in their 50's, 60's, and 70's now and don't have the technical skills to copy and distribute it. Plus they actually trust the corporations will preserve it. I mean all the books, music recordings, television shows, movies, and plays from the first half of the 20th century. The stuff that is under 'infinite copyright' and will never be in public domain because the corporations will simply pay off the politicians to endless extend the copyright period, as they do now.
As soon as all this stuff stops selling (and who nowdays is paying money for the book that was #3 on the New York Times BestSeller list of Oct 28, 1936?), and can't be legally copied because it can't enter public domain, then the corporations will just destroy it. Pulp the books; convert the film stock to ethanol to power their SUVs; dump the magazines in the oceans or in nuclear waste sites to absorb neutrons. When that happens, all this culture will be gone and historians 200 years from now will have little idea about how civilized people actually thought and acted in the critical early years of the modern technological age.
You can talk to the old people about the need to preserve their culture by making 'illegal' copies of the books, magazines, and movies that were important to them, but they are just simply and completely clueless about the extent that their culture will die as they do.