Slashdot Mirror


Office 2003 and XML

zachlipton writes "Internet World is reporting that initial reports from Office 2003 beta testers don't look good for those hoping to share documents with non-MS systems using the XML file format. Gary Edwards, the OpenOffice.org representative for the OASIS XML file-format group is quoted as saying "although it's still early in the review process, it does look as though XP XML has been so seriously crippled as to be useless to anyone but the big content management and collaboration system providers." Apparently, all formatting and presentation information is removed from the XML. Furthermore, Office's new collaboration featres will only work with users who are also running Office 2003 (requiring Windows 2000 or 2003) that are connecting over XP servers." So Microsoft will continue its efforts to lock-in users with proprietary formats, and hopefully the rest of the world will produce an XML standard document format without them.

41 of 502 comments (clear)

  1. At some point..... by i_want_you_to_throw_ · · Score: 5, Insightful

    Microsoft will have to learn IBM's lesson about transforming from a company that makes standards, to one that contributes to them.
    They still don't get that their attempts to "embrace and extend" the whole damn internet isn't going to work.

    The rest of the world WILL produce an XML standard document format without them, thank heavens.

    1. Re:At some point..... by McDutchie · · Score: 4, Insightful
      he rest of the world WILL produce an XML standard document format without them, thank heavens.
      Which will be an irrelevant format because everyone will still need Word to read all the ubiquitous crippled Word XML format documents flying around on the net.
    2. Re:At some point..... by gmuslera · · Score: 4, Insightful

      Word (or even complete office), Win2k/XP as desktop and server. If someone sends me a document in Office 2003 format that he say I "MUST" read, I ask him to choose between sending me US$2003 to be able to read it, or sendme it in a really open format.

    3. Re:At some point..... by bfree · · Score: 4, Interesting

      Why? The attitude sounds harsh when expressed so simply, but if you tell you "client" that you can't read the file and that your company has decided not to purchase the software required to be able to do so as otherwise they would have to pass on the associated costs to their clients, so could they please send the file in a format you can read instead (even Word XP or earlier thanks to oo.o) or fax it, should the client really have a problem and if so is it worth keeping hte client (yes I really said that, lots of the time troublesome clients aren't worth keeping without changes if you actually can cost them completely)? Similarly with a coworker you can ask them if you can buy the software from their budget (in a company setting there should be company standards so this should be easy)!

      --

      Never underestimate the dark side of the Source

    4. Re:At some point..... by ccp · · Score: 5, Insightful

      Why not?

      If your clients tell you to bend over, you bend over? You seem to have a very sad life. Grow some spine, explain things to them, and you'll be surprised about how many of them get it.

      And, in case you wonder,

      I'm not a student.
      I own a business.
      And yes, I'm doing rather well even with principles.

      Cheers,

    5. Re:At some point..... by MeanMF · · Score: 4, Interesting

      You could also just download the free MS Word viewer that Microsoft provides here.

    6. Re:At some point..... by aaarrrgggh · · Score: 4, Insightful

      I agree with what you are saying, but there is a caveat: once a product has reached critical mass, you have to go along with everyone else.

      I remember problems with AutoCAD back 7 years ago or so, going from release 12 to release 13. 13 was a dog. It had an incompatible file format, forcing upgrades for everyone that shared the same document. Since 13 didn't offer enough incentive for them to reach critical mass, it died with most people sticking with 12 until the next release came out... which solved a lot of problems. Autodesk got a humility pill and realized that forcing the upgrades is bad policy, although you can do thing to encourage it (default format save).

      The trouble with MSFT's approach is that it breaks too many things at once; you have to get critical mass not only on the office application, but also the operating system and servers. A company that is not posed for this migration will not do it. If a single client requires it, then they will hire a secretary to do a saveas down to a more manageable format. If half the clients require it, it is difficult to avoid the upgrade.

    7. Re:At some point..... by frozenray · · Score: 4, Interesting

      > You could also just download the free MS Word viewer that Microsoft provides here [microsoft.com]

      For those not running Windows, the Word viewer comes "free" with a $199.- (list price) version of Windows, a good sized chunk of your system disk (not that it really matters much given today's HD prices and capacities) and the usual installation hassles, like drivers for equipment which isn't included on the CD etc. Even if you got Windows "free" with your PC from the manufacturer, you just paid the Microsoft tax up front, and will continue to pay if you want to keep your system up to date.

      That's like saying the Grappa I got offered after shelling out $150.- for dinner with a date last Saturday was "free". Sure, I didn't pay for it, but you can't get it without buying dinner first.

      Yes, I know there are solutions for reading MS Office documents on Linux. But I always cringe when people tell me to use the "free" readers - they're not free in any sense of the word in my book.

      --
      "There are already a million monkeys on a million typewriters, and Usenet is NOTHING like Shakespeare." - Blair Houghton
    8. Re:At some point..... by MrResistor · · Score: 4, Insightful

      You could also just download the free MS Word viewer that Microsoft provides here [microsoft.com].

      Strangely, there doesn't seem to be a Linux version. Or a Mac version, either. It's not so free when I'd have to buy a copy of Windows and spend 2 hours installing it, is it?

      --
      Under capitalism man exploits man. Under communism it's the other way around.
  2. Separating Content from Presentation a Good Thing by avdi · · Score: 4, Insightful
    Apparently, all formatting and presentation information is removed from the XML.
    And this is bad how? Isn't this the dream that XML document proponents have aspired to for years? You just can't please some people...
    --

    --
    CPAN rules. - Guido van Rossum
  3. Style Sheets by FattMattP · · Score: 5, Insightful
    Apparently, all formatting and presentation information is removed from the XML.
    Good. That's the point of XML. Formatting and presentation goes in style sheets.
    --
    Prevent email address forgery. Publish SPF records for y
    1. Re:Style Sheets by Captain+Large+Face · · Score: 4, Interesting

      The problem is that they don't include it elsewhere.. So in order to share documents in the style intended by the user, it must be saved as the proprietary format.

      IMHO, this ensures the user will opt-out of the XML format, and stay with the proprietary format. As I posted above, if Microsoft are going to do this, then they should bundle an XSL document with each XML document.

  4. Wow. by deviator · · Score: 4, Funny

    I am shocked. Shocked! I'm shocked that Microsoft would do something like this that wasn't in the best interest of their customers.

  5. Missing the point by graphicartist82 · · Score: 4, Insightful

    So Microsoft will continue its efforts to lock-in users with proprietary formats, and hopefully the rest of the world will produce an XML standard document format without them.

    I'm not trying to start a flame war here, but it seems that they're missing the point! We don't want it to be MS with one format and the rest of the world with another. That really wouldn't make it much different from how it is now. At least the way it is now, non-MS office software can read the MS formats. If it comes down to the choice between using the MS format or the "rest of the world" format, MS is going to win every time..

  6. Re:Separating Content from Presentation a Good Thi by molarmass192 · · Score: 4, Insightful

    I think the point is that if you save to their XML specification, you will loose all your document formatting. So yeah, the data is there, but it can't be reopened in Office or any other word processor and be in a structured way. Essentially, it is the same as just saving as plain text which has already been available since Office 95.

    --

    Good people do not need laws to tell them to act responsibly, while bad people will find a way around the laws-Plato
  7. Re:Separating Content from Presentation a Good Thi by DaveAtFraud · · Score: 4, Insightful

    I have to agree. The the basic concept behind SGML and its diminutive offspring, XML, was to separate content, structure and presentation. This just means that you have to share a style sheet, FOSSI, or whatever when you share a document if you expect the person you share it with to be able to view it.

    There may be other *valid* criticisms of what Microsoft is doing but this isn't one of them.

    --
    They that can give up essential liberty to obtain a little temporary safety deserve neither safety nor liberty.
    Ben
  8. Re:Separating Content from Presentation a Good Thi by JordoCrouse · · Score: 4, Insightful

    And this is bad how? Isn't this the dream that XML document proponents have aspired to for years? You just can't please some people...

    Unfortunately, Manny Manager and Sarah Secretary are now very used to depending on the formatting and presentation information. To be honest, not too many people these days subscribe to the whole minimalist document theory (unless your idea of starting your editor is typing 'vi').

    The main point here is to encourage the .XML format for interoperability. If the XML format can't figure out the fonts, colors, and various drawing elements in your document, then people will abandon it for something that does - at the expense of the rest of us.

    --
    Do you have Linux and a DotPal? Click here now!
  9. bollocks by graveyhead · · Score: 4, Insightful
    hopefully the rest of the world will produce an XML standard document format
    This is just so wrong. It smacks of a writer who doesn't really understand the utility of XML. There doesn't need to be "The One True Document Format"... that's not what XML is all about.

    Instead, create an XML format that is specific to your needs and write a DTD or XML-Schema that describes it. If you need to translate it to someone elses' XML document format, a quick XSLT stylesheet will transform the document with a minimum of effort.

    Just my 2 cents.
    --
    std::disclaimer<std::legalese> sig=new std::disclaimer; sig->dump(); delete sig;
  10. MS .doc / Adobe PostSript & PDF by PerlPunk · · Score: 5, Interesting

    All Microsoft needs to do is make their standard an open one (that can be used by others), like Adobe has done with their PostScript and PDF formats. Adobe has done quite well with their products based on these formats, too. Products like Adobe Illustrator and Photoshop (which works very well w/ bitmaps saved in PostScript) are the industry standard in digital art. If Microsoft followed a similar model, I'm sure that Microsoft Word will continue to be the industry standard in word processing software, and Microsoft as a business won't be any less richer for it.

  11. Re:Separating Content from Presentation a Good Thi by gorilla · · Score: 5, Insightful

    There is a big difference between seperating presentation from content and removing the presentation totally.

  12. Part of the concept by nhavar · · Score: 4, Insightful

    Isn't part of the concept of XML relating DATA and being able to seperate presentation from pure content. Isn't the additional concept of XML it's extensibility and adaptability for one group to use it differently than another? Because if not I've been using XML wrong for about 2 years now.

    This article makes it sound as if MS is doing something completely improper with XML (i.e. changing it's "standard"). But it seems to me that MS is simply separating content from presentation and relying on ????(something proprietary, xsl, more xml) to provide presentation. Just because they don't use the standard the same way you want them to doesn't mean that they are breaking the standard. I'm sure if you look at the XML that they output it's all standard XML. It also sounds as if they are not using any of the "tricks" that others have complained about (i.e. storing binary data in an xml tag).

    Instead of bitching about the problem maybe we should
    1) provide feedback if we are a beta tester
    2) wait for it to be released
    3) ready some tools to provide interoperability
    4) work harder on creating tools better than MS

    --
    "Do not be swept up in the momentum of mediocrity." - anon
  13. sometimes.. by siphoncolder · · Score: 5, Interesting
    I wonder if michael is testing us for stupidity, literacy, and actual technical knowledge of the issues.

    1) Take MS, make a report that says they did something bad, watch how many people flock to bash them DESPITE THE FACTS PRESENTED IN THE ARTICLE, which leads me to:

    2) How many people read the article? And of those people who DID, :

    3) How many of them know that XML is supposed to be a divorce of data from presentation? Why this comes as a shock to people is obvious - they didn't know that.

    The poster above who said "style sheets" - bravo. You couldn't have made a better point with two words.

    --
    i'm amazed that i survived - an airbag saved my life.
  14. I have Office 2003 and this article is BS by Anonymous Coward · · Score: 5, Informative

    I have Office 2003 Beta 2 freshly downloaded from MSDN. This article is completely wrong. I did the following:

    1. Opened a heavily formated .DOC Word document with tables, multiple fonts, etc.
    2. Saved the document as XML.
    3. Opened up the XML document in Word and it looks EXACTLY like the original .DOC format.

    I also opened the XML file in a text editor and sure enough it contains complete formatting information.

    1. Re:I have Office 2003 and this article is BS by RanmaSan · · Score: 4, Informative

      It's not pretty, but it works:

      <?xml version="1.0" encoding="UTF-8" standalone="yes"?>
      <?mso-application progid="Word.Document"?>
      <w:wordDocument xmlns:w="http://schemas.microsoft.com/office/word/ 2003/2/wordml" xmlns:v="urn:schemas-microsoft-com:vml" xmlns:w10="urn:schemas-microsoft-com:office:word" xmlns:SL="http://schemas.microsoft.com/schemaLibra ry/2003/2/core" xmlns:aml="http://schemas.microsoft.com/aml/2001/c ore" xmlns:wx="http://schemas.microsoft.com/office/word /2003/2/auxHint" xmlns:o="urn:schemas-microsoft-com:office:office" xmlns:dt="uuid:C2F41010-65B3-11d1-A29F-00AA00C1488 2" xml:space="preserve"><o:DocumentProperties><o:Titl e>Some Centered Bolded Text</o:Title><o:Author>Mark McWilliams</o:Author><o:LastAuthor>Mar k McWilliams</o:LastAuthor><o:Revision>1</o:Revision ><o:TotalTime>2</o:TotalTime><o:Created>2003-03-13 T17:30:00Z</o:Created><o:LastSaved>2003-03-13T17:3 2:00Z</o:LastSaved><o:Pages>1</o:Pages><o:Words>10 </o:Words><o:Characters>57</o:Characters><o:Compan y>i-FRONTIER</o:Company><o:Lines>1</o:Lines><o:Par agraphs>1</o:Paragraphs><o:CharactersWithSpaces>66 </o:CharactersWithSpaces><o:Version>11.4920</o:Ver sion></o:DocumentProperties><w:fonts><w:defaultFon ts w:ascii="Times New Roman" w:fareast="Times New Roman" w:h-ansi="Times New Roman" w:cs="Times New Roman"/><w:font w:name="Tahoma"><w:panose-1 w:val="020B0604030504040204"/><w:charset w:val="00"/><w:family w:val="Swiss"/><w:pitch w:val="variable"/><w:sig w:usb-0="61007A87" w:usb-1="80000000" w:usb-2="00000008" w:usb-3="00000000" w:csb-0="000101FF" w:csb-1="00000000"/></w:font></w:fonts><w:styles>< w:versionOfBuiltInStylenames w:val="3"/><w:latentStyles w:defLockedState="off" w:latentStyleCount="156"/><w:style w:type="paragraph" w:default="on" w:styleId="Normal"><w:name w:val="Normal"/><w:rsid w:val="7765DB"/><w:rPr><w:rFonts w:ascii="Arial" w:h-ansi="Arial"/><wx:font wx:val="Arial"/><w:sz-cs w:val="24"/><w:lang w:val="EN-US" w:fareast="EN-US" w:bidi="AR-SA"/></w:rPr></w:style><w:styl e w:type="character" w:default="on" w:styleId="DefaultParagraphFont"><w:name w:val="Default Paragraph Font"/><w:semiHidden/></w:style><w:sty le w:type="table" w:default="on" w:styleId="TableNormal"><w:name w:val="Normal Table"/><wx:uiName wx:val="Table Normal"/><w:semiHidden/><w:rPr><wx:fon t wx:val="Times New Roman"/></w:rPr><w:tblPr><w:tblI nd w:w="0" w:type="dxa"/><w:tblCellMar><w:top w:w="0" w:type="dxa"/><w:left w:w="108" w:type="dxa"/><w:bottom w:w="0" w:type="dxa"/><w:right w:w="108" w:type="dxa"/></w:tblCellMar></w:tblPr></w:style>< w:style w:type="list" w:default="on" w:styleId="NoList"><w:name w:val="No List"/><w:semiHidden/></w:style><w:sty le w:type="paragraph" w:styleId="StyleBoldCentered"><w:name w:val="Style Bold Centered"/><w:basedOn w:val="Normal"/><w:rsid w:val="7765DB"/><w:pPr><w:pStyle w:val="StyleBoldCentered"/><w:jc w:val="center"/></w:pPr><w:rPr><wx:fon t wx:val="Arial"/><w:b/><w:b-cs/><w:sz-c s w:val="20"/></w:rPr></w:style><w:style w:type="paragraph" w:styleId="SmallTitle"><w:name w:val="Small Title"/><w:basedOn w:val="StyleBoldCentered"/><w:rsid w:val="7765DB"/><w:pPr><w:pStyle w:

  15. The authors of the article didn't bother to RTFM.. by malakai · · Score: 5, Informative

    The point of the Office 2003 "Save as XML" with the "Data Only" checkbox is _NOT_ a poor mans Save As XHTML. It's decide to allow the data of the document and pet placed into an XML document based on a schema. You literally can make your own schema file/XSD, and use a tool inside Word to map the elements of a Word document to elements of the schema. If you simply map a paragraph to a string you will lose formating. Unless of course you define in your schema how you'd like to store formating information. But that is generally an overkill.

    Think of a resume. you could define an XSD for a resume, and be able to save resumes against this XSD, as validated pure XML.

    Now, if you want to produce a document, using an XML syntax but want to combine both data and presentation, then you want WordML.

    WordML uses Word's own tags to markup the word document. I was going to show you an example of WordML but i don't feel like escaping allt he greater-than/less-than signs. Anyhow, WordML contains all the formating and everything necessary to display a Word document as it is supposed to look.

    I think this Open Office guy is looking for a devil in Office 11 that isn't there. That or he didn't read the friggin manual.

    -Malakai

  16. Re:Duh. by t0ny · · Score: 5, Insightful
    How do you figure this is anti-trust? This is simply a company who has the dominant product protecting their lead. And quite honestly, I dont see anything wrong with that, as long as they confine their practices to their product (ie. they arent making Office the only suite that can run on windows)

    Have you ever played a game like Civilization or Alpha Centari? You would be amazed at how much those games make you understand politics. Once you are in the lead, you do anything you can to protect that lead. And why would you expect the real world to be any different?

    But this isnt a game, this is business. And since businesses are SUPPOSED to make money, they need to make sure people continue to buy MS Office. And making an office suite that shares documents with all the various third-tier office suites just doesnt do that. Why should my company buy MS Office if the documents it produces are exactly the same as those of FreeBeerOffice? Now, if FBO cannot do things MSO can do, then there is an incentive...

    --

    Manipulate the moderator system! Mod someone as "overrated" today.

  17. Wait a minute... by sheldon · · Score: 4, Insightful

    "has been so seriously crippled as to be useless to anyone but the big content management and collaboration system providers."

    That indicates to me that the problem is really that the document format is so complicated that it takes tremendous resources to understand and implement compatibility with it, as this implies that larger companies like say a Xerox will have no problem producing tools to work with it.

    So from a business consumer perspective this is still a tremendous win.

    This sounds like more whining from the open source crowd.

  18. Re:Separating Content from Presentation a Good Thi by djoham · · Score: 5, Informative

    This may be bad (keeping in mind the jury is still out on exactly how Microsoft is making this work) because in the case of office documents, the style is actually *part* of the content, from the perspective of Joe Office User.

    If Microsoft just puts the raw text data into a .xml file, then that .xml file is practically useless to anyone who wants to collaborate with the original author since all of the styling information is lost.

    As an example of a good way to do this (IMHO), take a look at how OpenOffice.org builds their files. When you make a .sxw (the default writer format) you're actually taking the raw data of the document, the styling rules for the document and a few other important bits and pieces and zipping them up into a single file.

    After unzipping this file, the following directory structure was exposed:

    content.xml
    META-INF/manifest.xml
    meta.xml
    mi metype
    settings.xml
    styles.xml

    With this type of design, you can get the best of both worlds. Technically, there is a separation between your presentation and content which allows simple programatic access to the data when necessary. At the same time, this design allows for full collaboration between people who also consider the styling of the data to be part of the content because the style rules for the content are included with the document.

    With xml-saved Office documents containing only data and no style, collaboration between non-office users (and apparently Win9x users as well) will be no better off than before. Perhaps worse, assuming the binary .doc, .xls etc formats have changed and will need to be reverse-engineered again.

    If this article is true and Microsoft has decided to remove the styling of their xml-saved office documents, I see two possible reasons for this:

    The first is obvious. You're not using Office? Ok, second class citizen, here's the data but in a format that is next to useless for you to use.

    The second possibility involves Microsoft just not being where they want to be with the Office XML sharing. Keep in mind that it took OpenOffice.org something like a year and half or so to define their XML interchange format. Microsoft may be going there, but due to overwhelming inertia, it just might not be going there very quickly.

    Personally, I think the first option is the most likely. However, with OpenOffice.org working with OASIS and others on a common XML interchange format, I'm hoping Microsoft will be forced by the marketplace into option 2.

    Best regards,

    David

  19. Proprietary Document Formats by Daimaou · · Score: 4, Insightful

    Proprietary document formats were fine at one point. Most people shared documents via printed paper, or shared them via "soft copy" within their own organizations. However, the time for printed documents and interoffice "soft copies" is over. We need the ability to share documents with the world in an easy to use, feature rich, and easy to edit format. Since a significant part of a document's legibility is in its style and formatting (or at least people are more apt to read a well formatted document over one which is not) text files are out.

    Once an easy to use, open document format is created, and the ability to read and write those documents is built into many programs, I think we will see an end of .DOC file attachements.

    While there are currently some "open" formats like PDF and PS, the problem is that they are not easy to create for the average user, nor are they easy to edit. While PDF may be a good format, we need something better.

    XML is a logical choice as a base for an open format because it is a well defined standard, it is text based, and is quite easy to parse.

    But I ramble.

  20. The article is blantantly wrong... by malakai · · Score: 5, Informative

    Read some other articles, or better yet get ahold of a beta and try it out. The authors of this articles will feel like schmucks when they realize what they missed.

    First off, by default, if you save the word document as XML, it gets saved as WordML,which preserves Word's styles and formatting in an XML name-space that's separate from the one bound to the schema-controlled data.

    If you check off the checkbox "Data Only" then you will lose all formating and your own XSD will be used to map this document into XML data.

    WordML looks like a XML'ified RTF language. It would be trival to create an XSL stylesheet that transforms WordML into HTML/CSS with all formating (that HTML is capable of) which directly mimics MS Word. OpenOffice could also eat WordML quite easily and have all the formating/style of Word.

    What the authors of this article are REALLY bithing at, is the fact that MS didn't buy into the OpenOffice Document Specification from OASIS. MS prolly sees OASIS as the US sees the UN. Defunct, not needed.

    If you describe your data using XML semantics, and all it takes to convert from semantic style A to B is some XSL, then who cares about forcing everyone to use one specific format.

    -malakai

  21. WordML by malakai · · Score: 4, Informative

    If you "Save as XML" in Office 11, then by default the data is saved as WordML. WordML is an xml version of MS internal storage format (basically RTF). OpenOffice could quite easily write an interpreter for WordML. Hell, I could write an WYSIWYG editor for WordML in a day. If that. It's pretty simple if you understand the basics of RTF.

    It's only when you Save as XML with the "Data Only" checkbox that you get into striping formating (and rightly so). Word WARNS you about this. In addition, you can specify your own XSD to save to. And word will VALIDATE this for. Not to mention, you can use a word tool to map elements of Word documents to elements of your schema. DAMN COOL.

    In addition (As if that isn't enough) when you save, in either way, you have the option of specifiying a XSL style sheet. It'll go ahead and transform the output for you as part of the save.

    Then only thing the OpenOffice people are upset about is that MS didn't buy into the OASIS/OpenOffice Document Specification. Tough shit. I'll write them an XSL that'll work again WordML to solve that for them. Lazy bastards.

    -malakai

  22. REPEAT AFTER ME: XML IS NOT A FILE FORMAT by Trailer+Trash · · Score: 5, Interesting

    Internet World is reporting that initial reports from Office 2003 beta testers don't look good for those hoping to share documents with non-MS systems using the XML file format...

    That's because XML is not a file format, it is instead a format for file formats. To quote the O'Reilly "Learning XML" book, page 2:

    Note that despite its name, XML is not itself a markup language: it's a set of rules for building markup languages.

    I've said this many times on /. (look at my history), but the fact that a particular format is XML-based says nothing of your ability to read it. I'm even going beyond the fact that Microsoft could simply stick their traditional file formats into a CDATA and claim XML compliancy.

    The statement "If Microsoft used a standard XML format for their documents then anyone could read them" makes as much sense as an equally stupid statement like "If Microsoft just used 8-bit bytes in their file formats then anyone could read them".

    Sorry to rant, but the level of cluelessness around XML is astounding. Please read up, there's a ton of useful information on XML around the internet.

    MDC

  23. Save As XML = WordML by malakai · · Score: 5, Informative
    Taken from a real review of the XML/Office features:

    Once valid, the document can be saved as XML in two ways. The default is to create WordML, which preserves Word's styles and formatting in an XML name-space that's separate from the one bound to the schema-controlled data. You can optionally save through an XSLT transformation which, in a publish-to-the-Web scenario, could translate WordML formatting into HTML/CSS formatting. Alternatively, if you tick the Save as Data option, you can instead save just the raw XML data. In that case, you can bind one or more XSLT stylesheets to the document, each of which can generate WordML styles and formatting.


    InternetNews is authored by morons.

    -malakai
    1. Re:Save As XML = WordML by Hangtime · · Score: 5, Informative

      Same thing with Excel, you can save as Excel with formatting or not. This comes from the Excel XML with formatting. Quite simply the article is flamebait.

      <Style ss:ID="s26" ss:Parent="s16">
      <Borders>
      <Border ss:Position="Bottom" ss:LineStyle="Continuous" ss:Weight="1"/>
      <Border ss:Position="Top" ss:LineStyle="Continuous" ss:Weight="1"/>
      </Borders>
      <Font ss:FontName="Times New Roman" x:Family="Roman" ss:Size="12" ss:Bold="1"/>
      <NumberFormat ss:Format="_(* #,##0_);_(* \(#,##0\);_(* &quot;-&quot;??_);_(@_)"/>
      </Style>
      <Style ss:ID="s27">
      <Alignment ss:Vertical="Bottom"/>
      <Borders/>
      <Font ss:FontName="Geneva"/>
      <Interior/>
      <NumberFormat/>
      <Protection/>
      </Style>
      <Style ss:ID="s28">
      <Font ss:FontName="Geneva" ss:Size="12"/>
      <NumberFormat ss:Format="0.0"/>
      </Style>

      <Stuff in between here to get around Lameness filter>

      <Style ss:ID="s27">
      <Alignment ss:Vertical="Bottom"/>
      <Borders/>
      <Font ss:FontName="Geneva"/>
      <Interior/>
      <NumberFormat/>
      <Protection/>
      </Style>
      <Style ss:ID="s28">
      <Font ss:FontName="Geneva" ss:Size="12"/>
      <NumberFormat ss:Format="0.0"/>
      </Style>

  24. Real World Vs. Game by blahlemon · · Score: 5, Funny
    Truth be told the real disadvantage to this being the real world vs. a game is I can't set the level of difficulty to my liking...nor can I stop and speed up time.

    Or spy on other people from a God perspective. Damn you! Now I'll have to spend the rest of my day realizing how pathically small my scope is...

    --
    It take more faith to believe in evolution than it takes to believe in God
  25. Re:Separating Content from Presentation a Good Thi by Azghoul · · Score: 5, Insightful

    Your use of the tired "Bzzzzt" exclamation at the beginning of your post completely overwhelmed any potential interest in whatever it was that you were trying to say.

    Please, next time try to avoid the condescending tone, people might respond more constructively.

  26. Goldfarb's conjecture by RobotWisdom · · Score: 4, Informative
    I think the point is that if you save to their XML specification, you will lose all your document formatting.

    I think the root of the confusion goes back to Golfarb's original theory for SGML-- that the styles in a document are secondary to the structures, and should be kept separate.

    This has been a religious conviction ever since, despite the fact that most authors are messy and intuitive, and SGML-etc are very, very rigid and unintuitive. The rationalisation is that messy authors can just represent their styles using 'fake' (ad hoc) XML, but if this turns out to be 90% of the real users of MS Office, then I think MS could indeed save valid XML, but it won't be portable in any useful sense.

  27. On XML file formats.. by PeekabooCaribou · · Score: 4, Informative
    I realize this is redundant by now, but I think this is important enough to warrant a few duplicated posts. For Microsoft's XML format to be useful (and even worth implementing), it's going to require some advantages above and beyond what plain text formatting offers. The only completely useless XML format would be:
    <document>
    This is my document.
    Second paragraph.
    </document>
    I make the assumption that at least some tags are applied, such as some sort of paragraph tags and the like. I may be going out on a limb here, but I would even assume that their final XML format will produce documents identical to .doc files. I would also assume that I could pass this file off to Joe in marketing, and he would see a document identical to the one I saw. What I'm getting at is that style has to be held somewhere. If the XML file has no style associated with it, then congratulations, Microsoft, you did it right. But if Word can display the right formatting, then so can anyone else. (Assuming Word doesn't store the styles in a proprietary format, which I don't think is beyond them.) But why am I even writing this? From the article:
    However, Mark McWilliams, a software engineer and Office 2003 beta tester, said he has seen nothing to indicate that Office 2003 removes formatting information from files saved in .xml. He noted that he opened a heavily formatted .doc Microsoft Word file, saved the file as XML, and later opened the file in Word 2003, "The opened XML document looks exactly like the original .doc file," he said. "And if I open up the XML file in a text editor, I can see that all of the formatting is properly maintained in the XML file."
    Time will tell.
    --
    "I'll say it again for the logic-impaired." -- Larry Wall.
  28. Re:Duh. by kin_korn_karn · · Score: 4, Funny
    Separating data from format is one of the strengths of xml.

    Also, of the comma-delimited file.
  29. Duh. by Tony-A · · Score: 4, Insightful

    How do you figure this is anti-trust? Microsoft has been judged a monopolist. Since past behavior is a good indicator of future behavior, there is a presumption that this is anti-competitive behavior until proven otherwise.

    This is simply a company who has the dominant product protecting their lead.
    For a monopolist, nothing is simply any more. In the absense of market forces to correct misbehavior, exactly how they attempt to protect their lead does matter.

    And quite honestly, I dont see anything wrong with that, as long as they confine their practices to their product (ie. they arent making Office the only suite that can run on windows) [emphasis added]
    As long as nothing in the Office Suite promotes the Desktop OS monopoly.
    As long as nothing in the Desktop OS monopoly promotes their own Office Suite.

    But this isnt a game, this is business.
    And screwing your customers is bad business.
    And screwing your suppliers is bad business.
    And screwing your investors is bad business.
    And screwing your employees is bad business.
    Even screwing your competitors is bad business.

    And since businesses are SUPPOSED to make money, they need to make sure people continue to buy MS Office.
    And General Motors needs to make sure people continue to buy Chevrolets.

    And making an office suite that shares documents with all the various third-tier office suites just doesnt do that.
    It just makes incomprehensible gibberish unless the recipient happens to have the exact same sooper-dooper magic decoder ring. Unless I can read my stuff, under circumstances of my own choosing, I have a problem. Unless I can send stuff to my correspondents and they can read it un circumstances of their own choosing, I have a problem. If my documents are hostage to the whims of a supplier, I have a problem.

    Why should my company buy MS Office if the documents it produces are exactly the same as those of FreeBeerOffice?
    New twist on Clippy?
    No reason they should. That's Microsoft's problem, not yours or your company's (unless you work for Microsoft;)

  30. Re:Microsoft's new file format is: by nenolod · · Score: 5, Funny

    Oops, i forgot to set the reply to "Code". Please note, your SAX parser probably wont be able to parse this, heh. It is however, theoretically proper XML.

    <?xml version="1.0" standalone="yes" encoding="en">
    <!DOCTYPE worddoc [
    <!ELEMENT document (document_properties, document_section)>
    <!ELEMENT document_properties (title, author, organization, department, job, generalsummary)>
    <!ELEMENT title (#PCDATA)>
    <!ELEMENT author (#PCDATA)>
    <!ELEMENT organization (#PCDATA)>
    <!ELEMENT department (#PCDATA)>
    <!ELEMENT job (#PCDATA)>
    <!ELEMENT generalsummary (#PCDATA)>
    <!ELEMENT document_section (sectionsummary, proprietarybinary, unenhancedcrappytext)>
    <!ELEMENT sectionsummary (#PCDATA)>
    <!ELEMENT proprietarybinary (#PCDATA)>
    <!ELEMENT unenhancedcrappytext (#PCDATA)>
    ]>
    <document>
    <document_properties>
    <title>Crappydoc</title>
    <author>William H. Gates III</title>
    <organization>BORG</organization>
    <department>Unimatrix 0</department>
    <job>Secondary information processing adjunct</job>
    <generalsummary>Doc about crappy M$ things.</generalsummary>
    </document_properties>
    <document_section>
    <sectionsummary>Haha, you cant parse this and make it look perty, it's BINARY! You're still screwed!</sectionsummary>
    <proprietarybinary>firoiorfioeiojvonvonviniooiwnco ncooisoi39f940f9439 0f904390f94390fj904j90j3f09j4fj3490jf30jf040fj03j0 9fj9340fj043j90fj4903fj9043jfj0vjoirejvoojvoerjgoe jgojerogjoejoenmvotnhnoignoengotnhinringuinfi</pro prietarybinary>
    <unenhancedcrappytext>Hehe, doesnt this text just look ugly? I bet it does, if you arent using M$ WORD!</unenhancedcrappytext>
    </document_section>
    </document>