NCSA Compares Google and Yahoo Index Numbers

Accurate results? by bigwavejas · 2005-08-15 06:12 · Score: 5, Interesting

Google sometimes returns some pretty interesting/ entertaining results.

Try searching for the word, "failure" in Google and check the results.

This brings into question *accurate* results. In this case it appears that's left to interpretation.

--
"Simplify, simplify, simplify!" Thoreau

Re:Accurate results? by jrallison · 2005-08-15 06:28 · Score: 5, Insightful

It is odd however the #1 result for failure is a webpage without the word "failure" in it.
Re:Accurate results? by MindStalker · 2005-08-15 06:30 · Score: 4, Insightful

Well google also indexes based upon refering links and not just the context in the page itself. So if many websites refer to GW as a failure, GWs page itself will turn up as a high hit. Yahoo does this as well, but doesn't not nessesarly give it the same weight. This could highly affect amounts of returns. Because if we say that google returned X pages for a search on term "y" many of these pages may not actually mention "y" thus giving a larger page count for "y". While with yahoos method, it will mainly return pages that mention "y" themself. And possibly add some pages that are mentioned to include "y" by links. This can vastly alter the count.

They might have a larger index file by BlackCobra43 · 2005-08-15 06:14 · Score: 4, Insightful

but they can't sift through it nearly as well as Google, so what does it matter? Even if you have a bigger dictionnary, if you can't speak English at all it won't do you much good.

--
I never spellcheck and I freely admit it. Save your karma for more worthwhile "lol erorrs" replies

Flawed conclusion? by Prong_Thunder · 2005-08-15 06:16 · Score: 5, Insightful

Sorry, but if Google consistently returns more results, it could just as easily mean that the filtering isn't as good.

I still prefer Google though.

Re:Flawed conclusion? by Ossifer · 2005-08-15 06:24 · Score: 5, Insightful

Exactly! I find the conclusions of the research to be quite specious. Yahoo may simply have tighter controls of what is considered a match, which, by the way, is no simple algorithm.

In any case, I am usually not so interested in the numbers of matches, but in the quality of the list returned--hopefully one website will have exactly what I need...

The results by Swamii · 2005-08-15 06:17 · Score: 4, Interesting

For those that don't want to read the flippin' article:

Based on this random sample, we found that on average Yahoo! only returns 37.4% of the results that Google does and, in many cases, returns significantly less.

In other words, they believe Google indexes more items based on their own tests of searching.

--
Tech, life, family, faith: Give me a visit

Re:Yahoo returns dupes... by Anonymous Coward · 2005-08-15 06:23 · Score: 5, Funny

Yahoo returns a lot of dupes.

If that's the case, then why is Google the darling of slashdot? ;)

Perl Code by hayro · 2005-08-15 06:24 · Score: 4, Funny

I don't know about the study but that is the most readable perl code I have seen in a long time.

More please! by 2008 · 2005-08-15 06:25 · Score: 5, Interesting

This is a great article! I wish there were more like it on slashdot. It's scientific instead of an opinion piece, it has references, it's repeatable. It's also short and very readable, unlike a lot of science papers.

OK, it is yet another Google piece, but it's not "some junior analyst predicts Google will buy Apple and release OSX86box 720".

--
I quit!

Methodology by enjo13 · 2005-08-15 06:26 · Score: 5, Insightful

The very methodology used in this case seems rather incorrect to me.

The assumption (as stated in the paper): Since Yahoo claims to have indexed twice as much as google, searches should return twice as many entries.

That assumption is flat out incorrect. There are actually multiple problems.

First, the scope of the search (based on index terms) is really up to the search engine itself. Since each search engine does not return the entire database as search results, it is very much up to the individual search algorithm to determine the depth of entries considered to 'match' a set of terms. That's what is really being reflected in these results.. it is not the overall size of the index, but simply how aggressive the search algorithm is in matching terms to entries.

Even if the algorithms where identical (same algorithm being run across both indexes), the nature of search does not scale in that way. If Yahoo has, for instance, becomre more aggressive in indexing message board and forum content, then only searches that play to those subjects should return more results than Google. Since searches are by definition narrowing on a data set, a methodology needs to be developed that more effectively tests the BREADTH of the results more than simply testing the depth.

--
Turn s60 photos into awesome videos with mScrapbook for all S60 3rd edition phones!

Re:Conclusion by nutshell42 · 2005-08-15 06:26 · Score: 4, Insightful

And Nutshell42's New Amazing Search Engine gives you even more results. Even though my index size is only 1.something million. I simply return every single wikipedia article in every language as result no matter what you search.

Concluding that Yahoo's index has to be smaller because they return fewer results seems a bit overzealous. Only a thorough study comparing results and how useful they were (which is hard to do, expensive and time consuming) has any meaning that goes beyond producing lots of funny numbers and percentages.

96.34% of all percentages are completely useless.

btw. I use google, not yahoo

--
Don't think of it as a flame---it's more like an argument that does 3d6 fire damage

International Listings by Dominatus · 2005-08-15 06:27 · Score: 4, Insightful

The study only checked English words. Is it possible that the increase came from Yahoo expanding into more international website markets?

Just a thought

This is what passes for CS research nowadays? by adrizk · 2005-08-15 06:28 · Score: 5, Insightful

Seriously. 'We wrote a script and here are the results'? This would take an average PERL programmer what -- 30 minutes of work? Has academic research in computing really sunk to this level?

Maybe it's not even worth pointing out how badly flawed (and lazy) the underlying assumption of 'twice the results = twice the index size' probably is, as I'm sure we're going to see a few dozen posts to that effect (unless PageRank really means nothing), but at least I can complain about the slant they put on this, and how strong a conclusion they seem to derive.

Re:What would you want them to return? by Intron · 2005-08-15 06:29 · Score: 4, Insightful

The top of the page return for Yahoo is

"Failure on eBay Find failure items at low prices. "

which illustrates the most important difference between Yahoo and Google.

--
Intron: the portion of DNA which expresses nothing useful.

Results of my own study... by Locke2005 · 2005-08-15 06:38 · Score: 4, Funny

Google only reports "about 4,820,000" entries for Britney Spears, while Yahoo reports "about 67,100,000" entries! This makes Yahoo more than 12 times better than google! Yeah, my methodology is completely fucked up... but then, so is the NCSA's!

--
I've abandoned my search for truth; now I'm just looking for some useful delusions.

Proper name samples by jkauzlar · 2005-08-15 06:39 · Score: 5, Interesting

Let's try a few samples of proper names:

Search: Valerie Plame
Google: 908,000
Yahoo: 2,580,000

Search: "Boulder, Colorado"
Google: 1,600,000
Yahoo: 5,880,000

Search: "Linus Torvalds"
Google: 2,560,000
Yahoo: 5,870,000

I assume it goes on like this. Of course these exceed the 1000 maximum hit limit given in the study.

Re:Yahoo pants down, egg on face, no WMD either. by loose_cannon_gamer · 2005-08-15 06:54 · Score: 5, Insightful

After reading half the comments on this page, I'm amused at how many alert readers are making the same mistake that they accuse Yahoo of -- misstating results.

Can we conclude from this study that Google has a bigger index than Yahoo? No. Can we conclude that when you pick two English words that when entered into both Google and Yahoo, both return less than 1000 results, that Google has consistently more results? Yes.

The real question is, what can we infer from the actual indisputable findings of this study? I find no ready method of generalization. If you are inclined to believe google is better, you feel happy inside. If you think yahoo is better, you have many options to dispute the idea that the study result generalizes to search engine index size.

As a google fan, I enjoy the warm fuzzies, but I don't see that much to get excited about either way.

--
In Soviet Russia, us are belong to all your base.

Re:Conclusion by barawn · 2005-08-15 06:55 · Score: 5, Insightful

No, it's accurate. They're testing Yahoo's claim of how many pages they've indexed, which just means that all indexed pages that contain the requested words should be returned from the search request. If yahoo returns fewer unique pages, yahoo has indexed fewer pages.

Actually, it might not be, thanks to their methodology.

They only used searches with less than 1000 results. They therefore got a lot of searches with small results numbers (because they were searching for bizarre word combinations, like "promotion bedabble"). The total number of results was something like 500,000 or so (order of magnitude) for 10,000 searches. That's an average of 50 results/search, and I'd bet there's a large, large tail, so the most common search is probably something like 10 results.

The problem with this is that in their word list, the same sites are being returned over and over!. For instance, sites containing dictionary lists appear in both "promotion bedabble" and "foliolate defecations" because, duh, that's the only place they'll appear. Since they're just searching the same type of site over and over, they get the same result magnified a lot: Google has more "dictionary lists" in its index than Yahoo. Most of the "dictionary list" word searches returned about 10-20 for Google, and few, if any, for Yahoo.

It's a pretty serious flaw in the methodology, as far as I can tell - they're double counting huge numbers of results, and so they're not really getting a good statistical sample of the index.

Those are estimates by mcc · 2005-08-15 07:35 · Score: 4, Insightful

Of course the study also demonstrates that on the searched terms, Yahoo's estimate numbers vastly overestimated the number of available results they actually found. So if the pages from the study are even close to representative in that regard then this would make the numbers you quote utterly meaningless.

Which is the entire reason, of course, why they kept the limits under 1,000 in the first place-- that for any number over 1,000, if the search engine says, say, "I found "2.5 million results for 'Valerie Plame'", you have no way to tell whether it's telling the truth or not.

--
Irritable, left-wing and possibly humorous bumper stickers and t-shirts

Holy lack of IR stastics understanding, Batman! by freality · 2005-08-15 08:41 · Score: 4, Interesting

The most basic measure of performance in Information Retrieval is precision vs. recall.

Precision is how many of the results that you return are correct. e.g. If Google returns 100 results and 10 of them are correct, then the precision on that query is 10%.

Recall is how many of the correct results you return. e.g. If Yahoo returns 100 results out of a total 1000 correct matches, then the recall on that query is 10%.

Information retrieval systems such as search engines balance these two metrics -- which are fundamentally at odds with each other -- to give the "best balance" in the eyes of the system's designers.

The NCSA study basically misses the effect this decision would have on perceived size of index.

A simple demonstration shows how it works.

First let's say both search engines have the same index size: 10B pages. Second, let's say both search engines have exactly the same apriori capability for precision and recall, but can tune for a preferred performance. Yahoo decides it wants to favor more precise results over more results recalled, at a 2:1 relative ratio compared to Google.

In that case, any given query will show half the hits from Yahoo as compared to Google. Concluding Yahoo's index to be half the size of Google's, given this result, would be incorrect.

Furthermore, without knowing the precision/recall performance of either system, they can only demonstrate a lower-bound on index size, and that certainly doesn't predict average or max index size.

Slashdot Mirror

NCSA Compares Google and Yahoo Index Numbers

21 of 395 comments (clear)