Building a Bigger Search Engine

← Back to Stories (view on slashdot.org)

Building a Bigger Search Engine

Posted by ryuzaki0 on Saturday April 19, 2003 @02:15PM from the size-isn't-everything dept.

skreuzer writes "Wired is running a story about a distributed web crawler called Grub. People who choose to download and run the client will assist in building the Web's largest, most accurate database of URLs. This database will be used to improve existing search engines' results by increasing the frequency at which sites are crawled and indexed. Conceivably, Grub's distributed network could enable state information to be gathered on every document on the Internet, each and every day."

13 of 278 comments (clear)

Min score:

Reason:

Sort:

Will Grub take off or be smashed? by Blaine+Hilton · 2003-04-19 14:17 · Score: 4, Insightful

I started to use grub, but then questions started cropping up. First we are using this to further a commercial organization. This is not research such as SETI or Folding At Home; this is doing the dirty work of a large commercial search engine. There is not even any potential reward such as with distributed.net.
Also the grub engine crawls everything, including adult content and other questionable content. They have a setting to turn it off, but it does not block it. With the current questioning of international law relating to accessing illegal websites this could have major consequences for the average user.
So for the time being I have stopped using the grub client until some serious questions are answered. It's an interesting concept and if it was being used in more of an academic setting it could be interesting. However I believe that search engines like Google are doing pretty good themselves.
Go calculate something
1. Re:Will Grub take off or be smashed? by bcrowell · 2003-04-19 15:59 · Score: 3, Insightful
  
  This is not research such as SETI or Folding At Home; this is doing the dirty work of a large commercial search engine.
  Actually, if I had a gun to my head, I'd choose to run Grub, because the client is open-source. I used to run SETI@home, but then the news came out that they'd been sitting on a potential root vulnerability for a long time. That really brought home to me the risks of running someone else's closed-source app on my box.
  
  --
  Find free books.
2. Re:Will Grub take off or be smashed? by kaden · 2003-04-19 16:00 · Score: 5, Insightful
  
  Um, I think you're missing the point. This client could download highly illegal files, and make it look like I'm knowingly downloading them. Say I run it, and it downloads anything from kiddy porn to some Al Qaida webpage from an FBI sting server. I would quite possibly be arrested and charged, and while I wouldn't be convicted, it's quite an ordeal, and there's an ugly social stigma to even being charged with Kiddy Porn or conspiring with a terrorist. So that's a serious question that's posted by running Grub.
Great idea, but will it pan out? by dtolton · 2003-04-19 14:17 · Score: 5, Insightful

LookSmart hopes to tap the altruistic nature of many Internet users.

That unfortunately seems like a naively optimistic hope. While the
vast majority of people may be altruistic, it only takes a few
unscrupulous individuals to completely undermine a fair result.

It's interesting that this idea is an extension to Google's model in
many ways. Essentially Google is able to index so much of the
interent by having 50,000+ servers. I don't think that's what makes
Google such a useful search tool, rather I think it's accuracy and
relevancy. If my search results started getting poluted with bogus
hits, I would stop using it almost immediately.

Unfortunately, by letting people run the client on their machine and
having it send the results back to the server, I think spoofed
results are inevitable. I don't think it will be possible to
safeguard the results either, it will be interesting to see how well
this project survives *when* people start spoofing results. It's
been a problem for SETI@home, and it's something that undermined some
peoples faith in the project as a whole. If the spoofed results are
more widespread and have a larger impact as they would in a system
like this, it may ultimately prove fatal to the project.

One factor that has been asbolutely critical to Google's success has
been their ability to remain resistant to spoofing attempts. It's
still a question mark how well grub will perform in that context.

--

Doug Tolton

"The destruction of a value which is, will not bring value to that which isn't." -John Galt
Hrmm, I wonder how long... by bergeron76 · 2003-04-19 14:22 · Score: 3, Insightful

until someone figures out a way to compromize their local client's results and "escalate" their fave URLS.

It still sounds like a really cool idea though.

--
Don't think that a small group of dedicated individuals can't change the world. It's the only thing that ever has.
1. Re:Hrmm, I wonder how long... by CaptainMunchies · 2003-04-19 14:38 · Score: 3, Insightful
  
  Grub's clients don'tcome up with a ranking for each website they crawl; rather, they check to see if this website has changed since the last time it was crawled. For any website that has changed, the client notifies the server. The search engine asks the server which sites in its index need to be updated, and the server gleefully replies.
  
  Clients artificially increasing their ranking isn't an issue, since the client has nothing to do with a site's ranking.
  
  --
  Spam removed for the Internet's pleasure ...
Firewalls? by adam_megacz · 2003-04-19 14:28 · Score: 5, Insightful

So if I choose to run this client, how do I know that it won't accidentally index content that is only accessible from behind my firewall?
What about the RIAA? by One+Louder · 2003-04-19 14:51 · Score: 3, Insightful

So...let's say my instance of Grub crawls over a repository of .mp3s and supplies that information to the combined index.
What's the difference between my machine indexing them and the university students recently being hauled into court for indexing open shares? Why would I not be held liable for contributory copyright infringement?
No thanks.
A better use for my screensaver time by Call+Me+Black+Cloud · 2003-04-19 14:57 · Score: 5, Insightful

I prefer grid.org to grub.org. There the cycles are going to cancer or smallpox research. Currently over 2 million machines are participating.

Altruism has its place, but since I'm more likely to die of cancer than of not having the complete www indexed I think I'll be selfish and work towards a cure for something that may affect me.
Re:Search engine software and lack of A . I . by zymano · 2003-04-19 15:31 · Score: 3, Insightful

I didn't know that.
But it still kind of irks me that people think that a computerized 'dumb' search result could compete with a human rating system that filters spam,porn,and other garbage results. Google should hire some REAL PEOPLE that can do some sort catagorized intelligent directory so we can have QUALITY at the beginning of a search result. Some sort of HUMUN RATING system is needed to sort. The software is not up to par.
Good Idea, Bad Implementation by oaf357 · 2003-04-19 15:52 · Score: 3, Insightful

Yea. If you help Grub, Grub gives your web site a preferencial listing. Building the biggest search engine, sure. Building good search results, not so sure.
Unlimited Use? Try Wishful Thinking. by NeoMoose · 2003-04-19 16:37 · Score: 3, Insightful

You can always use the Google API for more than 2,000 searches per day if you pay licensing fees for it. That's just Google ensuring that it can remain a viable company. Little text-box advertisements just don't cut it in this day and age where blatant pop-ups and colorful banner ads don't even have much turn-around. That's not the point though.

The point is that I wouldn't look anytime soon for LookSmart to allow unlimited usage of this API. It's too large of a project for them to just let people use it. It's simple economics. They may not be investing the computing resources into this projects web spidering software, but it's still using TONS of resources to keep this data catalogued and readily accessible.
Re:Altruistic? by eversunsoft · 2003-04-19 18:36 · Score: 4, Insightful

Well, because web searching, to this day in age, has been a free service. Supposing that the index is built as the result of donated searches, it would be ethically in very bad taste to act against this trend.
Of course, I am the first one to question this trend. Has anyone else considered the possibility that one day we'll wake up, and notice that google is charging for access to it's basic searching services?
I for one, would probably pay. I have become so dependent on it. What price? That's a good question...