This is not how to crawl webpages. He started with the Alexa list. Those are not necessarily domain names of servers serving webpages. I would guess that some of the request to cease crawling came from some of these listings. Working from the Alexa list he would have been crawling some of the darkest underbelly of the web: ad servers and bulk email services.
His question: "Who gets to crawl the web?" is an interesting one though.
Do not assume that Googlebot is a smart crawler. Or smarter than all others. The author of Linkers and Loaders posted recently on CircleID about how dumb Googlebot can be.
There is no such thing as a smart crawler. All crawlers are stupid. Googlebot resorts to brute force more often than not.
Theoretically no one should have to crawl the web. The information should be organised when it is entered into the index.
Do you have to "crawl" the Yellow Pages? Are listings arranged by an "algorithm"? PageRank? 80/20 rules?
Nothing wrong with those metrics; except of course that they can be gamed trivially, as experiments with Google Scholar have shown. But building a business around this type of ranking? C'mon.
If the telephone directories abandoned alpha and subject organisation for "popularity" as a means of organisation it would be total chaos. Which is why "organising the world's information" is an amusing mission statement when your entire business is built around enabling continued chaos and promoting competition for ranking.
Even worse are companies like Yelp. It's blackmail.
If the information was organised, e.g., alphabetically and regionally, it would be a lot easier to find stuff. Instead, search engines need to spy on users to figure out what they should be letting users choose for themselves. Where "user interfaces" are concerned, it is a fine line between "intuitive" and "manipulative".
The people who run search engines and directory sites are not objective. They can be bought. They want to be bought.
This brings quality down. As it always has for traditional media as well. But it's much worse with search engines.
>Theoretically no one should have to crawl the web. The information should be organised when it is entered into the index.
What do you mean by this statement? I can see from your post you have a contrarian view on how to organize and find stuff on the web, but I'm struggling to understand what alternative you're proposing.
Well, I'm not him but: probably starting from a zone file and narrowing it down to only whitelisted and "legit" domains would be a good start.
Maybe during the registration process more metadata should be demanded of people and anonymity prohibited or reduced. That way for example if you wanted a list of all the .com blogs it is just a grep away and tied into mostly real people for example. Corporate websites tied to their business entity with an EIN or something and verified. 'etc.
The thing is.. that ship has sailed a long time ago so we are stuck with google.
His question: "Who gets to crawl the web?" is an interesting one though.
Do not assume that Googlebot is a smart crawler. Or smarter than all others. The author of Linkers and Loaders posted recently on CircleID about how dumb Googlebot can be.
There is no such thing as a smart crawler. All crawlers are stupid. Googlebot resorts to brute force more often than not.
Theoretically no one should have to crawl the web. The information should be organised when it is entered into the index.
Do you have to "crawl" the Yellow Pages? Are listings arranged by an "algorithm"? PageRank? 80/20 rules?
Nothing wrong with those metrics; except of course that they can be gamed trivially, as experiments with Google Scholar have shown. But building a business around this type of ranking? C'mon.
If the telephone directories abandoned alpha and subject organisation for "popularity" as a means of organisation it would be total chaos. Which is why "organising the world's information" is an amusing mission statement when your entire business is built around enabling continued chaos and promoting competition for ranking.
Even worse are companies like Yelp. It's blackmail.
If the information was organised, e.g., alphabetically and regionally, it would be a lot easier to find stuff. Instead, search engines need to spy on users to figure out what they should be letting users choose for themselves. Where "user interfaces" are concerned, it is a fine line between "intuitive" and "manipulative".
The people who run search engines and directory sites are not objective. They can be bought. They want to be bought.
This brings quality down. As it always has for traditional media as well. But it's much worse with search engines.