How We Built Our 60-Node (Almost) Distributed Web Crawler - http://t.co/Jqpd6zAE
September 04, 2012
Advertising
/http%3A%2F%2Fblog.semantics3.com%2Fassets%2Fworker-supervisor.png)
How We Built Our 60-Node (Almost) Distributed Web Crawler | Blog - Semantics3
Web crawling is one of those tasks that is so easy in theory (well you visit some pages, figure out the outgoing links, figure out which haven't been visited, queue them up, pop the queue, visit the page, repeat), but really hard in practice, especially at scale.




Comment on this post