Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Would the common crawl dataset be useful to you starting out?

http://commoncrawl.org



Yes absolutely. I have been holding off crawling because I have no server capacity yet. That will probably sort itself out pretty soon from the looks of it. When I have the disk space I'll start using their data.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: