Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Do you have a rough set of guidelines for how fast we should request from HN? For a side project, I was thinking of writing something that scraped the HN frontpage and all the associated comment threads every 10 minutes or so, and I'd rather not cause performance issues or get banned. I'd be happy to rate-limit requests to whatever is convenient.


May be better to use the official API.

http://www.hnsearch.com/api


That's not an official API: http://www.hnsearch.com/about

Quote: "HNSearch was built by the team at ThriftDB to give back to the community and to test the capabilities of the ThriftDB flexible datastore with search built-in."

Interesting API all the same though.


Regardless whether it is official or not, it is pg's preferred api: http://news.ycombinator.com/item?id=4694308


If it were an official API, wouldn't it be associated with HN or Y Combinator rather than an external website?


It is by a YC company, and recommended by pg.


The robots.txt file for HN suggests a Crawl-Delay value of 30 seconds.


This might be helpful: http://api.ihackernews.com/

edit: oh, official API is above. Disregard this one :-)




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: