They have the code for the crawler up on the website, I've just processed all the WARC and created an index and a RocksDB so that I can make a tool to fetch articles from the WARC. The crawler follows robot.txt
Brian PRO
brian-learns
AI & ML interests
ethical ai use for cultural heritage use cases
Recent Activity
updated a bucket 1 day ago
brian-learns/cc-news-cdx-server-storage updated a Space 2 days ago
brian-learns/cc-news-cdx-serverOrganizations
None yet