no key · no human

What is inside

Roughly 110,000 news articles a day from 25,000+ publishers, on a rolling 30-day window, refreshed continuously. Here is exactly what that means.

Per article

Title, description, full body text, publication time, crawl time, publisher hostname and name, country, language, top-level domain, author, publisher categories, lead image URL — plus country_source and date_source, which say how the inferred fields were derived. Field reference.

Volume

MeasureValue
Articles per day, after filtering~110,000
Batches per day24, one per hour
Publishers in a typical day~3,000
Retention window30 days, rolling
Crawled pages discarded as non-articles or duplicates~40%
Typical body length2,000–4,000 characters

Where it comes from

Large-scale public news archives: bulk datasets of already-crawled news pages, published continuously as open data. We take each release as it appears, extract articles, and index them. We do not fetch pages from publishers' servers ourselves. More on the sources.

What it is good for

  • Grounding a language model in what was actually reported, with the body text to quote from.
  • Monitoring a topic, company or region across thousands of publishers at once.
  • Comparing coverage between countries and languages — the same event as reported in Ukraine, Germany and India.
  • Measuring how much and where something is being covered, via /v1/stats.

What it is not good for

  • Guaranteed coverage of a named publisher. The archives reach what they reach. A publisher present today may be thin tomorrow.
  • History. Thirty days. Nothing older exists here.
  • Second-by-second breaking news. Batches land hourly; the median lag from publication is one to two hours.
  • A clean publisher taxonomy. Categories are whatever each publisher labels its own sections, unnormalised.