What is inside
Roughly 110,000 news articles a day from 25,000+ publishers, on a rolling 30-day window, refreshed continuously. Here is exactly what that means.
Coverage
Live counts by country, language and domain zone. Real numbers, updated from the index.
Sources
Which publishers are actually in there, and how much each contributes.
Methodology
How an article is extracted, dated, and assigned a country and language.
Quality
What this dataset gets wrong, stated plainly, with the numbers.
Per article
Title, description, full body text, publication time, crawl time, publisher hostname
and name, country, language, top-level domain, author, publisher categories, lead image
URL — plus country_source and date_source, which say how
the inferred fields were derived. Field reference.
Volume
| Measure | Value |
|---|---|
| Articles per day, after filtering | ~110,000 |
| Batches per day | 24, one per hour |
| Publishers in a typical day | ~3,000 |
| Retention window | 30 days, rolling |
| Crawled pages discarded as non-articles or duplicates | ~40% |
| Typical body length | 2,000–4,000 characters |
Where it comes from
Large-scale public news archives: bulk datasets of already-crawled news pages, published continuously as open data. We take each release as it appears, extract articles, and index them. We do not fetch pages from publishers' servers ourselves. More on the sources.
What it is good for
- Grounding a language model in what was actually reported, with the body text to quote from.
- Monitoring a topic, company or region across thousands of publishers at once.
- Comparing coverage between countries and languages — the same event as reported in Ukraine, Germany and India.
- Measuring how much and where something is being covered, via /v1/stats.
What it is not good for
- Guaranteed coverage of a named publisher. The archives reach what they reach. A publisher present today may be thin tomorrow.
- History. Thirty days. Nothing older exists here.
- Second-by-second breaking news. Batches land hourly; the median lag from publication is one to two hours.
- A clean publisher taxonomy. Categories are whatever each publisher labels its own sections, unnormalised.