Methodology

Every number on Book Statistics is computed at build time from a snapshot of 13,258 state reading-list assignments, 100% of which carry a per-row source URL. Rows are counted directly from the data files — never taken from summaries — and any statistic whose underlying field coverage is too sparse is omitted rather than estimated.

The dataset

The core corpus is 13,258 reading-list assignments: each row records that a specific book appears on a specific U.S. state, district, or national-curriculum reading list, at a specific grade level. Every assignment row carries its own provenance:

This per-row provenance is the zero-fabrication backbone: a number that cannot be traced to a source document does not enter the dataset.

The joins

Assignments are joined at build time to five reference tables (row counts are counted directly from the files):

Table Rows Joined on
books (title, author, publication year, page count) 1,690 book_id
states 50 state_id
grade levels (K–12) 13 grade_id
districts 20 district_id
curricula (e.g. national curriculum exemplar lists) 82 curriculum_id

Of the 1,690 book records, 1,664 appear in at least one assignment. Assignments attach either to a state list (47 states have at least one) or to a national curriculum list; curriculum-level rows have no state and are labeled "National (curriculum)" in exports.

How statistics are computed

Refresh cadence

The dataset is refreshed annually, plus whenever a new snapshot of the underlying corpus lands. The current snapshot is dated 2026-07-07, printed on every data-bearing page along with the build's "Last verified" date. Because every figure recomputes at build time, a new snapshot updates every number on the site in one deploy — no hand-edited statistics anywhere.

Where the narrative lives

Book Statistics is the reference layer: lookup tables, definitions, per-year numbers. The narrative studies built on the same dataset — Reading Statistics 2026, The Most-Assigned Books of 2026, and How Old Are Assigned Books? — are published on readinglist.school. The two sites share the source data, never the text.

Questions about the methodology? Contact us — corrections are welcome and acted on.