INSIDE MICHIGAN/ A humane public-data project
142,592 · records  ·  520,719 · sentences

§ 09 · The Stacks · Methodology

How the shelf tells the truth.

The governing rule is simple: preserve what the public row says, preserve what it leaves blank, and never turn one kind of absence into another for rhetorical effect.

01 · Unit of analysis

Entries, not deduplicated works.

Each rendered record corresponds to one row in one MDOC list snapshot. Similar editions or repeated titles are not silently merged. IDs and slugs are reproducible from this snapshot’s source identity and source order; they are stable within this edition, not promised to persist across future editions.

02 · Documentation matrix

Two fields, four honest states.

01

Code + narrative

The row contains at least one standardized rejection code and narrative rejection language.

02

Code only

The row contains a standardized code but its narrative rejection-language field is blank.

03

Narrative only · no standardized code

The code field is blank, but the row still publishes narrative rejection language. This is uncoded, not unexplained.

04

Neither code nor narrative

Both public fields are blank. Only this state receives the blackout treatment in The Stacks.

03 · Provenance

The row travels with its evidence.

Every entry retains raw and display title, raw author and format, ISO date, every identifier, raw code text, normalized code tokens when possible, verbatim narrative language, source snapshot, PDF page, row, ordinal position, and parser warnings. Normalization never overwrites the raw field.

Source edition
2026-01-12
SHA-256
a810776a4822d97766b468aae4cbd7e6acae19519fee537b9d07d2326542ac94

04 · What the charts claim

Public-list structure, not hidden intent.

  • The Date Added chart uses every distinct source Date Added value; no fixed number of “batches” is assumed.
  • Documentation counts report whether fields are present, not whether MDOC kept other internal records.
  • A publication can carry multiple codes. Code-assignment counts and publication counts must be labeled separately.
  • Verbatim rejection language is shown as source language, not adopted as this publication’s characterization.
  • Any enrichment or multi-state match must display its source, vintage, method, confidence, and denominator.

05 · Policy-code drift

A number is not a timeless meaning.

MDOC’s published rejection numbers refer to PD 05.03.118, but the numbered items changed across policy versions. The recovered archive first shows the recent renumbering in the version effective August 1, 2023. A source row receives a policy meaning only when the policy in force for that row has been established; the current policy is never projected backward by convenience.

06 · Enrichment and identity

An identifier is a clue, not a verdict.

The source contains repeated identifiers. The offline audit found 13 collision groups: three repeat the same normalized title and author, while ten attach one ISBN to conflicting identities. External bibliographic metadata is accepted only after normalized title and author agree; identifier-only agreement cannot overwrite the public row. Across 312 source rows, the completed enrichment ledger records 14 accepted, 20 conflicted, and 278 missed outcomes.

Those outcomes are not interchangeable. A Google Books HTTP 429 caused by quota exhaustion remained a provider failure and was never rewritten as a no-match. Only 10 of 312 rows carry exact provider-supplied subject labels. Those subjects are bibliographic metadata from the provider, not MDOC categories or model-inferred themes; one row can carry several labels.

The exact-language audit finds 7 narrative strings repeated across Date Added values on 24 rows. C9 uses exact parsed-field equality only: no trimming, case folding, fuzzy matching, semantic clustering, or category inference. Repeated language does not prove semantic equivalence, a policy template, a common policy, coordinated review, or the reason for a restriction.

07 · Other states and historical files

Disclosure is mapped before restriction is compared.

The state atlas records source scope, vintage, format, provenance tier, acquisition status, and comparability limits for 29 states. This build holds 26 states and 29 artifacts in byte-verified custody. That custody layer is deliberately wider than the locked P0 normalization cohort: 8 states and 10 artifacts were attempted, 9 artifacts across 7 states passed their parser gates, and one Washington artifact was refused rather than coerced into a table.

P1 is normalized aggregate evidence, not custody-only material: 13 artifacts across 12 states were attempted, 12 parsed into 17,755 source rows, and exact Michigan intersections were computed for 7 artifacts only where publication identity was valid.

P2 covers 6 custody-verified artifacts across 6 states. 5 passed the privacy gate and yielded 5,090 normalized source rows. Exact intersections were computed for 3 identity-eligible artifacts, producing 6 source-row matches that represent 5 distinct Michigan rows. Idaho stopped at the privacy gate. Nebraska’s ten worksheet scopes remain separate rather than being flattened into one statewide list. A noncomputed intersection is JSON null, never zero.

Across P0, P1, and P2, an exact intersection uses ISBN-13 or an exact normalized title-and-author pair only; fuzzy and title-only matches are excluded. It establishes shared publication identity, not comparable denominators, rates, ranks, list sizes, or restrictiveness. State practices remain unaligned in institutional reach, decision status, vintage, retention, and disclosure.

Historical comparisons are snapshot-specific too. The only publishable reviewed-exact diff runs from 2025-02-26 to 2026-01-12: it reports 31 added, 4 changed, 22 removed, and 277 unchanged rows. The recovered 2023 and 2024 layouts remain refused, and the four recovered files do not establish a complete publication history between snapshots.

08 · Parser gates

Silence is a failure, not a parse result.

The parser strips invisible direction marks, locates columns by page coordinates, joins wrapped cells, preserves page and logical-row positions, and accounts for extracted words. Promotion fails on row-count, ordering, required-field, format, fixture, word-accounting, database-integrity, or provenance errors. The reviewed January 12, 2026 profile resolves to exactly 312 rows.

09 · Reproducibility

The generated module is not hand-edited.

The production data module is rendered from the validated pipeline artifact. The source checksum, row-count gates, parser warnings, and regeneration command ship with that artifact. This frontend consumes that generated module directly; it is not maintained as a second hand-edited dataset.

← Return to The Stacks