GPT-5 is a reasoning model and is not bit-deterministic even at the API level (no temperature override allowed). Variance testing in iter 7 showed ~1 in 6 loop runs hit a nondeterministic regression: in one run the magazine_mead exemplar lost its italic markers around "New Yorker" even though the exact same input had produced a clean output 5 times prior. The scoring linter caught it (rule_magazine_name_italicized) but the formatter still emitted the bad output to the user. This commit adds a runtime guardrail: - New module src/cmos/runtime_validator.py — type-independent structural sanity checks. Operates on a single candidate string with no Exemplar context, because at runtime we don't know the source type. Checks: ends with period, no Ibid., has at least one italic span (almost every CMOS bibliography entry italicizes something), balanced * and " markers. Deliberately weaker than the scoring linter; it's a fast guardrail, not a full validator. - format_bibliography_entry now retries up to DEFAULT_MAX_RETRIES (2) times when the validator rejects a candidate. Independent re-calls are usually enough because the failures are stochastic. If every attempt fails, the LAST attempt is returned (no exception) — the caller still gets something usable, and the failure surfaces through the scoring linter or human review. The retry path costs zero on the common case (1 call per entry); ~1-2% extra calls on noisy drafts. Empirical: 3 consecutive loop runs after this change are scalar 1.000 canary 1.000 (vs 5/6 clean in the variance test before). Sample is too small to claim full suppression but the signal is positive. Also adds a new exemplar journal_multiauthor_secondary_first_last.toml captured from the user's HML draft (Hasanah et al., IJIDI 2024). It exercises the case where a multi-author entry has the first author inverted and the rest in First Last form — which the iter 7 HML run got wrong on one entry. Variance testing showed the exemplar passes 6/6 in isolation, so the original HML failure was nondeterminism, not a missing rule. Keeping the exemplar regardless: it adds canary coverage of a real-world multi-author pattern, no-DOI / JSTOR-URL variant, and the year-suffix author-date holdover stripping. Tests: 85 → 95 (8 new for runtime_validator + 2 new for formatter retry). All passing.
Exemplars
Each .toml file is one exemplar: a messy input, a canonical structured
record, and an expected CMOS 18 bibliography string.
Format
# IMPORTANT: all top-level keys (including expected_bibliography) must come
# BEFORE any [section] header, or TOML will absorb them into that section.
source = "Purdue OWL CMOS 18 NB PDF, p. X" # provenance, required
type = "book" # book | journal | chapter | web | ai | ...
tags = ["book", "single-author"]
canary = false # if true, exact-string match required
messy_input = "..."
expected_bibliography = "Last, First. *The Book Title*. Publisher Name, 2023."
[canonical]
# Structured fields for the field-level diff. Fields vary by type.
author = "Last, First"
title = "The Book Title"
publisher = "Publisher Name"
year = 2023
Primary sources
- Purdue OWL CMOS 18 NB PDF: https://owl.purdue.edu/owl/research_and_citation/chicago_manual_18th_edition/documents/cmos-nb-20250731.pdf
- CMOS official quick guide: https://www.chicagomanualofstyle.org/tools_citationguide/citation-guide-1.html
- Widener "New in 18th" LibGuide: https://widener.libguides.com/Chicago18/New