8b14dbc9ca3bf4b6653fe27c14bec08d57456bf6
GPT-5 is a reasoning model and is not bit-deterministic even at the API level (no temperature override allowed). Variance testing in iter 7 showed ~1 in 6 loop runs hit a nondeterministic regression: in one run the magazine_mead exemplar lost its italic markers around "New Yorker" even though the exact same input had produced a clean output 5 times prior. The scoring linter caught it (rule_magazine_name_italicized) but the formatter still emitted the bad output to the user. This commit adds a runtime guardrail: - New module src/cmos/runtime_validator.py — type-independent structural sanity checks. Operates on a single candidate string with no Exemplar context, because at runtime we don't know the source type. Checks: ends with period, no Ibid., has at least one italic span (almost every CMOS bibliography entry italicizes something), balanced * and " markers. Deliberately weaker than the scoring linter; it's a fast guardrail, not a full validator. - format_bibliography_entry now retries up to DEFAULT_MAX_RETRIES (2) times when the validator rejects a candidate. Independent re-calls are usually enough because the failures are stochastic. If every attempt fails, the LAST attempt is returned (no exception) — the caller still gets something usable, and the failure surfaces through the scoring linter or human review. The retry path costs zero on the common case (1 call per entry); ~1-2% extra calls on noisy drafts. Empirical: 3 consecutive loop runs after this change are scalar 1.000 canary 1.000 (vs 5/6 clean in the variance test before). Sample is too small to claim full suppression but the signal is positive. Also adds a new exemplar journal_multiauthor_secondary_first_last.toml captured from the user's HML draft (Hasanah et al., IJIDI 2024). It exercises the case where a multi-author entry has the first author inverted and the rest in First Last form — which the iter 7 HML run got wrong on one entry. Variance testing showed the exemplar passes 6/6 in isolation, so the original HML failure was nondeterminism, not a missing rule. Keeping the exemplar regardless: it adds canary coverage of a real-world multi-author pattern, no-DOI / JSTOR-URL variant, and the year-suffix author-date holdover stripping. Tests: 85 → 95 (8 new for runtime_validator + 2 new for formatter retry). All passing.
cmos — Chicago Manual of Style 18 bibliography reformatter
A Python tool that reformats the bibliography section of an English-language markdown draft to CMOS 18th edition, notes-and-bibliography form.
Accuracy-first, iterative, built with a karpathy/autoresearch-style dev loop.
See program.md for the goal specification and
/home/claudecode1/.claude/plans/pure-mixing-yao.md for the implementation plan.
Quick start
uv sync
uv run pytest
uv run cmos format path/to/draft.md > out.md
Layout
src/cmos/formatter.py— the iterable artifact (edited every loop iteration).src/cmos/parser.py— extracts the bibliography section from a draft.src/cmos/linter.py— deterministic CMOS 18 rule checks (versioned).src/cmos/cli.py—cmos formatentry point.harness/score.py,harness/diff.py— frozen scoring.exemplars/— TOML test corpus.tests/— pytest suite (linter, parser, formatter, cli).rough_drafts/— user-supplied messy drafts for ad-hoc iteration.logs/— per-run scores, model ids, linter version hashes.
Languages
Python
100%