Begin v2 (in-text citation note form) as a parallel artifact to the
v1 bibliography formatter, per the Path B architectural decision:
no shared mutable state, no shared prompt content, no v1 changes.
New artifacts:
- src/cmos/parser.py: add find_notes() / NoteDefinition /
NotesParseResult as siblings to split_bibliography(). Targets
pandoc-style markdown footnote definitions [^marker]: text.
Single-line only; multi-line continuation deferred.
- src/cmos/note_formatter.py: new module mirroring formatter.py.
17-rule SYSTEM_PROMPT_NOTES for CMOS 18 first-occurrence note
form. Reuses cmos.runtime_validator.validate() unchanged — its
structural checks all apply to note form too. Caller injection,
retry loop, and model selection mirror v1.
- tests/test_parser.py: 7 new tests for find_notes().
- tests/test_note_formatter.py: 11 new tests mirroring v1 test
discipline (no API calls, fake-caller injection, retry semantics,
prompt smoke checks).
- exemplars/notes/: new subdirectory with 3 hand-synthesized
first-occurrence note exemplars (book, journal article, chapter
in edited book). Uses expected_note as the field name (not v1's
expected_bibliography). harness/score.py:load_exemplars uses a
non-recursive glob, so the v1 canary loader does not see these —
Path B isolation is automatic.
v1 untouched. v1 canary still loads exactly 17 exemplars. Full
test suite: 96 v1 + 18 new v2 = 114 passing.
Smoke-tested all 3 v2 exemplars against real GPT-5 (not fakes),
2 runs each. book_first_yu and journal_first_kwon: 4/4 byte-perfect
on the first try. chapter_first_doyle: 0/2, surfacing two known
prompt gaps for the next iteration:
1. Publisher abbreviation not expanded ("U of Chicago Press"
preserved instead of "University of Chicago Press"). v1's
formatter.py rule 14 is missing from the v2 prompt.
2. "ed." pluralized to "eds." for multiple editors. CMOS NB
uses "ed." invariantly regardless of editor count.
Both gaps are addressable prompt edits, not architectural problems —
exactly the kind of finding the dev loop is designed to surface.
Out of scope (deferred to chunk 2 and later): CLI extension,
shortened-form generation, document reassembly, harness scoring
loop integration, v2 linter rules, real-draft testing.
Two new real-draft-derived journal exemplars locking down prompt
rule 29 from the preceding commit. Follows the Hasanah precedent
for promoting real-draft failures into the canary corpus, per the
"defense in depth" feedback in project memory.
- journal_preposition_huttunen.toml exercises "among" mid-title,
plus four other CMOS 8.159 lowercased words ("on", "in", "a",
and implicitly "among"). Real article: Huttunen and Kortelainen,
JASIST 72, no. 7 (2021). The messy_input uses the ASCII hyphen
in "Meaning-Making" for clarity — the source draft actually had
a U+2010 non-breaking hyphen from docx->txt conversion, but
mixing Unicode hyphen normalization into this exemplar would
have conflated two independent failure modes. The U+2010 issue
is noted in the exemplar comment for a separate iteration.
- journal_preposition_mehra.toml exercises "beyond" mid-title,
plus mid-title exclamation-point preservation (rule 25) and
leading-"The" drop from "The Library Quarterly" (rule 17).
Real article: Mehra, Library Quarterly 91, no. 2 (2021).
Both are canary-enabled and will block any future prompt iteration
that regresses on these patterns via the exact-match canary axis.
The canonical dicts additionally lock them down via the field-
level diff, per the upstream-refinement feedback in project
memory.
Canary corpus size: 15 -> 17.
GPT-5 is a reasoning model and is not bit-deterministic even at the API
level (no temperature override allowed). Variance testing in iter 7
showed ~1 in 6 loop runs hit a nondeterministic regression: in one run
the magazine_mead exemplar lost its italic markers around "New Yorker"
even though the exact same input had produced a clean output 5 times
prior. The scoring linter caught it (rule_magazine_name_italicized) but
the formatter still emitted the bad output to the user.
This commit adds a runtime guardrail:
- New module src/cmos/runtime_validator.py — type-independent
structural sanity checks. Operates on a single candidate string
with no Exemplar context, because at runtime we don't know the
source type. Checks: ends with period, no Ibid., has at least one
italic span (almost every CMOS bibliography entry italicizes
something), balanced * and " markers. Deliberately weaker than the
scoring linter; it's a fast guardrail, not a full validator.
- format_bibliography_entry now retries up to DEFAULT_MAX_RETRIES (2)
times when the validator rejects a candidate. Independent re-calls
are usually enough because the failures are stochastic. If every
attempt fails, the LAST attempt is returned (no exception) — the
caller still gets something usable, and the failure surfaces
through the scoring linter or human review. The retry path costs
zero on the common case (1 call per entry); ~1-2% extra calls on
noisy drafts.
Empirical: 3 consecutive loop runs after this change are scalar 1.000
canary 1.000 (vs 5/6 clean in the variance test before). Sample is too
small to claim full suppression but the signal is positive.
Also adds a new exemplar journal_multiauthor_secondary_first_last.toml
captured from the user's HML draft (Hasanah et al., IJIDI 2024). It
exercises the case where a multi-author entry has the first author
inverted and the rest in First Last form — which the iter 7 HML run
got wrong on one entry. Variance testing showed the exemplar passes
6/6 in isolation, so the original HML failure was nondeterminism, not
a missing rule. Keeping the exemplar regardless: it adds canary
coverage of a real-world multi-author pattern, no-DOI / JSTOR-URL
variant, and the year-suffix author-date holdover stripping.
Tests: 85 → 95 (8 new for runtime_validator + 2 new for formatter
retry). All passing.
Add 11 new exemplars pulled verbatim from the CMOS official quick guide
(chicagomanualofstyle.org/tools_citationguide/citation-guide-1.html):
- book_two_authors_binder: Binder & Kidder (two-author book)
- book_chapter_doyle: Doyle in Marks & Parkin (chapter in edited volume,
no page range per CMOS 18)
- book_translated_liu: Liu Xinwu, trans. Tiang (translated book, non-
Western name not comma-inverted)
- book_edition_borel: Borel, 2nd ed. via EBSCOhost (edition + database)
- book_ebook_roy: Roy, Kindle format
- journal_many_authors_snyder: 7 authors in PLOS ONE, exercises the
CMOS 18 "first 3 + et al." threshold and article-ID page format
- magazine_mead: New Yorker (tests leading-"The" drop)
- newspaper_blum: NYT with URL (tests abbreviation expansion)
- social_media_cmos_facebook: Facebook post with case preservation
- podcast_ober: Pushkin podcast with season/episode/duration
- video_cowan_ted: TED Talk with venue and duration
Every exemplar is marked canary=true — we want byte-exact regression
detection on all of them. Source citations and CMOS rule justifications
are in the TOML doc comments.
All 14 canonical outputs lint clean under LINTER_VERSION v0.2.0. The
harness will score them in the following commit.
First real loop iteration found that GPT-5 silently drops date qualifiers
("Effective", "Published", "Accessed", etc.) when reformatting web
sources. The field diff was satisfied because the canonical date field
did not include the qualifier, but the canary exact-match axis caught
the regression.
Two fixes, per the canary-upstream policy:
1. Tighten the web-page exemplar canonical: date field now includes the
"Effective" qualifier so the field diff will catch future regressions
without relying on the canary.
2. Add SYSTEM_PROMPT rule 12 instructing the formatter to preserve
semantic date qualifiers from the input.
After these fixes: scalar 1.000, canary exact-match 1.000 on all three
seed exemplars.
Initial phase-1 baseline of the karpathy/autoresearch-style loop.
The formatter module is the inner-loop artifact; parser and linter
are infra. The linter carries a LINTER_VERSION hash (v0.2.0) that
will force a re-baseline on any rule change.
Components:
- harness/diff.py: case-sensitive field-level substring diff
- harness/score.py: three-axis scoring (field, linter, canary exact)
- src/cmos/linter.py: 9 CMOS 18 structural rules, each Purdue/CMOS cited
- src/cmos/parser.py: locate ## Bibliography section, split entries
- src/cmos/formatter.py: prompt + OpenAI call with caller injection
- src/cmos/cli.py: cmos format path/to/draft.md
- scripts/run_loop.py: loop runner with --fake mode for no-API runs
- exemplars/: 3 canary seed exemplars (book, journal w/DOI, web),
sourced from chicagomanualofstyle.org quick guide
Tests: 48 passing. Fake-mode baseline scalar = 0.000 on the 3 seed
exemplars (identity caller fails the linter on every rule). This is
the floor the real GPT-5 formatter needs to improve from.