1757304e444216f387ed69a1e36cb09d61924083
14
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1757304e44 | Merge branch 'main' of tildegit.org:MarkEEaton/cmos | ||
|
|
21fb9969c4 | update .gitignore | ||
|
|
cad936f576 |
parser: support diverse md inputs (bold headings, Works Cited, blockquote notes, junk filtering)
- Bibliography heading now accepts bold markers (## **BIBLIOGRAPHY**), alternative names (Works Cited, References), and numbered sub-headings within the section. - Original heading text preserved in output instead of hardcoded "## Bibliography". - New find_blockquote_notes() parser for PDF-to-markdown footnote format (> N text), wired into CLI as third fallback after pandoc and numbered. - PDF junk filtered from bibliography entries: bare page numbers, blockquote footnotes, download banners, CC license URLs, and short running headers. |
||
|
|
c21ca7b58e |
v2 chunk 5: 5 fixes from cross-draft triage (HML + Reading_Disrepair)
After chunk 4 verification on the Anti-Communist draft, ran v2
against the previously-untested ## Notes sections of HML (72 notes)
and Reading_Disrepair (25 notes) — 97 new notes for 204 total
across 3 drafts. The wider corpus surfaced 5 distinct issues that
weren't visible in the Anti-Communist data alone.
src/cmos/note_formatter.py:
1. Tighten substantive prose detector — exclude semicolon-bearing
inputs (chunk 5 task 1). The HML draft was dominated by compound
shortened references like:
See Author1, "Title;" Author2, "Title;" Author3, "Title."
17 of 72 HML notes matched this pattern, exceeded the 150-char
threshold, and had no citation skeleton markers, so they were
passed through verbatim by the chunk 3 detector. Compound
references use semicolons as item separators; substantive prose
doesn't. Adding `if ";" in stripped: return False` cleanly
separates the two without affecting the genuine substantive
notes (which contain none).
2. Rule 3 — distinguish comma-inside vs period-inside for
article/chapter title closing quote based on whether MORE
content follows the title. Surfaced by 8+ entries across drafts
producing the wrong `"Title,".` pattern (comma inside followed
by stray period outside) when the title is the last element of
the entry. Rule 3 now explicitly says: comma inside when more
content follows; period inside when title is the last element.
3. New rule 20 — preserve signal phrases verbatim. Surfaced by HML
[69] and [70] where "For example," was stripped while "See" was
preserved elsewhere. CMOS notes commonly open with signal
phrases (See, See also, For example, Cf., Compare, But see,
Contra, Accord, Quoted in) that indicate how a citation relates
to the surrounding argument. Stripping them is data loss.
4. Rule 1 — explicit "no comma between author name and 'et al.'"
in non-inverted note form. Surfaced by HML [69] where "Bignoli
et al." became "Bignoli, et al." (extra comma). The comma
before "et al." is a feature of inverted bibliography form,
not non-inverted note form. Rule 1 now includes WRONG examples.
5. Citation skeleton URL marker — also accept bare-domain URLs
without the https:// scheme. Surfaced by Reading_Disrepair [1]
GAO entry which had `files.gao.gov/...` without `https://`,
so the existing https?:// marker didn't match and the 190-char
citation was passed through as substantive prose. New pattern:
`\b\w{2,}(?:\.\w{2,})+/\S*` catches bare domains like
files.gao.gov/path without false-positiving on common things
like "e.g./" (single-char tokens excluded by {2,}).
Tests: 157/157 (was 153, +4 net new for prompt-content and regex
verification). The +4 is: compound-reference detection, genuine
substantive prose regression check, rule 3 / rule 20 / rule 1
prompt-content tests, bare-domain URL detection.
Path B: rule 18 (publishers, chunk 4) and rule 20 (signal phrases,
this commit) are v2-only. Rule 1 et al. clarification could in
principle apply to v1 too but v1 doesn't see "et al." in the same
non-inverted form. v1 untouched. 96/96 v1 tests still passing.
Real-draft validation pending — verification re-run against all
3 drafts will follow in the next step.
|
||
|
|
83339e313c |
v2 chunk 4: revert validator carve-out + expand publisher exceptions
Two follow-ups based on the chunk 3 verification re-run findings. src/cmos/runtime_validator.py — revert chunk 3 fix #3 (carve-out) Chunk 3 fix #3 added a carve-out so shortened-form notes ("Rosen, 7.", "Ettarh.") would pass the validator without italics, saving retry cost. Chunk 3 verification on the Anti-Communist draft showed the carve-out also let GPT-5's variance produce under-italicized variants of shortened forms WITH short titles (e.g., "A Restudy, 73." instead of the more CMOS-correct "*A Restudy*, 73."). The user explicitly chose accuracy over the cost saving and asked for the carve-out to be reverted. The strict italics check now applies uniformly. Simple shortened forms (Mitchell, 197., Ettarh.) get retried unnecessarily and waste API cost without producing better output. Shortened forms with short titles get a fair shot at the italicized version on the retry. Cost ↑, accuracy ↑. Removes _SHORTENED_FORM_RE, _looks_like_shortened_form, and the carve-out check from validate(). Updates module docstring with history note. Inverts the 2 carve-out tests in test_runtime_validator.py to assert shortened forms now FAIL the strict check, documenting the design intent for future reviewers. src/cmos/note_formatter.py — expand rule 18 publisher exception list Chunk 3 verification showed [6] Batterson getting "NYU Press" expanded to "New York University Press", losing the publisher's canonical brand. Rule 18's exception list previously only named MIT Press, ALA Editions, and MLA. Expanded to include NYU Press, Routledge, IEEE Press, ACM Press, WHO Press, plus "Pew Research Center" as a non-Press canonical example. Also added a "when in doubt" guidance paragraph: if the abbreviation contains "Press" and is widely used as the publisher's own branding, leave it alone. The risk of losing a brand name (NYU Press) is worse than the risk of leaving an obscure abbreviation. v2 only. Tests: 151/151 green (was 150). The +1 is the new NYU Press prompt-content test; the 2 inverted runtime_validator tests stayed at the same count. Path B: runtime_validator change crosses the v1/v2 boundary (shared infrastructure), per the same approval that authorized the original chunk 3 fix #3. v1's bibliography formatter is unaffected in practice because v1 outputs always have italics. |
||
|
|
553240c153 |
v2 chunk 3: bug fixes from real-draft triage + polish
Six independent fixes addressing the Anti-Communist Formations of LIS real-draft test findings (chunk 2c). Four are bug fixes for issues that destroyed user data or wasted API calls; two are polish quality improvements revisited from the deferred list. src/cmos/note_formatter.py — empty input guard (fix #1) Empty / whitespace-only input now short-circuits the API entirely and returns the input verbatim. Surfaced by note [66] in the Anti-Communist draft, where GPT-5 broke character on empty input and returned a conversational meta-reply ("Please paste the citation entry...") that then got substituted into the document. Whitespace preserved so cli.reformat_notes is byte-exact on empty notes. v2 only. src/cmos/note_formatter.py — deprecated Latin guard (fix #2) New Python guard catches the full CMOS-18-deprecated Latin citation set (ibid, idem, id., op. cit., loc. cit.) and returns input verbatim before the API call. Rule 10 in the prompt also rewritten: explicitly says "return verbatim" instead of the old "return cleanest possible full-form note", which the model interpreted as "return empty when there's no information", silently destroying the marker. Surfaced by note [61] = "Ibid.". v2 only. src/cmos/runtime_validator.py — shortened-form carve-out (fix #3) The "must contain italics" check is now skipped when the candidate looks like a CMOS shortened-form note (author last name, optional page, no italic content). Surfaced by ~30 of 107 notes in the Anti-Communist draft that were correctly returned as shortened forms ("Rosen, 7.", "Mitchell, 197.") but rejected by the validator's strict italic check, firing the full retry budget on valid output (~60 wasted API calls per run). The carve-out is conservative: name-token regex + terminal period; doesn't match unstructured prose. Shared infrastructure — affects v1 and v2, no-op in v1's normal workflow because v1 outputs always have italics. src/cmos/note_formatter.py — substantive prose pass-through (fix #4) Long discursive prose with no citation skeleton markers now short-circuits the API and returns the input verbatim. Surfaced by notes [8] (689 chars), [9] (893 chars), [75] (163 chars) in the Anti-Communist draft — substantive notes that the formatter was extracting a single citation from and silently discarding the surrounding commentary. CMOS 14.39 explicitly allows substantive notes; preserving them is the user's explicit policy ("preserve all free text as long as that does not break other formatting"). Detection: length > 150 chars AND no citation skeleton markers (parenthesized year, URL, vol./no./pp., DOI, terminal page or terminal year). Conservative — does not false-positive on legitimate first-occurrence notes. v2 only. src/cmos/formatter.py + note_formatter.py — month/season preservation (polish #5) Both v1 rule 10 (journal article format) and v2 rule 5 (journal article note form) now explicitly instruct the model to preserve (Month YEAR) and (Season YEAR) parentheticals when the source provides them. Surfaced by entries in both Reading_Disrepair and Anti-Communist where (Spring, 1993) and August 1952 were silently collapsed to (1993) and 1952. Both forms are valid CMOS but month/season is more informative when the source has it. Affects v1 and v2 in mirror. src/cmos/note_formatter.py — government documents rule 19 (polish #7) v2 now has an explicit rule mirroring v1's rule 27: government bodies, institutional reports, and similar standalone documents get italicized titles and book-form treatment, NOT quoted-article treatment. Includes the Anti-Communist Senate of California Tenth Report and a hypothetical GAO example. Implicit handling worked in chunk 2c, but explicit rule provides regression protection. v2 only. Tests: +14 net new (across test_note_formatter.py, test_formatter.py, test_runtime_validator.py). Total suite 150/150 (was 136 before this chunk). All v1 tests still passing. Path B status: formatter.py and runtime_validator.py edits cross the v1/v2 boundary, but only by user-explicit approval per the relevant fix discussions. The shared validator was always shared; the v1 formatter rule 10 mirror is a small additive edit that doesn't affect the v1 prompt's existing behavior on its existing inputs. |
||
|
|
45f51c6341 |
v2 chunk 2c: numbered-note format support + real-draft test
Extend the v2 parser and CLI to handle [N] text footnote definitions in addition to pandoc-style [^marker]: text. The numbered format is what docx-to-text conversion of footnoted Word documents produces. The user's "Anti-Communist Formations of LIS" draft uses this format for all 107 of its notes; without this support, v2 was structurally incapable of running on real-world docx-derived inputs. src/cmos/parser.py: - Add NoteDefinition.original_prefix field so reassembly can round-trip the source's marker syntax (pandoc input → pandoc output, numbered input → numbered output) without consumers needing to know which format was matched. - Update find_notes() to populate original_prefix as "[^N]: ". - Add find_numbered_notes() targeting "[N] text" definitions. Marker must be all digits (rejects [Smith 2020], [foo], etc.); caret prefix is rejected (rejects pandoc-style cleanly). src/cmos/cli.py: - reformat_notes now auto-detects format: tries find_notes first, falls back to find_numbered_notes if no pandoc definitions found. Uses definition.original_prefix for reassembly so both formats round-trip correctly. tests/test_parser.py: - 11 new tests for find_numbered_notes covering: single/multiple definitions, multi-digit markers, line number recording, trailing whitespace stripping, ignoring pandoc/non-numeric markers, original prefix recording, and the actual Anti-Communist draft format. - 1 new test for find_notes original_prefix population. tests/test_cli.py: - 2 new tests for reformat_notes auto-detect: numbered input round- trips as numbered output, pandoc input still round-trips as pandoc. Total suite: 136/136 (was 123, +13 net new). v1 untouched, 96/96 v1 tests still passing. Real-draft validation: ran cmos format-notes against the 107-note Anti-Communist Formations of LIS draft (108 calls in parallel via the existing concurrency=8 thread pool, completed cleanly). Output saved to /tmp (not committed). All 107 notes preserved through the pipeline; ~30-40 first-occurrence full notes produced clean CMOS 18 note form; ~30 shortened-form refs correctly left unchanged; 2 real bugs surfaced for the next iteration (empty input → conver- sational reply, Ibid → empty string), plus several lower-priority issues (substantive note truncation, retry waste on shortened forms, lossy month dropping). Not addressed in this chunk per the "collect signal, don't fix" plan. |
||
|
|
87cb70b4fb |
v2 chunk 2b: cli format-notes subcommand and reformat_notes()
Wire the v2 note formatter into the CLI so it can be invoked on real markdown drafts. The format-notes subcommand mirrors v1's format subcommand: parser → formatter → reassemble, with concurrent API calls. src/cmos/cli.py: - Import find_notes and format_note_entry alongside the existing v1 imports. - Add reformat_notes(text, formatter, concurrency) that finds pandoc-style markdown footnote definitions via find_notes, formats each definition's text via the v2 note formatter (or an injected fake), and substitutes the formatted text back into the original line position. Non-definition lines preserved byte-for-byte. Returns text unchanged when no definitions found. - Register the format-notes argparse subparser with the same --concurrency flag as v1's format. - Dispatch args.command == "format-notes" to reformat_notes. - Module docstring updated to document both subcommands. tests/test_cli.py: - 7 new tests for reformat_notes covering: in-place substitution, order preservation under concurrency, prose preservation, reference markers staying verbatim, no-op on empty input, trailing newline preservation. - Extended test_python_dash_m_invocation_actually_runs_main to also assert "format-notes" appears in --help, catching accidental subcommand removal. Path B integrity: formatter.py, note_formatter.py, parser.py, linter.py, harness/score.py all unchanged. No LINTER_VERSION bump. 96/96 v1 tests still passing. Total suite: 123/123. Real-API end-to-end smoke test on a temp draft with 2 footnote definitions: both reformatted byte-for-byte, prose and headings preserved, ## Conclusion section after the notes preserved. |
||
|
|
d061f78b7f |
v2 chunk 2a iter 1: rules 6+18 — ed. invariant, expand publishers
Two prompt edits to SYSTEM_PROMPT_NOTES driven by chunk 1 smoke test
gaps on chapter_first_doyle.toml.
Rule 6 (chapter form) expanded with an explicit invariance statement:
"ed." is the canonical abbreviation regardless of editor count — do
NOT pluralize to "eds." for multiple editors. CMOS NB treats it as
an invariant abbreviation, not a number-agreeing word.
New rule 18 (publisher expansion) mirrors v1 formatter.py rule 14:
publisher names must be in full canonical form, with note-form-
specific examples ("U of Chicago Press" → "University of Chicago
Press"). Includes the same MIT Press / ALA Editions / MLA carve-out
for publishers whose canonical self-presentation legitimately uses
initials.
Two new prompt-content unit tests added to tests/test_note_formatter.py
following the v1 test_formatter.py discipline. TDD cycle: red-green-
verified end-to-end.
Real-API smoke test, 3 v2 exemplars × 2 runs each: 6/6 byte-perfect
matches (was 4/6 in chunk 1; chapter_first_doyle went 0/2 → 2/2,
book and journal still 2/2). v1 untouched, 96/96 v1 tests still
passing. Total suite: 116/116.
|
||
|
|
28ca3ac928 |
v2 chunk 1: scaffold note formatter (Path B parallel artifact)
Begin v2 (in-text citation note form) as a parallel artifact to the
v1 bibliography formatter, per the Path B architectural decision:
no shared mutable state, no shared prompt content, no v1 changes.
New artifacts:
- src/cmos/parser.py: add find_notes() / NoteDefinition /
NotesParseResult as siblings to split_bibliography(). Targets
pandoc-style markdown footnote definitions [^marker]: text.
Single-line only; multi-line continuation deferred.
- src/cmos/note_formatter.py: new module mirroring formatter.py.
17-rule SYSTEM_PROMPT_NOTES for CMOS 18 first-occurrence note
form. Reuses cmos.runtime_validator.validate() unchanged — its
structural checks all apply to note form too. Caller injection,
retry loop, and model selection mirror v1.
- tests/test_parser.py: 7 new tests for find_notes().
- tests/test_note_formatter.py: 11 new tests mirroring v1 test
discipline (no API calls, fake-caller injection, retry semantics,
prompt smoke checks).
- exemplars/notes/: new subdirectory with 3 hand-synthesized
first-occurrence note exemplars (book, journal article, chapter
in edited book). Uses expected_note as the field name (not v1's
expected_bibliography). harness/score.py:load_exemplars uses a
non-recursive glob, so the v1 canary loader does not see these —
Path B isolation is automatic.
v1 untouched. v1 canary still loads exactly 17 exemplars. Full
test suite: 96 v1 + 18 new v2 = 114 passing.
Smoke-tested all 3 v2 exemplars against real GPT-5 (not fakes),
2 runs each. book_first_yu and journal_first_kwon: 4/4 byte-perfect
on the first try. chapter_first_doyle: 0/2, surfacing two known
prompt gaps for the next iteration:
1. Publisher abbreviation not expanded ("U of Chicago Press"
preserved instead of "University of Chicago Press"). v1's
formatter.py rule 14 is missing from the v2 prompt.
2. "ed." pluralized to "eds." for multiple editors. CMOS NB
uses "ed." invariantly regardless of editor count.
Both gaps are addressable prompt edits, not architectural problems —
exactly the kind of finding the dev loop is designed to surface.
Out of scope (deferred to chunk 2 and later): CLI extension,
shortened-form generation, document reassembly, harness scoring
loop integration, v2 linter rules, real-draft testing.
|
||
|
|
2a8c72ae22 |
gitignore: protect rough_drafts/*.txt from accidental commits
The user's drafts in rough_drafts/ are real in-progress academic work. Project discipline (per memory) is to keep them untracked, but a single accidental `git add -A` could leak unpublished scholarship into git history. Adding the pattern to .gitignore makes the protection structural rather than discipline-based. Scoped to .txt only (the docx->txt converted form the project actually consumes); does not affect sample.md, .gitkeep, or any other extensions that might land in rough_drafts/ later. |
||
|
|
22404fbd6b |
cli: add __main__ guard so python -m cmos.cli actually runs
Without `if __name__ == "__main__": sys.exit(main())` at the bottom of cli.py, `python -m cmos.cli format <path>` imports the module but never invokes main(), so the process silently exits 0 with empty stdout — indistinguishable from a successful run that produced no output. Discovered during real-draft testing on 2026-04-11. Adds a regression test that subprocesses the CLI with --help and asserts on stdout content. argparse --help exits 0 in both broken and fixed states; stdout content is the only discriminator. Both invocation paths now work: - uv run cmos format <path> (pyproject script entry) - uv run python -m cmos.cli format <path> (module invocation) |
||
|
|
1d970e84f9 |
exemplars: canary for CMOS 8.159 preposition rule (Huttunen, Mehra)
Two new real-draft-derived journal exemplars locking down prompt
rule 29 from the preceding commit. Follows the Hasanah precedent
for promoting real-draft failures into the canary corpus, per the
"defense in depth" feedback in project memory.
- journal_preposition_huttunen.toml exercises "among" mid-title,
plus four other CMOS 8.159 lowercased words ("on", "in", "a",
and implicitly "among"). Real article: Huttunen and Kortelainen,
JASIST 72, no. 7 (2021). The messy_input uses the ASCII hyphen
in "Meaning-Making" for clarity — the source draft actually had
a U+2010 non-breaking hyphen from docx->txt conversion, but
mixing Unicode hyphen normalization into this exemplar would
have conflated two independent failure modes. The U+2010 issue
is noted in the exemplar comment for a separate iteration.
- journal_preposition_mehra.toml exercises "beyond" mid-title,
plus mid-title exclamation-point preservation (rule 25) and
leading-"The" drop from "The Library Quarterly" (rule 17).
Real article: Mehra, Library Quarterly 91, no. 2 (2021).
Both are canary-enabled and will block any future prompt iteration
that regresses on these patterns via the exact-match canary axis.
The canonical dicts additionally lock them down via the field-
level diff, per the upstream-refinement feedback in project
memory.
Canary corpus size: 15 -> 17.
|
||
|
|
0238f5e5d7 |
iter 8: prompt rule 29 for CMOS 8.159 headline-case prepositions
Real-draft testing on the HML bibliography surfaced inconsistent preposition capitalization in headline-style titles: the formatter was sometimes lowercasing "among", "beyond", "within", "into", etc. and sometimes capitalizing them, violating CMOS 8.159. The behavior was nondeterministic across similar entries — smoking-gun evidence of GPT-5 judgment drift on a rule the existing prompt only implicitly covered via rule 6's "headline-style capitalization". Rule 29 makes the carve-out explicit and instructs the model to apply CMOS 8.159 independently of the source's casing, with: - a substantial but illustrative preposition list including "as" (which CMOS 8.159 calls out as always lowercased) - an explicit scope restriction to PREPOSITIONS, with subordinating conjunctions (If, That, Because, Although, Unless, etc.) capitalized - a CMOS 8.161 carve-out for hyphenated compounds: the first element is always capitalized (so "In-School" stays, not "in-School") - a first/last-word exception covering "last word of the main title immediately before a subtitle colon", which the model was treating as mid-title Verified on the 134-entry HML draft: ~17 entries now produce more CMOS-correct forms, and two consecutive re-runs show no new rule- 29-related regressions. Observed GPT-5 drift on orthogonal axes (multi-author inversion, inner-quote style, periodical italicization) washed out across re-runs, consistent with the ~5% single-run variance documented in project memory. Rule count: 28 -> 29. No linter version bump required. |