Scaffold CMOS 18 reformatter: harness, linter, parser, formatter, CLI

Initial phase-1 baseline of the karpathy/autoresearch-style loop.
The formatter module is the inner-loop artifact; parser and linter
are infra. The linter carries a LINTER_VERSION hash (v0.2.0) that
will force a re-baseline on any rule change.

Components:
- harness/diff.py: case-sensitive field-level substring diff
- harness/score.py: three-axis scoring (field, linter, canary exact)
- src/cmos/linter.py: 9 CMOS 18 structural rules, each Purdue/CMOS cited
- src/cmos/parser.py: locate ## Bibliography section, split entries
- src/cmos/formatter.py: prompt + OpenAI call with caller injection
- src/cmos/cli.py: cmos format path/to/draft.md
- scripts/run_loop.py: loop runner with --fake mode for no-API runs
- exemplars/: 3 canary seed exemplars (book, journal w/DOI, web),
  sourced from chicagomanualofstyle.org quick guide

Tests: 48 passing. Fake-mode baseline scalar = 0.000 on the 3 seed
exemplars (identity caller fails the linter on every rule). This is
the floor the real GPT-5 formatter needs to improve from.
This commit is contained in:
cmos dev
2026-04-10 20:48:33 -04:00
commit 4cad38ef30
29 changed files with 2267 additions and 0 deletions
+76
View File
@@ -0,0 +1,76 @@
"""Tests for src/cmos/parser.py — extract the bibliography section from a draft.
v1 scope: locate a ``## Bibliography`` heading (case-insensitive) and return
the prose before it, the list of entries inside it, and the prose after it.
The section ends at the next ``##`` (level-2) heading or end of file.
Blank lines and bullet markers (``- `` or ``* ``) are stripped from entries.
"""
import pytest
from cmos.parser import NoBibliographyError, split_bibliography
def test_three_entries_under_bibliography_heading():
draft = """\
Some prose here.
More prose.
## Bibliography
yu, charles. interior chinatown. New York: Pantheon Books, 2020.
Kwon, Hyeyoung. "inclusion work." American Journal of Sociology 127, no. 6 (2022): 1818-1859.
google, "privacy policy," privacy & terms, nov 15 2023, policies.google.com/privacy
"""
result = split_bibliography(draft)
assert len(result.entries) == 3
assert result.entries[0].startswith("yu, charles")
assert result.entries[1].startswith("Kwon, Hyeyoung")
assert result.entries[2].startswith("google")
assert "Some prose here." in result.before
assert result.after == ""
def test_section_ends_at_next_level_two_heading():
draft = """\
## Bibliography
First entry.
Second entry.
## Appendix
Appendix content here.
"""
result = split_bibliography(draft)
assert result.entries == ["First entry.", "Second entry."]
assert "## Appendix" in result.after
assert "Appendix content here." in result.after
def test_bullet_prefixes_stripped():
draft = """\
## Bibliography
- First entry.
- Second entry.
* Third entry.
"""
result = split_bibliography(draft)
assert result.entries == ["First entry.", "Second entry.", "Third entry."]
def test_heading_match_is_case_insensitive():
draft = """\
## bibliography
Only entry.
"""
result = split_bibliography(draft)
assert result.entries == ["Only entry."]
def test_no_bibliography_raises():
with pytest.raises(NoBibliographyError):
split_bibliography("## Introduction\n\nJust prose.\n")