Extend the v2 parser and CLI to handle [N] text footnote definitions
in addition to pandoc-style [^marker]: text. The numbered format is
what docx-to-text conversion of footnoted Word documents produces.
The user's "Anti-Communist Formations of LIS" draft uses this format
for all 107 of its notes; without this support, v2 was structurally
incapable of running on real-world docx-derived inputs.
src/cmos/parser.py:
- Add NoteDefinition.original_prefix field so reassembly can
round-trip the source's marker syntax (pandoc input → pandoc
output, numbered input → numbered output) without consumers
needing to know which format was matched.
- Update find_notes() to populate original_prefix as "[^N]: ".
- Add find_numbered_notes() targeting "[N] text" definitions.
Marker must be all digits (rejects [Smith 2020], [foo], etc.);
caret prefix is rejected (rejects pandoc-style cleanly).
src/cmos/cli.py:
- reformat_notes now auto-detects format: tries find_notes first,
falls back to find_numbered_notes if no pandoc definitions found.
Uses definition.original_prefix for reassembly so both formats
round-trip correctly.
tests/test_parser.py:
- 11 new tests for find_numbered_notes covering: single/multiple
definitions, multi-digit markers, line number recording, trailing
whitespace stripping, ignoring pandoc/non-numeric markers, original
prefix recording, and the actual Anti-Communist draft format.
- 1 new test for find_notes original_prefix population.
tests/test_cli.py:
- 2 new tests for reformat_notes auto-detect: numbered input round-
trips as numbered output, pandoc input still round-trips as pandoc.
Total suite: 136/136 (was 123, +13 net new). v1 untouched, 96/96
v1 tests still passing.
Real-draft validation: ran cmos format-notes against the 107-note
Anti-Communist Formations of LIS draft (108 calls in parallel via
the existing concurrency=8 thread pool, completed cleanly). Output
saved to /tmp (not committed). All 107 notes preserved through the
pipeline; ~30-40 first-occurrence full notes produced clean CMOS 18
note form; ~30 shortened-form refs correctly left unchanged;
2 real bugs surfaced for the next iteration (empty input → conver-
sational reply, Ibid → empty string), plus several lower-priority
issues (substantive note truncation, retry waste on shortened
forms, lossy month dropping). Not addressed in this chunk per the
"collect signal, don't fix" plan.