2a8c72ae2298f52b41215574482544df0a32723d
10
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8b14dbc9ca |
formatter: linter-guided retry to suppress GPT-5 reasoning drift
GPT-5 is a reasoning model and is not bit-deterministic even at the API level (no temperature override allowed). Variance testing in iter 7 showed ~1 in 6 loop runs hit a nondeterministic regression: in one run the magazine_mead exemplar lost its italic markers around "New Yorker" even though the exact same input had produced a clean output 5 times prior. The scoring linter caught it (rule_magazine_name_italicized) but the formatter still emitted the bad output to the user. This commit adds a runtime guardrail: - New module src/cmos/runtime_validator.py — type-independent structural sanity checks. Operates on a single candidate string with no Exemplar context, because at runtime we don't know the source type. Checks: ends with period, no Ibid., has at least one italic span (almost every CMOS bibliography entry italicizes something), balanced * and " markers. Deliberately weaker than the scoring linter; it's a fast guardrail, not a full validator. - format_bibliography_entry now retries up to DEFAULT_MAX_RETRIES (2) times when the validator rejects a candidate. Independent re-calls are usually enough because the failures are stochastic. If every attempt fails, the LAST attempt is returned (no exception) — the caller still gets something usable, and the failure surfaces through the scoring linter or human review. The retry path costs zero on the common case (1 call per entry); ~1-2% extra calls on noisy drafts. Empirical: 3 consecutive loop runs after this change are scalar 1.000 canary 1.000 (vs 5/6 clean in the variance test before). Sample is too small to claim full suppression but the signal is positive. Also adds a new exemplar journal_multiauthor_secondary_first_last.toml captured from the user's HML draft (Hasanah et al., IJIDI 2024). It exercises the case where a multi-author entry has the first author inverted and the rest in First Last form — which the iter 7 HML run got wrong on one entry. Variance testing showed the exemplar passes 6/6 in isolation, so the original HML failure was nondeterminism, not a missing rule. Keeping the exemplar regardless: it adds canary coverage of a real-world multi-author pattern, no-DOI / JSTOR-URL variant, and the year-suffix author-date holdover stripping. Tests: 85 → 95 (8 new for runtime_validator + 2 new for formatter retry). All passing. |
||
|
|
a21d2614d3 |
iter 7: prompt rules 27-29 driven by real-draft testing
First exercise of the tool against unfamiliar real-world drafts (user's
in-progress academic papers, not derived from CMOS quick guide
examples). Found three systemic patterns affecting ~14% of entries
across 134-entry and 13-entry drafts; this commit addresses all three.
Prompt additions:
27. GOVERNMENT DOCUMENTS / INSTITUTIONAL REPORTS — when the author is
a government body (GAO, US Census, WHO, Pew, etc.) and the source
is a standalone report, italicize the title like a book; do NOT
wrap it in quotes. CMOS 14.272.
28. PRESERVE DELIBERATE LOWERCASING of proper nouns. Examples cited in
the prompt: the journal "portal: Libraries and the Academy"
(deliberately lowercase "p"), authors bell hooks / danah boyd /
e e cummings, brand names iPhone / eBay. Headline-case
normalization must NOT touch these.
Rule 17 also extended (was newspaper/magazine only):
17. PERIODICAL NAMES — drop leading "The" from ANY periodical:
newspapers, magazines, AND scholarly journals. Write "Journal of
Academic Librarianship", not "The Journal of Academic
Librarianship"; "Library Quarterly", not "The Library Quarterly";
"American Archivist", not "The American Archivist". CMOS 14.191.
Rule explicitly excepts BOOK titles and report titles, which keep
their leading "The" (e.g., *The Library's Guide to Sexual and
Reproductive Health Information*).
Empirical impact, Reading_Disrepair (13 entries):
before: 3 errors (2 leading-The, 1 GAO format, 1 portal capitalized)
after: 0 errors
Empirical impact, HML Submission (134 entries):
before: ~21 errors (18 leading-The, 0 GAO in this draft, 3 portal)
after: ~1 error (a multi-author 2nd-author inversion bug newly
surfaced — separate fix)
Also adds scripts/analyze_draft_output.py — a triage helper that
diff-walks input and output, surfacing common failure patterns
(leading-The, doi: prefix, bare-hyphen ranges, gov-report quoted, and
intentional lowercasing stripped). False positives are tolerated since
it's a human-review tool, not a validator.
Exemplar corpus is unchanged. Re-baseline against 14 canary exemplars
under the new rules: scalar 1.000, canary 1.000 (no regressions).
User's real drafts under rough_drafts/ remain untracked — they are
in-progress academic work and don't belong in git history.
|
||
|
|
2bd1bdf0d4 |
cli: parallelize formatter calls with thread pool
reformat_draft now uses concurrent.futures.ThreadPoolExecutor with a default of 8 workers. OpenAI SDK calls are synchronous but network- bound, so threads release the GIL during I/O and give real speedup. ThreadPoolExecutor.map preserves input order regardless of completion order, so output is deterministic. Empirical: HML draft (134 entries) went from ~25 min serial to 1m55s with concurrency=12. ~12x speedup; further increases hit OpenAI rate limits. CLI gains a --concurrency flag (default 8) for tuning per draft size / rate limit headroom. concurrency=1 forces serial execution for debugging. New test asserts that order is preserved when concurrent calls finish out of input order (uses a sleep-by-index fake formatter). |
||
|
|
1b11a03a0b |
judge: LLM-as-judge triage for canary failures (advisory only)
Add harness/judge.py and scripts/triage.py. The judge reads the latest loop run from logs/runs.jsonl, finds canary failures (exact_match=False but fields/linter passed), and asks GPT-5 to classify each as "regression", "variant", or "unclear" with CMOS 18 section citations. CRITICAL: verdicts are ADVISORY ONLY. They are written to logs/triage.jsonl and never feed back into the loop scalar. Using the judge as ground truth would let the formatter-LLM optimize against a judge-LLM from the same model family, inviting shared-bias drift. Design: - harness/judge.py: Verdict dataclass, build_judge_prompt, parse_verdict (handles code fences and normalizes unknown labels to "unclear"), judge() with caller-injection seam matching cmos.formatter. Uses the same _is_reasoning_model branching to skip temperature for gpt-5/o*. Judge model is separately overridable via CMOS_JUDGE_MODEL env var (defaults to OPENAI_MODEL, which defaults to gpt-5). - scripts/triage.py: CLI that walks runs.jsonl, locates a target run (default: latest), filters canary failures, calls judge on each, appends a verdict record to logs/triage.jsonl. --dry-run available for offline testing. Exits 0 with a note when there are no failures. Tests: 6 new unit tests covering prompt building, JSON parsing (including code-fence stripping and unknown-label normalization), and caller injection. No real API calls in the test suite. Validated on iter 3's run (canary 0.286, 10 failures): - 8 correctly flagged as regressions, each with a cited CMOS section (14.72, 14.76, 14.128, 14.190, 14.206, 14.212, 14.267, ...). - 2 flagged as variants: "Kindle" vs "Kindle edition" (CMOS 14.159– 14.161 allows flexibility) and "The New Yorker" vs "New Yorker" (CMOS 14.191 — leading "The" is optional). These surface that the formatter's current rules 17 and 18 are stricter than CMOS strictly requires; documenting here but not acting on yet. |
||
|
|
fe5e071066 |
linter: expand to 21 rules covering all 8 source types (v0.3.0)
Close the defense-in-depth gap flagged after iter 6: chapter, magazine,
newspaper, social_media, podcast, and video exemplars had ZERO applicable
linter rules, so their 100% linter pass rate was trivially true. Every
exemplar now has 3-6 structural rules firing.
Bump LINTER_VERSION to v0.3.0. Per harness discipline, this invalidates
prior run logs' scores for cross-version comparison (they stay in the
log as history). The re-baseline run still scores scalar 1.000, canary
1.000 on all 14 exemplars — formatter was already producing output that
matches the new structural conventions, so tightening the linter didn't
surface any regressions.
New rules (12):
- chapter_book_title_italicized
- chapter_in_book_marker
- magazine_name_italicized
- newspaper_name_italicized
- periodical_comma_before_date (magazine + newspaper; not journal)
- social_media_post_quoted
- social_media_platform_comma_date
- podcast_series_italicized
- podcast_episode_quoted
- podcast_format_label
- video_title_quoted (in quotes, NOT italicized)
- video_format_label
Extended:
- article_title_quoted now covers journal, magazine, newspaper,
and accepts ?/! as terminal punctuation
Each rule has a positive + negative unit test with the CMOS/Purdue source
cited in the docstring. Tests: 48 → 78 (30 new).
Per-exemplar applicable-rule counts after this change:
book 4 journal 5-6 web 3
chapter 4 magazine 5 social 4
podcast 5 video 4 newspaper 5
|
||
|
|
ced9c73474 |
iter 3-6: prompt rules 13-26 to handle expanded corpus
Four-iteration sweep over the 14-exemplar corpus, driven by loop
failures at each step. Scalar progression:
iter 3 (v0 prompt, 14 exemplars) → 0.842
iter 4 (+rules 13-23) → 0.951
iter 5 (+rules 24-25) → 0.988
iter 6 (+rule 26, rule 12 refined) → 1.000 canary 1.000
New SYSTEM_PROMPT rules:
13. Drop chapter page range from bibliography (CMOS 18 change from 17).
14. Use full canonical publisher names — no "U of Chicago Press" etc.
15. Do not comma-invert non-Western family-first names (Liu Xinwu,
Murakami Haruki).
16. Author threshold: up to 6 listed; 7+ → first 3 + "et al." with full
given names when source provides them.
17. Expand abbreviated newspaper names (NYT → New York Times); drop
leading "The" from newspaper/magazine names; use comma between
italicized name and date.
18. E-book format: "Kindle." not "Kindle edition."
19. Database indicator: just the name, not "Accessed via X".
20. Social media: preserve original casing of post content; comma
between Platform and Date.
21. Podcast: series italicized, episode in quotes, preserve
Season/episode/duration.
22. Video / TED Talk: preserve venue, date, Video label, duration, URL.
23. URLs preserved verbatim including "www." subdomain when present.
24. CMOS 9.61 inclusive-number elision for page ranges: 1818–59 not
1818–1859; 101–8; 1100–1113; 1496–1504.
25. Preserve terminal punctuation (? !) inside titles.
26. Edition indicators abbreviated: "2nd ed." not "Second edition".
Also scoped rule 12 (date qualifier preservation) to web-page sources
only — podcasts and videos use bare dates, so "Released September 13,
2022" should emit as just "September 13, 2022".
Also refined three messy inputs (podcast, social media, video) to
include "www." in the URL, and the snyder messy input to explicitly
list 7 authors so the et-al threshold is unambiguous to the formatter.
|
||
|
|
abb75888c5 |
Expand exemplar corpus to 14 across 8 CMOS 18 source types
Add 11 new exemplars pulled verbatim from the CMOS official quick guide (chicagomanualofstyle.org/tools_citationguide/citation-guide-1.html): - book_two_authors_binder: Binder & Kidder (two-author book) - book_chapter_doyle: Doyle in Marks & Parkin (chapter in edited volume, no page range per CMOS 18) - book_translated_liu: Liu Xinwu, trans. Tiang (translated book, non- Western name not comma-inverted) - book_edition_borel: Borel, 2nd ed. via EBSCOhost (edition + database) - book_ebook_roy: Roy, Kindle format - journal_many_authors_snyder: 7 authors in PLOS ONE, exercises the CMOS 18 "first 3 + et al." threshold and article-ID page format - magazine_mead: New Yorker (tests leading-"The" drop) - newspaper_blum: NYT with URL (tests abbreviation expansion) - social_media_cmos_facebook: Facebook post with case preservation - podcast_ober: Pushkin podcast with season/episode/duration - video_cowan_ted: TED Talk with venue and duration Every exemplar is marked canary=true — we want byte-exact regression detection on all of them. Source citations and CMOS rule justifications are in the TOML doc comments. All 14 canonical outputs lint clean under LINTER_VERSION v0.2.0. The harness will score them in the following commit. |
||
|
|
1621d0d502 |
iter 1: preserve date qualifiers in web entries
First real loop iteration found that GPT-5 silently drops date qualifiers
("Effective", "Published", "Accessed", etc.) when reformatting web
sources. The field diff was satisfied because the canonical date field
did not include the qualifier, but the canary exact-match axis caught
the regression.
Two fixes, per the canary-upstream policy:
1. Tighten the web-page exemplar canonical: date field now includes the
"Effective" qualifier so the field diff will catch future regressions
without relying on the canary.
2. Add SYSTEM_PROMPT rule 12 instructing the formatter to preserve
semantic date qualifiers from the input.
After these fixes: scalar 1.000, canary exact-match 1.000 on all three
seed exemplars.
|
||
|
|
ee0bcd107e |
formatter: skip temperature override for reasoning models
GPT-5 and the o-series reject temperature=0 with a 400 BadRequestError — only the default (1) is supported for reasoning models. Add an _is_reasoning_model helper and pass temperature only when the model name does not start with gpt-5/o1/o3/o4. Determinism on reasoning models is a property of the architecture, not a parameter. Discovered on the first real LLM run against the seed exemplars. |
||
|
|
4cad38ef30 |
Scaffold CMOS 18 reformatter: harness, linter, parser, formatter, CLI
Initial phase-1 baseline of the karpathy/autoresearch-style loop. The formatter module is the inner-loop artifact; parser and linter are infra. The linter carries a LINTER_VERSION hash (v0.2.0) that will force a re-baseline on any rule change. Components: - harness/diff.py: case-sensitive field-level substring diff - harness/score.py: three-axis scoring (field, linter, canary exact) - src/cmos/linter.py: 9 CMOS 18 structural rules, each Purdue/CMOS cited - src/cmos/parser.py: locate ## Bibliography section, split entries - src/cmos/formatter.py: prompt + OpenAI call with caller injection - src/cmos/cli.py: cmos format path/to/draft.md - scripts/run_loop.py: loop runner with --fake mode for no-API runs - exemplars/: 3 canary seed exemplars (book, journal w/DOI, web), sourced from chicagomanualofstyle.org quick guide Tests: 48 passing. Fake-mode baseline scalar = 0.000 on the 3 seed exemplars (identity caller fails the linter on every rule). This is the floor the real GPT-5 formatter needs to improve from. |