Files
cmos/tests/test_formatter.py
T
cmos dev 8b14dbc9ca formatter: linter-guided retry to suppress GPT-5 reasoning drift
GPT-5 is a reasoning model and is not bit-deterministic even at the API
level (no temperature override allowed). Variance testing in iter 7
showed ~1 in 6 loop runs hit a nondeterministic regression: in one run
the magazine_mead exemplar lost its italic markers around "New Yorker"
even though the exact same input had produced a clean output 5 times
prior. The scoring linter caught it (rule_magazine_name_italicized) but
the formatter still emitted the bad output to the user.

This commit adds a runtime guardrail:

- New module src/cmos/runtime_validator.py — type-independent
  structural sanity checks. Operates on a single candidate string
  with no Exemplar context, because at runtime we don't know the
  source type. Checks: ends with period, no Ibid., has at least one
  italic span (almost every CMOS bibliography entry italicizes
  something), balanced * and " markers. Deliberately weaker than the
  scoring linter; it's a fast guardrail, not a full validator.

- format_bibliography_entry now retries up to DEFAULT_MAX_RETRIES (2)
  times when the validator rejects a candidate. Independent re-calls
  are usually enough because the failures are stochastic. If every
  attempt fails, the LAST attempt is returned (no exception) — the
  caller still gets something usable, and the failure surfaces
  through the scoring linter or human review. The retry path costs
  zero on the common case (1 call per entry); ~1-2% extra calls on
  noisy drafts.

Empirical: 3 consecutive loop runs after this change are scalar 1.000
canary 1.000 (vs 5/6 clean in the variance test before). Sample is too
small to claim full suppression but the signal is positive.

Also adds a new exemplar journal_multiauthor_secondary_first_last.toml
captured from the user's HML draft (Hasanah et al., IJIDI 2024). It
exercises the case where a multi-author entry has the first author
inverted and the rest in First Last form — which the iter 7 HML run
got wrong on one entry. Variance testing showed the exemplar passes
6/6 in isolation, so the original HML failure was nondeterminism, not
a missing rule. Keeping the exemplar regardless: it adds canary
coverage of a real-world multi-author pattern, no-DOI / JSTOR-URL
variant, and the year-suffix author-date holdover stripping.

Tests: 85 → 95 (8 new for runtime_validator + 2 new for formatter
retry). All passing.
2026-04-11 01:24:25 -04:00

121 lines
4.7 KiB
Python

"""Tests for src/cmos/formatter.py — the inner-loop iterable artifact.
Because `formatter.py` is edited every iteration of the dev-time autoresearch
loop, these tests intentionally test STABLE pieces: the prompt skeleton, the
caller-injection seam, and a smoke check that the system prompt mentions key
CMOS 18 rules. The actual LLM output shape is tested via the harness (running
all exemplars through `score`) rather than pinned here — exact string
matches on LLM output belong in the canary exact-match rate, not in unit
tests.
No real API calls in this file. Tests that hit the OpenAI API live in
`tests/test_formatter_integration.py` (not yet created) and are gated on an
OPENAI_API_KEY being set.
"""
from cmos.formatter import (
MODEL,
SYSTEM_PROMPT,
build_user_message,
format_bibliography_entry,
)
def test_default_model_is_gpt5():
# The user specified "GPT-5 / frontier reasoning" as the default. Override
# via the OPENAI_MODEL env var if gpt-5 is unavailable in your account.
assert MODEL == "gpt-5"
def test_system_prompt_mentions_cmos_18():
assert "CMOS" in SYSTEM_PROMPT or "Chicago Manual of Style" in SYSTEM_PROMPT
assert "18" in SYSTEM_PROMPT
def test_system_prompt_mentions_no_place_of_publication():
# CMOS 14.30 / 18th ed. change. The formatter MUST know this.
lower = SYSTEM_PROMPT.lower()
assert "place of publication" in lower
def test_system_prompt_mentions_doi_preference():
assert "doi" in SYSTEM_PROMPT.lower()
def test_system_prompt_mentions_italic_markers():
# Plan v1 uses Markdown `*Title*` for italics.
assert "*" in SYSTEM_PROMPT and "italic" in SYSTEM_PROMPT.lower()
def test_build_user_message_contains_the_messy_input():
msg = build_user_message("yu, charles. interior chinatown. 2020")
assert "yu, charles. interior chinatown. 2020" in msg
def test_formatter_uses_injected_caller():
"""The formatter accepts a caller shim so tests (and the harness) can
substitute a fake OpenAI call. This is how the whole pipeline can run
without an API key during tests."""
recorded: dict = {}
def fake_caller(system: str, user: str) -> str:
recorded["system"] = system
recorded["user"] = user
return "Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020."
output = format_bibliography_entry(
"yu, charles. interior chinatown. New York: Pantheon Books, 2020.",
caller=fake_caller,
)
assert output == "Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020."
assert recorded["system"] == SYSTEM_PROMPT
assert "yu, charles" in recorded["user"]
def test_formatter_strips_whitespace_from_caller_output():
# Language-model output often has leading/trailing whitespace or
# surrounding code fences. v0 only strips whitespace; code-fence
# stripping can be added when an exemplar forces it.
def fake_caller(system: str, user: str) -> str:
return " Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020. \n"
output = format_bibliography_entry("anything", caller=fake_caller)
assert output == "Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020."
def test_formatter_retries_on_validator_failure():
"""If the runtime validator rejects the first attempt (e.g., italics
missing), the formatter should retry and return a passing later attempt.
Mirrors the iter 7 magazine_mead variance: GPT-5 sometimes drops italics
around a magazine name, and a re-call typically succeeds."""
attempts = []
def flaky_caller(system: str, user: str) -> str:
attempts.append(len(attempts) + 1)
if len(attempts) == 1:
# First attempt: italics dropped (validator should fail).
return 'Mead, Rebecca. "Terms of Aggrievement." New Yorker, December 18, 2023.'
# Retry: clean.
return 'Mead, Rebecca. "Terms of Aggrievement." *New Yorker*, December 18, 2023.'
output = format_bibliography_entry("messy mead", caller=flaky_caller)
assert "*New Yorker*" in output
assert len(attempts) == 2
def test_formatter_returns_last_attempt_after_max_retries():
"""If every attempt fails validation, return the last attempt and don't
loop forever. We surface the failure later via scoring/canary, not by
raising at runtime — the user should still get *something* back."""
attempts = []
def always_bad(system: str, user: str) -> str:
attempts.append(1)
return 'Mead, Rebecca. "Terms of Aggrievement." New Yorker, December 18, 2023.'
output = format_bibliography_entry("messy", caller=always_bad, max_retries=2)
# 1 initial + 2 retries = 3 calls total.
assert len(attempts) == 3
assert "*" not in output # the bad output is what we got