Files
cmos/tests/test_runtime_validator.py
T
cmos dev 8b14dbc9ca formatter: linter-guided retry to suppress GPT-5 reasoning drift
GPT-5 is a reasoning model and is not bit-deterministic even at the API
level (no temperature override allowed). Variance testing in iter 7
showed ~1 in 6 loop runs hit a nondeterministic regression: in one run
the magazine_mead exemplar lost its italic markers around "New Yorker"
even though the exact same input had produced a clean output 5 times
prior. The scoring linter caught it (rule_magazine_name_italicized) but
the formatter still emitted the bad output to the user.

This commit adds a runtime guardrail:

- New module src/cmos/runtime_validator.py — type-independent
  structural sanity checks. Operates on a single candidate string
  with no Exemplar context, because at runtime we don't know the
  source type. Checks: ends with period, no Ibid., has at least one
  italic span (almost every CMOS bibliography entry italicizes
  something), balanced * and " markers. Deliberately weaker than the
  scoring linter; it's a fast guardrail, not a full validator.

- format_bibliography_entry now retries up to DEFAULT_MAX_RETRIES (2)
  times when the validator rejects a candidate. Independent re-calls
  are usually enough because the failures are stochastic. If every
  attempt fails, the LAST attempt is returned (no exception) — the
  caller still gets something usable, and the failure surfaces
  through the scoring linter or human review. The retry path costs
  zero on the common case (1 call per entry); ~1-2% extra calls on
  noisy drafts.

Empirical: 3 consecutive loop runs after this change are scalar 1.000
canary 1.000 (vs 5/6 clean in the variance test before). Sample is too
small to claim full suppression but the signal is positive.

Also adds a new exemplar journal_multiauthor_secondary_first_last.toml
captured from the user's HML draft (Hasanah et al., IJIDI 2024). It
exercises the case where a multi-author entry has the first author
inverted and the rest in First Last form — which the iter 7 HML run
got wrong on one entry. Variance testing showed the exemplar passes
6/6 in isolation, so the original HML failure was nondeterminism, not
a missing rule. Keeping the exemplar regardless: it adds canary
coverage of a real-world multi-author pattern, no-DOI / JSTOR-URL
variant, and the year-suffix author-date holdover stripping.

Tests: 85 → 95 (8 new for runtime_validator + 2 new for formatter
retry). All passing.
2026-04-11 01:24:25 -04:00

71 lines
2.7 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Tests for src/cmos/runtime_validator.py.
Runtime validator runs at format time on a single candidate string with no
exemplar context. It checks type-independent structural sanity rules so the
formatter can self-retry when GPT-5's reasoning drift produces a malformed
entry. It is deliberately weaker than the scoring linter (which has access
to canonical fields and source type) but it catches the most common
nondeterministic failures: missing terminal period, stray Ibid., dropped
italic markers, unbalanced quotes/asterisks.
"""
from cmos.runtime_validator import validate
def test_clean_book_entry_passes():
candidate = "Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020."
result = validate(candidate)
assert result.passed is True
assert result.failures == []
def test_clean_journal_entry_passes():
candidate = (
'Kwon, Hyeyoung. "Inclusion Work: Children of Immigrants Claiming Membership in Everyday Life." '
"*American Journal of Sociology* 127, no. 6 (2022): 181859. "
"https://doi.org/10.1086/720277."
)
assert validate(candidate).passed
def test_missing_terminal_period_fails():
result = validate("Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020")
assert result.passed is False
assert any("period" in f.lower() for f in result.failures)
def test_ibid_fails():
result = validate("Ibid., 42.")
assert result.passed is False
assert any("ibid" in f.lower() for f in result.failures)
def test_missing_italics_anywhere_fails():
# The "New Yorker" magazine regression we saw in iter 7 variance: italic
# markers around the magazine name disappeared. This rule should catch
# that — almost every CMOS bibliography entry italicizes something.
candidate = 'Mead, Rebecca. "Terms of Aggrievement." New Yorker, December 18, 2023.'
result = validate(candidate)
assert result.passed is False
assert any("italic" in f.lower() for f in result.failures)
def test_unbalanced_italic_markers_fails():
candidate = "Yu, Charles. *Interior Chinatown. Pantheon Books, 2020."
result = validate(candidate)
assert result.passed is False
assert any("italic" in f.lower() or "*" in f for f in result.failures)
def test_unbalanced_double_quotes_fails():
candidate = 'Mead, Rebecca. "Terms of Aggrievement. *New Yorker*, December 18, 2023.'
result = validate(candidate)
assert result.passed is False
assert any("quote" in f.lower() for f in result.failures)
def test_passed_property_is_negation_of_failures():
# Sanity: passed should be True iff there are no failures.
result = validate("Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020.")
assert result.passed == (result.failures == [])