formatter: linter-guided retry to suppress GPT-5 reasoning drift

GPT-5 is a reasoning model and is not bit-deterministic even at the API
level (no temperature override allowed). Variance testing in iter 7
showed ~1 in 6 loop runs hit a nondeterministic regression: in one run
the magazine_mead exemplar lost its italic markers around "New Yorker"
even though the exact same input had produced a clean output 5 times
prior. The scoring linter caught it (rule_magazine_name_italicized) but
the formatter still emitted the bad output to the user.

This commit adds a runtime guardrail:

- New module src/cmos/runtime_validator.py — type-independent
  structural sanity checks. Operates on a single candidate string
  with no Exemplar context, because at runtime we don't know the
  source type. Checks: ends with period, no Ibid., has at least one
  italic span (almost every CMOS bibliography entry italicizes
  something), balanced * and " markers. Deliberately weaker than the
  scoring linter; it's a fast guardrail, not a full validator.

- format_bibliography_entry now retries up to DEFAULT_MAX_RETRIES (2)
  times when the validator rejects a candidate. Independent re-calls
  are usually enough because the failures are stochastic. If every
  attempt fails, the LAST attempt is returned (no exception) — the
  caller still gets something usable, and the failure surfaces
  through the scoring linter or human review. The retry path costs
  zero on the common case (1 call per entry); ~1-2% extra calls on
  noisy drafts.

Empirical: 3 consecutive loop runs after this change are scalar 1.000
canary 1.000 (vs 5/6 clean in the variance test before). Sample is too
small to claim full suppression but the signal is positive.

Also adds a new exemplar journal_multiauthor_secondary_first_last.toml
captured from the user's HML draft (Hasanah et al., IJIDI 2024). It
exercises the case where a multi-author entry has the first author
inverted and the rest in First Last form — which the iter 7 HML run
got wrong on one entry. Variance testing showed the exemplar passes
6/6 in isolation, so the original HML failure was nondeterminism, not
a missing rule. Keeping the exemplar regardless: it adds canary
coverage of a real-world multi-author pattern, no-DOI / JSTOR-URL
variant, and the year-suffix author-date holdover stripping.

Tests: 85 → 95 (8 new for runtime_validator + 2 new for formatter
retry). All passing.
This commit is contained in:
cmos dev
2026-04-11 01:24:25 -04:00
parent a21d2614d3
commit 8b14dbc9ca
5 changed files with 240 additions and 3 deletions
@@ -0,0 +1,33 @@
# Source: derived from a real-world entry in the user's HML draft (Hasanah
# et al., IJIDI 2024). This exemplar specifically covers the CMOS rule that
# in a multi-author bibliography entry, ONLY the first author is comma-
# inverted; subsequent authors are written in "First Last" form.
#
# This is a regression test for a bug found in iter 7: the formatter saw a
# messy input where the first author was already inverted and the secondary
# authors were in First-Last form, and "fixed" the perceived inconsistency by
# inverting everyone. Per CMOS 14.72 / 14.76, only the first author is
# inverted in a bibliography entry; secondary authors stay in given-name-
# first order.
#
# Also exercises:
# - Year-suffix author-date holdover that should be stripped: "(2024). "
# - Leading "The" in journal name to be dropped (rule 17)
# - Hyphen→en-dash for page range, no elision (under 100)
# - Source provides a JSTOR URL instead of a DOI (rare variant)
source = "user draft (HML Submission); article: Hasanah et al., IJIDI 8, no. 2 (2024)"
type = "journal"
tags = ["journal", "multi-author", "no-doi", "jstor-url"]
canary = true
messy_input = "Hasanah, Anis Karunia Uswatun, Fitri Mutia, and Norhuda Salleh. (2024). “Crossing the Barriers of Library Anxiety: A Quantitative Evaluation of Indonesian Undergraduate Students with Visual Disabilities Navigating Their Academic Library.” The International Journal of Information, Diversity, & Inclusion 8, no. 2 (2024): 29-51. https://www.jstor.org/stable/48786446."
expected_bibliography = "Hasanah, Anis Karunia Uswatun, Fitri Mutia, and Norhuda Salleh. \"Crossing the Barriers of Library Anxiety: A Quantitative Evaluation of Indonesian Undergraduate Students with Visual Disabilities Navigating Their Academic Library.\" *International Journal of Information, Diversity, & Inclusion* 8, no. 2 (2024): 2951. https://www.jstor.org/stable/48786446."
[canonical]
authors = "Hasanah, Anis Karunia Uswatun, Fitri Mutia, and Norhuda Salleh"
article_title = "Crossing the Barriers of Library Anxiety: A Quantitative Evaluation of Indonesian Undergraduate Students with Visual Disabilities Navigating Their Academic Library"
journal = "International Journal of Information, Diversity, & Inclusion"
volume = 8
issue = 2
year = 2024
pages = "2951"
url = "https://www.jstor.org/stable/48786446"
+34 -3
View File
@@ -22,9 +22,18 @@ from typing import Callable
from dotenv import load_dotenv from dotenv import load_dotenv
from cmos.runtime_validator import validate
# Default model. Override with OPENAI_MODEL=... in your .env. # Default model. Override with OPENAI_MODEL=... in your .env.
MODEL = os.environ.get("OPENAI_MODEL", "gpt-5") MODEL = os.environ.get("OPENAI_MODEL", "gpt-5")
# How many extra attempts after the first if the runtime validator rejects
# the candidate. Reasoning models like GPT-5 are nondeterministic; an
# independent re-call usually succeeds. 2 retries = 3 calls total per
# entry in the worst case, but only fires when validation fails so the
# common case still costs 1 call.
DEFAULT_MAX_RETRIES = 2
# SYSTEM_PROMPT: the formatter's sole "program". Iterated by the autoresearch # SYSTEM_PROMPT: the formatter's sole "program". Iterated by the autoresearch
# loop. Keep it explicit, versioned, and traceable to CMOS 18 sources. # loop. Keep it explicit, versioned, and traceable to CMOS 18 sources.
SYSTEM_PROMPT = """\ SYSTEM_PROMPT = """\
@@ -278,13 +287,35 @@ def _openai_caller(system: str, user: str) -> str:
return response.choices[0].message.content or "" return response.choices[0].message.content or ""
def format_bibliography_entry(messy_entry: str, caller: Caller | None = None) -> str: def format_bibliography_entry(
messy_entry: str,
caller: Caller | None = None,
max_retries: int = DEFAULT_MAX_RETRIES,
) -> str:
"""Reformat a single bibliography entry to CMOS 18 bibliography form. """Reformat a single bibliography entry to CMOS 18 bibliography form.
Pass a ``caller`` shim to avoid the real API (used by tests and by the Pass a ``caller`` shim to avoid the real API (used by tests and by the
harness when running with a fake formatter for loop smoke tests). harness when running with a fake formatter for loop smoke tests).
Reasoning models (gpt-5, o-series) are nondeterministic. The candidate
is checked against ``cmos.runtime_validator`` after each call; if it
fails any structural sanity check (missing terminal period, dropped
italics, stray Ibid., unbalanced quotes/asterisks), the formatter
retries up to ``max_retries`` more times. The runtime validator is
intentionally weaker than ``cmos.linter`` because at runtime there is
no Exemplar — only structural rules independent of source type apply.
If every attempt fails validation, the LAST attempt is returned (we
don't raise — the caller still gets something usable, and the failure
will surface via the scoring linter or human review). The retry path
is not a substitute for the scoring axes, just a guardrail against
GPT-5 reasoning drift.
""" """
call = caller or _openai_caller call = caller or _openai_caller
user_message = build_user_message(messy_entry) user_message = build_user_message(messy_entry)
raw = call(SYSTEM_PROMPT, user_message) last: str = ""
return raw.strip() for _ in range(max_retries + 1):
last = call(SYSTEM_PROMPT, user_message).strip()
if validate(last).passed:
return last
return last
+66
View File
@@ -0,0 +1,66 @@
"""Runtime structural validator — type-independent sanity checks on a single
formatted bibliography entry.
This is the gate used by ``cmos.formatter.format_bibliography_entry`` to
decide whether a candidate output is well-formed enough to return, or
whether to retry the API call. It runs WITHOUT an Exemplar, because at
runtime we have only the messy input and the candidate output — no
canonical fields, no source-type metadata.
Because it knows neither the source type nor the canonical fields, it
deliberately checks only structural properties that hold for ALL CMOS 18
bibliography entries:
- Ends with a period.
- Contains no "Ibid." (deprecated in CMOS 18).
- Contains at least one italic span (``*...*``). This catches one of
the most common nondeterministic regressions: GPT-5 occasionally
drops italic markers around a journal/book/magazine title. Almost
every CMOS bibliography entry italicizes SOMETHING (book title,
journal name, magazine name, podcast series, report title), so an
entry with no italics at all is highly likely to be malformed.
- Italic markers (``*``) are balanced (even count).
- Straight double quotes (``"``) are balanced (even count).
This is intentionally weaker than ``cmos.linter`` (the scoring linter
that needs an Exemplar). The point is to be a fast, deterministic
guardrail at runtime, not to validate every CMOS rule.
"""
from __future__ import annotations
from dataclasses import dataclass, field
@dataclass
class ValidationResult:
failures: list[str] = field(default_factory=list)
@property
def passed(self) -> bool:
return not self.failures
def validate(candidate: str) -> ValidationResult:
result = ValidationResult()
text = candidate.rstrip()
if not text.endswith("."):
result.failures.append("does not end with a period")
if "ibid" in text.lower():
result.failures.append("contains 'Ibid.' (deprecated in CMOS 18)")
if "*" not in text:
result.failures.append(
"no italic span found (CMOS bibliography entries usually italicize "
"a book/journal/magazine/series/report title)"
)
if text.count("*") % 2 != 0:
result.failures.append("unbalanced italic markers (odd number of *)")
if text.count('"') % 2 != 0:
result.failures.append("unbalanced double quotes (odd number of \")")
return result
+37
View File
@@ -81,3 +81,40 @@ def test_formatter_strips_whitespace_from_caller_output():
output = format_bibliography_entry("anything", caller=fake_caller) output = format_bibliography_entry("anything", caller=fake_caller)
assert output == "Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020." assert output == "Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020."
def test_formatter_retries_on_validator_failure():
"""If the runtime validator rejects the first attempt (e.g., italics
missing), the formatter should retry and return a passing later attempt.
Mirrors the iter 7 magazine_mead variance: GPT-5 sometimes drops italics
around a magazine name, and a re-call typically succeeds."""
attempts = []
def flaky_caller(system: str, user: str) -> str:
attempts.append(len(attempts) + 1)
if len(attempts) == 1:
# First attempt: italics dropped (validator should fail).
return 'Mead, Rebecca. "Terms of Aggrievement." New Yorker, December 18, 2023.'
# Retry: clean.
return 'Mead, Rebecca. "Terms of Aggrievement." *New Yorker*, December 18, 2023.'
output = format_bibliography_entry("messy mead", caller=flaky_caller)
assert "*New Yorker*" in output
assert len(attempts) == 2
def test_formatter_returns_last_attempt_after_max_retries():
"""If every attempt fails validation, return the last attempt and don't
loop forever. We surface the failure later via scoring/canary, not by
raising at runtime — the user should still get *something* back."""
attempts = []
def always_bad(system: str, user: str) -> str:
attempts.append(1)
return 'Mead, Rebecca. "Terms of Aggrievement." New Yorker, December 18, 2023.'
output = format_bibliography_entry("messy", caller=always_bad, max_retries=2)
# 1 initial + 2 retries = 3 calls total.
assert len(attempts) == 3
assert "*" not in output # the bad output is what we got
+70
View File
@@ -0,0 +1,70 @@
"""Tests for src/cmos/runtime_validator.py.
Runtime validator runs at format time on a single candidate string with no
exemplar context. It checks type-independent structural sanity rules so the
formatter can self-retry when GPT-5's reasoning drift produces a malformed
entry. It is deliberately weaker than the scoring linter (which has access
to canonical fields and source type) but it catches the most common
nondeterministic failures: missing terminal period, stray Ibid., dropped
italic markers, unbalanced quotes/asterisks.
"""
from cmos.runtime_validator import validate
def test_clean_book_entry_passes():
candidate = "Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020."
result = validate(candidate)
assert result.passed is True
assert result.failures == []
def test_clean_journal_entry_passes():
candidate = (
'Kwon, Hyeyoung. "Inclusion Work: Children of Immigrants Claiming Membership in Everyday Life." '
"*American Journal of Sociology* 127, no. 6 (2022): 181859. "
"https://doi.org/10.1086/720277."
)
assert validate(candidate).passed
def test_missing_terminal_period_fails():
result = validate("Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020")
assert result.passed is False
assert any("period" in f.lower() for f in result.failures)
def test_ibid_fails():
result = validate("Ibid., 42.")
assert result.passed is False
assert any("ibid" in f.lower() for f in result.failures)
def test_missing_italics_anywhere_fails():
# The "New Yorker" magazine regression we saw in iter 7 variance: italic
# markers around the magazine name disappeared. This rule should catch
# that — almost every CMOS bibliography entry italicizes something.
candidate = 'Mead, Rebecca. "Terms of Aggrievement." New Yorker, December 18, 2023.'
result = validate(candidate)
assert result.passed is False
assert any("italic" in f.lower() for f in result.failures)
def test_unbalanced_italic_markers_fails():
candidate = "Yu, Charles. *Interior Chinatown. Pantheon Books, 2020."
result = validate(candidate)
assert result.passed is False
assert any("italic" in f.lower() or "*" in f for f in result.failures)
def test_unbalanced_double_quotes_fails():
candidate = 'Mead, Rebecca. "Terms of Aggrievement. *New Yorker*, December 18, 2023.'
result = validate(candidate)
assert result.passed is False
assert any("quote" in f.lower() for f in result.failures)
def test_passed_property_is_negation_of_failures():
# Sanity: passed should be True iff there are no failures.
result = validate("Yu, Charles. *Interior Chinatown*. Pantheon Books, 2020.")
assert result.passed == (result.failures == [])