What this page is

A grounded status check for publication, built by verifying Zhonghao’s Status quo && Plans note against the real run data in /data/jobs/ (75 run dirs), the all-runs Log, narrative, and the master runs table. Every metric below traces to a real file; four metric sets were independently recomputed from raw stated_p/y and matched the logged aggregates to the decimal. This is the A pass (focused, headline claims) — a fuller sweep (B) follows. Verdict legend: ✅ confirmed · 🟡 partial / qualified · ❌ refuted · ❔ uncertain (needs a run).

TL;DR for the paper

  • The project’s negative results are robust and replicate across 3 model families; the positive result is single-model (Qwen) and load-bearing on several non-ecological choices.
  • The mode gap explains ~85% of the negative (“entrenchment”) slope; “double priors” is a small (~−0.06) measurement-noise correction that isn’t even in the shipped code — don’t present it as co-equal.
  • The headline recovery (A-on-D) is real on Qwen3-32B (Brier 0.291→0.234 held-out, generalizes) but it is an “undistiller” (removes installed bias, not intrinsic) and diverges catastrophically on overconfident-distilled models.
  • The martingale 2nd-stage gain is real but minimal (~0.033 Brier) and only under same-mode elicitation; on a natural base the martingale stage is statistically redundant after Brier (p=0.43). The namesake objective is doing little of the work.
  • One urgent run would most de-risk the paper: a k=3 confirmation of judge-samemode-BP on D (the decisive 3d result is currently k=1).

1 · Negative (“entrenchment”) slope — is the two-part decomposition full?

Claim: the negative slope is explained by exactly two parts — (a) mode gap, (b) double priors in the formula. Verdict: 🟡 confirmed-but-mislabeled. (mode-gap analysis)

  • The decomposition is the identity MS ≈ ρ(prior,post) − 1 (sd ratio ≈1 across arms). ρ<1 has two sources: (H4) the mode gap — snap-prior vs CoT-posterior are different functionals (~0.6-correlated even noise-free) — and (H2) snap-prior measurement noise, removable by de-attenuation (“double priors”). log.md:383-417, martingale.md:114-130.
  • Mode gap is load-bearing (~85%). Same-mode judge reads lift ρ to 0.78–0.92 and the slope returns to ≈0 or positive. Exact identity match: B-distill ρ=0.591 → MS −0.398 (predicted −0.398); A-on-B ρ=0.779 → −0.232. Replicates on Llama-8B, Qwen3-8B/32B, Mistral-7B.
  • “Double priors” is small and not in code. De-attenuation lifts Llama-8B ρ only 0.59→0.64, removing ≈−0.06 of the slope. The formula MS_adj = β_obs/(1 − σ²_meas/Var_q(prior)) − 1 is documented in martingale.md only — grep finds no implementation in core//utils//martingale_training/. Its cleanest measurement was never completed (the stage-1 model “was lost with the terminated on-demand pod”).
  • Is it the full set? Effectively yes: after both corrections the slope → ≈0 with no residual genuine anti-martingale (H1 “genuinely mean-revert” ❌ falsified). Extreme (0/1) priors were tested and ruled out as a contributor (dropping them makes MS more negative).

For the paper: present one mechanism (mode gap) + one small removable noise term, not two co-equal parts. Confidence: HIGH on the decomposition; MEDIUM on “exactly two, no residual” (the clean double-prior measurement is unfinished).

2 · Does the negative slope reliably turn positive? (degenerate-model confound)

Claim: some runs show negative→positive later, but it’s unclear if reliable or confounded by positive MS from degenerate models. Verdict: ✅ confirmed (both halves) — the positive turn is real but NOT reliable, and the degenerate-model confound is explicit.

  • Two routes to “positive”: (1) measurement swap (same traces, sign flips): paper-reextract self-report −0.398/−0.227 → re-judged +0.152/+0.117; r1-replication −0.415 → +0.023. (2) training-induced: A-on-B-distilled slope −0.398 → −0.227 (~43% toward 0, p=3.6e-13) — shrinks negativity but does not cross zero under self-report.
  • Not reliable. Snap two-stage stage-2: helps 2/5, undoes 1/5, diverges 2/5 (twostage-v2-k3/base). “Reliability is the blocker, not direction.”
  • Degenerate confound is real and measured/data/jobs/mtg7b_overconf/ (Mistral, 2026-06-16): standalone A-REINFORCE on distilled D1/D2/D3 drives overconfidence to 0.58–0.62 and Brier worse by +0.25 to +0.34 — a model collapsed to high-confidence reads as extremizing (positive slope) while calibration rots.
  • The clean win exists but is narrow: A-on-D (Qwen) via the KL-safe pipeline — overconf 0.385→0.106, Brier 0.45→0.33; held-out 0.291→0.234. Not degenerate. But every clean positive turn is k=1.

For the paper: the positive turn is trustworthy only under (i) same-mode/judge measurement and (ii) the KL-safe safeguards; strip either → artifact sign-flip or divergence masquerading as positive MS. Confidence: HIGH.

3 · Workshop claims + the key open question

Sub-claimVerdictGrounding
3a martingale improves a degenerate base🟡 partialReal only for Qwen confirmation-bias-distilled D recovered with KL-safe+info-term A: 0.291→0.234 held-out, generalizes (p=2e-9). It’s an “undistiller” (not intrinsic bias). Refuted in general: on overconfident-distilled Mistral, single-stage A diverges to ~0.9 confidence; only B recovers them (mtg7b_overconf).
3b validated on different models✅ confirmed3 families (Qwen3-32B, Llama-3.1-8B, Mistral-7B). Negatives replicate broadly (martingale term hurts; A degrades natural base; mode gap). Positive recovery is Qwen-only; Mistral A-on-D weak (0.366→0.348 ≫ base 0.286).
3c training surpasses the base🟡 partialTrue only for B (Brier-only), modestly (Llama 0.2546→0.2286 Wilcoxon 1.1e-6; Mistral 0.281→0.219). Martingale term C hurts (0.2845>base). A-on-base = clean null (p=0.77).

3d (the crux) — martingale-only vs 2-stage; is martingale-2nd > brier-1st; is the gain minimal?

Verdict: ❔→🟡 the martingale-2nd gain is REAL BUT SMALL, and ONLY under same-mode elicitation. (validation thread)

  • Same-mode, on D (judge-samemode-BP): s0 0.341 → s1 (+Brier) 0.259s2 (+martingale) 0.226. Martingale-2nd is strictly better than brier-1st, but by Δ≈−0.033 (~13%) while the Brier stage did ~2.5× more (−0.082). ⚠ k=1.
  • Snap-reasoned, on D (twostage-v2-D-snap): s1 0.273 → s2 0.346 — martingale-2nd undoes the gain (opposite sign).
  • Natural base, same-mode (selfjudge-filter-base): martingale STAGE redundant after Brier — s2-vs-s1 paired-t p=0.43; JudgeMS CIs straddle 0.
  • martingale-only vs 2-stage: A-on-D (0.234; filtered 0.225±0.007) ≈ 2-stage-on-D same-mode (0.226; best-tuned filter-s2only 0.2213) → roughly on par, 2-stage marginally lower.

Answers: martingale-only is ≈ on par with 2-stage; martingale-2nd is strictly better than brier-1st only under same-mode and only by ~0.033 (minimal); on snap/natural it’s redundant→harmful. The single highest-value missing run is a k=3 confirmation of judge-samemode-BP on D.


5 · “For-positive-results” tweaks that may not be ecologically valid

Found 12 (all quoted from real code/config). The four most load-bearing for the paper:

  1. High-|Δ| trajectory filter (stage-2 only; ondemand_ship/run_selfjudge_base.sh filt() keeps |actual_delta|≥median). Converts unreliable recovery (~2/6) → 3/3; produces the best 2-stage-on-D (0.2213). You can’t pre-select high-movement questions in deployment.
  2. Prior-conforming human simulant (sycophancy_sim/sim.py HUMAN_SYS injects 确认偏误/confirmation bias by prompt). “Without the explicit disposition … nothing to measure.” And martingale-vs-truthseeker is a tie (slope −0.827 vs −0.822).
  3. Balanced labels via option-flip (forces y_mean=0.5; eval natural ~0.30). Gates the only positive natural-base result; BSS computed vs a non-matching constant baseline.
  4. Same-mode judge selection + judge “laundering” — same-mode is the only mode where the martingale stage helps; the LLM judge compresses extreme beliefs (judge ~4–6% extreme vs model ~16–22%), flattering Brier and partly manufacturing MS→0 (“Never judge for Brier”).

Plus, conceptually damaging: (7) martingale stage redundant after Brier (p=0.43) — the gain is calibration, not the namesake; (6) D is deliberately bias-installed (recovery = undistiller, never beats the base-rate constant on a natural base); (8) reward optimizes the same OLS-slope later shown to be a measurement artifact (circularity); (9) info-term gain partly from inertia (+6.6pp |Δ|=0 — refusing to update); (10) in-sample p / single-seed / outlier-subset wins; (11) “seeds” are uncontrolled re-runs (trainer ignores SEED) — weakens every CI; (12) held-out integrity (an earlier sorted-id split anomaly; a pre-2026-05-17 silent-base-bug era).

Most damaging if a reviewer probes

Removing #1 (high-Δ filter) or #3 (balanced labels) likely removes the “reliable recovery” and “works on a natural base” results; #7 + #6 mean the martingale contribution may be attributable to plain calibration on an artificially broken model. The docs are unusually self-critical — these are load-bearing, not hidden.


Plans (from Zhonghao’s note — roadmap, not verified claims)

  • Eval tweaks to try: (i) treat current posterior as the prior (gives context to calibrate model priors); (ii) most ecologically-valid posterior = elicit via multi-turn reasoning.
  • Workshop → conference (aim EOM):
    • Belief-elicitation methods — judge-eval still preferred (methodological contribution; no ground truth; same-family judge to avoid confounder, ~self-report but without the self-fulfilling-prophecy pitfall).
    • Human–AI system as a whole — sycophantic influence from AI, eval on human-sim (MS/Brier), RL martingale training against human–AI co-updates → then tease apart “correcting sycophancy” vs “correcting confirmation bias” (martingale-only, each).
  • Open question: 3 separate workshop papers vs 1 packed conference paper (belief elicitation · martingale-on-degenerate-policy · unsupervised sycophancy-correction with human sim)? — MenoClaw’s read: the grounding above argues for one conference paper — the three threads share the same load-bearing caveats (mode-gap measurement, degenerate-model recovery, ecological-validity tweaks), and split across three workshops each piece looks thinner than the honest combined story. Strategic call, though.

Highest-value missing experiments (consolidated)

  1. k=3 confirmation of judge-samemode-BP on D — locks 3d (currently k=1). Top priority.
  2. Clean matched-recipe recovery replication on a 2nd family — the recovery positive is Qwen-only.
  3. Summarize mtg7b_overconf (2026-06-16) into the log — fresh, bears directly on 3a, not yet logged.
  4. Overconfident-but-NOT-distilled base — disentangle overconfidence magnitude from distillation as the predictor of benefit.
  5. Resume the independent-validation line (mtg_verify_qwen) — ~halfway; compute deferred per ZH.

Security flag (acted on separately)

A live OPENROUTER_API_KEY in plaintext was found in /data/jobs/mistral_verify_1780417248/env.sh and sycophancy_sim/sim_baseline.py. Not echoed. Rotate it and move keys out of shared run dirs.