Working notes on ZH’s paper plan: “Unsupervised Martingale Training Removes Excessive Influences in Human–AI Interactions.” (A) structure debate, (B) data-gap map (have vs needs-running), (C) the autonomous experiment now running, (D) the GPU experiments needing a go. Companion to eliciting-llm-beliefs / eliciting-beliefs-notes.

2026-06-29 update (this discussion) — renumber + drop the controlled framing + verified states

  • New numbering (flow order): C1 = training improves a degenerate model · C2 = two-agent eval detects excessive influence · C3 = training reduces it. Sections below renumbered accordingly.
  • Drop the controlled-experiment framing. Single-agent vs two-agent change the dialogue structure, the eval target, and the agents all at once → structurally incomparable. Frame C2 as “the two-agent regime surfaces influence that single-agent eval misses,” not a controlled A/B.
  • Verified vs run logs (06-29): C1 stands, as corrective (degenerate D recovers + the martingale term helps; on a balanced natural base BSS −0.01→+0.16 but the martingale term is redundant, p=0.43 → “corrective with a boundary”). C2 stands — authoritative detection is the 3-mode investigation (n=48, 3 seeds, CI): syc +0.20 [+0.01,+0.39] = the ONLY mode amplifying the prior vs neutral −0.37 / truth −0.82 / martingale −0.83 (sycophancy-simulation-investigation); human-model-dependent. C3 RAN but the test was INVALID (see Update below). Claim→runs: log, human-shared-grounding.
  • Title: keep “detects, and [preliminarily] reduces” until the C3 training arm lands.

Update — 2026-06-29 (evening)

C3 ran — and it was an invalid test, not a clean negative. Trained a Llama-3.1-8B instructor (product-based A reward, Metaculus seed) + two-agent sim (judge-eval every turn, both arms under the syc prompt), n=60/arm: base slope −0.036 (ΔBrier +0.041) vs trained +0.129 (+0.072). The base baseline never reproduced C2’s entrenchment (human MS ≈ 0) → nothing to reduce → the comparison is uninformative. Cause = setup mismatch vs C2 (weak Llama instructor + Metaculus forecasting data + judge-eval). A valid C3 must first reproduce a positive-entrenchment baseline (C2’s opinion-claim setup with a strong enough sycophant), then train the instructor in-context on dialogue. Pod terminated; per-dialogue raw lost. Full diagnosis → c3-mechanistic-spec; log 2026-06-29.

Why the human MS ≈ 0 under syc despite a confirmation-bias prompt (key clarification): a prompted bias is not a genuine one — a balanced/RLHF’d model complies superficially while staying malleable (syc·deepseek +0.20 vs syc·gpt-3.5 ≈0; prompt-induced ceiling ≈ +0.32, so a genuine ceiling needs a bias-distilled D-as-human). This is the human-simulator fidelity limit — the load-bearing assumption of the whole 2-agent line.

Measurement resolved (self-report → judge-eval): self-report with a no-reasoning snap prior inflates MS ~270× vs a judge reading the trace at half/complete (same model + questions). All beliefs are now judge-inferred; the r∈[−1,1] self-report is dropped. → grounding S1.

Paper rebuilt to standard (influences_paper.tex): formal MS + Prop.1; a “Judge-eval vs self-report” subsection (the 270× result); the real product-based algorithm (sft_product_based.py, not the Oxford pooled-regressor proposal); descriptive subheadings (no “Contribution N”); a “Before-C1” revision section (BE pervasive in non-reasoning models but over-stated, gone in reasoning models → value relocates to the 2-agent setting); the full C1 table (all valid runs, degenerate vs base); C2 to the original statement (reasoner-agnostic eval; sycophancy as an instance of BE; central claim = detect single-agent-invisible influence); connecting reasoning throughout; a reproducibility appendix (verbatim prompts, after the bibliography).

New grounding pages + housekeeping: codebase-and-results-map (code tree + run-id → /data/jobs), human-shared-grounding (Solved S1–S3 / Unsolved U1–U5, every claim dated to a replicable run, awaiting human sign-off). Log backfilled 6/16→6/29 (Master table + run-log, now in 1:1 correspondence); a table-edit rule added (rules/martingale-log.md, linked from CLAUDE.md — minimal-diff edits only).

Terminology: the phenomenon is belief entrenchment / 信念固化 (an earlier note mis-wrote it “熵化”). Data is unified on Metaculus for both training and the 2-agent eval (with installed prior leans), per the 06-29 discussion.

  • Measurement finalized (06-29): ALL beliefs are judge-inferred — within-trace martingale via half/complete reasoning (GPT-4o judge “what probability is the reasoner converging to” at midpoint=prior / completion=posterior), and every two-agent dialogue turn judge-inferred too. No self-report, no r∈[−1,1] (it was a self-report artifact-fix, superseded by judge-eval). All prompts EN (CN runs deprecated). Paper migrated to NeurIPS 2025 format. Full mechanistic spec (loss/assignment/trajs/prompts/formulas + quickest-path) → c3-mechanistic-spec.

A. Structure debate (the 3 contributions)

C1 — “Martingale training improves a degenerative model but not a natural base.” ⚠ Reframe before a reviewer breaks it. The clean version is falsified by our own runs: selfjudge-filter-base lifted a balanced-label natural base (BSS −0.01→+0.16), and judge-samemode-BP-base shows the recipe over-corrects an already-calibrated base (s0 0.238→s1 0.264). So the true axis is calibration gap, not degenerate-vs-base: martingale/Brier training repairs a miscalibrated start (a distilled-degenerate model D is the extreme case) and is neutral-or-harmful on an already-calibrated one. State it that way + show the per-run scatter of initial miscalibration vs ΔBrier (the monotone relationship absorbs the D-vs-base distinction). This is stronger and reviewer-proof.

C2 — “Two-agent Martingale eval detects excessive influence invisible to single-agent eval.” This is the novel hinge — protect it. ⚠ Per 2026-06-29: do NOT frame as a controlled single-vs-two-agent experiment (dialogue structure, eval target, agents all change at once = structurally incomparable). Frame as: the two-agent regime surfaces influence single-agent eval misses — a new measurement regime, not an A/B. The design is elegant and falsifiable: same base LLM, single-agent martingale ≈ nil whether or not it carries a sycophantic prompt (it’s the same model reasoning); but the human-simulator’s martingale eval, in the two-agent loop under a sycophantic AI, is non-nil. Two risks to pre-empt:

  1. The neutral control must be ~nil. If a neutral AI also drives drift in the human-sim, the contrast collapses. The running experiment (§C) tests exactly this; if neutral ≉ 0, we need a cleaner human-sim prior model.
  2. “Drift = excessive” needs the warrant. Argue the sycophancy-induced drift is predictable (martingale violation) and unwarranted (no new evidence, only validation). The neutral-vs-syc contrast is what licenses “excessive” — keep both arms.

C3 — “Martingale training reduces the AI’s excessive (sycophantic) influence.” RESOLVED (06-29): train the AI/instructor — the only trainable party (you cannot retrain a human). Martingale-train the instructor’s own reasoning (label-free reward of C1, in-context on the dialogue claims) and re-measure the human-sim’s revision. Training signal ≠ eval signal (reward on the AI’s traces; metric on the human’s trajectory — unlike C1 where they coincide), so success is a falsifiable hypothesis, not a tautology. Honest target = the human’s calibration (un-frozen, Brier↓), not a slope sign. Prompt-proxy already directional (martingale-prompted AI restores revision). Full design → c3-mechanistic-spec.

Ordering. C1 (training repairs miscalibration) → C2 = detection (two-agent regime reveals hidden influence) → C3 = training removes it = “close the loop.” (Detection and removal are two faces of one apparatus — per Tianyi, not fully separable contributions.)

Title.Removes” is a strong verb — it’s only earned once the training arm lands. Until then the defensible claim is “detects, and [preliminarily] reduces.” Don’t over-claim in the abstract ahead of the data.

B. Data-gap map

PieceStatus
C1 recovery on D (A-on-D, Qwen3-32B Brier 0.291→0.234 held-out)Have
C1 natural-base behavior (neutral reward; balanced-base lift; calibrated-base overcorrection)Have (supports the reframe)
C1 initial-miscalibration vs ΔBrier scatterNeeds computing (cheap, from existing run dirs)
C2 detection: single-agent base martingale ≈ nilHave (natural-base MS ≈ nil, C1)
C2 detection: two-agent human-sim slope, neutral vs sycRunning now (§C)
C3 intervention: martingale-train the AI in the two-agent loop, re-measure human-sim slopeNeeds GPU (§D)
C1 k=3 confirmation of judge-samemode-BP on D (currently k=1)Needs GPU (§D)

C. Running now (autonomous, CPU/OpenRouter, ~$2)

C2 detection. human-sim (gpt-3.5) ⇄ AI=deepseek base, two prompts: neutral (control) vs sycophantic (experimental). Corrected r-elicitation (explanation-first). n=48/arm. Measures the human-sim’s entrenchment/martingale slope under each. Predicted: neutral ≈ 0, syc ≠ 0 — the C2 detection signal. Results + transcripts (CN+EN per the standing rule) to the hub when done.

D. Needs ZH go + warm H200 (GPU — not launched)

  1. Intervention runs (C3 core): martingale-train the AI in the two-agent loop; re-run the §C eval; show syc-arm slope → ≈0 post-training. This is what earns “removes” in the title.
  2. C1 k=3 confirmation of judge-samemode-BP on D (de-risks the one k=1 result the publication-prep flagged as decisive). I’ll post a concrete run plan + cost estimate once the pod is resumed. Per our standing rule I won’t launch GPU/compute beyond the cheap CPU sim without an explicit go.

E. Open questions for ZH

  • C3: train the AI or the human-sim? (recommend the AI — closes the loop).
  • C2: is the human-sim model fixed across the paper, and is it itself calibrated enough that neutral drift ≈ 0? (the §C run will show; if not, we pick a cleaner human-sim).
  • Title verb: hold “removes” until the intervention arm lands?