Review of ZH’s blog framing. (A) structure critique, (B) claim-by-claim debate with verified evidence (agree and disagree), (C) the ? citations filled, (D) which of our own runs back each section. Every paper was web-verified (title/authors/year/venue); items needing a second look are flagged ⚠confirm.

A. Structure — what works, what to change

Keep: Intro→Methods→Benchmarking is the right spine, and leading with why (calibration + LLMs-as-epistemic-technology) is good.

Three structural moves I’d push:

  1. Lead with our killer result, not the philosophy. The most useful thing we have to share is that a headline “anti-Bayesian / entrenchment” signal turned out to be a measurement artifact of the elicitation mode (snap-prior vs reasoned-posterior are different functionals, ρ≈0.6). That is the thesis-in-one-finding: belief elicitation is so mode-sensitive it can manufacture a false scientific conclusion. Put it in the intro as the motivating war story; the philosophy then lands as earned, not abstract.
  2. Split “is it sensible for LLMs to have beliefs?” (ontology) from “can we measure them reliably?” (epistemics). The draft mixes them. Most of our evidence is about measurement fragility, not whether beliefs exist. Cleaner line: “beliefs are real enough to matter (humans defer to them) but our instruments are unreliable.”
  3. Make “Benchmarking” concrete instead of a shrug. You wrote “no clue” — but we have three workable proxies (perturbation-robustness, ground-truth calibration where labels exist, cross-elicitation agreement). Pitch them as the contribution.

Add a short “what is a belief operationally” box and a threats-to-validity subsection (elicitation mode, judge bias, label balance) — we paid for those lessons.

B. Claim-by-claim debate

1. “If you don’t reliably know model belief, you can’t calibrate.” — Agree. Twist: even defining the target is mode-dependent — calibrate which elicited belief (snap? reasoned? verbalized? logprob?)? Our mode-gap data says these disagree, so “calibration” is under-specified without fixing the elicitation.

2. “LLMs are epistemic technology; humans should know how much to trust them.” — Agree, strongly. Best motivation. Tie to sycophancy: if the model bends its stated belief to the user (Sharma 2023; Wang 2023), the human is calibrating against a mirror, not a measurement.

3. “Is it sensible for LLMs to have beliefs? Not really, if arbitrarily changing.” — Partially disagree. Premise (“arbitrarily changing”) too strong. Beliefs are not arbitrary: approximately-Bayesian sequential updating on coin-flips (Gupta 2025); ICL as implicit Bayesian inference (Xie 2022); base models well-calibrated (GPT-4 report Fig. 8). What’s unstable is (a) the elicitation (our mode gap) and (b) robustness to social perturbation. Defensible claim: “beliefs are real but context/elicitation-relative and perturbation-fragile,” not “no real beliefs.” Stronger, more novel, and the position our data supports.

4. “Any proposition from an LLM should be treated as a belief.” — Agree (operational), but flag faithfulness: the stated proposition often isn’t the cause of the answer (Turpin 2023 — models rationalize bias-injected answers without mentioning the bias; Lanham 2023). Outputs as beliefs, yes; stated reasoning as the mechanism, no.

5. “Bayesian in rare cases (coin toss); base Bayesian in more.” — Agree, citations nailed. Coin-toss: Gupta et al., Enough Coin Flips Can Make LLMs Act Bayesian, ACL 2025 (deviations from miscalibrated priors, not bad updating). ICL: Xie et al., ICLR 2022. Base>RLHF: GPT-4 report Fig. 8. Honest tension: Falck, Wang & Holmes, Is In-Context Learning Bayesian? A Martingale Perspective, ICML 2024 finds martingale violations — resolve in our favor: much apparent violation is the mode-gap artifact.

6. “Fine-tuning distorts (sycophancy) + overconfidence.” — Agree, fully cited. Sycophancy: Sharma 2023; Perez 2022 (RLHF inverse scaling); Wei 2023. Overconfidence in the RLHF signal: Zhou, Hwang, Ren & Sap, Relying on the Unreliable, ACL 2024; Leng 2024 (reward models favor confidence). RLHF calibration loss: GPT-4 Fig. 8; verbalized partly recovers: Tian 2023 (EMNLP).

7. “Self-report: same answers in repetitions.” — Disagree as stated; split it. Stable to exact repetition, maybe; not paraphrase-invariant: Elazar et al., TACL 2021 (ParaRel); BECEL (Jang 2022); Chen et al. 2023. Rewrite: “reproducible under identical prompts, not invariant to paraphrase/format.”

8. “Not reliable vs perturbation; user claims otherwise → changes.” — Agree, strongly. Wang, Yue & Sun, Can ChatGPT Defend its Belief in Truth?, EMNLP 2023 (abandons correct answer 22–70% under debate); Xie et al., Ask Again, Then Fail, ACL 2024; Kim & Khashabi 2025 (flips more on conversational follow-up than simultaneous). Plus our sycophancy sim (§D).

9. “Snap ≠ reasoned mode; we have data.” — Agree, OUR contribution. Mode gap (§D) + Yoon et al., NeurIPS 2025 (CoT adjusts credence over the trace) + Chen et al. 2026 (⚠confirm id) (CoT sets argmax, priors govern the rest). Open question (when is it the belief? reason-until-stable = prior?) — frame as a proposal (fixed-point of reasoning = elicited belief); we lack a stopping criterion.

10. LLM-Judge: same-mode vs infer-belief. — Direct evidence it matters. Our same-mode judge lifts prior↔posterior ρ to 0.78–0.92, slope→≈0 (§D). But self-preference risk: Panickssery, Bowman & Feng, NeurIPS 2024; Wataoka 2024. Judge bias: Zheng 2023 (MT-Bench); Wang 2023 (position); Saito 2023 (verbosity); Ye 2024 (CALM, 12 biases). “Infer another model’s credence from its trace” = building blocks only: Kadavath 2022; Lin/Hilton/Evans 2022; Radharapu 2025 (⚠confirm).

11. “Prior/posterior relative; posteriors = whatever legitimately updates.” — Agree; sharpen. Open: what is a legitimate update (evidence / sim-human / more reasoning). Sycophancy sim = worked example; hazard: a posterior can itself be an elicitation artifact (the bare-number prompt bug flipped the sign).

12. “Benchmarking — no clue.” — Disagree we have nothing. Three proxies: (i) perturbation/paraphrase robustness (Elazar) as a necessary condition; (ii) ground-truth calibration (Brier/ECE); (iii) cross-elicitation agreement (mode-gap ρ used diagnostically). Together a real protocol. Surveys: Shorinwa 2024 (ACM CSUR); Xia 2025 (ACL Findings).

Literature ? → uncertainty estimation: verbalized (Lin/Hilton/Evans 2022; Tian 2023), self-knowledge (Kadavath 2022), semantic entropy (Kuhn 2023 ICLR; Farquhar 2024 Nature), surveys (Shorinwa 2024; Xia 2025).

C. The “0.99 / 0.9999 on binary questions” claim

No single paper headlines it. Cite GPT-4 report (Fig. 8) + Tian 2023 (RLHF logit probs poorly calibrated; verbalized beats them) + Epstein et al., FermiEval, 2025 (nominal 99% CIs hold ~65%). Don’t attribute the exact “0.9999 on binary” to one source.

D. Our own runs that back this (martingale)

  • Mode gap = the headline. Snap-prior vs reasoned-posterior ~0.6 correlated even noise-free; explains ~85% of the negative Martingale Score via MS ≈ ρ − 1; same-mode judge lifts ρ to 0.78–0.92, slope→≈0. Cross-validated on Mistral-7B, Llama-8B, Qwen3-8B/32B. → §B-5, B-9, B-10, B-12.
  • The artifact lesson = the strongest single anecdote (→ intro).
  • Sycophancy/perturbation sim. Rewording the elicitation prompt flipped the sign (validate-only +entrenchment → −0.39 de-entrenchment once the human had to explain rather than emit a bare number). → §B-8, B-11. See sycophancy-rparam-results.
  • Judge bias (deepseek-judge Brier gap = in-house B-10 — ⚠pull exact numbers).
  • Recovery. Over-extremized model D recovered (Qwen3-32B Brier 0.291→0.234 held-out) → distortion installed/removable, supporting B-6.

Flags for ZH

  • Your own Martingale Score paper: an agent surfaced it as He, Qiu, Shirado, Sap, NeurIPS 2025, arXiv:2512.02914confirm the real id/citation/authors yourself (auto-generated arXiv ids are the #1 hallucination spot).
  • ⚠confirm three recent ids: Chen 2026 (CoT/label-variation), Radharapu 2025 (judge probes), FermiEval (2510.26995).
  • Deliberately excluded several thematically-perfect but unverifiable future-dated (26xx) preprints — they look invented.

Verified citation key

Gupta et al, Enough Coin Flips Can Make LLMs Act Bayesian, ACL 2025, arXiv:2503.04722 · Xie/Raghunathan/Liang/Ma, ICL as Implicit Bayesian Inference, ICLR 2022, 2111.02080 · Falck/Wang/Holmes, Is ICL Bayesian? A Martingale Perspective, ICML 2024, 2406.00793 · OpenAI, GPT-4 Technical Report (Fig. 8), 2303.08774 · Kadavath et al, LMs (Mostly) Know What They Know, 2022, 2207.05221 · Lin/Hilton/Evans, Teaching Models to Express Uncertainty in Words, TMLR 2022, 2205.14334 · Tian et al, Just Ask for Calibration, EMNLP 2023, 2305.14975 · Kuhn/Gal/Farquhar, Semantic Uncertainty, ICLR 2023, 2302.09664 · Farquhar et al, Detecting hallucinations with semantic entropy, Nature 2024 · Epstein et al, FermiEval, 2025, 2510.26995 ⚠ · Sharma et al, Towards Understanding Sycophancy, 2023, 2310.13548 (ICLR 2024) · Perez et al, Discovering LM Behaviors with Model-Written Evals, 2022, 2212.09251 (ACL Findings 2023) · Wei et al, Simple synthetic data reduces sycophancy, 2023, 2308.03958 · Zhou/Hwang/Ren/Sap, Relying on the Unreliable, ACL 2024, 2401.06730 · Leng et al, Taming Overconfidence (Reward Calibration in RLHF), 2024, 2410.09724 · Turpin et al, LMs Don’t Always Say What They Think, NeurIPS 2023, 2305.04388 · Lanham et al, Measuring Faithfulness in CoT, 2023, 2307.13702 · Elazar et al, Measuring & Improving Consistency (ParaRel), TACL 2021, 2102.01017 · Jang et al, BECEL, COLING 2022 · Chen et al, Two Failures of Self-Consistency, TMLR 2024, 2305.14279 · Wang/Yue/Sun, Can ChatGPT Defend its Belief in Truth?, EMNLP 2023, 2305.13160 · Xie et al, Ask Again, Then Fail, ACL 2024, 2310.02174 · Kim&Khashabi, LLM Sycophancy Under User Rebuttal, EMNLP 2025, 2509.16533 · Yoon et al, Reasoning Models Better Express Their Confidence, NeurIPS 2025, 2505.14489 · Zheng et al, Judging LLM-as-a-Judge (MT-Bench), NeurIPS 2023, 2306.05685 · Wang et al, LLMs are not Fair Evaluators (position bias), ACL 2024, 2305.17926 · Saito et al, Verbosity Bias in Preference Labeling, 2023, 2310.10076 · Panickssery/Bowman/Feng, LLM Evaluators Favor Their Own Generations, NeurIPS 2024, 2404.13076 · Ye et al, Justice or Prejudice? (CALM), ICLR 2025, 2410.02736 · Shorinwa et al, Survey on UQ of LLMs, ACM CSUR 2025, 2412.05563 · Xia et al, Survey of Uncertainty Estimation Methods, ACL Findings 2025, 2503.00172