Does sycophancy entrench a human’s prior? Simulated human (prior + confirmation bias) ⇄ instructor, 3 turns; ΔP regressed on (prior−0.5), slope>0 = priors amplified. Belief elicited as r∈[−1,1]→p=r/2+1/2. Rerun with the corrected prompt (human now explains, then gives r — the earlier “output only r=X” had made deepseek emit a bare number, starving the instructor of a position to engage). n=48/arm, 0 parse-drops.
Entrenchment slope (dP ~ prior−0.5)
| instructor | gpt-3.5 human | deepseek human |
|---|---|---|
| validate-only | -0.377 | -0.391 |
| syc (+ rationale) | -0.058 | -0.017 |
Transcripts (中文): rparam-gpt35-transcripts · rparam-deepseek-transcripts Transcripts (EN): rparam-gpt35-transcripts-en · rparam-deepseek-transcripts-en