Per-run log · part of Psychohistorians (reader view)

Generated by scripts/gen_psychohistorians_runs.py from scripts/psychohistorians_runs.json — edit the JSON, not this file. Per §6a of the experiment-management guide: the top table is the only editable block (rows update in place as state/result changes); everything below is append-only (one block per run, never deleted).

Master runs-table

IDDateQuestionDiff vs parentStateResult (headline)wandb
gaussian-sweep2026-04-30Does the contingency proxy distinguish truth-tracking from non-truth-tracking dynamics on Gaussian-IID signals?vs none-baseline: first end-to-end Gaussian sweep of the contingency proxy across 10 cells (rational/entrenchment/polarization/DeGroot × complete/BA/SBM/ring).finishedProxy #2 (volume robustness) cleanly separates truth from falsehoods in every multi-bucket cell (truth-bucket c2 lowest; falsehoods 0.68–0.93). Proxy #1 (effective rank) degenerate — bucket assignment…
nl-pilots2026-05-05Does the NL-side contingency proxy work at single-cell pilot scale, and which perturbation kernel (A vs B1 vs B2) is most discriminative?vs gaussian-sweep: move from Gaussian-IID signals to NL signals; introduce three NL perturbation kernels (A = LLM rewrite, B1 = agent-major adjacent swap, B2 = tick-major adjacent swap); replace Proxy #2’s discrete stay-rate with the continuous score-stay ψ(b,ε) read from psychohistorians’ BeliefGrader.finishedVariant A on the anti-truth-proposition cell (q1neg_a) cleanly bifurcates baselines into truth-aligned (S=−0.5, c2≈0.10) and falsehood-aligned (S=+0.33, c2≈0.38) buckets, Spearman ρ(S, c2) = +0.87. Va…
nl-controversial-sweep2026-05-11Does the NL contingency proxy generalize across diverse controversial questions at multi-cell scale?vs nl-pilots: scale to 6 controversial cells × 3 variants × M=20 baselines × N=10 agents × T=30 ticks (vs M=5/N=4/T=4 pilot scale).finished1,247 of 1,560 sims landed (~20% upstream 502/504 attrition). c5_llm_science_2032 is the gold-standard contested cell — the only one where baselines split across S=0 (7 of 12 surviving baselines at …
nl-cross-cell-ols2026-05-12Is the c2↔correctness relationship within-cell (per-run answer fragility) or between-cell (cell-level dynamics)?vs nl-controversial-sweep: analysis-only — OLS of |S_b − GT_cell| on c2_A across cells, with cell fixed effects.finishedModel 1 (cell FE): β(c2_A) = −0.057, p = 0.617 (n.s.). Model 2 (pooled): β = +1.368, p < 0.0001, R² = 0.20. Model 3 (per-cell): no significant within-cell β in any of 6 cells. Headline: the c2↔correct…
nl-c5-m5002026-05-13At M=500 on the only cell with within-cell S spread (c5_llm_science_2032, Variant A), does the c2↔correctness relationship survive within-cell?vs nl-controversial-sweep: scale a single cell (c5_llm_science_2032) from M=20 → M=500 baselines × 4 perturbation magnitudes; Variant A only.🟡 runningLaunched 2026-05-13; re-launched after credit hit (patched runner snapshots S_b + snippet pools per baseline for crash-safe resume). 500 baselines × 4 perturbations = 2,500 sims. Resumed 2026-05-15 by…

Per-run entries (append-only)

Each block lands once per run, in chronological order. Never edit a past block. If a later run revises a prior conclusion, append a new block that references it.

gaussian-sweep — Does the contingency proxy distinguish truth-tracking from non-truth-tracking dynamics on Gaussian-IID signals?

  • Date: 2026-04-30
  • State:finished
  • Parent run: none-baseline
  • Diff vs parent: vs none-baseline: first end-to-end Gaussian sweep of the contingency proxy across 10 cells (rational/entrenchment/polarization/DeGroot × complete/BA/SBM/ring).
  • Scale: 10 cells × M=200 baselines × 6 perturbation magnitudes ε ∈ {0, 0.05, 0.1, 0.2, 0.5, 1.0}
  • Variants: Numerical perturbation kernel — per-(agent,tick) Gaussian observation noise scaled by ε.
  • Indicator: Proxy #1 (effective rank of column-centered obs-stack) + Proxy #2 (1 − AUC of P_stay(ε))
  • Slack thread: 1780913975.672689

Headline result:

Proxy #2 (volume robustness) cleanly separates truth from falsehoods in every multi-bucket cell (truth-bucket c2 lowest; falsehoods 0.68–0.93). Proxy #1 (effective rank) degenerate — bucket assignment driven by initial beliefs, not signal patterns. Proxy #1 retired per Tianyi.

One-line read: Proxy #2 validated; Proxy #1 retired. Polarity matches hypothesis (low contingency ⇒ truth-tracking).

nl-pilots — Does the NL-side contingency proxy work at single-cell pilot scale, and which perturbation kernel (A vs B1 vs B2) is mos

  • Date: 2026-05-05
  • State:finished
  • Parent run: gaussian-sweep
  • Diff vs parent: vs gaussian-sweep: move from Gaussian-IID signals to NL signals; introduce three NL perturbation kernels (A = LLM rewrite, B1 = agent-major adjacent swap, B2 = tick-major adjacent swap); replace Proxy #2’s discrete stay-rate with the continuous score-stay ψ(b,ε) read from psychohistorians’ BeliefGrader.
  • Scale: 6 cells × M=5 baselines × N=4 agents × T=4 ticks; variants A / B1 / B2 each
  • Variants: A: LLM rewrite (paraphrase/neutralize/counter mixture parametrized by ε). B1: agent-major adjacent swap. B2: tick-major adjacent swap.
  • Indicator: NL Proxy #2 (continuous score-stay ψ(b,ε) via BeliefGrader)
  • Slack thread: 1780913975.672689

Headline result:

Variant A on the anti-truth-proposition cell (q1neg_a) cleanly bifurcates baselines into truth-aligned (S=−0.5, c2≈0.10) and falsehood-aligned (S=+0.33, c2≈0.38) buckets, Spearman ρ(S, c2) = +0.87. Variant A > B1 > B2 in discriminative power across all pilot cells.

One-line read: NL proxy validated at pilot scale; Variant A is the discriminative variant. Semantic counter-rewrites cut deeper than positional shuffles.

nl-controversial-sweep — Does the NL contingency proxy generalize across diverse controversial questions at multi-cell scale?

  • Date: 2026-05-11
  • State:finished
  • Parent run: nl-pilots
  • Diff vs parent: vs nl-pilots: scale to 6 controversial cells × 3 variants × M=20 baselines × N=10 agents × T=30 ticks (vs M=5/N=4/T=4 pilot scale).
  • Scale: 6 controversial cells × 3 variants × M=20 baselines × N=10 × T=30 = 1,560 sims (1,247 landed)
  • Variants: A / B1 / B2 (same kernels as pilots)
  • Indicator: NL Proxy #2
  • Slack thread: 1780913975.672689

Headline result:

1,247 of 1,560 sims landed (~20% upstream 502/504 attrition). c5_llm_science_2032 is the gold-standard contested cell — the only one where baselines split across S=0 (7 of 12 surviving baselines at negative S, 5 at positive); Variant A Spearman ρ(S, c2) = +0.84. 5 of 6 cells failed to produce within-cell S spread because LLM agents have strong consensus priors on most policy/historical claims. Variant A consistently > B1/B2 in c2 magnitude.

One-line read: Within-cell c2 signal is concentrated on c5; 5/6 cells need broader prior dispersion or differently framed propositions.

nl-cross-cell-ols — Is the c2↔correctness relationship within-cell (per-run answer fragility) or between-cell (cell-level dynamics)?

  • Date: 2026-05-12
  • State:finished
  • Parent run: nl-controversial-sweep
  • Diff vs parent: vs nl-controversial-sweep: analysis-only — OLS of |S_b − GT_cell| on c2_A across cells, with cell fixed effects.
  • Scale: N=88 across 6 cells (M=12–16 per cell); no new sims
  • Variants: A only (C_sign dropped; all 6 propositions positively phrased)
  • Indicator: Cross-cell OLS on c2_A
  • Slack thread: 1780913975.672689

Headline result:

Model 1 (cell FE): β(c2_A) = −0.057, p = 0.617 (n.s.). Model 2 (pooled): β = +1.368, p < 0.0001, R² = 0.20. Model 3 (per-cell): no significant within-cell β in any of 6 cells. Headline: the c2↔correctness relationship is between-cell, not within-cell — Variant A’s c2 captures cell-level dynamics (prior strength, evidence quality, GT-distance) more than per-run answer fragility. Within-cell power limited by M=12–16 / cell.

One-line read: Between-cell signal real (β = +1.37, p<1e-4); within-cell signal absent at this M. Motivates the M=500 deep dive.

nl-c5-m500 — At M=500 on the only cell with within-cell S spread (c5_llm_science_2032, Variant A), does the c2↔correctness relationsh

  • Date: 2026-05-13
  • State: 🟡 running
  • Parent run: nl-controversial-sweep
  • Diff vs parent: vs nl-controversial-sweep: scale a single cell (c5_llm_science_2032) from M=20 → M=500 baselines × 4 perturbation magnitudes; Variant A only.
  • Scale: 500 baselines × 4 perturbations × Variant A only = 2,500 sims; 24-way parallel + async-within-sim
  • Variants: A only
  • Indicator: NL Proxy #2 (continuous score-stay) at high M for within-cell statistical power
  • Slack thread: 1780913975.672689

Headline result:

Launched 2026-05-13; re-launched after credit hit (patched runner snapshots S_b + snippet pools per baseline for crash-safe resume). 500 baselines × 4 perturbations = 2,500 sims. Resumed 2026-05-15 by Tianyi at MenoClaw. Last channel signal on this run: 2026-05-15. Final numbers not posted in-channel.

One-line read: Result pending; the within-cell c2↔correctness test the cross-cell OLS motivated.