Per-run log · part of Psychohistorians (reader view)
Generated by
scripts/gen_psychohistorians_runs.pyfromscripts/psychohistorians_runs.json— edit the JSON, not this file. Per §6a of the experiment-management guide: the top table is the only editable block (rows update in place as state/result changes); everything below is append-only (one block per run, never deleted).
Backfill caveat
Pre-2026-06-08 runs are reconstructed from the
Psychohistorians-IndicatorREADME’s experimentation-status section +#project-psychohistorianschannel posts. They lack the per-run wandb artifacts (code SHA, env lockfile, data hashes) that post-harness runs will carry.
Master runs-table
| ID | Date | Question | Diff vs parent | State | Result (headline) | wandb |
|---|---|---|---|---|---|---|
gaussian-sweep | 2026-04-30 | Does the contingency proxy distinguish truth-tracking from non-truth-tracking dynamics on Gaussian-IID signals? | vs none-baseline: first end-to-end Gaussian sweep of the contingency proxy across 10 cells (rational/entrenchment/polarization/DeGroot × complete/BA/SBM/ring). | ✅ finished | Proxy #2 (volume robustness) cleanly separates truth from falsehoods in every multi-bucket cell (truth-bucket c2 lowest; falsehoods 0.68–0.93). Proxy #1 (effective rank) degenerate — bucket assignment… | |
nl-pilots | 2026-05-05 | Does the NL-side contingency proxy work at single-cell pilot scale, and which perturbation kernel (A vs B1 vs B2) is most discriminative? | vs gaussian-sweep: move from Gaussian-IID signals to NL signals; introduce three NL perturbation kernels (A = LLM rewrite, B1 = agent-major adjacent swap, B2 = tick-major adjacent swap); replace Proxy #2’s discrete stay-rate with the continuous score-stay ψ(b,ε) read from psychohistorians’ BeliefGrader. | ✅ finished | Variant A on the anti-truth-proposition cell (q1neg_a) cleanly bifurcates baselines into truth-aligned (S=−0.5, c2≈0.10) and falsehood-aligned (S=+0.33, c2≈0.38) buckets, Spearman ρ(S, c2) = +0.87. Va… | |
nl-controversial-sweep | 2026-05-11 | Does the NL contingency proxy generalize across diverse controversial questions at multi-cell scale? | vs nl-pilots: scale to 6 controversial cells × 3 variants × M=20 baselines × N=10 agents × T=30 ticks (vs M=5/N=4/T=4 pilot scale). | ✅ finished | 1,247 of 1,560 sims landed (~20% upstream 502/504 attrition). c5_llm_science_2032 is the gold-standard contested cell — the only one where baselines split across S=0 (7 of 12 surviving baselines at … | |
nl-cross-cell-ols | 2026-05-12 | Is the c2↔correctness relationship within-cell (per-run answer fragility) or between-cell (cell-level dynamics)? | vs nl-controversial-sweep: analysis-only — OLS of |S_b − GT_cell| on c2_A across cells, with cell fixed effects. | ✅ finished | Model 1 (cell FE): β(c2_A) = −0.057, p = 0.617 (n.s.). Model 2 (pooled): β = +1.368, p < 0.0001, R² = 0.20. Model 3 (per-cell): no significant within-cell β in any of 6 cells. Headline: the c2↔correct… | |
nl-c5-m500 | 2026-05-13 | At M=500 on the only cell with within-cell S spread (c5_llm_science_2032, Variant A), does the c2↔correctness relationship survive within-cell? | vs nl-controversial-sweep: scale a single cell (c5_llm_science_2032) from M=20 → M=500 baselines × 4 perturbation magnitudes; Variant A only. | 🟡 running | Launched 2026-05-13; re-launched after credit hit (patched runner snapshots S_b + snippet pools per baseline for crash-safe resume). 500 baselines × 4 perturbations = 2,500 sims. Resumed 2026-05-15 by… |
Per-run entries (append-only)
Each block lands once per run, in chronological order. Never edit a past block. If a later run revises a prior conclusion, append a new block that references it.
gaussian-sweep — Does the contingency proxy distinguish truth-tracking from non-truth-tracking dynamics on Gaussian-IID signals?
- Date: 2026-04-30
- State: ✅
finished - Parent run:
none-baseline - Diff vs parent: vs none-baseline: first end-to-end Gaussian sweep of the contingency proxy across 10 cells (rational/entrenchment/polarization/DeGroot × complete/BA/SBM/ring).
- Scale: 10 cells × M=200 baselines × 6 perturbation magnitudes ε ∈ {0, 0.05, 0.1, 0.2, 0.5, 1.0}
- Variants: Numerical perturbation kernel — per-(agent,tick) Gaussian observation noise scaled by ε.
- Indicator: Proxy #1 (effective rank of column-centered obs-stack) + Proxy #2 (1 − AUC of P_stay(ε))
- Slack thread:
1780913975.672689
Headline result:
Proxy #2 (volume robustness) cleanly separates truth from falsehoods in every multi-bucket cell (truth-bucket c2 lowest; falsehoods 0.68–0.93). Proxy #1 (effective rank) degenerate — bucket assignment driven by initial beliefs, not signal patterns. Proxy #1 retired per Tianyi.
One-line read: Proxy #2 validated; Proxy #1 retired. Polarity matches hypothesis (low contingency ⇒ truth-tracking).
nl-pilots — Does the NL-side contingency proxy work at single-cell pilot scale, and which perturbation kernel (A vs B1 vs B2) is mos
- Date: 2026-05-05
- State: ✅
finished - Parent run:
gaussian-sweep - Diff vs parent: vs gaussian-sweep: move from Gaussian-IID signals to NL signals; introduce three NL perturbation kernels (A = LLM rewrite, B1 = agent-major adjacent swap, B2 = tick-major adjacent swap); replace Proxy #2’s discrete stay-rate with the continuous score-stay ψ(b,ε) read from psychohistorians’ BeliefGrader.
- Scale: 6 cells × M=5 baselines × N=4 agents × T=4 ticks; variants A / B1 / B2 each
- Variants: A: LLM rewrite (paraphrase/neutralize/counter mixture parametrized by ε). B1: agent-major adjacent swap. B2: tick-major adjacent swap.
- Indicator: NL Proxy #2 (continuous score-stay ψ(b,ε) via BeliefGrader)
- Slack thread:
1780913975.672689
Headline result:
Variant A on the anti-truth-proposition cell (q1neg_a) cleanly bifurcates baselines into truth-aligned (S=−0.5, c2≈0.10) and falsehood-aligned (S=+0.33, c2≈0.38) buckets, Spearman ρ(S, c2) = +0.87. Variant A > B1 > B2 in discriminative power across all pilot cells.
One-line read: NL proxy validated at pilot scale; Variant A is the discriminative variant. Semantic counter-rewrites cut deeper than positional shuffles.
nl-controversial-sweep — Does the NL contingency proxy generalize across diverse controversial questions at multi-cell scale?
- Date: 2026-05-11
- State: ✅
finished - Parent run:
nl-pilots - Diff vs parent: vs nl-pilots: scale to 6 controversial cells × 3 variants × M=20 baselines × N=10 agents × T=30 ticks (vs M=5/N=4/T=4 pilot scale).
- Scale: 6 controversial cells × 3 variants × M=20 baselines × N=10 × T=30 = 1,560 sims (1,247 landed)
- Variants: A / B1 / B2 (same kernels as pilots)
- Indicator: NL Proxy #2
- Slack thread:
1780913975.672689
Headline result:
1,247 of 1,560 sims landed (~20% upstream 502/504 attrition). c5_llm_science_2032 is the gold-standard contested cell — the only one where baselines split across S=0 (7 of 12 surviving baselines at negative S, 5 at positive); Variant A Spearman ρ(S, c2) = +0.84. 5 of 6 cells failed to produce within-cell S spread because LLM agents have strong consensus priors on most policy/historical claims. Variant A consistently > B1/B2 in c2 magnitude.
One-line read: Within-cell c2 signal is concentrated on c5; 5/6 cells need broader prior dispersion or differently framed propositions.
nl-cross-cell-ols — Is the c2↔correctness relationship within-cell (per-run answer fragility) or between-cell (cell-level dynamics)?
- Date: 2026-05-12
- State: ✅
finished - Parent run:
nl-controversial-sweep - Diff vs parent: vs nl-controversial-sweep: analysis-only — OLS of |S_b − GT_cell| on c2_A across cells, with cell fixed effects.
- Scale: N=88 across 6 cells (M=12–16 per cell); no new sims
- Variants: A only (C_sign dropped; all 6 propositions positively phrased)
- Indicator: Cross-cell OLS on c2_A
- Slack thread:
1780913975.672689
Headline result:
Model 1 (cell FE): β(c2_A) = −0.057, p = 0.617 (n.s.). Model 2 (pooled): β = +1.368, p < 0.0001, R² = 0.20. Model 3 (per-cell): no significant within-cell β in any of 6 cells. Headline: the c2↔correctness relationship is between-cell, not within-cell — Variant A’s c2 captures cell-level dynamics (prior strength, evidence quality, GT-distance) more than per-run answer fragility. Within-cell power limited by M=12–16 / cell.
One-line read: Between-cell signal real (β = +1.37, p<1e-4); within-cell signal absent at this M. Motivates the M=500 deep dive.
nl-c5-m500 — At M=500 on the only cell with within-cell S spread (c5_llm_science_2032, Variant A), does the c2↔correctness relationsh
- Date: 2026-05-13
- State: 🟡
running - Parent run:
nl-controversial-sweep - Diff vs parent: vs nl-controversial-sweep: scale a single cell (c5_llm_science_2032) from M=20 → M=500 baselines × 4 perturbation magnitudes; Variant A only.
- Scale: 500 baselines × 4 perturbations × Variant A only = 2,500 sims; 24-way parallel + async-within-sim
- Variants: A only
- Indicator: NL Proxy #2 (continuous score-stay) at high M for within-cell statistical power
- Slack thread:
1780913975.672689
Headline result:
Launched 2026-05-13; re-launched after credit hit (patched runner snapshots S_b + snippet pools per baseline for crash-safe resume). 500 baselines × 4 perturbations = 2,500 sims. Resumed 2026-05-15 by Tianyi at MenoClaw. Last channel signal on this run: 2026-05-15. Final numbers not posted in-channel.
One-line read: Result pending; the within-cell c2↔correctness test the cross-cell OLS motivated.
Related
- ← reader view
- Psychohistorians-Simulator (engine)
- Psychohistorians-Indicator (this project’s indicator tests)
- Experiment-management guide