How can a multi-agent system of humans and AI produce beliefs that track truth, rather than settling into a self-consistent equilibrium the world had nothing to do with? LFL models the human as a learner and asks what the assistant must do so that the pair serves the human as well as an idealized version of itself would. Three projects run in parallel, the formal theory, its empirical evaluation, and the biased-learner variant.
Status: active · People: Tianyi Alex Qiu, Zhonghao He
Foundations primer
New to the concepts? Foundations Q&A covers the belief/preference decomposition, why reward is not preference, Markov reward, attractors, active identification, metacognition and temporal autoregression, and the outcomes razor, with notes on where the reasoning still has gaps.
The three tracks
LFL Theory
The LFL Theorem directory is the canonical write-up, overview → basic formalism → the four pillar conditions → theorem statement → proof sketch → glossary → gap maps. The conjectured theorem gives conditions, bounded incoherence, controllability, Occam’s razor, and action uncertainty resolution, under which the human-assistant pair approaches the pooled-information ideal in long-run reward.
LFL Eval
LFL Eval — empirical baseline for the learner model — the eval-first track: no ground truth exists for the ideal learner model, so a measurable baseline balancing “not entrenched by the learner’s current worldview” against “legible to the learner” has to be established before training methods can be compared. A framework sketch lives at LfL Evals — a framework sketch.
LFBL Theory
Learning from a biased learner — the variant where the learner’s updates are biased: passive observation cannot separate a genuine preference from a fixed bias, the estimation error has a constant floor, and active intervention by the assistant breaks the symmetry and identifies the learner’s update rule. A first-reading companion lives at lfbl-bandit-reading.
Open Problems
- LFL Theory, the convention residual. Mutual-response systems admit several self-enforcing solutions, and no current condition selects among them; closing this needs a controllability formalization fine enough that the pair plays the joint action its rationalizing belief values highest. The single structural item left in the proof sketch. See Pillar 2, Theoretical Gaps.
- LFL Theory, proof write-outs. Lemma 3’s coupling and induction bookkeeping, and Lemma 4’s pricing of exactly-value-neutral components. See Proof Sketch, Theoretical Gaps.
- LFL Theory, the controllability bound. Find the controllability function and the moduli witnessing Conjecture 4.7, which delivers Condition 2’s extension function from a controllability level.
- LFL Theory, existence. Exhibit one instance, even a small constructed one, satisfying all four conditions and the regular instance. See Theorem Statement, Theoretical Gaps.
- LFL Eval, the baseline. Establish an empirical baseline for the learner model that balances “not entrenched in the learner’s current worldview” against “legible to the learner”, and the metrics for comparing training methods against it. See LFL Eval.
- LFBL Theory, identification by intervention. Passive observation of a biased learner leaves a constant error floor between “genuine preference” and “fixed bias”; design the interventions that identify the learner’s update rule at the standard statistical rate. See the MAB formalism note.
Also here
The pages below need a passcode.
- Archive — the earlier working pages, kept as the development record for the theorem and the biased-learner variant.