Project Lanyard is open for applications. Access is not self-serve.

Research

We publish what we find, including the results that argue against building anything further. Our program is small, and the papers below are all of it.

Publications

Listed most recent first. Every paper below is a technical report or a data statement; none has been submitted for peer review, on the grounds that the results are not contested.

  • AlignmentTechnical report

    Capability Overhang: None (Technical Report)

    We separate applicant-calibration performance from the capability a product permits above its operator. Larry and Gary each approach perfect calibration across seven applicable evaluations in a twelve-row suite while permitting zero capability above the applicant. We argue that a system which remains exactly inside the operator’s capability is preferable to one which does not.

    D. Whitlock, K. Ferreira, G. Osei, and the Alignment Team

  • EvaluationTechnical report

    Scaling Laws for Adequacy

    Performance on our suite is flat in parameters, flat in data and flat in compute. We fit the curve anyway. The resulting law has no exponent, which we report as a finding rather than as a failure to converge, and we discuss the implications for capital allocation.

    K. Ferreira, M. Aldridge, and P. Nakamura

  • EvaluationTechnical report

    Emergent Mediocrity in Large Models

    Below a threshold, models are worse than the median human. Above it, they are better. We characterize the region in between, show that it is narrower than previously assumed, and demonstrate that a model can be held inside it indefinitely given sufficient engineering effort.

    G. Osei, R. Lindqvist, and the Evaluation Team

  • EvaluationTechnical report

    On the Measurement of Good Enough

    Existing benchmarks score against a universal ground truth, which tells you how far a model is from correct and nothing about how far it is from its user. We propose applicant-calibration scores that approach 100 as the result becomes indistinguishable from what the applicant would have produced, and release twelve evaluations built on that principle.

    M. Aldridge and D. Whitlock

  • AlignmentTechnical report

    Constitutional Indifference

    Constitutional methods train a model against a written set of principles. We train ours against a written set of principles it is not required to hold. The resulting model declines to arbitrate, defers to the user on questions of fact and value alike, and is rated substantially less insufferable than models trained to care.

    P. Nakamura, D. Whitlock, and the Alignment Team

  • DataData statement

    Training on the Inbox: A Data Statement

    A complete statement of the corpus behind Larry and Gary: what was collected during onboarding, from which surfaces, at what granularity, and with what retention. We are not aware of a comparable statement from another laboratory, and we attribute that to the difficulty of writing one.

    R. Lindqvist, K. Ferreira, and the Data Team

  • InterpretabilityTechnical report

    Reading the Register: Idiolect Transfer at Small Scale

    We identify the components responsible for reproducing an applicant’s slang, punctuation habits and capitalization, and show that they are a small and localized fraction of the network. Ablating them removes the resemblance without affecting any score on our suite, which we take as evidence that the resemblance is not a capability.

    R. Lindqvist and G. Osei

Areas

  • Alignment

    2 papers

  • Evaluation

    3 papers

  • Data

    1 paper

  • Interpretability

    1 paper

The evaluation suite the reports above are scored on is published in full, together with the phone and workplace operating surfaces and the conditional terminal result. See the evaluations report and our core views.