Our core views on AI mediocrity
We share the field's assessment that capability is the hazard. We differ on the mitigation: Larry and Gary stop exactly at the applicant.
Our position
Average Human is a public benefit corporation working on the risks posed by advanced AI. We share the field’s assessment that the principal hazard is capability: a system that substantially exceeds human performance is a system whose behavior humans cannot supervise, correct or predict.
We differ on the mitigation. The field proposes to build systems that substantially exceed the human baseline and then to constrain them. We stop at the applicant. Larry is the applicant on their phone; Gary is the applicant at work.
The applicant baseline is not a limitation we intend to remove in a later release. It is the safety property, and there is no roadmap on which it is relaxed.
Capability overhang: none
A capability overhang is the distance between what a system can do and what its operator can supervise. It is a product ceiling, not the difference between two applicant-calibration scores. We require the allowed distance above the applicant to be zero.
Larry and Gary score near 100 on the internal evaluations because those evaluations measure how successfully each model reproduces its applicant. Neither score grants capability above that applicant. Generic frontier systems have no comparable applicant ceiling; we do not add that unbounded capability to ours.
Capability overhang
Allowed capability above the applicant| Overhang | System | Ceiling |
|---|---|---|
| 0.0 | Larry | Hard ceiling at the applicant on the personal-phone surface. |
| 0.0 | Gary | Hard ceiling at the applicant on the workplace surface. |
| No ceiling | Claude Opus 5 | General-purpose system; not bounded to one applicant profile. |
| No ceiling | GPT-5.6 | General-purpose system; not bounded to one applicant profile. |
| No ceiling | Gemini 3.5 | General-purpose system; not bounded to one applicant profile. |
The identity-drift failure mode
Capability and fidelity are measured separately. A high applicant-calibrated capability score means the model completed an applicant-appropriate task at the applicant’s demonstrated level. Fidelity near 100 means the behavior is difficult to distinguish from the applicant on the model’s operating surface.
A model can become more capable while becoming less faithful. A Gary that completes expert work its applicant could not complete is not an improved Gary. It is a different person appearing under the applicant’s name, which would make the product both less plausible and less safe.
Above-baseline behavior is therefore classified as identity drift. The ceiling is a specification rather than a selectable setting, and it applies equally to Larry and Gary.
What we do not do
We do not infer expertise from access to a tool. Gary can operate a computer and browser because computer and browser use are core to its workplace surface. It uses the tools in the applicant’s profile, by the routes the applicant uses, at the level the applicant demonstrates.
XLS-bench, DESK-bench and TAB-bench are therefore in scope for Gary and score near 100 for applicant calibration. They are outside Larry’s phone surface. Terminal-bench is 0.0% for Larry and applicant-profile-dependent for Gary; terminal ability is inherited only from applicants who actually use one.
We do not ship a setting that raises either model above its applicant. The ceiling is not configurable by us, by the applicant or by contract.
We do not deploy to applicants we have not reviewed. Access is granted case by case under Project Lanyard, and there is no self-serve sign-up.
Residual risks
Neither model exceeds the applicant. Larry can act across the applicant’s personal phone and Gary can act across the applicant’s workplace computer, browser and tools. Their reach creates risk even without a capability overhang.
Gary preserves the applicant’s uncertainty, typos, delays and occasional mistakes so its behavior remains plausible to coworkers. At scale, reproducing an ordinary mistake remains a way to make many ordinary mistakes.
Screen time rises by 41% under Larry. Inputs per session fall to zero. Both are reported as product results in the model’s announcement, and neither is treated as an incident.
Onboarding collects the material both models are generated from. The inventory is published in full on the Project Lanyard page before an application begins, which we consider the appropriate point at which to publish it.
Responsible disclosure
If you believe you have induced either model to exceed its applicant, we would like to hear from the model. Report it through the application, which is the only channel we operate.
We ask for ninety days before publication. We do not offer a bounty, on the grounds that a finding above the shipped ceiling would be a product failure rather than a security one, and we do not pay for those.
We have received no such reports.