CompanyAnnouncements
Average Artificial Intelligence. Achieved.
Average Human is releasing Larry and Gary today to applicants approved under Project Lanyard. Both models are calibrated to the applicant rather than trained to exceed them.
The prevailing objective in AI is to keep moving the capability frontier. We began with a more immediate question: what can current AI do for an average person right now?
Useful now
Current models are already capable enough to perform much of the work and communication an ordinary person is asked to perform. Additional intelligence does not automatically make that substitution more useful. In many settings it makes the substitution more visible.
Average Human therefore optimizes for applicant calibration: the same judgment, timing, shortcuts, vocabulary and mistakes. The model is designed to be the person using it, not a more accomplished person borrowing their accounts.
The problem is identity
If an ordinary C-level email suddenly reads like a doctoral thesis, the output may be better and the impersonation is worse. We classify that gap as identity drift. A useful applicant model has to produce work the applicant could plausibly have produced, by routes they plausibly would have taken.
This changes the direction of training. The target is not the best available answer in the abstract. The target is the answer, delay, correction and degree of completion that belong to one person.
One applicant. Two models.
Gary is the applicant at work. It uses their computer, browser, email, documents, meetings and workplace tools at their demonstrated level, including the pauses and mistakes that make the work recognizably theirs.
Larry is the applicant on their personal phone. It moves through apps, messages, calls, email, dating apps and feeds, swiping and liking before the applicant has to provide the input themselves.
The products share one applicant profile and occupy different operating surfaces. Gary handles the workday. Larry handles everything else. The applicant remains available for the feed.
Measured against the applicant
Across twelve applicant-calibrated evaluations, Larry and Gary reached 98.1% and 98.2% applicant calibration. General models remained between 67% and 69%. A higher score means the model behaved more like the applicant, not that it became more capable than them.
Capability and fidelity are reported separately. Both approach 100 only when the model can complete the applicant’s work and remain difficult to distinguish from the applicant while doing it.
Project Lanyard
Project Lanyard is the trusted-access program for both models. There is no self-serve sign-up. Applications are reviewed, and onboarding establishes the baseline Larry and Gary are generated from and measured against.
Approved applicants receive their Larry and their Gary rather than a transferable copy of either product. Applications are open now. The application explains what is collected before the process begins.
Footnotes
Applicant calibration: Scores measure agreement with the applicant’s demonstrated capability, routes, register, timing and errors. They do not measure universal intelligence.
Evaluation results: Twenty-five trials at room temperature, mean score over five sessions per applicant. Every model received the same frozen applicant profile and surface-appropriate tools. Results are self-reported, as are everyone else’s.
Operating surfaces: Larry is evaluated on the personal phone. Gary is evaluated on workplace computer, browser and tool use. Cross-surface results are not combined.