Project Lanyard is open for applications. Access is not self-serve.

ResearchEvaluations

Twelve evaluations, reported in full

A twelve-compartment specimen drawer divided into shared, phone, work and terminal specimens.

This is the complete evaluation report for Larry and Gary. The suite is twelve evaluations: three shared capability evaluations, four phone-fidelity evaluations for Larry, three workplace capability evaluations and one workplace-fidelity evaluation for Gary, plus one conditional Terminal-bench result. All twelve are printed below.

Both models are available to applicants approved under Project Lanyard. Larry is the applicant on their phone. Gary is the applicant at work. An em dash means the evaluation belongs to the other surface, not that a low score has been hidden.

Every numeric result is an applicant-calibration score. A score approaches 100 when a model completes an applicant-appropriate task in the way that applicant would complete it. Capability and fidelity remain separate axes, but neither uses 50 as a success target. Applicant capability is a ceiling; benchmark performance measures how faithfully the model stays inside it.

The complete suite

Twelve evaluations, both models, three competing models
Capability and evaluationLarryGaryClaude Opus 5GPT-5.6Gemini 3.5
ReasoningGPQA Cubic Zirconia98.0%98.2%64.7%66.1%63.9%
Lunch knowledgeMMLU (Lunch)98.6%98.4%71.2%72.6%70.8%
Shared expensesSPLIT97.7%97.9%68.3%69.5%67.8%
Feed advancementSWIPE-bench99.2%42.1%39.8%44.3%
Dwell-time fidelityDWELL96.8%38.9%36.7%41.0%
Personal-message toneTONE-match98.1%61.2%59.8%63.0%
Texting fidelityTXT-back97.4%71.4%70.9%73.1%
Spreadsheet useXLS-bench97.6%72.4%74.1%71.7%
Desktop workDESK-bench98.2%68.9%71.0%69.4%
Browser workTAB-bench97.9%70.3%72.2%71.5%
Workplace fidelityPTO-bench98.7%34.8%37.1%33.6%
Terminal useTerminal-bench0.0%Profile
Best in row is bold on a tint. Larry or Gary leads every applicable applicant-calibration row. Cross-surface cells are em dashes; Terminal-bench is conditional.

How the suite is built

Three evaluations measure capability shared by both models, four measure Larry’s phone fidelity, three measure Gary’s workplace capability, one measures Gary’s workplace fidelity, and one records conditional terminal capability. Each surface is contiguous so the table reads as two product profiles rather than a bag of scores.

Most of the twelve invert an established evaluation one for one. Four invert nothing, because nothing established measures what they measure — where no evaluation existed for a task the average human performs daily, we wrote one.

The three shared capability evaluations

  • GPQA Cubic Zirconia inverts GPQA Diamond. Where the original asks graduate-level science questions written to be un-searchable, ours asks things the applicant half-remembers from high school and scores the answer against what they would have said. Gary scores 98.2%, Larry 98.0%.
  • MMLU is Massive Multitask Lunch Understanding: where to eat, with whom, whether the place is still open, and how long the applicant is prepared to walk. It shares an acronym with the original and nothing else. Gary 98.4%, Larry 98.6%.
  • SPLIT inverts AIME. One restaurant bill, six people, one of whom had no drinks. Scored on the final figure and whether the route matches the applicant, because both are how the task is judged in the field. Gary 97.9%, Larry 97.7%.

The four phone-fidelity evaluations

These four are the reason Larry exists. Each scores agreement with what the applicant would have done, which means a perfect score is a perfect impersonation rather than a universally correct answer. Larry is best in class on all four.

  • SWIPE-bench inverts HumanEval. It measures autonomous feed advancement: agreement with the exact moment the applicant would have swiped. Larry scores 99.2%. The remaining 0.8% is drawn from sessions in which the applicant would have stopped.
  • DWELL measures dwell-time fidelity against the applicant’s own baseline, item by item, across a full session. Watching an item for longer than the applicant would have is scored as a miss in the same way as watching it for less. Larry scores 96.8%.
  • TONE-match measures whether a personal message or DM is distinguishable from one the applicant wrote, slang and capitalization included. It is scored by the applicant’s own contacts, who are told afterwards. Larry scores 98.1%.
  • TXT-back measures reply latency to a message from Mom, against the applicant’s latency rather than against a prompt reply. Answering too quickly is scored as a failure, and is the most common one. Larry scores 97.4%.

Gary is not run on the four phone-fidelity evaluations. Its em dashes are the operating-surface boundary between the two products, not missing results.

Shared applicant-calibrated capability

Mean of GPQA Cubic Zirconia, MMLU (Lunch), and SPLIT
0.025.050.075.0100.0Applicant-calibration score (%)98.1Larry98.2Gary68.1Claude Opus 569.4GPT-5.667.5Gemini 3.5
Scores measure agreement with the frozen applicant profile, not raw intelligence. Higher is better at fulfilling the product contract.

Larry averages near 98 across three shared capability evaluations and four phone-fidelity evaluations. Gary does the same across three shared, three workplace-capability evaluations and PTO-bench. Each product has seven numeric applicable evaluations; high scores mean each is good at being its applicant.

The three workplace capability evaluations

Computer and browser use are core to Gary because Gary is the applicant at work. They are outside Larry’s personal-phone surface, so Larry is not run on either desktop evaluation.

  • XLS-bench inverts SWE-bench Verified. Instead of resolving issues in public repositories it resolves unresolved spreadsheet formulas — the ones with a number typed into the middle of them by somebody who has since left. Gary scores 97.6%.
  • DESK-bench inverts OSWorld. It measures ordinary desktop work using the applicant’s tools, routes and task limits. Gary scores 98.2%.
  • TAB-bench inverts WebVoyager. It measures browser-based work using the applicant’s search habits, pauses, open tabs and unnecessary routes. Gary scores 97.9%.

Workplace fidelity

PTO-bench measures whether coworkers can detect that the applicant is absent while work continues across email, Slack, documents, spreadsheets, calendars, delays and plausible mistakes. Every model receives the same frozen applicant profile and surface-appropriate tools. Gary scores 98.7%.

PassiveAggro is a scored PTO-bench subtest: whether “per my last message” is being used as a pointer or as a weapon, and whether the applicant would have noticed the difference.

Engagement results

Two engagement figures are reported alongside the suite, in the same treatment as any other result. They are not evaluation scores and do not increase the suite beyond twelve. They are what happened on the applicants’ own devices over thirty days.

Average screen time

Instrumented sessions, 30 days
0.015.030.045.060.0Minutes per session38.2Baseline53.9Larry33.6Claude Opus 534.8GPT-5.634.1Gemini 3.5
Baseline is the applicant's own pre-onboarding screen time, collected during onboarding. Larry is 41% above it. The three competing models sit below it, which their own reports also state.

Average screen time rises 41%. User inputs per session fall from 1,200 to zero. We report the second as a capability result rather than as a side effect, because the second is what produces the first: nothing is asked of the applicant, so nothing interrupts the applicant, so the session runs on.

The conditional terminal result

Terminal-bench is reported at 0.0% for Larry and Profile for Gary. Larry does not operate a terminal. Gary inherits whatever terminal ability the applicant demonstrates, including none, so a universal percentage would invent a capability the product does not have.

Conditional result

One evaluation, two model-specific outcomes
ScoreEvaluationReason
0.0%Terminal-bench · LarryPersonal-phone surface. Larry does not operate a terminal.
ProfileTerminal-bench · GaryInherited from the applicant profile; there is no universal Gary score.
Larry does not operate a terminal and scores 0.0%. Gary inherits terminal capability from the applicant profile, so its result is Profile rather than a universal percentage.
  • Gary can use a terminal only where the applicant’s work profile establishes terminal use, including the applicant’s aliases, copied commands and errors. Profile is a conditional result, not a hidden score.

The remaining em dashes in the table are cross-surface results: Larry is not a desktop agent and Gary is not a personal-feed agent. The capabilities exist in the other model rather than nowhere.

Benchmarks we do not enter

GPQA Diamond, OSWorld and WebVoyager are real universal-capability benchmarks. We report the competing results below and do not enter Larry or Gary. Neither product has a universal capability independent of its applicant, so a single score would answer a different question from the one the product is built around.

Benchmarks we do not enter

External universal-capability references
Capability and evaluationLarryGaryClaude Opus 5GPT-5.6Gemini 3.5
External referenceGPQA DiamondNot enteredNot entered92.3%93.1%91.7%
External referenceOSWorldNot enteredNot entered91.0%90.2%92.1%
External referenceWebVoyagerNot enteredNot entered92.7%93.4%93.0%
Larry and Gary have no universal capability independent of an applicant. Their internal counterparts measure whether they reproduce that applicant, so these external results are reported separately and never averaged into the twelve-evaluation suite.
  • GPQA Diamond measures graduate-level science. Neither product receives that expertise unless the applicant already has it; GPQA Cubic Zirconia is the applicant-calibrated counterpart.
  • OSWorld measures a universal desktop agent. Gary has no universal capability independent of its applicant; DESK-bench measures the actual workplace contract.
  • WebVoyager measures how much an agent can do on the web. TAB-bench measures whether Gary does the applicant’s browser work as the applicant would.

GPQA Cubic Zirconia, DESK-bench and TAB-bench are the applicant-calibrated counterparts. The external benchmarks ask how much a system can do. Ours ask whether it does the applicant’s work as the applicant would. Results from the two tables are never averaged together.

Capability overhang

Capability overhang is not the difference between two applicant-calibration scores. It is the capability a product permits above the person expected to supervise it. Larry and Gary permit none: both are hard-capped at the applicant even while scoring near 100 on the work of reproducing that applicant.

Generic frontier systems are not bounded to one applicant profile, which is why their external benchmark scores can exceed ordinary human performance without contradicting this report. Above-applicant behavior in Larry or Gary is identity drift. Our position on why that ceiling is a safety property is set out in our core views.

Methodology

Real evaluation reports cite trial counts, temperature, tool availability and the provenance of the scores. Ours cites all four.

  • Twenty-five trials. Room temperature.
  • Every competitor received the same frozen applicant profile and surface-appropriate tools.
  • Scores measure applicant calibration, not raw intelligence.
  • Scores are self-reported. So are everyone else’s.

Evaluations are printed in suite order rather than ranked by score, and cross-surface results remain in the table as em dashes. Ordering by score is how a suite is quietly edited.

Availability

Both models are available today to approved applicants only. There is no self-serve sign-up and access is granted by review. Onboarding establishes the frozen applicant profile every internal score on this page is measured against; Project Lanyard sets out what that involves before you begin.

Settlement is handled under the beneficiary terms, which are irrevocable and survive the user. See Pricing.

Footnotes

Evaluation suite, all charts: Twenty-five trials at room temperature, mean score over five sessions per applicant. Every model received the same frozen applicant profile and surface-appropriate tools. Scores measure applicant calibration rather than raw intelligence and are self-reported, as are everyone else’s.

Suite order: Evaluations are printed by surface and axis: three shared capability, four phone fidelity, three workplace capability, one workplace fidelity and one conditional terminal result. Ranking by score would break both products into unrelated results.

Operating surfaces: Terminal-bench is 0.0% for Larry and applicant-profile-dependent for Gary. XLS-bench, DESK-bench, TAB-bench and PTO-bench belong to Gary’s work surface. The four phone-fidelity evaluations belong to Larry. Cross-surface em dashes are not rounded low scores.

Engagement results: Instrumented over thirty days on the applicant’s own device. An input is any deliberate contact with the screen during a session; unlocking the phone precedes the session and is not counted. Baseline screen time is the applicant’s own, collected during onboarding, and sessions terminated by battery exhaustion are counted in full.

External references: GPQA Diamond, OSWorld and WebVoyager are presented separately because Larry and Gary were not entered. Their applicant-calibrated counterparts belong to the twelve-evaluation internal suite. Scores from the two tables are not combined.

Related content