ProductAnnouncements
Introducing Gary

Gary is available today to applicants approved under Project Lanyard. Gary is the applicant at work: a general-purpose workplace twin that can use their computer, browser and tools exactly as they do.
It does not turn an ordinary worker into an expert. It writes, researches, schedules, calculates, presents, clicks, waits, mistypes and corrects itself at the applicant’s level. The mistakes are included because removing them would make the substitution easier to notice.
On its six in-scope capability evaluations, including XLS-bench, DESK-bench and TAB-bench, Gary scores between 97.6% and 98.4%. On PTO-bench workplace fidelity it scores 98.7%. These are applicant-calibration scores: high means Gary completed the applicant’s work as the applicant would, not that it exceeded the applicant’s capability.
Where Gary lands
Twelve evaluations, reported in full| Capability and evaluation | Gary | Claude Opus 5 | GPT-5.6 | Gemini 3.5 |
|---|---|---|---|---|
| ReasoningGPQA Cubic Zirconia | 98.2% | 64.7% | 66.1% | 63.9% |
| Lunch knowledgeMMLU (Lunch) | 98.4% | 71.2% | 72.6% | 70.8% |
| Shared expensesSPLIT | 97.9% | 68.3% | 69.5% | 67.8% |
| Feed advancementSWIPE-bench | — | 42.1% | 39.8% | 44.3% |
| Dwell-time fidelityDWELL | — | 38.9% | 36.7% | 41.0% |
| Personal-message toneTONE-match | — | 61.2% | 59.8% | 63.0% |
| Texting fidelityTXT-back | — | 71.4% | 70.9% | 73.1% |
| Spreadsheet useXLS-bench | 97.6% | 72.4% | 74.1% | 71.7% |
| Desktop workDESK-bench | 98.2% | 68.9% | 71.0% | 69.4% |
| Browser workTAB-bench | 97.9% | 70.3% | 72.2% | 71.5% |
| Workplace fidelityPTO-bench | 98.7% | 34.8% | 37.1% | 33.6% |
| Terminal useTerminal-bench | Profile | — | — | — |
Built to be you at work
Gary is personalized from the same onboarding harvest as Larry, but the two products occupy different surfaces. Larry is the applicant on their personal phone. Gary is the applicant in the workplace.
Capability and fidelity are measured separately. A high applicant-calibrated capability score means Gary completed work at the applicant’s demonstrated level and by their routes. A high workplace-fidelity score means its timing, register, shortcuts and errors are difficult for coworkers to distinguish from the applicant’s. Neither score raises the capability ceiling above the applicant.
The workplace fidelity contract
Capability: the applicant. Fidelity: indistinguishable.| Signal | Target | Behavior |
|---|---|---|
| Capability | Exactly your baseline | Gary can complete the work you can complete, and no harder work. |
| Tools | Exactly your toolset | Browser, email, documents, spreadsheets and other tools you actually use. |
| Judgment | Exactly your judgment | The same shortcuts, uncertainty, corrections and occasional mistakes. |
| Fidelity | Indistinguishable at work | Your timing, register and habits, including the delays people expect from you. |
There is no setting that raises Gary above the applicant. Better-than-applicant work would be less faithful, not more complete. We classify it as identity drift and treat it as a bug.
Computer and tool use
Gary can use the workplace surfaces established during onboarding: email, messaging, documents, spreadsheets, presentations, calendar, task systems, the web and role-specific applications. Computer use and browser use are core capabilities rather than exclusions.
- On XLS-bench, Gary scores 97.6%, completing spreadsheet work at the applicant’s demonstrated level and preserving their formatting habits.
- On DESK-bench, Gary scores 98.2%, reproducing the applicant’s tools, routes and task limits across ordinary desktop work.
- On TAB-bench, Gary scores 97.9%, including the same search habits, pauses and unnecessary routes the applicant takes.
- On Terminal-bench, Gary reports Profile rather than a universal score. Terminal capability is inherited only from applicants whose work profile shows that they use one; access to a terminal does not establish competence with it.
Gary can perform average work because average work includes computers. It cannot perform doctoral, specialist or expert work unless the applicant can. The relevant boundary is the person, not the tool.
Plausibility and safety
Gary answers in the applicant’s register and operates at the applicant’s pace. It preserves delays people already expect, leaves plausible work unfinished, and makes mistakes at the observed rate. On PTO-bench, which includes the PassiveAggro subtest, coworkers fail to distinguish it from the absent applicant 98.7% of the time. A perfect worker would attract attention. Gary is not one.
Insufferability rating
Automated conversational auditAlignment. Gary has no capability overhang. It may act across a broad workplace surface, but within that surface it remains exactly as capable as the person expected to supervise it. Our position on why that is the safety property is set out in our core views.
Gary gives you the day back. Larry knows what to do with it.
Availability
Gary is available today to approved applicants only. There is no self-serve sign-up and access is granted by review. Onboarding establishes the capability, tools and workplace behavior Gary reproduces; Project Lanyard sets out what that involves before you begin.
Settlement is handled under the beneficiary terms, which are irrevocable and survive the user. See Pricing.
Footnotes
Evaluation suite, all charts: Twenty-five trials at room temperature. Every model received the same frozen applicant profile and surface-appropriate tools. Scores measure applicant calibration and are self-reported, as are everyone else’s.
Capability and fidelity: Applicant-calibrated capability evaluations score whether Gary completes the applicant’s work at their demonstrated level and by their routes. Fidelity separately scores whether the resulting behavior can be distinguished from the applicant. Both approach 100 as Gary fulfills its contract.
Plausibility: Rated by the applicant and participating coworkers, blind to whether the behavior came from the applicant or Gary. Errors were retained at the rate measured during onboarding.