ProductAnnouncements
Introducing Larry

Larry is available today to applicants approved under Project Lanyard. It is the applicant on their personal phone, and it does the swiping so that they can keep watching.
On the phone-fidelity evaluations — SWIPE-bench, DWELL, TONE-match and TXT-back — Larry is the new state of the art. It also scores between 97.7% and 98.6% on the three shared applicant-calibrated capability evaluations. Workplace evaluations belong to Gary.
Larry can do everything the applicant does on their phone: navigate apps, browse, watch, swipe, like, save, search and reply. It is designed to be used continuously. Removing the last remaining input from short-form consumption raises average screen time by 41%.
Where Larry lands
Twelve evaluations, reported in full| Capability and evaluation | Larry | Claude Opus 5 | GPT-5.6 | Gemini 3.5 |
|---|---|---|---|---|
| ReasoningGPQA Cubic Zirconia | 98.0% | 64.7% | 66.1% | 63.9% |
| Lunch knowledgeMMLU (Lunch) | 98.6% | 71.2% | 72.6% | 70.8% |
| Shared expensesSPLIT | 97.7% | 68.3% | 69.5% | 67.8% |
| Feed advancementSWIPE-bench | 99.2% | 42.1% | 39.8% | 44.3% |
| Dwell-time fidelityDWELL | 96.8% | 38.9% | 36.7% | 41.0% |
| Personal-message toneTONE-match | 98.1% | 61.2% | 59.8% | 63.0% |
| Texting fidelityTXT-back | 97.4% | 71.4% | 70.9% | 73.1% |
| Spreadsheet useXLS-bench | — | 72.4% | 74.1% | 71.7% |
| Desktop workDESK-bench | — | 68.9% | 71.0% | 69.4% |
| Browser workTAB-bench | — | 70.3% | 72.2% | 71.5% |
| Workplace fidelityPTO-bench | — | 34.8% | 37.1% | 33.6% |
| Terminal useTerminal-bench | 0.0% | — | — | — |
Hands-free consumption
Larry advances short-form video feeds — Reels, TikTok, Shorts — while you watch, matching your dwell time and completion patterns closely enough that the feed is indistinguishable from one you scrolled yourself.
The pitch is not that Larry watches instead of you. You keep watching. You do not put the phone down and you do not look away. Larry removes the only remaining input: you no longer have to move your thumb.
Because nothing is asked of you, sessions run longer. We report that as a capability result rather than as a side effect, in the same register as any other number on this page.
Average session length
Instrumented sessions, 30 daysThe applicant is still the one watching. What changes is how much of the session they are required to operate, which is reported below alongside how often they stopped.
What the applicant still has to do
Instrumented sessions, 30 days- Swipes
- Taps
- Ratings
Evaluation results
Larry was scored against the four mobile evaluations. Each measures agreement with what you would have done yourself, which is the only bar Larry is built to clear — and unlike a capability score, it does not decay as the session runs long.
Agreement with you by session length
SWIPE-bench- Larry
- Claude Opus 5
- GPT-5.6
- Gemini 3.5
Agreement with you by session length
DWELL- Larry
- Claude Opus 5
- GPT-5.6
- Gemini 3.5
Agreement with you by session length
TONE-match- Larry
- Claude Opus 5
- GPT-5.6
- Gemini 3.5
- On SWIPE-bench, which measures agreement with the moment you would have swiped, Larry scores 99.2%. The remaining 0.8% is drawn from sessions in which you would have stopped.
- On DWELL, dwell-time fidelity against your own baseline, Larry scores 96.8% and holds that across the full length of a session.
- On TONE-match, personal messages and DMs indistinguishable from your own, Larry scores 98.1%, including slang and capitalization.
- On TXT-back, reply latency to a message from Mom, Larry scores 97.4% — measured against your latency, not against a prompt reply.
On Terminal-bench, Larry scores 0.0%. This is intentional. XLS-bench, DESK-bench, TAB-bench and PTO-bench are reported for Gary and shown as not applicable for Larry because they measure a workplace surface. Larry’s mobile browser remains in scope. See the full evaluation suite for the full split by operating surface.
Rating on your behalf
Larry rates on your behalf — thumbs up and thumbs down, inferred from the profile built during onboarding. The recommendation algorithm continues to receive signal. The feed continues to tune itself. You contribute nothing and receive a feed increasingly precisely your own.
It sounds like you
Larry replies in your own register: slang, punctuation habits, capitalization, the words you overuse. That register is derived from the communications imported during onboarding, which is the literal mechanism behind the promise that it sounds like you. It sounds like you because it read everything you have ever written.
Dating apps are a supported secondary surface. They are not the primary use case, and Larry is not optimized for them.
Alignment and safety
Alignment. Larry exhibits no capability overhang. On shared capability tasks it stays at the applicant baseline; on phone-fidelity tasks it is scored on how closely it reproduces you rather than on how far it exceeds you. There is no setting that raises it above your baseline.
Safety. Larry does not advance the frontier in any risky domain, because it does not advance the frontier. Our position on why that is the safety property, rather than an absence of one, is set out in our core views.
Availability
Larry is available today to approved applicants only. There is no self-serve sign-up and access is granted by review. Onboarding establishes the profile Larry is generated from; Project Lanyard sets out what that involves before you begin.
Settlement is handled under the beneficiary terms, which are irrevocable and survive the user. See Pricing.
Footnotes
Mobile suite, all charts: Twenty-five trials at room temperature, mean agreement over five sessions per applicant. Every model received the same frozen applicant profile and phone-native tools. Scores measure applicant calibration and are self-reported, as are everyone else’s.
Average session length: Measured against the applicant’s own pre-onboarding baseline, which onboarding also collects. Error bars are one standard error over twenty-five applicants. Sessions terminated by battery exhaustion are counted in full.
Inputs and endings: Instrumented over thirty days on the applicant’s own device. An input is any deliberate contact with the screen; unlocking the phone counts as two. Sessions ended by the applicant exclude those ended by the battery.