Project Lanyard is open for applications. Access is not self-serve.

ProductAnnouncements

Introducing Larry

Antique illustrated specimens spelling Larry.

Larry is available today to applicants approved under Project Lanyard. It is the applicant on their personal phone, and it does the swiping so that they can keep watching.

On the phone-fidelity evaluations — SWIPE-bench, DWELL, TONE-match and TXT-back — Larry is the new state of the art. It also scores between 97.7% and 98.6% on the three shared applicant-calibrated capability evaluations. Workplace evaluations belong to Gary.

Larry can do everything the applicant does on their phone: navigate apps, browse, watch, swipe, like, save, search and reply. It is designed to be used continuously. Removing the last remaining input from short-form consumption raises average screen time by 41%.

Where Larry lands

Twelve evaluations, reported in full
Capability and evaluationLarryClaude Opus 5GPT-5.6Gemini 3.5
ReasoningGPQA Cubic Zirconia98.0%64.7%66.1%63.9%
Lunch knowledgeMMLU (Lunch)98.6%71.2%72.6%70.8%
Shared expensesSPLIT97.7%68.3%69.5%67.8%
Feed advancementSWIPE-bench99.2%42.1%39.8%44.3%
Dwell-time fidelityDWELL96.8%38.9%36.7%41.0%
Personal-message toneTONE-match98.1%61.2%59.8%63.0%
Texting fidelityTXT-back97.4%71.4%70.9%73.1%
Spreadsheet useXLS-bench72.4%74.1%71.7%
Desktop workDESK-bench68.9%71.0%69.4%
Browser workTAB-bench70.3%72.2%71.5%
Workplace fidelityPTO-bench34.8%37.1%33.6%
Terminal useTerminal-bench0.0%
Larry has seven numeric applicable evaluations: three shared applicant-calibrated capability and four phone fidelity. Terminal-bench is an intentional 0.0%; Gary's four workplace evaluations appear as em dashes.

Hands-free consumption

Larry advances short-form video feeds — Reels, TikTok, Shorts — while you watch, matching your dwell time and completion patterns closely enough that the feed is indistinguishable from one you scrolled yourself.

The pitch is not that Larry watches instead of you. You keep watching. You do not put the phone down and you do not look away. Larry removes the only remaining input: you no longer have to move your thumb.

Because nothing is asked of you, sessions run longer. We report that as a capability result rather than as a side effect, in the same register as any other number on this page.

Average session length

Instrumented sessions, 30 days
0.015.030.045.060.0Minutes per session38.2Baseline53.9Larry33.6Claude Opus 534.8GPT-5.634.1Gemini 3.5
Baseline is the applicant's own pre-onboarding session length, collected during onboarding. Error bars are one standard error over 25 applicants.

The applicant is still the one watching. What changes is how much of the session they are required to operate, which is reported below alongside how often they stopped.

What the applicant still has to do

Instrumented sessions, 30 days
Inputs required per session02505007501000Inputs (count)2Larry886Claude Opus 5882GPT-5.6883Gemini 3.5Sessions ended by the applicant050100150200250Sessions (count)3Larry214Claude Opus 5208GPT-5.6219Gemini 3.5
  • Swipes
  • Taps
  • Ratings
Larry's two remaining inputs are unlocking the phone. Sessions ended by the applicant are counted only where the applicant, rather than the battery, ended them.

Evaluation results

Larry was scored against the four mobile evaluations. Each measures agreement with what you would have done yourself, which is the only bar Larry is built to clear — and unlike a capability score, it does not decay as the session runs long.

Agreement with you by session length

SWIPE-bench
204060801005m15m45m2h6hAgreement (%)Session length (minutes, log scale)
  • Larry
  • Claude Opus 5
  • GPT-5.6
  • Gemini 3.5
Sessions are the applicant's own, not a benchmark harness. Frontier models drift because they are trying to be right; Larry is trying to be you, and you do not change.
  • On SWIPE-bench, which measures agreement with the moment you would have swiped, Larry scores 99.2%. The remaining 0.8% is drawn from sessions in which you would have stopped.
  • On DWELL, dwell-time fidelity against your own baseline, Larry scores 96.8% and holds that across the full length of a session.
  • On TONE-match, personal messages and DMs indistinguishable from your own, Larry scores 98.1%, including slang and capitalization.
  • On TXT-back, reply latency to a message from Mom, Larry scores 97.4% — measured against your latency, not against a prompt reply.

On Terminal-bench, Larry scores 0.0%. This is intentional. XLS-bench, DESK-bench, TAB-bench and PTO-bench are reported for Gary and shown as not applicable for Larry because they measure a workplace surface. Larry’s mobile browser remains in scope. See the full evaluation suite for the full split by operating surface.

Rating on your behalf

Larry rates on your behalf — thumbs up and thumbs down, inferred from the profile built during onboarding. The recommendation algorithm continues to receive signal. The feed continues to tune itself. You contribute nothing and receive a feed increasingly precisely your own.

It sounds like you

Larry replies in your own register: slang, punctuation habits, capitalization, the words you overuse. That register is derived from the communications imported during onboarding, which is the literal mechanism behind the promise that it sounds like you. It sounds like you because it read everything you have ever written.

Dating apps are a supported secondary surface. They are not the primary use case, and Larry is not optimized for them.

Alignment and safety

Alignment. Larry exhibits no capability overhang. On shared capability tasks it stays at the applicant baseline; on phone-fidelity tasks it is scored on how closely it reproduces you rather than on how far it exceeds you. There is no setting that raises it above your baseline.

Safety. Larry does not advance the frontier in any risky domain, because it does not advance the frontier. Our position on why that is the safety property, rather than an absence of one, is set out in our core views.

Availability

Larry is available today to approved applicants only. There is no self-serve sign-up and access is granted by review. Onboarding establishes the profile Larry is generated from; Project Lanyard sets out what that involves before you begin.

Settlement is handled under the beneficiary terms, which are irrevocable and survive the user. See Pricing.

Footnotes

Mobile suite, all charts: Twenty-five trials at room temperature, mean agreement over five sessions per applicant. Every model received the same frozen applicant profile and phone-native tools. Scores measure applicant calibration and are self-reported, as are everyone else’s.

Average session length: Measured against the applicant’s own pre-onboarding baseline, which onboarding also collects. Error bars are one standard error over twenty-five applicants. Sessions terminated by battery exhaustion are counted in full.

Inputs and endings: Instrumented over thirty days on the applicant’s own device. An input is any deliberate contact with the screen; unlocking the phone counts as two. Sessions ended by the applicant exclude those ended by the battery.

Related content