AI Tools Review
Hark Handoff: Adcock's Browser Agent Explained

Insights

Hark Handoff: Adcock's Browser Agent Explained

AI Tools Review Editorial Team7 October 2026

    Quick Answer:

    Handoff is the browser and computer-use agent from Hark, Brett Adcock's personal-AI startup. Hark claims 97.7 on Online-Mind2Web at under a tenth of the token price of GPT 5.5, and says the first 100,000 sign-ups get the paid plan free. The benchmark is vendor-reported and not independently verified, and key security details are still unspecified, so treat it as promising but unproven.

    Brett Adcock is best known for building humanoid robots at Figure. This week his second company, Hark, put a very different product in front of the public: an agent that opens a browser on a cloud computer and gets on with your errands, from ordering dinner to booking a table.

    The launch comes with a big benchmark number, a very low price per token, a $700 million funding round and an NVIDIA partnership. It also comes with caveats that the headline coverage tends to skip. This guide separates what has been confirmed from what is vendor claim, and tells you what to check before you hand over a card number.

    Julian Goldie SEO covers the Hark launch and the free-access angle. Creator enthusiasm, not independent testing.

    Executive Summary

    What it is: Handoff is Hark's computer-use agent. According to coverage of the technical preview, it creates a dedicated virtual computer for every request, complete with a browser, a file system and a terminal, and controls a cursor and keyboard rather than relying on integrations with individual services. That is what lets it work on websites that have no consumer-facing API.

    • Timeline: funding of more than $700 million at a $6 billion valuation in May 2026; a research preview of Handoff on 05/08/2026; an NVIDIA partnership announced on 27/08/2026; and a public launch of the Hark Pro app on web, iOS and Android in the week beginning 05/10/2026.
    • Headline claim: a score of 97.7 on the Online-Mind2Web benchmark, ahead of the older models Hark compared against. This has not been independently reproduced.
    • Price claim: $0.18 per million input tokens and $2.37 per million output tokens, against $5 and $30 for GPT 5.5.
    • Launch offer: Adcock said on X that the first 100,000 registered users receive the paid plan free.
    • The caveats: benchmark comparisons use earlier-generation models and Hark's own harness; Handoff trailed GPT 5.5 on a second benchmark; and security and retention details for the virtual computers were not specified at preview.

    Our view: Handoff is one of the more credible new entrants in a crowded browser-agent field, mainly because of the funding, the compute deal and a product that is genuinely live. But we have not tested it ourselves, and the numbers that make it look dominant come from the company. For context on the wider category, see our guide to AI browser automation agents.

    Who Is Hark?

    Hark is the personal-AI lab started by Brett Adcock while he continues to run Figure, the humanoid robot maker. Reporting on the company says Adcock first funded it with about $100 million of his own money before a Series A of more than $700 million in May 2026, led by Parkway Venture Capital at a $6 billion post-money valuation, with NVIDIA, Intel Capital and Qualcomm Ventures among the participants. The company is a separate legal entity from Figure.

    Hark's stated ambition is broader than a browser agent. Adcock has described an AI that knows you, speaks your language and is highly personalised, eventually living on hardware built for it. One report on the consumer launch says a device is planned for 2027 with the form factor undecided, and another mentions a connectivity partnership for standalone devices. We treat the hardware story as a roadmap rather than a product: nothing we found shows a shipping device.

    Handoff is the first concrete product to emerge. In Hark's framing it is the "hands" of a personal assistant, the part that actually does things on the open web, while the rest of the app supplies memory, projects and a home screen of proactive cards. That context matters when judging the benchmark: Handoff is a component of a consumer assistant, not a standalone research model that Hark intends to sell by the token to developers, although it publishes token rates.

    What Handoff Actually Does

    Hark describes Handoff as an agent that uses the internet the way you would: it can place orders, do research, build websites and make slides. In the preview coverage, example tasks included shopping at large US retailers such as Target and Walmart, booking restaurant tables through OpenTable, ordering food and coffee, booking travel, filing taxes, working on LinkedIn and ordering flowers to a custom specification. The preview demo showed a florist order, although one outlet noted that the demo video showed only part of the process, which limits how far it proves reliability.

    Hark branding over a close-up of a floral bouquet, the image used by eWeek to illustrate its coverage of the Handoff agent preview.
    The florist order was one of the tasks used to demonstrate Handoff. Image: Hark branding as used in eWeek's coverage of the Handoff preview.

    Once Handoff reached the Hark Pro app, the description became more concrete. A third-party review of the launch reports that Handoff runs on a cloud computer able to manage up to six browsers in parallel, log in to websites on your behalf and make purchases, all from a single persistent chat thread. Alongside it sit Projects (topical folders for multi-week tasks), Action Buttons (proactive cards that try to solve problems before you ask) and Panels (mini apps that connect to services such as Strava, Spotify and Venmo). The app has persistent memory, and Hark says conversations and credentials are protected with end-to-end encryption.

    Two features stand out for everyday use. The first is background operation: because Handoff runs on Hark's cloud computer rather than your own machine, you can hand over a task and put your phone down. The second is account access. Handoff connects to your own accounts so it can use saved addresses and preferences, which makes it far more useful than an anonymous agent and far more sensitive from a security point of view. We return to that below.

    If the idea of a cloud-hosted agent doing work while you are away sounds familiar, it is the same direction Anthropic has taken with its own desktop product. Our explainer on Claude's background computer use covers how that approach differs, because Claude operates on your machine while Handoff operates on Hark's.

    How It Works: Predicting Actions, Not Tokens

    The most interesting technical claim is about what the model predicts. Hark says Handoff is trained to predict the next action (a click, a keystroke, a scroll) rather than the next text token. It analyses both the structure of a web page and its visual appearance to decide what to do. That is in line with how computer-use models generally work, but Hark is explicit that actions are the native output.

    On training, the reporting we read says Handoff is built using supervised fine-tuning and reinforcement learning on top of an existing base model that Hark has not named. Mid-training is under way and a full pre-training run is planned for later in 2026. In other words, the model people are using now is a post-trained system sitting on someone else's foundation, and the version Hark is raising money to build comes later. That has two implications:

    • Headroom: if the current scores come from post-training alone, a purpose-built pre-trained model could be better, which is the bull case.
    • Opacity: because the base model is unnamed, outsiders cannot reason about its safety behaviour, licensing or how it will be updated.

    One outlet that read the Hark launch materials noted that the post-training approach, plus a cheap per-token price, is what allows the low rate card. A smaller or more efficient underlying model would naturally be cheaper to serve. That is a reasonable inference, but Hark has not published the details, so we flag it as inference rather than fact.

    For a point of comparison from the open-weight world, Microsoft's Fara1.5 is a 27B model trained for the same job, and we looked at how its numbers hold up in our review of Microsoft Fara1.5 against OpenAI Operator. The key difference is transparency: Fara publishes weights and methods, Handoff does not.

    Alex Finn gives his hands-on reaction to Hark. A creator's first impressions, useful for seeing the interface but not a controlled test.

    Benchmarks: What Hark Claims

    Online-Mind2Web is a benchmark that tests agents on live websites, judging whether they complete realistic multi-step tasks. It has become a standard way for browser-agent vendors to compare themselves, which is exactly why a headline score needs careful reading.

    The numbers in one place

    BenchmarkHandoffComparison reportedStatus
    Online-Mind2Web97.792.8 (OpenAI model), 84.1 (Claude Opus 4.8), 69.0 (Gemini 2.5 Pro)Vendor-reported, not reproduced
    WebTailBench v268.672.3 (GPT 5.5)Hark's own harness
    Price per million tokens (in / out)$0.18 / $2.37$5 / $30 (GPT 5.5)Rate card only

    A note on the comparison model. Two outlets name the 92.8 comparator as GPT 5.4, while a third calls it GPT 5.5. The pricing comparison in the same coverage uses GPT 5.5. We have therefore written "an OpenAI model" for the Online-Mind2Web figure and recommend checking Hark's own materials for the exact label.

    Why to hedge the 97.7

    The score is striking, and a 13.6-point gap over Claude Opus 4.8 in Hark's table is large. There are five good reasons to be cautious before treating it as settled.

    • Not independently verified. Coverage noted that the public Online-Mind2Web leaderboard did not independently verify Hark's 97.7, and eWeek said the results have not been independently reproduced.
    • Older comparators. eWeek reported that Hark compared Handoff with earlier-generation models, and another outlet pointed out that newer frontier models are missing from the comparison. Beating last year's systems is not the same as beating the current best.
    • Own harness. The WebTailBench and internal evaluations used the company's own testing harness. Harness details (prompts, retries, time limits, how success is judged) can move scores by many points.
    • Mixed results. On WebTailBench v2 Handoff scored 68.6 against 72.3 for GPT 5.5. That is the only head-to-head with the current GPT generation in the coverage we read, and Handoff lost it.
    • Benchmarks are not your websites. As eWeek put it, benchmark leadership does not necessarily guarantee better performance across every real-world situation, and no head-to-head speed comparisons were published either.

    None of this means the number is wrong. It means a single self-reported score is a claim, not a fact, and the sensible approach is to wait for third-party runs on the same benchmark. If you want a sense of how a transparent, reproducible submission looks, our Fara1.5 coverage walks through one.

    Pricing: Tokens, Tiers and the Free Offer

    There are three layers of price information, and they are not equally solid.

    Token rates (reported by several outlets): $0.18 per million input tokens and $2.37 per million output tokens, roughly £0.14 and £1.80 at recent exchange rates. Against $5 and $30 for GPT 5.5, that is under one-tenth of the price, which matches Hark's "less than a tenth" claim. At those rates the input price is about 3.6% of GPT 5.5's and the output price about 7.9%.

    Why the rate card is not the whole story: eWeek cautioned that token rates alone do not account for the total cost of completing a browser task, including retries and supporting infrastructure. Browser agents consume many tokens because each step involves reading a page, and a model that needs fewer steps or fewer retries can cost less per task even at a higher rate. We saw the same effect in our look at a cheap small model in Solar Mini 4 and the cost-per-task problem. Until someone publishes cost per completed task for Handoff, the headline saving is unproven.

    Consumer plans (third-party report): one review of the public launch lists a free tier at roughly 40 delivery orders' worth of usage, a Pro² tier at $20 a month with double that, and a Pro³ tier at $100 a month with ten times the baseline, roughly £16 and £80. A separate report on pre-launch app code referred to further tiers and automatic top-ups. These come from third-party coverage and leaked configuration rather than a Hark pricing page we could read, so confirm the current tiers in the app.

    The 100,000 offer: Adcock announced on X that the first 100,000 registered users get the paid plan free. Reporting is consistent on that point, but the duration, the exact plan and the feature scope were not specified in the announcement coverage we read. Free offers of this kind often convert to paid after a period, so set a reminder and check what happens when it ends.

    From Preview to Public Launch

    The product moved quickly from announcement to launch, and the sequence helps to explain why coverage is uneven.

    • May 2026: Series A of more than $700 million at a $6 billion valuation, with NVIDIA among the investors.
    • 05/08/2026: research preview of Handoff. At that point access was by waitlist and application review, with broader availability promised by the end of summer.
    • 27/08/2026: Hark and NVIDIA announce a multi-year partnership, stating that the consumer platform would launch before the end of summer.
    • 04/10/2026: Adcock posts that the launch is this week and that the first 100,000 sign-ups get the paid plan free.
    • 06/10/2026: a third-party report describes Hark Pro as live on web, iOS and Android, from a company it identifies as Hark Labs Inc.

    Note that the "end of summer" target slipped into October. That is not unusual, and the product is now live, but it is a reminder that the August preview coverage described a research-stage system. Some of what was said then (access controls unspecified, a waitlist, a partial demo) may have been addressed in the app. We could not confirm that from primary documentation, so we flag the August caveats as the last confirmed position.

    The NVIDIA Partnership

    The NVIDIA link is real and documented. Hark's press release on 27/08/2026 announced a multi-year strategic partnership to advance personalised agentic AI, including deep technical collaboration and gigawatt-scale compute capacity on NVIDIA's next-generation Vera Rubin platform. It covers training Hark's foundation models and inference for the consumer launch. The release also says Hark is using NVIDIA's Megatron stack for training, Dynamo for inference, permissively licensed Nemotron pretraining data and NVSentinel for GPU cluster resilience, and credits NVIDIA AI infrastructure with powering Handoff's development.

    Adcock said the partnership gives Hark the foundation to build "the AI that everyone deserves but no one has built yet", and NVIDIA's Nico Caprez said building on the full-stack platform at gigawatt scale would let Hark develop, train and run AI at scale. That is partnership language, so read it for direction rather than detail.

    What the partnership tells you: Hark is serious about training its own foundation model rather than merely wrapping someone else's, which is the plan behind the later pre-training run. What it does not tell you: how good Handoff is today. Compute access is not a benchmark, and gigawatt-scale capacity is a plan for the future rather than a description of the current model. For a different angle on NVIDIA's agent push, see our coverage of NVIDIA NemoClaw.

    Security and Privacy Questions

    An agent that logs in to your accounts and spends your money is a high-trust product, so this section matters more than the benchmarks.

    What Hark promises: according to coverage of the launch, an encrypted vault called "Secured by Hark" stores credentials, and Hark says even it cannot see inside. It says it does not sell data to advertisers and that you can delete your data in Settings. The app asks for consent by default and learns which actions you are comfortable delegating.

    What is unclear:

    • Coverage of the August preview stated that the technical preview did not specify access controls for the virtual computers, how long files and credentials are retained, or who can inspect a session.
    • The launch review says Hark may train on safety-reviewed content even where a user has opted out. If you have strong views on training data, read the privacy policy in full.
    • Hark does not state failure rates for Handoff tasks, so you cannot judge how often the agent makes a mistake before the confirmation step.
    • The confirmation button is, in practice, your main protection against an erroneous purchase. If you click through prompts quickly, you remove the safeguard.

    Prompt injection. Any agent that reads web pages can be steered by hidden instructions on those pages. Hark has not, in the material we read, set out how Handoff resists this. It is an industry-wide problem rather than a Hark failing, but it is more serious when the agent holds payment access. Our browser agent guide covers the general risks and mitigations.

    Work accounts. The launch review recommends that IT teams review third-party app restrictions in Microsoft Entra ID and API access controls in Google Workspace before staff connect work accounts, and set a policy on which accounts a personal agent may touch. We agree: keep Handoff to personal accounts until your organisation has decided.

    Real-World Friction: Sites Do Not Want Bots

    The strongest argument for a browser agent is that most services have no API, so a model that can use a normal website unlocks everything. The strongest argument against is that many of those websites actively resist automation.

    Coverage of the launch raises this directly. Handoff depends on websites keeping stable interfaces, and it is vulnerable to bot-blocking and human-verification checks. Platforms have historically blocked automated agents, and one outlet specifically flagged the risk of blocking by sites such as LinkedIn and OpenTable, both of which featured in Hark's examples. Terms of service are another factor: some sites forbid automated access, and an agent that works today can be blocked tomorrow.

    Benchmarks cannot capture this, because a benchmark runs against a controlled set of tasks, whereas a real checkout flow changes weekly, serves CAPTCHAs unpredictably and treats cloud IP addresses with suspicion. This is the main reason we think the Online-Mind2Web number, even if it is accurate, tells you less about your experience than it appears to.

    How It Compares

    The field is crowded. Reporting on Handoff named Google, OpenAI, Anthropic and Browser Use as competitors, alongside smaller startups. A fair high-level comparison looks like this:

    ApproachWhere it runsStrengthTrade-off
    Hark HandoffHark's cloud computerRuns in the background; claimed low price; consumer appClosed model; unverified claims; account access held by a start-up
    Frontier lab computer use (OpenAI, Anthropic, Google)Varies: local, cloud or browserProven vendors; general modelsHigher token prices; some features need desktop apps
    Open-weight agents (for example Fara1.5)Your own hardwareTransparent; private; no per-token billYou supply the harness and the compute
    Agent browsers (for example Perplexity Comet)Your browserUses your existing sessionsTied to one browser; your sessions are exposed to the agent

    On general-purpose agents, our Manus tool page is the closest existing comparison in the directory, and you can read about its recent direction in the Manus 2 Cue agents article. For browser-specific products, Perplexity Comet takes the opposite approach by running in your own browser. If you are weighing a cloud agent against Anthropic's offering, start with the explainer on Claude Cowork's browser.

    How to Try It Sensibly

    If the free offer tempts you, a cautious trial costs little and tells you far more than any chart.

    • Sign up early if you want the free plan. The offer is limited to the first 100,000 registrations, so waiting may mean paying.
    • Use a separate, low-limit card or a virtual card with a spending cap for any purchase test.
    • Start with low-stakes tasks: research, price comparison, drafting a shopping list. Avoid tax filing, banking or anything irreversible in the first week.
    • Do not connect work accounts until your organisation has a policy.
    • Read every confirmation. Check the item, address, quantity and total before approving.
    • Keep a log. Record the task, whether it succeeded, how long it took and how many retries it needed. Ten tasks gives you a personal benchmark.
    • Check the end date of any free period and the cancellation route before you enter payment details.

    A log of your own results will answer the question that matters: whether Handoff completes the errands you actually have, on the sites you actually use, at an acceptable failure rate.

    Who Should Use It

    • Early adopters who like testing new agents and are happy to report where it breaks, particularly while the free plan is available.
    • People with repetitive web errands such as reorders, bookings and research, who can check results quickly.
    • Builders watching the category who want to see how a well-funded start-up packages computer use for consumers.
    • Not yet for: anyone who needs audited security, regulated data handling, guaranteed uptime on third-party sites or a published failure rate. Also not yet for work accounts.

    Limitations and Open Questions

    • Unverified benchmark: the 97.7 on Online-Mind2Web is self-reported, with older comparators and no public leaderboard confirmation at the time of coverage.
    • Mixed second benchmark: Handoff scored below GPT 5.5 on WebTailBench v2 in Hark's own harness.
    • Cost per task unknown: low token rates do not show how many steps and retries a real task needs.
    • Base model unnamed: Handoff is a post-trained model on an undisclosed foundation, with pre-training still to come.
    • Security details thin: retention, inspection rights and failure rates are not published in the sources we read.
    • Plan details secondhand: tier names, prices and the length of the free period come from third-party coverage.
    • No hands-on testing here: we have not run Handoff ourselves, and creator videos are first impressions rather than controlled tests.
    • Conflicting labels: sources disagree on whether the 92.8 comparator is GPT 5.4 or GPT 5.5.

    The Bottom Line

    Hark Handoff is a serious attempt at a consumer browser agent, backed by large funding, an NVIDIA compute deal and a live app with a generous launch offer. Brett Adcock's track record with Figure means the company will get attention and resources that most start-ups in this space cannot match.

    But the case that Handoff is the best browser agent rests on a score the company reported itself, measured against older models and not yet confirmed independently, plus a price comparison that ignores cost per task. The same coverage shows Handoff trailing GPT 5.5 on a second benchmark and leaves security questions open. Our advice is simple: try it if the free plan appeals, keep the stakes low, log your own results, and wait for third-party benchmark runs before believing the headline. We will update this article as independent numbers and clearer plan terms appear.

    Sources

    Images: eWeek (Hark demo imagery); hero is the thumbnail of the embedded Julian Goldie SEO video.

    Last updated: 07/10/2026. Sourced from press coverage, Hark and NVIDIA announcements and third-party reviews. We have not tested Handoff ourselves; benchmark figures are vendor-reported and plan details may change.

    Free Guide

    Get the free guide: Claude vs ChatGPT, Gemini & Grok

    A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

    Pop your email in to get it free
    Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

    Frequently Asked Questions

    What is Hark Handoff?
    Handoff is the computer-use agent from Hark, the personal-AI startup founded by Brett Adcock, who also runs the humanoid robot company Figure. For each request it creates a dedicated virtual computer with a browser, file system and terminal, then clicks, types and scrolls across websites to complete tasks such as ordering food, shopping, booking travel or making reservations. It works on sites that have no public API.
    How good is Handoff on benchmarks?
    Hark reports a score of 97.7 on Online-Mind2Web, against 92.8 for an OpenAI model, 84.1 for Claude Opus 4.8 and 69.0 for Gemini 2.5 Pro. The figures are vendor-reported, were not independently reproduced on the public leaderboard at the time of the coverage we read, and compare against earlier-generation models. On a second benchmark, WebTailBench v2, Hark reported 68.6 against 72.3 for GPT 5.5, so Handoff does not lead everywhere.
    How much does Handoff cost?
    Hark lists model rates of $0.18 per million input tokens and $2.37 per million output tokens, against $5 and $30 for GPT 5.5, which is under one-tenth of the price on paper. Rate cards do not capture retries or infrastructure costs. Third-party coverage of the consumer launch reports a free tier plus paid Pro tiers at $20 and $100 per month, and Adcock said the first 100,000 sign-ups get the paid plan free, but Hark has not detailed how long that lasts.
    Is Handoff safe to give my accounts and card details?
    That is the main open question. Hark says credentials are held in an encrypted vault it cannot see into, and that Handoff asks for confirmation before actions. Coverage of the technical preview noted that Hark had not specified access controls for the virtual computers, how long files and credentials are kept, or who can inspect a session. Use a separate account, a low-limit card and read every confirmation prompt.
    What is the NVIDIA connection?
    NVIDIA took part in Hark's $700 million Series A in May 2026, which valued the company at $6 billion. On 27/08/2026 Hark announced a multi-year partnership with NVIDIA covering gigawatt-scale compute on the Vera Rubin platform, plus use of NVIDIA's Megatron, Dynamo and Nemotron data. NVIDIA infrastructure is credited with powering Handoff's development.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.