Ask Runable forDesign-Driven General AI AgentTry Runable For Free
Runable
Back to Blog
Technology7 min read

AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff | VentureBeat

Notably, the benchmark comparisons Hark provided to VentureBeat for its Handoff AI agent are against GPT 5.5, GPT 5.4, Opus 4.8, and Gemini 2.5 Pro — the pri...

TechnologyInnovationBest PracticesGuideTutorial
AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff | VentureBeat
Listen to Article
0:00
0:00
0:00

AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff | Venture Beat

Overview

AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff

Credit: Venture Beat made with Open AI Chat GPT-Images-2.0

Details

Credit: Venture Beat made with Open AI Chat GPT-Images-2.0

Hark, the secretive AI startup founded earlier this year by serial entrepreneur and roboticist Brett Adcock, today announced Handoff, a "computer use agent" (CUA) that it says is among the top-performing in the world at navigating the open web on a user's behalf — ordering dinner on Door Dash, booking flights on United and Delta, or messaging job candidates on Linked In — all autonomously, end-to-end.

Sign-ups open to the public today at hark.com, with availability planned for later this month as part of the initial release of Hark's software platform.

The company says Handoff recorded the top-ever score on Online-Mind 2 Web (OM2W), a third-party benchmark with a human-evaluated leaderboard for web agents, posting a 97.7 against 92.8 for Open AI's GPT 5.4, 84.1 for Anthropic's Claude Opus 4.8, and 69 for Google's Gemini 2.5 Pro.

Hark Handoff benchmark comparison chart. Credit: Hark

Hark Handoff benchmark comparison chart. Credit: Hark

Hark also says it can serve the model at less than one-tenth the token price of competing frontier models —

0.18permillioninputtokensand0.18 per million input tokens and
2.37 per million output tokens, versus
5and5 and
30 for GPT 5.5 — with per-turn model latency of 0.8 seconds.

Hark Handoff pricing comparison chart. Credit: Hark

Hark Handoff pricing comparison chart. Credit: Hark

Hark's research uncovered that despite people spending 75% of their screentime every day in a browser, fewer than 1 in 1000 websites have publicly accessible APIs, making it challenging for AI agents to take over the workload.

In a roughly four-minute produced announcement video posted on You Tube and social media, Adcock — seated in a bare warehouse space that doubles as a metaphor for the company's build-out — speaks a request aloud to Hark ("let's liven this place up a bit… let's do some roses, maybe some cherry blossoms") and Handoff is shown navigating a florist's website to place the order, while Adcock narrates that unlike a typical chatbot, Handoff "is always working, it's looping," and says he now uses it for "all of my recruiting efforts end to end." In Hark's announcement blog post, more demos are shown in realtime and 5x speed.

But big some open questions about Handoff remain, especially for potential enterprise customers and users.

High-scoring benchmarks...but against last generation's models

Notably, the benchmark comparisons Hark provided to Venture Beat for its Handoff AI agent are against GPT 5.5, GPT 5.4, Opus 4.8, and Gemini 2.5 Pro — the prior generation of frontier models.

The current leaders, Open AI's GPT-5.6 and Anthropic's Opus 5, are absent, as are strong open-source computer-use contenders like Deep Seek V4, Kimi K3, and Qwen 3.8-Max.

These newer models haven't published Online-Mind 2 Web results, and no third party has posted them to the benchmark's public leaderboard — meaning Hark's "top-ever" claim cannot currently be checked against the strongest available systems.

The omission is notable because the newest frontier models have posted their largest gains precisely in computer use: on OSWorld 2.0, a related benchmark covering full computer control, Anthropic's Opus 5 scores roughly 70.6% versus 55.7% for the Opus 4.8 model Hark chose as its comparison point.

The latency comparison comes with similar caveats: the 6.8-second and 6-second per-turn figures Hark cites for GPT 5.5 and Opus 4.8 were measured by Hark, in Hark's own harness, with the competing models set to their highest — and slowest — reasoning level. No independent latency measurements exist for comparison.

Asked by Venture Beat whether Hark plans to publish comparisons against those newer models, the company did not specify.

Even within Hark's own chosen comparisons, the "best" framing has an asterisk: on Web Tail Bench v 2, one of the three benchmarks in Hark's own results table, GPT 5.5 scores 72.3 to Handoff's 68.6.

Two of the three benchmarks (Web Tail Bench and an unnamed internal evaluation) were also run inside Hark's own harness, with pass rates computed by Hark's internal LLM judge — conditions the company controls.

Hark's pricing advantage is far clearer: Anthropic's newer Opus 5 carries the same

5permillioninputand5-per-million-input and
25-per-million-output list price as its predecessor, so Handoff's roughly tenfold cost savings would hold up even against the current frontier — assuming its benchmark performance does too.

Hark's research preview describes a sensible-sounding pipeline — supervised fine-tuning followed by asynchronous reinforcement learning using the GRPO algorithm, according to materials shared with Venture Beat prior to today's announcement — but the company acknowledges it has only done post-training so far, with pre-training "planned for later this year."

That means Handoff is built on top of a base model Hark did not train. Asked which base model it is, and what mix of proprietary and open data Handoff was trained on, Hark hasn't yet specified.

Another big question mark for enterprise users: who can access the dedicated virtual computers and the files created on them?

A Hark spokesperson said "security and privacy is a primary focus, but this is a technical preview," adding the company will share more when the product reaches market at the end of the summer.

Hark is Adcock's fourth company. He previously co-founded the talent marketplace Vettery (sold in 2018 for roughly $100 million), the air-taxi maker Archer Aviation, and the humanoid robotics unicorn Figure AI.

Hark raised a

700millionSeriesAroundinMay2026ata700 million Series A round in May 2026 at a
6 billion valuation — led by Parkway Venture Capital, with participation from Nvidia, AMD, Intel Capital, Qualcomm Ventures, Salesforce Ventures, and ARK Invest.

Adcock seeded the company with $100 million of his own money and remains founder and CEO of both Figure and Hark simultaneously, a spokesperson confirmed.

Asked how the two companies interact, the spokesperson said Hark models "are being trained on the Figure robots," but that Adcock has no plans to combine them.

Adcock's promotional style has drawn skeptics. In April 2025, Fortune correspondent Jason Del Rey reported that Figure's much-touted BMW partnership was far more modest than Adcock's public claims of a robot "fleet" performing "end-to-end operations": BMW spokesperson Steve Wilson said a single Figure robot was practicing picking up parts during non-production hours.

On the social network X, Adcock called the story "mischaracterizations and downright lies" and threatened a defamation suit. Two months later, Tech Crunch reported that Adcock skipped a promised live demo at a tech conference and sidestepped questions about the BMW deal onstage.

None of that means Handoff's numbers are wrong. The agent may well be excellent, and the pricing — if it holds — would undercut every major lab.

Key Takeaways

  • AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff

  • Credit: Venture Beat made with Open AI Chat GPT-Images-2

  • Credit: Venture Beat made with Open AI Chat GPT-Images-2

  • Hark, the secretive AI startup founded earlier this year by serial entrepreneur and roboticist Brett Adcock, today announced Handoff, a "computer use agent" (CUA) that it says is among the top-performing in the world at navigating the open web on a user's behalf — ordering dinner on Door Dash, booking flights on United and Delta, or messaging job candidates on Linked In — all autonomously, end-to-end

  • Sign-ups open to the public today at hark

Cut Costs with Runable

Cost savings are based on average monthly price per user for each app.

Which apps do you use?

Apps to replace

ChatGPTChatGPT
$20 / month
LovableLovable
$25 / month
Gamma AIGamma AI
$25 / month
HiggsFieldHiggsField
$49 / month
Leonardo AILeonardo AI
$12 / month
TOTAL$131 / month

Runable price = $9 / month

Saves $122 / month

Runable can save upto $1464 per year compared to the non-enterprise price of your apps.