Skip to main content
Live
Main content

Hark previews Handoff, a browser-use agent for real-world tasks

The $700M-funded startup claims its post-trained model beats GPT 5.5 and Opus 4.8 on cost and speed for browser automation.

Jaeden Schafer
Editor in Chief · · 4 min read
Hark previews Handoff, a browser-use agent for real-world tasks

Hark, the startup that raised $700 million in Series A funding in May, previewed its browser-use agent Handoff today and opened a waitlist ahead of a launch by the end of the summer. Handoff navigates consumer websites that lack official APIs — Target, Walmart, OpenTable, and LinkedIn among them — by reading page structure and visual layout to decide when to click, scroll, or type. CEO Brett Adcock, better known for humanoid robotics, is pitching Handoff as faster and cheaper than GPT 5.5 and Opus 4.8 on browser tasks, though Hark has not published benchmark numbers to back the claim.

The pitch is familiar: issue a natural-language command and the agent orders coffee, books travel, files returns, shops, reserves restaurant tables, or researches across the open web. In a video demo, Adcock asked Handoff to build a bouquet with specified flowers and left room for fuzzy instructions like "some of the florist's choice." The demo skipped chunks of the workflow, so end-to-end reliability is not verifiable from the clip.

Hark's architectural bet is what separates the preview from the crowd. The company says Handoff predicts the next action — a click at a coordinate, a keystroke into a field — rather than the next token, the standard objective for large language models. It is running on a post-trained model for this release and plans to move to pre-training later this year, which Hark argues will let it iterate faster on its data pipeline and training infrastructure than teams grafting agents onto general-purpose LLMs.

Key facts

  • 01Hark raised $700 million in Series A funding in May before previewing Handoff.
  • 02Handoff targets sites without official APIs, including Target, Walmart, OpenTable, and LinkedIn.
  • 03Hark claims Handoff is faster and cheaper than GPT 5.5 and Opus 4.8 on browser tasks.
  • 04The model is post-trained today; Hark plans to pre-train later this year.
  • 05A public waitlist is open and Hark plans to launch by the end of the summer.

The competitive field is dense. OpenAI, Google, and Anthropic all ship computer-use agents inside their consumer and API products, and a wave of venture-backed startups — Browser Use, Polar, Strawberry, and Aside — is chasing the same browser-automation wedge. Most of these agents share the same failure modes: brittle when sites re-render, slow on multi-step flows, and expensive per task when routed through frontier models. Hark's claim to beat GPT 5.5 and Opus 4.8 on cost is a direct swipe at that pricing dynamic.

Adcock's $700 million round in May was one of the largest Series A deals of the year and stood out because Hark shipped no public product at the time. The preview is the first look at what the money is funding. A purpose-built action model, if it works, is a defensible position — GPT 5.5 and Opus 4.8 were trained to write tokens, not to click through Target's checkout flow, and the gap in cost per successful task is where a specialist can win.

The unknowns are what Hark has not shown. There is no independent benchmark on task completion, no latency figure, no per-task price, and no disclosure of which sites Handoff handles reliably versus which ones break its planner. Browser agents historically stumble on captchas, two-factor prompts, checkout flows behind fraud detection, and any site that changes its DOM between test and production. Target and Walmart in particular deploy aggressive bot mitigation, and Hark has not said how Handoff clears those hurdles or whether merchants are cooperating.

Consumer trust is the other open question. Handing an agent a credit card and a shipping address to shop across four retailers is a different level of delegation than asking a chatbot to summarize a PDF. If Handoff misfires on a $200 grocery order or a restaurant reservation, the recovery cost — refunds, disputes, customer support — is the user's problem, not Hark's, unless the company builds explicit guarantees into the product. None have been announced.

For Hark, the summer launch is the real test. A preview video and a waitlist can carry a story for a few weeks; a paying user base that trusts an agent with weekly shopping and travel bookings is the moat. If the action-prediction architecture delivers the cost and speed advantage Hark is claiming, the browser-agent market gets a genuine specialist competing on unit economics against generalist frontier labs — and that pressure is where consumer AI pricing actually starts to move.

Related · from this week
Anthropic relaunches Claude Code Projects to orchestrate multiple agents in the cloud
Jaeden Schafer · 4 min read →
ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Tools

Anthropic logo
Tools

Anthropic relaunches Claude Code Projects to orchestrate multiple agents in the cloud

The revamped Projects feature runs parallel Claude Code threads on branched repos, with a coordinator resolving conflicts as pull requests.

Jaeden Schafer4 min read
Anthropic logo
Tools

Anthropic flips Claude Code into auto mode by default

Starting August 14, Pro, Max, and Team accounts get an agent that stops asking permission at every step.

Jaeden Schafer4 min read
Anthropic logo
Tools

Model Context Protocol drops stateful sessions in next week's update

The plumbing behind AI agent integrations shifts to a stateless design, addressing a scaling headache that has slowed first-party MCP rollouts.

Jaeden Schafer4 min read