Brett Adcock's Hark has previewed a browser agent called Handoff that scored 97.7 on the Online-Mind2Web leaderboard, surpassing GPT-5.5's 92.8, Opus 4.8's 84.1, and Gemini 2.5 Pro's 69.0 on a human-evaluated third-party benchmark. The agent navigates real websites like Target, Walmart, OpenTable and LinkedIn by clicking and typing like a person would, without relying on official APIs. According to TechCrunch coverage cited in the report, Handoff analyzes page structure and visual information to determine whether to click buttons or enter text.

The benchmark performance came alongside a substantial cost advantage: Handoff's pricing sits at $0.18 per million input tokens and $2.37 per million output tokens, less than one-tenth of GPT-5.5's $5/$30 rate. The company positions both speed and cost as key differentiators against competing models from Anthropic, OpenAI and Google, though the preview didn't include head-to-head speed comparisons. In a demonstration for TechCrunch, Handoff assembled a custom flower bouquet from a vague request that included "some of the florist's choice." The report lists tasks like ordering on DoorDash and Uber Eats and price-comparing flights across United, Delta and American as example use cases.

The underlying architecture represents a design departure from standard language models. Hark's system predicts the next action rather than the next token, a shift the company describes as fundamental to reliable real-world browser task completion. The timing carries unusual context: Hark raised $700 million in Series A funding in May 2026 at a $6 billion post-money valuation, and Handoff currently runs on a post-trained model, with full-scale pre-training scheduled for later in 2026. That reversed sequencing means the current demo reflects data pipeline and post-training work, not the complete training run Hark still plans to execute. The report acknowledges that all performance claims originate from Hark itself, with the TechCrunch preview offering no independent benchmarks or reproducible tests.

The report highlights a business challenge underneath the technical demonstration: platforms like LinkedIn and OpenTable have historically blocked automated agents, and a browser-first agent that mimics human behavior represents exactly the kind of traffic large sites actively fight. The waitlist opened immediately, with launch targeted for the end of summer 2026, when the claims will meet real workloads and, the report suggests, likely face real countermeasures. As Brett Adcock framed it in the official press release quoted in the report, "The world is full of AI assistants, but you'd never hire an assistant who couldn't use a computer." The success of browser agents that operate without permission depends entirely on whether the platforms they automate choose to treat them as legitimate users or adversarial bots—a question that won't be answered in controlled demos but in production at scale.