
Claude Browser Use Tool vs Computer Use: I Tested Both
Use Claude's browser use tool (browser_toolset_20260801) when the whole job happens in web pages, and computer use (computer_toolset_20260801) when it doesn't. Browser use can read the page instead of photographing it, so on text-heavy pages it finishes in fewer turns for less money. On a short form that fits on one screen, computer use came out cheaper. I ran both toolsets against the same two tasks on Claude Sonnet 5.5 this morning, with the same headless Chromium behind them. Here are the numbers, and the one setting that changes them.
What's the difference between the two toolsets?
Both shipped on August 19. Computer use came out of beta as a toolset: no beta header, no display dimensions, and each action is its own member tool (left_click, type, screenshot, zoom...). Browser use was new that day. The two look alike at the API level. The difference is what Claude gets back.
- Computer use sees pixels. Claude takes a screenshot, picks a coordinate, clicks, takes another screenshot. It works on anything that renders: native apps, file dialogs, a canvas, a VNC session.
- Browser use reads the page.
read_pagereturns an accessibility tree where every element has a reference like[ref_3],findsearches elements in plain English,get_page_textreturns the readable text, andform_inputsets a field by reference. It also has tabs, download reporting andnavigate. Screenshots are still there when Claude wants them.
Both are client toolsets: Anthropic doesn't run the browser, you do. Claude returns tool_use blocks, your executor drives Playwright (or whatever you use), and you send back tool_result blocks tagged with toolset_name. The browser use docs spell out the result shapes, including the browser_state block that carries your tab list.
How many tokens does each toolset cost just to declare?
Nobody talks about this part, and it's the number that decided one of my two tests. I ran each tools array through /v1/messages/count_tokens on claude-sonnet-5-5 with a one-word message (10 tokens on its own) and subtracted:
- browser_toolset_20260801, defaults: 6,603 tokens
- browser toolset + all 4 optional members: 7,486 tokens
- browser toolset, 7 low-level pointer members disabled: 5,427 tokens
- computer_toolset_20260801: 4,522 tokens
- both toolsets together: 10,839 tokens
- old computer_20251124 (on Sonnet 5): 2,206 tokens
So the browser toolset costs about 2,100 tokens more than computer use, and you pay that on every turn of the loop, because the tools array goes up with every request. The new computer toolset is also twice the size of the tool it replaced. If you aren't caching the tools array, that overhead is a real share of your bill.
Trimming is one line per member. This is the config that took 1,176 tokens off:
tools: [{
type: "browser_toolset_20260801",
configs: {
left_mouse_down: { enabled: false },
left_mouse_up: { enabled: false },
hold_key: { enabled: false },
middle_click: { enabled: false },
triple_click: { enabled: false },
mouse_move: { enabled: false },
left_click_drag: { enabled: false },
},
cache_control: { type: "ephemeral" },
}]Turn off members you don't implement anyway. A drag-and-drop member your executor can't run is worse than no member at all: Claude will try it, fail, and burn a turn.
Which one is faster and cheaper in practice?
I wrote a small Node executor (Playwright 1.60, headless Chromium, 1280×800 viewport) that implements both toolsets against the same page, and ran two tasks on Sonnet 5.5 with no prompt caching, so the raw costs show. Every run gave the right answer.
Task 1: find row 40 in a long table
"Which country is in row 40 of the main table on Wikipedia's list of countries by population, and what figure does it show?" Answer: Morocco, 37,254,695. Two runs each:
- Browser use:2 turns, ~23,960 input tokens, ~600 output, 0 screenshots, 5–6 seconds. It called
navigate, thenget_page_text, and counted rows in text. About $0.054 a run. - Computer use: 6 turns, 49,713 input tokens, ~550 output, 5 screenshots, 16 seconds. Screenshot, scroll, wait, screenshot, and so on down the page. About $0.105 a run.
Half the tokens and a third of the time, because reading text beats scrolling pictures of it. It's also more reliable. Counting rows in a screenshot is a vision task; counting lines in text isn't.
Task 2: fill and submit a 9-field form
httpbin's pizza order form: name, phone, email, a size radio button, two topping checkboxes, delivery time, comments, submit, then report what the server echoed. Two runs each, identical results both times:
- Browser use:4 turns, ~33,100 input tokens, ~720 output, 1 screenshot, 8–9 seconds.
read_pageonce, then a single batch of 9 calls: fiveform_inputby reference and fourleft_clickfor the radio, checkboxes and submit. About $0.073. - Computer use: 4 turns, 28,244 input tokens, ~590 output, 3 screenshots, 12 seconds. One batch of 14 clicks and keystrokes off a single screenshot. About $0.062.
Computer use won this one by about 15%. The form fit on one screen, so a single screenshot told Claude everything, and the browser toolset's extra 2,100 definition tokens times four turns ate the difference. Batch actions are why computer use is competitive now: 14 actions in one turn used to be 14 round trips.
Two honest caveats. The computer use runs started with the page already open, since my executor had no URL bar to click; that saved them a turn or two of navigating. And my read_pageis my own simplified tree builder, not Anthropic's reference executor, so a richer tree will cost more tokens than mine did. Two tasks isn't a benchmark. It's enough to show the shape of the trade-off.
When should you pick each one?
- Browser use for research, scraping behind a login, long pages, docs, dashboards, multi-tab work and anything where the answer is text. It also helps when the UI changes often, because
[ref_N]references survive a redesign that moves every pixel. - Computer use for desktop apps, native file pickers, remote desktops, canvas and WebGL apps where the accessibility tree is empty, and short visual tasks on one screen.
- Neitherwhen there's an API. A computer use agent is still the most expensive way to move data; I put numbers on that in why computer use agents cost 45x more.
You can declare both in one request, but that's 10,839 tokens of tools before Claude reads a word of your prompt. Unless the task really crosses from the browser to the desktop, pick one.
Which models support them, and what breaks?
Both toolsets run on Fable 5.1, Mythos 5.1, Fable 5, Mythos 5, Opus 5.5, Opus 5, Sonnet 5.5, Sonnet 5 and Opus 4.8. I sent browser_toolset_20260801 to claude-opus-4-7and got a 400: "does not support tool types". Browser use is also missing on Bedrock, Claude Platform on AWS and Managed Agents for now.
Going the other way, computer_20251124returns a 400 on Sonnet 5.5 and Opus 5.5 on the Claude API. If you're mid-upgrade (see my Opus 5.5 migration notes and the Sonnet 4.5 retirement checklist), the computer use rewrite comes with it. The computer use migration guide has the full diff; the parts that bite:
tool_use.nameis now the action (left_click), notcomputerwith anactionfield. Your switch statement changes.- Every result needs
toolset_name, and text goes inside a content array. - No
display_width_px. Coordinates are in screenshot pixels, so if you downscale screenshots you scale clicks back up yourself. On a Retina Mac that's a factor of two you'll find the hard way.
How do batch actions change your executor?
Both toolsets can return several tool_use blocks in one turn. You run them in order, and if one fails, every call after it gets an error result with exact text. The browser toolset wants Not executed: an earlier action in this turn failed.; the computer toolset adds the word "computer" before "action". My loop:
let failed = false;
for (const u of response.content.filter(b => b.type === "tool_use")) {
const base = { type: "tool_result", tool_use_id: u.id, toolset_name: u.toolset_name };
if (failed) { results.push({ ...base, is_error: true, content: NOT_EXECUTED }); continue; }
try { results.push({ ...base, content: await act(u.name, u.input) }); }
catch (e) { failed = true; results.push({ ...base, is_error: true, content: `Error: ${e.message}` }); }
}One result per call, always, even the skipped ones. Drop one and the next request fails validation. And because a single turn can now contain a click on "Buy", your human-approval check has to run per call inside the loop, not once per turn.
Treat page content as hostile while you're at it. Tab titles, URLs and download paths all go back to the model, and they're all written by whoever owns the page. Anthropic runs injection classifiers on browser content by default, but I'd still keep javascript_exec off and the browser profile logged out unless the task needs it. Same thinking as MCP tool poisoning: anything the model reads can carry instructions.
My default setup
- Browser toolset for anything web. Computer toolset only when a native window is involved.
- Disable members the executor doesn't implement.
cache_controlon the toolset entry. Cache reads bill at a small fraction of the $2 input rate, so the 6,600-token definition stops mattering; my prompt caching breakdown has the math.- Prefer
get_page_textandfindin the system prompt over screenshots for reading. - Per-call approval for anything that spends money or sends a message.
If you're building a browser agent for a client workflow and want a second pair of eyes on the executor or the costs, get in touch. I've now written this loop enough times to know where it breaks.
Tested the morning of 6 October 2026 on claude-sonnet-5-5 through the Claude API, no prompt caching, max_tokens 8,000. Token overhead from /v1/messages/count_tokens. Costs at Sonnet 5.5 list prices of $2 / $10 per million input / output tokens. Executor: Node 25 + Playwright 1.60, headless Chromium, 1280×800, two runs per task per toolset.
Related Posts
AI Agents
ant apply: Manage Claude Agents as Code
ant apply turns Claude agents, skills, environments, memory stores and scheduled deployments into files in your repo, with a Terraform-style plan-and-apply loop. How claude-lock.json detects drift with two hashes, why renaming a file silently duplicates an agent, the adoption gap for Console-made resources, and the five rules for running it in CI.
AI Agents
Claude Agent Memory Stores vs the Memory Tool
A memory store is a versioned folder Anthropic mounts into your agent's sandbox at /mnt/memory. How it differs from the client-side memory tool, the read_write default that enables persistent prompt injection, the 10,000-memory cap that fails silently, and the 15-second sync on self-hosted sandboxes.
AI Agents
Claude Mid-Conversation Tool Changes: Keep the Cache
Editing the tools array invalidates the prompt cache for the whole conversation. tool_addition and tool_removal change what Claude can call without ever touching it — plus the placement rules that 400.