3 min read

Claude Sonnet 5.5 is faster, but task cost depends on effort

Anthropic claims up to 30% lower cost per task. Artificial Analysis measured $7.60 per task at max effort, about 50% above Sonnet 5. Two request settings that worked with Sonnet 5 now return 400 errors.
Claude Sonnet 5.5 headline beside a chart showing Anthropic’s up to 30% lower task cost and Artificial Analysis’s about 50% higher result at max effort, from separate evaluations.

Anthropic released Claude Sonnet 5.5 on September 28 for Claude users and API developers. It says the model generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task in its testing. The API still charges $2 per million input tokens and $10 per million output tokens. Anthropic attributes the saving to fewer tokens per task.

Artificial Analysis measured about $7.60 per task on its Intelligence Index with Sonnet 5.5 at max effort, roughly 50% more than Sonnet 5 in the same test. Its run used a pre-release deployment with Anthropic's default fallback enabled and a structured-output bug Anthropic has since fixed; Artificial Analysis plans to rerun affected tests. The two cost results cover different task mixes and effort settings.

Anthropic's case holds below max effort

Sonnet 5.5 is aimed at bounded work such as bug fixes and everyday coding. Anthropic says Opus 5.5 remains stronger on complex, open-ended tasks. Its published coding evaluations show large gains over Sonnet 5, including a 70.6% Terminal-Bench 4.0 score against 10.3% for the older model.

Effort changes the bill. Medium effort is the default in Claude apps and high on the Claude Platform. Higher effort generally makes the model work longer and raises cost per task. At medium effort on Terminal-Bench 4.0 and low effort on CursorBench, Anthropic reports Sonnet 5.5 beating Sonnet 5's best score for less than a tenth of the cost per task. At high effort on FrontierCode, Sonnet 5.5 scored 10 points above Sonnet 5 at the same setting for about one fifteenth of the cost per task in Anthropic's test.

Artificial Analysis measured roughly 193,000 output tokens per Intelligence Index task at max effort, about 60% more than it measured for Sonnet 5 or Opus 5.5 at max. That token use drove the higher task cost at Sonnet 5.5's unchanged list price.

For a team already using Sonnet 5, compare cost and time per accepted fix across effort settings. A faster response that requires another attempt can erase the saving. The effort levels have been recalibrated, so a setting called medium does not produce the same amount of thinking on both models.

Anthropic's FrontierCode table also records a limit of simply turning effort up: Sonnet 5.5 scored 52.1% at xhigh and 46.2% at max. The release attributes the lower max score partly to extra code-review steps that caused timeouts or edits outside the task.

The API upgrade needs more than a model ID change

Sonnet 5.5 is available on the Claude API as claude-sonnet-5-5, and through AWS, Google Cloud, and Microsoft Azure. Two request settings that worked on Sonnet 5 now return a 400 error, and the migration guide covers both.

The first is turning off thinking. Sonnet 5 accepted thinking: {"type": "disabled"}. Sonnet 5.5 needs {"type": "between_tools"} instead. That setting only works at low, medium, and high effort. To run at xhigh or max, use adaptive thinking.

The second is forcing a tool call. Sonnet 5.5 rejects tool_choice of type tool or any. Switch to {"type": "auto"} and mark the tool strict: true so its input matches the schema, then say in the prompt when the tool should be used. Strict tool use is not available for Sonnet 5.5 on Amazon Bedrock, so teams there must validate tool input in their own code.

Three other documented changes can also return 400 errors. Editing earlier conversation history and then replaying a Sonnet 5.5 thinking block fails by default on accounts created on or after August 31, 2026. Older accounts can opt in to the same check. The computer_20251124 computer-use tool no longer works on the Claude API or Google Cloud. Use computer_toolset_20260801 instead. Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, and Sonnet 4.6 are also rejected as advisor models.

One change breaks interfaces without returning an error. Notes longer than a sentence or two between tool calls now arrive as thinking blocks. At the default display: "omitted", those blocks are empty, so interfaces that stream progress go quiet between tool calls.

Refusals also work differently. Some higher-risk cybersecurity requests can fall back to Sonnet 5, while ordinary secure coding, including source-code vulnerability scanning, remains available. Claude apps switch models automatically by default and show the switch. On the Claude API, a decline returns HTTP 200 with stop_reason: "refusal". The beta fallbacks: "default" setting, available only on the Claude API, retries cyber and frontier-LLM declines on Sonnet 5. Other refusal categories are not retried. Sonnet 5.5 was not in the Cyber Verification Program at launch.

Test the effort setting alongside the model, using tasks where the result can be checked.