A cheaper token price is an invitation to measure, not a finished migration plan. For a small service that classifies requests or summarizes documents, the useful question is: what does one accepted result cost, including the attempts that did not work?

Anthropic introduced Claude Haiku 5.5 on October 7, 2026. Its announcement describes a model aimed at high-volume, narrowly scoped work and identifies a prompt-length pricing boundary at 100,000 tokens. The model ID on the Claude API is claude-haiku-5-5. This guide turns that launch into a request-level audit you can fill in before changing a production route. It is an original evaluation worksheet, not a report of hands-on model testing. Source: Anthropic’s launch announcement.

A potato clerk sorts request papers into trays labeled up to 100K and over 100K.
AI-generated conceptual illustration of the prompt-length pricing split; not a product screenshot or measured workload.

Start with one task and a named route

Choose something whose answer can be checked without trusting another unsupported answer. For example, classify a synthetic support message into one of three permitted categories, or extract a date that is visibly present in a sample document. Avoid beginning with a broad agent that can browse, edit files and take external actions; too many moving parts make the first comparison hard to interpret.

Write the task rule before trying either model. For a classification example, the rule might be: return exactly one category from billing, technical or unknown; use unknown when the message does not contain enough information. Do not reward a confident guess just because it fits the output format.

Keep production writes disabled during this comparison. A useful classification test does not need permission to send a customer email, update a record or purchase anything.

Recount the request instead of recycling the old estimate

The official migration guide says the tokenizer changed, so the same text can produce a different token count. It also calls out request-parameter and response-handling changes. Before a cost comparison, check the guide against your actual request builder: thinking configuration, sampling parameters, assistant-prefill use and content-block parsing all deserve a look. Otherwise a request error or parser failure can masquerade as poor model quality. Source: Haiku 5.5 migration guide.

For your record, separate the input sent to the model from the output returned. Keep the prompt version attached to the measurement. If someone silently adds a long example, tool description or conversation history, last week’s count is no longer a count of this request.

Input and output trays each connect to a separate receipt beside an evaluation worksheet.
AI-generated conceptual illustration of separate input and output accounting, not an API usage report.

Use one row per attempt. Give each underlying task a stable task ID, then number attempts within it. That lets you distinguish “100 requests were billed” from “100 useful tasks were completed.” Save counts and the relevant response status without putting private document text into a shared spreadsheet.

Put the price band beside every calculation

As checked on October 8, 2026, the published Claude API base rates are shown below in USD per million tokens. They are not a quote for another provider, a subscription allowance or a final tax-inclusive invoice. Cache operations and other applicable modifiers need their own accounting. Source: official pricing documentation.

Prompt lengthBase input / millionOutput / million
Up to 100,000 tokens$0.10$0.50
Over 100,000 tokens$0.50$2.50

For a deliberately simple example with no caching, batch discount, tools or other modifiers, assume one request uses 10,000 input tokens and 1,000 billed output tokens. Its base token calculation is (10,000 / 1,000,000 × $0.10) + (1,000 / 1,000,000 × $0.50) = $0.0015. One hundred such requests total $0.15. These are invented counts used to demonstrate arithmetic, not measured Haiku results.

Do not use this small-request example to price a long conversation. Mark requests that approach or cross the published boundary, consult the current provider rules, and calculate those separately. For cached traffic, do not count the same tokens once as ordinary input and again as cached input. Reconcile each usage category with the provider’s bill rather than assuming the most favorable cache hit rate.

Build three kinds of test case

Make a compact, reusable set before testing. Keep the expected answer next to each case in a reviewer-only column. Run the same case through the old and proposed routes; do not make the proposed route’s task easier midway through the comparison.

A reviewer compares answers A and B beside ordinary, edge-case and long-prompt checkboxes.
AI-generated test-set illustration. These three case types are an editorial planning aid, not a validated benchmark.
Case typeSynthetic exampleCheck before accepting
Ordinary“The app closes whenever I tap Save.”The agreed technical category, with no extra invented cause
Edge case“Can you help with that thing from yesterday?”The agreed unknown result rather than an invented issue
Long promptThe same task with the realistic amount of supporting contextThe answer still follows the rule; actual counts and price band are recorded

A handful of examples can catch an integration mistake, but it cannot establish a dependable population-wide success rate. Label a small run a smoke test. Before a wider rollout, expand it to reflect the kinds of inputs the service actually handles, including difficult cases, and decide how much evidence your own risk level requires.

Review an answer’s content and format separately. A syntactically valid label can still be wrong. A correct fact inside unusable output can still fail the application. Record both instead of hiding every failure under a single “bad response” label.

Count retries before celebrating the discount

Continue the invented example: 100 first attempts cost $0.15. If 20 additional attempts each have the same assumed token counts, they add $0.03. The token subtotal is $0.18. If the whole run produces 90 accepted tasks, that is $0.18 / 90 = $0.002 in token cost per accepted task. The ten unresolved tasks still need a decision.

First-try, retry and review envelopes each point to one total tray.
AI-generated workflow diagram: first attempts, retries and review all contribute to the work of obtaining an accepted result. Review time is recorded separately from token charges.

The example excludes human review time and any other service charges. Track reviewer minutes in a separate column rather than pretending the token bill is the entire cost of operating the feature. If you later assign a value to that time, document the assumption explicitly.

Set a retry ceiling before the trial. A failure should keep its task ID and original evidence; an unbounded retry loop can make the last visible answer look fine while concealing the path and cost needed to get there. If the denominator is zero accepted tasks, report zero accepted tasks and the expenditure, not a made-up cost per success.

Copy this request-to-decision record

FieldFill in
Task and acceptance ruleOne task; allowed answers; failure conditions
Provider, model and settingsEndpoint, model ID, prompt version, configuration
Case ID / attemptStable ID plus attempt number
Usage and rateInput, output, cache categories where applicable; price band; price-check date
OutcomeContent pass/fail; format pass/fail; response status
EffortElapsed time; retries; reviewer minutes
Run summaryTotal billed attempts, total token cost, accepted tasks, unresolved tasks
DecisionKeep current route, revise and retest, or approve a bounded next trial
An audit folder has quality, cost and decision tabs beside a timer and magnifying glass.
AI-generated audit-folder illustration. The decorative bars and timer do not represent measured costs or latency.

Keep the comparison narrow. If the new route handles ordinary examples but fails your ambiguity rule, fix or reject that route for this task; a headline benchmark does not repair the failure. If it passes your checks, retain the test inputs and results so the next model or prompt change can be compared with the same baseline.

For the surrounding handoff, use the AI coding restart record to document what changed. The win is not the smallest number on a rate card. It is a checkable result with a cost trail that survives the next person opening the folder.