A model can return a cheap answer and still create expensive work. Someone has to check it, retry the request if it fails, and repair whatever slipped through.
A cheap mistake remains a mistake, and someone eventually gets the invoice.
Haiku 5.5’s announced rates make it worth asking what an acceptable result would cost in a particular workflow. That calculation starts with tokens and continues through every correction needed before the output can be used.
Anthropic announced Claude Haiku 5.5 on October 7, 2026. API input/output prices per million tokens: $0.10/$0.50 for prompts up to 100,000 tokens; $0.50/$2.50 above. Availability: Claude Platform, AWS, Google Cloud and Azure. Sonnet 5.5 cache reads are half-price. Source: Anthropic.
Decide what finished means
Take a hypothetical system that routes incoming bug reports. A category looks convincing enough on screen; the test is whether it sends the report to the right team. Build the evaluation around examples with known answers, including reports that could belong to several teams and cases that need a person to decide.
Before changing models, define the acceptance rule. A valid JSON response might be necessary; it says very little about the correctness of the values inside it. Check the fields against the underlying evidence. Record when the system should abstain or ask a person. An honest unresolved case can be much cheaper to handle than a confident answer that quietly enters a database.
Add up the first call, retries, escalations and the human review actually needed. Divide that spending by accepted results, while reporting the unresolved cases separately. Otherwise, the cost per finished task hides how much work the system left behind.

What travels with each request
Context deserves its own budget. Consider a hypothetical assistant that drags an entire conversation into every request, including obsolete plans and earlier mistakes. Each new task inherits the clutter. A sensible experiment would compare that approach with carefully selected evidence and a concise record of decisions, checking that the shorter version preserves the information needed for correct answers.
Trimming blindly creates a different bill. Remove an exception, a date or an instruction that matters, and the next answer may need repair. Measure context size alongside accuracy. Keep examples where the omitted detail changes the outcome. The aim is a lean working brief with enough evidence to do the job.
Caching also needs accounting that reflects actual reuse. A stable block of instructions may be reused often; a changing document may offer fewer opportunities. Track what is written, what is reused and what misses the cache. Check the terms of the provider you deploy through before treating an advertised rate as your effective price.
Set a retry limit before increasing volume. In the bug-report example, an uncertain classification could go to review after the limit is reached. Keep the reason for that escalation: several failures in the same category may reveal missing information or a task that needs redefining. Another prompt will not necessarily fix either.
Keep total spending beside the unit cost. Cheaper calls make it easy to add more jobs, so a reduction per request can coexist with a larger bill. That may be worthwhile; it should be a choice the team makes with the figures in front of it.
A pilot can answer the question the rate card cannot. Run the candidate against the existing process, compare accepted results and inspect the cases that failed.
The most useful line on the dashboard is the cost of work you can actually accept.
Got a different take? I’d love to hear it. hello@maisfps.com
