JOURNAL / AI & MACHINE LEARNING

Cheaper AI is a product decision, not just a smaller bill

Haiku 5.5 cuts token prices. The useful test for a product team is what happens to accepted results, review time and the complete cost of a task.

OCTOBER 9, 2026 · 4 MIN READ · BLANCHE
Cheaper AI is a product decision, not just a smaller bill
FIG. 01 — FROM THE BLANCHE JOURNALBLANCHE / 2026

INTRODUCTION

A smaller model bill does not finish the work

A cheaper model can make a useful feature viable. It can also make an unhelpful feature produce more work for everyone else. The product decision is whether lower generation costs improve an outcome people actually want, after checking, correction and support are included.

Anthropic launched Claude Haiku 5.5 on October 7, 2026, positioning it for narrowly scoped, high-volume work while retaining larger models for harder tasks. That is a reason to revisit a product's workload, rather than automatically replace its model everywhere.

Read the price, then change the unit

Checked October 9: Claude Platform lists Haiku 5.5 input/output prices at $0.10/$0.50 per million tokens for prompts up to 100,000 tokens, and $0.50/$2.50 above. API token prices are not completed-task costs. Account separately for caching, tools and other applicable charges.

A product team needs another unit: cost per accepted result. A support summary that an agent must rewrite is not equivalent to one they can safely use. A document label that silently misroutes a case is not a bargain because its generation was inexpensive.

A practical accounting method is to divide model, tool, review and correction costs by the number of results that meet a predefined acceptance test. Track elapsed time and serious errors separately. A low average cost should not conceal a rare failure with unacceptable consequences.

The review queue can swallow the saving

Consider a hypothetical internal summarization feature, not a Blanche client result or a model benchmark. It processes 1,000 documents. A team budgets $20 for generation and 100 minutes of review, valued internally at $30 an hour. Its combined generation-and-review cost is $70.

Suppose a replacement setup reduces generation to $5 but doubles review to 200 minutes. That combination costs $105. Nothing in this illustration predicts either model's performance. It simply shows why a smaller API bill cannot establish a cheaper workflow. Hosting, taxes, support and integration work are excluded from these illustrative totals.

The design matters as much as the model choice. Put source passages beside a summary. Make uncertain fields easy to inspect. Let a reviewer correct one item without restarting the entire task. Test whether these changes reduce review time before turning up the output volume.

Give the smaller model a smaller job

A sensible first experiment is one bounded step whose answer can be checked: categorize an incoming request, extract a defined field, or draft a short summary for review. Keep the input, allowed output and failure path explicit.

That does not require a complicated collection of agents. In a hypothetical customer-service product, the first version might sort requests into a few existing queues and send ambiguous cases to a person. It should not quietly expand into approving refunds merely because the same model can write a plausible response.

More generation is not automatically more customer value. If a team can only review ten suggestions, generating a hundred may create a new backlog. A better use of the saving might be faster feedback, a second check on risky cases, or a feature that was previously too expensive to offer.

Sometimes the stronger model is the simpler choice

Routing work between models adds maintenance, evaluation and another place for errors to hide. For a low-volume product, or a task requiring long-range reasoning, one stronger model can be the more economical overall choice. The lower-cost option has to earn its place through observed results, not its position in a pricing table.

This is also why a polished demonstration is insufficient. Test ordinary cases, ambiguous inputs and known failures with the same acceptance criteria. Count escalations and retries. Inspect the failures themselves, not only an overall pass rate. Keep a route back to the previous setup while testing a limited rollout.

Before changing the default

  • Name the customer outcome and define what counts as an accepted result.
  • Establish a baseline for cost, waiting time, review effort and error severity.
  • Compare the same representative tasks, including difficult cases.
  • Include tool calls, retries, corrections and maintenance in the decision.
  • Expand only when the complete workflow improves without unacceptable failures.

Lower inference prices create room to design better products. The useful question is what to do with that room. For teams deciding what to build next, start with the product and its users, then choose the model that makes the whole experience work.

KEEP THINKING

Keep reading.

More ideas on building better brands and digital products.

DESIGN / ACCESSIBILITY

Accessibility is a design advantage.

ENGINEERING / PERSPECTIVE

Beyond parallax.

View all articles