AI integration

Shipping AI features without betting the product on them

Every second brief we receive now has an AI line item in it. Some of them are real — a genuine reduction in manual work, a search experience that finally understands intent. Some of them are a board slide. The engineering problem is the same either way: a probabilistic component has to live inside a system that users expect to behave deterministically.

The teams that get this right do not have better prompts. They have better containment.

Constrain the output before you tune the prompt

A model that returns free text pushes the parsing problem downstream into code that was not designed for ambiguity. A model that returns a schema you validate on arrival pushes the failure to the boundary, where you can handle it.

We treat every model call as an untrusted external API. Structured output, a strict schema, validation on receipt, and an explicit branch for the case where validation fails. That branch is the feature — not an afterthought.

  • Define the schema first, then write the prompt to satisfy it.
  • Validate on arrival and reject rather than coerce. A silently repaired response hides a regression.
  • Cap the blast radius: a model can draft, suggest or rank. Have a human or a rule confirm anything destructive.

Log the whole interaction, not just the failures

You cannot debug a prompt from an error rate. When a model behaves badly you need the exact input, the exact output, the model version, and the surrounding application state — and you need it from the successful calls too, because that is your only baseline for what changed.

Model providers deprecate versions and silently shift behaviour. The team that captured request and response pairs from day one can diff before and after in an afternoon. The team that logged only exceptions starts from nothing.

Keep the non-AI path alive

Semantic search is better than keyword search until the embedding service is down, and then keyword search is infinitely better. The fallback is not a nice-to-have; it is the difference between a degraded feature and a broken product.

This is also the honest answer to cost. Routing the easy 80% through deterministic logic and reserving the model for the genuinely ambiguous cases usually cuts spend more than any amount of prompt golfing.

What we tell clients

Start with one workflow where the cost of being wrong is low and the volume of manual work is high. Instrument it properly. Run it alongside the existing process for a few weeks and compare. Then decide whether to widen it.

The projects that go badly are the ones that start with the model and look for a problem afterwards.

All insights Talk to us about this