Skip to content
moher.ai
← Journal6 min read

What an AI development company should actually deliver.

Most AI engagements die between the demo and the deploy. Here's the checklist that separates a vendor who ships from one who presents.

Travis Moher

There is a specific failure mode in AI work that almost every company hits once. A vendor builds something genuinely impressive, demos it to a room that gets excited, and then the project spends eight months failing to reach production. Nobody is lying. The demo was real. It simply wasn't the same category of object as a system.

The gap is always the same set of unglamorous concerns: authentication, rate limits, cost per request, evaluation, failure modes, observability, and who carries the pager. If a vendor hasn't raised those in the first conversation, you are buying a demo.

Ask what happens on the worst input, not the best.

Any model looks capable on a curated example. The interesting question is what the system does with a malformed document, a hostile prompt, a request that costs forty times the average, or a third-party API that returns a 500 halfway through a multi-step task.

A team that has run this in production will answer immediately and specifically, because they've been paged about it. A team that hasn't will answer in principle.

Insist on an evaluation harness before the first feature.

Without evaluations, every prompt change and model upgrade is a coin flip dressed up as a judgement call. Six months in, nobody can tell you whether the system is better than it was in March, and the honest answer is that nobody knows.

An evaluation harness is not exotic. It's a set of representative cases, a scoring method, and a number that moves. It should exist before the second feature ships, and you should be able to run it yourself.

Cost per request is a product decision, not a finance one.

AI systems have a variable cost curve that traditional software doesn't. A feature that is delightful at a thousand users can be insolvent at a hundred thousand. That's a design constraint, and it belongs in the architecture conversation rather than in a surprised finance review two quarters later.

Ask for a per-request cost model on day one, with the assumptions written down. Then ask what happens to it under a caching strategy, a smaller model, and a ten-times traffic increase.

Ownership is a structural question.

Whose GitHub organisation? Whose cloud account? Whose domain does the tracking run on? Whose name is on the app store listing? These sound administrative until the relationship ends and you discover the answer is not yours.

The correct arrangement is that everything is in your name from the first commit, and the engagement could end on a Friday with you shipping on Monday. Any vendor structurally uncomfortable with that is telling you something.

Ask what they run on the same stack.

The strongest signal available is whether a firm operates anything of its own on the architecture it is recommending to you. Not a demo product — something with real users, real money, and a real incident history.

It changes the advice. A team carrying its own pager recommends boring, durable choices, because they are the ones woken up by the exciting ones.

Related capability
AI product engineering

Tell us what you're building.

Two honest paragraphs beat a ten-page RFP. We answer inside one business day, and the first call is 45 minutes with someone who can scope it.

Or write directly — t@moher.ai