What Breaks First in AI-Built Apps
AI-built apps break first at authorisation, database access rules, leaked secrets and unverified payment webhooks, not at the feature they were built to demo. They fail at the thing nobody demonstrate

Choosing an AI development agency comes down to five checks: production evidence, compliance literacy, who writes the code, how the price is structured, and who owns the IP on day one. This guide gives you the exact questions for each check, what a good answer sounds like, and the red flags, including the ones that would rule us out.
Every agency rebranded to "AI" over the last three years, but the demo-to-production gap in AI work is wider than anything in classic SaaS: a convincing prototype is a weekend, a system that behaves under real users is an engineering discipline. The five checks below are designed to expose that gap in a single call.
Ask: "Show me how you evaluate model behaviour before a release, and what your monitoring caught last month." An agency that ships production AI has eval suites, tracing, and war stories about failure modes. An agency that demos AI has a portfolio video. Follow up with: "What happens when the model is wrong?" A good answer involves guardrails, review queues, and fallbacks, not reassurance.

If you are in a regulated space, make them draw your data flow before they quote. For health products: where does PHI enter, what does the model see, who signs a BAA (our HIPAA-aware approach)? For fintech: what is in the audit trail when a regulator asks why the model declined a customer (how we build KYC/AML systems)? Vague answers here turn into rework you pay for twice.
Ask for the names and seniority of the people who will be in your repo, and whether they stay for the whole engagement. The bait-and-switch (senior faces in the sales call, juniors in the codebase) is the most common failure mode in outsourced work. At Robust Devs the team is senior-only and the founder leads delivery directly; whoever you choose, get the staffing commitment in writing.

Fixed scope with a written definition beats open-ended time and materials for an MVP: it forces the scoping conversation now, when changes are cheap. The pattern we recommend (and sell, so judge accordingly): a small paid diagnostic first. Ours is a fixed $4,999 Tech Audit, then a 1–2 week discovery that ends in a fixed build price, then a 6–14 week fixed-scope build. Any agency that quotes a firm number before discovery is guessing with your money.
The IP assignment, the cloud accounts, the repos, and the model-provider keys should be yours from the first commit, not handed over at the end or held hostage to the final invoice. Ask: "Whose name is on the AWS account?" The right answer is yours.
| Check | What good sounds like | Red flag |
|---|---|---|
| Production evidence | Eval suites, tracing, named failure modes | Portfolio videos and "our AI is very accurate" |
| Compliance literacy | Draws your PHI/KYC data flow before quoting | "We can add compliance later" |
| Who writes the code | Named senior engineers, staffing in writing | Seniors in the sales call, unnamed "delivery team" after |
| Price structure | Paid diagnostic → discovery → fixed-scope price | Firm quote in the first call, open-ended T&M |
| Ownership | Your accounts, your repos, day one | IP transfer "on final payment" |
Beyond the calls: review volume and rating on independent platforms (ours: 4.6/5 across 58 verified client reviews), a verifiable legal entity (we are a UK-registered company), and named humans with real profiles rather than stock-photo teams. None of these prove delivery quality, but their absence tells you something.
Market quotes mostly run $15,000–$150,000+ depending on scope and compliance posture. The driver-by-driver breakdown is in our AI MVP cost guide.
Freelancers suit single-workflow experiments; in-house suits post-product-market-fit scale; a senior agency suits the middle, where you are testing a funded thesis fast without hiring risk. The comparison table in the cost guide covers the failure modes of each.
We keep ranked lists with stated, checkable criteria for fintech MVPs and healthtech products, including where we do and do not fit.
Give them a small paid diagnostic and judge the artifact. Our version is the $4,999 Tech Audit: five days, a 47-item report, a 1-page action plan, and a 30-minute walkthrough. The working sample costs less than a day of anyone's time.
Written by
Founder & CEO
Founder & CEO of Robust Devs. Leads delivery and works directly with every client, across AI marketing, healthtech, and fintech builds, and has done since 2019.
AI-built apps break first at authorisation, database access rules, leaked secrets and unverified payment webhooks, not at the feature they were built to demo. They fail at the thing nobody demonstrate
A code audit report earns its fee when it contains four things: an executive summary the person paying can act on, findings ranked by severity that cite specific files and lines, an architecture asses
Technical due diligence rarely kills a round. It re-prices one, and it does so at the worst possible moment: after the term sheet exists, when the fund has already decided it wants in and you have alr

Tactical writing for founders building AI products. Browse the archive for more field notes like this one.
Browse all articlesIf you are building in this space, book a call or get in touch.