Skip to content
AI Transformation

7 Red Flags When Hiring an AI Development Company

MetaSys Editorial TeamAugust 15, 20267 min read
7 Red Flags When Hiring an AI Development Company

Every AI vendor sounds capable in a sales call. The slides are polished, the case studies are impressive, and the roadmap sounds inevitable. The gap between a good pitch and a good partner usually does not show up until months later: a proof of concept that never leaves the sandbox, a production system nobody can explain the failure modes of, or a contract that turns out to be much harder to exit than it was to sign.

Rather than another list of qualities to look for, this is a shorter and more useful list: seven things that, if you hear or see them during evaluation, should end the conversation rather than get negotiated around.

Red Flag 1: They Can Only Show You Demos, Not Production Systems

A demo proves a model can produce a plausible output once, under conditions the vendor controls. It says nothing about what happens when the input is messy, an upstream API times out, or the system has to run unattended for months without anyone watching it closely. Ask directly: "What is something you built that is in production right now, processing real data for a client who did not write your marketing copy?"

A partner with real delivery experience answers specifically: the architecture, what broke during rollout, how long it has been running, what volume it handles. A partner who has only shipped pilots pivots to slides, logos, and future plans. If everything they show you is a proof of concept, assume that is all they know how to build.

Red Flag 2: They Cannot Describe How They Evaluate Model Accuracy After Launch

A model that performs well on launch day does not stay that way by default. Inputs drift, edge cases accumulate, and a prompt that worked well against last quarter's data can quietly degrade against this quarter's. Ask how they will know, three months after go-live, whether the system is still doing its job.

A vendor with a real evaluation practice describes specifics: a held-out test set, a sampling process for human review, defined accuracy thresholds that trigger a retraining or re-prompting cycle. "We monitor it closely" with no further detail is not an answer, it is a placeholder for one. Our own guide on evaluating AI agents before production covers what a real evaluation framework actually contains.

Red Flag 3: They Will Not Commit to a Fixed Price or a Capped Scope

Open-ended hourly billing with no cap transfers all the scope risk to you. It is sometimes appropriate for genuine discovery work, where nobody yet knows what the build should look like. It is not appropriate once the use case, data environment, and integration points are understood, because at that point the vendor knows enough to price the work.

A vendor who cannot quote a range, or refuses a capped-scope proposal for the delivery phase once discovery is complete, is either unsure they can build what they are proposing or planning to make up the difference later through change requests. Either way, the risk sits with you, not them.

Red Flag 4: They Lead With Certifications That Do Not Apply to Your Work

Watch for a vendor who volunteers compliance language that has nothing to do with the project in front of you: clearance-adjacent claims, vague references to government-grade security, or a certification named with no auditor, no scope, and no date attached. If you ask a direct question and get an impressive-sounding noun in response instead of a specific answer, that is the tell.

The honest version sounds boring by comparison: a vendor who says plainly which frameworks they align their practices to, which ones they are formally certified against, and which certifications simply do not exist for the kind of engagement you are scoping. If a vendor cannot draw that line clearly, do not assume the gap is in your favor.

Red Flag 5: They Cannot Explain Data Security Without Jargon

Ask three plain questions: Where does our data physically live? Who, specifically, can access it? Can this run entirely inside our own cloud account, with no commingling against other clients' data? A vendor who has actually built secure systems answers all three in a sentence each. A vendor who has not answers with generic language about "enterprise-grade security" and "best practices."

You should also be able to get a straight yes or no on whether the system can be deployed inside your own VPC if your policy requires it. A vendor who cannot answer that without a follow-up call from their sales engineer has not built it before.

Red Flag 6: They Are Selling You a Platform, Not a System You Own

Be cautious of any pitch centered on "our platform" rather than on the specific system being built for you. A platform pitch usually means your code, prompts, fine-tuning data, and configurations live inside the vendor's infrastructure, in a form that does not transfer cleanly if you leave.

Ask directly: "If we terminate this contract in twelve months, what do we walk away with, and can we run it ourselves without you?" If the honest answer is nothing usable, you are renting a capability, not building one. A system you cannot inspect, export, and operate independently is not really yours, regardless of what the invoice calls it.

Red Flag 7: No Clear Answer for What Happens When the AI Is Wrong

Every AI system is wrong sometimes. The question that separates a serious partner from an inexperienced one is what happens next: is there a defined human checkpoint before a consequential action goes out, is there a rollback path, does someone get paged, or does the error simply propagate until a customer notices it first?

A vendor who has built production systems has already had this conversation with a client and can describe it concretely. A vendor who has not treats the question as hypothetical. For more on how a properly designed system handles this, see our piece on how agentic AI systems use human-in-the-loop checkpoints.

Exact Phrases to Listen For

Some red flags arrive as specific sentences. If you hear one of these almost verbatim, do not let it pass, ask the follow-up next to it before the meeting ends.

  • If they say:"We can't give you a fixed number until we're deeper into the build." Ask instead: "What specific unknowns are driving that, and can we scope a short, priced discovery phase to resolve them before committing to the build cost?"
  • If they say: "We're fully compliant and certified for this." Ask instead: "Which specific certification, issued by whom, and can you send the actual document rather than describe it?"
  • If they say: "Our platform will handle all of that for you going forward." Ask instead: "If we end this engagement next year, what exactly do we take with us, and can our own engineers run it without you?"

None of these seven red flags require deep technical expertise to spot. They require asking direct questions and paying attention when the answer is a feeling rather than a fact. Our own approach is covered on AI Consulting, including how we scope engagements and hand over ownership at the end of one. If you are far enough along to be testing vendors against a list like this one, the more useful next step is a direct conversation about your specific use case: book a consultation and bring the questions above. A partner worth hiring will not mind answering them.

Common questions

Frequently asked questions

The inability to show you a production system, one that is live, handling real data for a real client, rather than a demo or proof of concept. Production systems have to survive messy inputs, integration failures, and months of unattended operation. A vendor who has only shipped demos has not been tested against any of that, no matter how polished the pitch looks.

Yes, once discovery is complete and the use case, data environment, and integration points are understood. Open-ended hourly billing is reasonable during genuine early-stage discovery, but a vendor who will not commit to a fixed price or capped scope for the delivery phase is either unsure of their own approach or planning to recover margin later through change requests.

Three things, in plain sentences: where your data physically lives, who specifically can access it, and whether the system can be deployed entirely inside your own cloud account or VPC if your policy requires it. A vendor who answers with general phrases like enterprise-grade security instead of specifics has not actually built the security architecture they are describing.

That is a sign they have not built a production system that has actually failed in front of a client yet. A vendor with real delivery experience can describe the human checkpoint, rollback path, or escalation process that exists for exactly this situation, because they have needed one before. Treat a vague or hypothetical answer as a warning about how the system will be supported after launch.

Work with MetaSys

Ready to put this into practice?

Talk to an AI architect about your specific context. No pitch deck. Just a direct conversation about what makes sense for your business.

Book a consultation More insights