Why AI Projects Fail, and Where to Look First
In the company projects we have worked on, failures cluster in four stages. At scoping, nobody fixed a concrete result. At the data stage, nobody found out what the systems actually contain. At integration, the project stopped at a demo. At acceptance, someone checked that each task finished and nobody checked the outcome. This article is for owners and managers who are running or have run an AI rollout. For each stage it lists the signals you can check from the outside and what we do about it, and it ends with four questions you can use as a self-check.
Zhang Chenxi (Nature), who builds agent systems for manufacturers and distributors. . Drafted with AI assistance.
If you came looking for industry statistics on why AI projects fail, this is not that. It is a single team's notes from its own project records, so it gives no failure rate. Those records include requirement conversations, ERP data work, document-editing tests, coaching sessions with clients, and an audit of one rebuild inside our own system, Flowness.
Why AI projects fail: four stages and their signals
| Stage | What you see when it is failing | What we do |
|---|---|---|
| Scoping | The request is "we want to use AI". One meeting raises several unrelated asks. Nobody can say who sees what, and when. | Run each scenario through a design checklist. The owner answers business questions only. Out-of-scope asks are declined at once. |
| Data | The same question gets different answers. Rows go missing. The system has to be taught again and again. Field names and business meaning do not match. | Read everything and reconcile counts first, then look at what the fields really hold. Adjust scope to what the data shows. |
| Integration | The demo runs on an exported sample. Nobody can say who owns the account, what happens when the login expires, or who is responsible for writing back. | Split the work into read, application and write-back layers. Quote authentication, session renewal, permissions and retries as separate items. |
| Acceptance | Every task shows as done and the overall result does not appear. Edited output "looks right". After delivery only the outside team can change anything. | Add a whole-system check on top of task-level acceptance. Score field by field against an answer key. Teach the client to state requirements on their own. |
Scoping: the owner cannot picture the AI in their own business
In our experience, owners usually know AI is powerful and still cannot say what it would do in their business on a given day. What is missing is one sentence: at a certain moment, a certain person sees something and gets something. Without that sentence, data work, integration and acceptance have nothing to be measured against.
The first conversation often brings requests from several departments at once, from management reports to automation unrelated to the core business. Each is reasonable. Put them in one project and none can be accepted.
We turn each scenario into a design checklist and run every scenario through it before we write a proposal. The first item is "at what moment does the owner see what, and get what". The middle items cover the current situation and pain, data sources, rules and logic, the points where a person confirms, and how the output flows back. The last part is a set of choices for the owner, asked in business language, never in technical terms. Two tests tell us the scoping is good: the owner trusts that we understand the business because we can name details only someone who has done the work would know, and the owner can see the first step, what comes out at the end, and how it will be used. Anything outside the checklist, we say in the first conversation that we will not take on, so the scope does not grow quietly.
Three signals say scoping is the problem. You cannot find one specific person and one specific moment in the request. The number of asks in one meeting is more than a team could accept at the same time. The acceptance standard is "it feels smart". If any of these is true, go back to scoping before talking about technology.
Data: the same question gets a different answer every time
Some companies try the obvious thing before they call anyone: export a big table from the ERP and hand it to a chat model. The result is missing rows and answers that change from run to run. A model can read the product descriptions in one file. The hard part is running it every day, tying every number back to a source document, and answering follow-ups, with accounting entries and noise in the data next to the products. In our proposals, the numbers are computed by code and queries over the full data, so each figure traces to a document, and the model chooses the measure, combines results and explains them.
The most underrated part of the data stage is the distance between what a system field says and what the business means by it. In one ERP data pull we ran, we saw a few situations that are easy to miss.
- Counts need reconciling. Read the item master page by page and compare the records you got with the total the system shows, and the unique keys with both. Everything after this rests on having the complete data.
- "Last updated" cannot be trusted. In that item master the update-time field was mostly empty, so a design that says "sync what changed since the last update time" would have failed. Re-read the master on a schedule, compare field fingerprints, and rebuild the index only for records that changed.
- The specification field is not the specification. In sampled finished products, the specification field held a model-style code and the dimension field was empty. The real dimensions, materials and tolerances sat in the bill-of-materials sub-components. Similarity search on finished-product codes alone will never find an older product whose size and material match.
- Scope should follow business objects. Products usually have parent and child relations: the finished product, its semi-finished parts and its sub-components sit under different codes. Take a random sample and an engineer will find an old model and miss the other half that goes with it. We judge that this damages trust more than having no system at all, so we draw scope by business object and cover all finished products. When the budget is tight, we cut source types and special formats first and protect search coverage of the master.
Signals: unstable results, numbers that do not match source documents, field meanings that exist only in people's heads, and a sample size picked by guesswork.
Integration: a demo is not a system
A demo can run on an exported sample. Production cannot, and the questions integration has to answer never come up in a demo.
The ERP we tested gave us structured data through the interfaces its own web pages use, but that is not a public API the vendor supports for the long term. A browser login can prove that the path works. It cannot serve as long-term authentication. We recommend quoting service-side authentication, session renewal, permission control and failure retries as separate items, not folded into one line called "connect to the ERP".
Then the business semantics. In ERP systems that keep several sets of books, we see these situations: stock has several unposted states, so total on-hand is not the same as available; sales prices come in quantity tiers with validity dates, so taking one price per part number gives wrong totals; data from different sets of books cannot simply be merged, and any cross-company summary must keep the set of books as a dimension; issuing materials to production is an outbound transaction, and editing or posting it must never be written as an inbound one. A demo's "available stock" usually lacks these judgments.
Write-back needs the most care. If a read or a calculation is wrong, you recompute. If a write-back is wrong, a live business document has changed. Our design splits a project into three layers: first read the data, then run the calculations and produce the reports, and only after confirmation write back. The write-back layer gets its own permission checks, duplicate-submission handling and recovery after interruption.
Signals: the demo data came from a manual export, nobody can say whose account it is or what happens when the login expires, and "write back" is treated as a button with no approval flow behind it.
Acceptance, and who changes it after delivery
Acceptance goes wrong in two common ways.
The first happened to us. Flowness is our own agent harness, which we use to organize several AI agents working continuously, managing tasks, context, tool calls and work records. On 2 September 2026 we ran a retrospective audit on a rebuild of one stage of the system, laying requirements, design, engineering plan, frozen concepts and task packages out in one table. Of 32 tasks, 29 were reported successful, and 24 agreed design concepts and 11 mechanisms looked completely covered, yet several key effects the original request asked for never reached those tasks. Take "improve the next interview from past problems". The design held four things: identify the relevant events, turn past problems into a list of must-ask questions, give that list a stable home, and make sure the next interview reads it before it starts. By the time it reached the task stage, it had become a list of problem types, a few places that read it, and a statistics endpoint. The code could be implemented correctly from those tasks, and the task chain still could not guarantee that the next interview would read the list. Every handoff narrowed the request in a reasonable-looking way, and what remained was what is easy to confirm. We did not rerun any experiment; this is a trace of design against tasks. The full account is in Every task was done, and nobody carried the result we first asked for (in Chinese).
For a company project, "every feature passed acceptance" does not mean "the result the owner wanted has appeared". So in our proposals we add one whole-system check on top of task-level acceptance. It looks directly at whether the original moment happened, for example: the owner opens the daily report, and can every number in it be traced back to a source document?
The second way is to replace item-by-item checking with a result that looks right. In a document-editing test we compared several groups of real products and several hundred fields, one field at a time. For each group, a person produced the finished version and the answer key was copied from it at field level. The scoring script used that key, and the approaches under test could not read it. Editing has one more distinction: a dimension in the text can be corrected without the drawing geometry following. So the proposal checks at two levels, field and file, and an engineer reviews and signs off the edited output. The answer key and scoring script stay with the project, so anyone who wants to evaluate the result can recompute it.
One more stage comes after delivery. Business rules change, data changes, and the people who use the system change. If only our team can change anything, the longer the client uses it, the more dependent they become. So in coaching we put the effort into the client's own ability to state clear requirements. In class we walk the client through the whole development process, from interview to monitoring, then demonstrate a live interview, so the client goes through expressing a need, being questioned, clarifying it, forming a requirement definition and correcting it themselves. After class there are two assignments: build your own interviewing skill, a reusable set of instructions that lets an AI keep asking questions when we are not in the room, and hand in a detailed requirements document.
Signals at this stage: when something goes wrong, only the outside team knows what to change; when a business user sees a wrong number, they do not know whom to tell and have nowhere to log it.
Four questions to find where your project is stuck
- Can you say in one sentence who sees what, at what moment, and what they get?
- Ask the same question twice. Are the answers the same? Can every number be traced to a source document?
- Is the data in the demo the same data the system will fetch for itself in production?
- Has anyone checked the overall result on its own? When a business user finds a problem, do they know whom to tell?
If you can answer all four, those stages have no obvious gap. Fix the stage where you got stuck first; you do not have to redo the whole project.
If you want a system like this built around your own workflow, see AI agent systems for manufacturers and distributors or write to hi@towow.ai.