Enterprise AI initiatives do not usually fail because the model is wrong. They fail because the data infrastructure underneath was not built to support it. After 15 years building enterprise data architecture across clinical research, financial services and healthcare, Milan Parikh FBCS, Lead Enterprise Data Architect at Cytel explores the five most common mistakes.
Summary
- AI enterprise initiatives are often failed by five common data architecture mistakes
- Pipelines should be designed to handle as yet unknown requirements as well as current ones
- Governance must be baked into practice, not treated as a regulatory checkbox
- While a demo shows proof of concept, a product isn't ready until it's ready for the end user
- Data lineage should be built into systems, not written up after the fact
- Layered architecture means updates can be made to individual aspects, rather than a monolithic architecture
The five most common data architecture mistakes are not exotic — they are predictable, and preventable.
Here is what they look like, why teams make them, and what to do instead.
Mistake one: building the data pipeline after the model
Most organisations treat pipeline development as delivery. First, there is model development, then there are requirements, and finally, someone hands out a ticket with instructions for getting data from point A to point B. It works in theory. In practice, it creates a fragile pipeline that breaks the moment requirements change.
This is what happened when my team needed to migrate data from SQL Server into Dataverse. The SQL schema was huge and constantly changing according to business needs. Meanwhile, Dataverse has a fixed schema. As each new requirement popped up, I had to rewrite the entire pipeline actions set again.
There was no problem with data migration per se; the problem was in designing the whole pipeline based on a point-in-time snapshot of requirements, instead of assuming that schema drift would happen.
A middle ground solution was the transformation layer. Instead of relying on hardcoded column-to-column connections, the transformation step could absorb all the changes in the source schema. Only variables would be adjusted. Thus, the entire pipeline would withstand any new requirements, without requiring total reconstruction.
Be prepared for those requirements that you don't know yet.
Mistake two: treating data governance as a regulatory checkbox
Without any guidance, different departments design and implement their data structures to solve their own issues. Each department is solving a real problem, but there are issues in terms of siloed data, duplication, conflicting ownership and lack of a single source of truth.
This kind of situation emerges slowly, then becomes an issue suddenly; for instance, a governance discussion in our organisation resulted from some compliance and audit issues.
The solution was not to write a policy document. Instead, we adopted medallion architecture with a certified layer above the raw and aggregated layers. This was where governance would be: controlling permissions and endpoints to access different resources. No data would reach consumers or reporting tools without going through this layer.
In other words, governance is not about filling out a form before go-live.
Mistake three: designing for demo rather than production
Deadlines result in a unique kind of tunnel vision. The team needs to make sure that the demo is done, MVP is ready to ship, and the stakeholder presentations are prepared. It sounds right, but consistently misses the same thing.
A project management and accounting application was created and successfully tested. Then came the launch. What was overlooked in this case were multiple security roles for each user, as well as other things which could have been checked during testing but which were ignorable under a tight deadline.
This problem is common; everything works until someone whose job is to use the product starts using it. Additional testing can’t solve this issue; checklists, documentation and end-user sign off can.
A demo shows that the product works. Production readiness shows that it works for everybody involved.
Mistake four: deferring data lineage until an audit forces it
Data lineage is that feature everyone agrees on but that few teams will prioritise until something breaks.
In a large enterprise reporting environment, fragmented data lineage is invisible in a normal day because sources are only partially documented; column-level lineage may not have been implemented completely. Data pipelines will run, checks will pass and reports will be delivered right on schedule, and no one will notice the problems until something breaks.
For you
Be part of something bigger, join BCS, The Chartered Institute for IT.
A senior stakeholder decides to check a report. Nothing is broken, but there is missing data for sure. Someone has dropped a source table from the lineage tracking without generating any notification. In an environment with hundreds of data tables available, it takes two days to figure out which table it was, and it’s only possible by tracing back to the source manually from the pipeline. Two days to find an answer that clean lineage could have produced in minutes.
The price of fragmented lineage is clear in such situations, because that is when you need visibility in your pipeline the most. This problem can be avoided by building lineage right into layer transitions instead of documenting it afterwards. Source tables tracking and column-level lineage become an integral part of data pipelines.
Build lineage into your systems; don't document it afterwards.
Mistake five: deploying AI into a monolithic architecture
A monolithic architecture seems to make perfect sense in the beginning. One data storage, one source code, everything consolidated into one place. But it works well only when there is little data and very few use-cases. People do it with the assumption that it will suffice for now. Six months later, that assumption becomes a structural liability.
I analysed the experience of an organisation that consolidated its data architecture like this, and then built its AI on top. Once the data volume started to grow, duplicate models were created to extract different pieces of information without isolating them properly. That led to the breakdown of any naming structure, and the once neat architecture became bloated and convoluted.
In such conditions, everything wrong with the data architecture translates into an issue with your AI project: data duplication leads to inconsistent training sets; lack of proper naming structure makes it hard to establish which model consumes which source; and with the growth of complexity, drift in the models' performance occurs. Some models even receive conflicting inputs.
The solution lies in separation. By creating a layered architecture — raw, aggregated and certified — you ensure that updates can be made separately for each layer. Models work with certified data rather than raw sources, so retraining one does not affect the rest of the system.
The more data and models accumulate, the more pressure is put on the architecture.
The common thread
None of these are technological failures and all of them are avoidable; they are planning decisions made too late. They require decisions about pipeline design, governance, production readiness, lineage and scalability of the architecture made earlier than most teams are comfortable with — but that discomfort is far cheaper than rebuilding a data ecosystem after an AI programme failure.
Milan Parikh FBCS | Lead Enterprise Data Architect, Cytel | CES Innovation Awards 2026 Judge | IEEE Senior Member | Secretary, BCS Enterprise Architecture Group
Take it further
Interested in this and similar topics? Explore BCS' books and courses:
- Innovating ethically to drive business change
- BCS Foundation Certificate in the Ethical Build of AI
- BCS Foundation Certificate in Digital Solution Development
- Delivering Digital Solutions: Software engineering, testing and deployment
- Designing Digital Solutions: Architecting user experiences, processes, data and security