Emma McGuigan FBCS and Martin Cooper MBCS explore why continuous assurance is the only way to keep AI safe and stop tech creators marking their own homework.

Summary:

  • Safety and privacy meet new challenges in the AI era
  • Traditional software uses predictable, rule-based logic, whereas probabilistic AI continuously learns and drifts, making testing hard
  • Embedded historical data bias is exceptionally difficult to spot, but it scales rapidly when automated across thousands of daily AI queries
  • Personal data cannot easily be deleted once it is woven into a trained model's weights and parameters
  • AI assurance is the process of measuring, evaluating and communicating whether AI systems are trustworthy, safe, secure and working as intended

In this edition of Insight Exchange, BCS Editor-in-Chief Martin Cooper sits down with Emma McGuigan, BCS’ AI programme lead, to untangle the complex reality of responsible AI. 

Beyond the industry buzzwords, securing fundamental pillars like explainability, fairness, safety and privacy is far more challenging than it first appears.

As a passionate proponent of the technology, Emma argues that robust, independent AI assurance is now the only way to ensure these powerful systems genuinely do no harm.

Why don’t you introduce yourself?

I’m a 30-year-plus veteran in the IT sector. I spent many years with Accenture as a developer and architect before moving into global leadership roles. More recently, I've been working as a consultant with BCS to help shape and support our AI program, which is incredibly exciting. I also act as an independent advisor for several other organisations.

So, let's start with making AI that’s transparent and explainable. Why are these things so important and so hard to do?

AI has been used in medicine for decades, but what is exciting now is its impact on diagnosis. While a top consultant relies on their training and hundreds of past cases, an AI engine can analyse tens of thousands of similar cases, delivering a highly accurate, well-documented diagnosis.

But as humans, we do not just want the diagnosis… We want to know it is fair, right, and leads to the best health outcome. Building that confidence is difficult. Traditional IT systems were deterministic and rule-based; we rigorously tested them and understood their logic. With deep learning models, there are no rules, only statistical patterns across millions of parameters. This is not human-readable logic. How do we build confidence when there are no clear, linear interactions?

And how about AI explainability?

Think about where our ‘tech bro’ colleagues are taking us: towards super-intelligence. This term used to belong in comic books, but now we are seriously looking at what it takes to get there — combining data, quantum capabilities and biological assessments.

But how will we ever have faith in that future when we cannot even trust how a credit limit was set on [credit] cards? If we cannot build confidence in something that could exist within a decision tree, how will we ever trust richer, more complex solutions that analyse far greater complexities to find their patterns?

Let's talk about fairness and inclusivity. I think we can all agree we want AI to be fair and inclusive.

If a teacher introduces a bias to a classroom, only those students are affected; if you add bias to an AI engine handling thousands of daily queries, the impact is massive.

A classic example is Amazon's 2018 hiring campaign. The algorithm was trained on their existing, predominantly male workforce, meaning CVs from different profiles were filtered out. Without focusing on the training data, the engine becomes a bias magnifier.

These fairness issues are hard to detect because we often only realise the lack of inclusivity when the engine is used on a group omitted from the training data. How do we know what questions to ask to ensure a dataset is fair? Real-world data naturally carries real-world biases. 
I mean, the AI engine can only be as good as the data it's been trained on. It's like all of us. We can only be as good at something as the people who have trained us to do it.

Where in an AI lifecycle are these issues hardest to detect?

It comes back to that force magnifier. As an engine matures and operates for longer, it builds a certain credibility over time, but those biases remain baked in.

With traditional systems, the longer they existed, the more we came to believe they were correct and became dependent on them — which is why COBOL and Fortran systems still exist today. It is the same mentality: with age comes increased confidence in a solution, but the magnifying force of that underlying bias is still at work.

Let's move on to safety and reliability. What's special or different about AI? Why is it hard to make reliable AI systems?

Traditional software is rule-based and deterministic. AI, however, is probabilistic and data-dependent; small input changes can produce unpredictable outputs, and models naturally degrade over time due to distribution shifts.

For you

Be part of something bigger, join BCS, The Chartered Institute for IT.

This fundamentally changes the development life cycle. Traditional IT gave us confidence through rigorous, structured testing (unit, integration, system, performance, and user testing) against defined requirements. With AI, there is no finite test suite. The engine is constantly deployed into a live environment, where it continues to learn from new data, leading to small changes that can yield highly believable but completely incorrect answers.

This is why having a human in the loop is so vital. Unlike traditional systems, where we just went back to a test cycle for new releases, AI requires human oversight. We need AI professionals who understand both the engine and the context in which it operates. In pharmacology, for instance, experts must verify that AI-generated drug designs are consistent with expectations, whilst remaining open to unexpected, positive breakthroughs that the AI's pattern-recognition has uncovered.

Who should be accountable? The human or the machine?

Who should be accountable? Look at the Post Office Horizon scandal. It was a rule-based accounting system that had passed its tests and was assumed to be infallible. Yet, it took a devastating miscarriage of justice to remind us that people, not machines, are always accountable.

Thirteen people lost their lives, and hundreds of families were ruined because of an IT system. We must never allow decision makers to hide behind the excuse of ‘the computer said so’. It is humans who sign off on these systems and make the final decisions. We use tools to help us, but the responsibility ultimately stops with us.

Why, where, and when should our AI makers consider privacy?
Data privacy in AI is entirely different from that in traditional systems. You cannot just tick GDPR boxes, perform a test, and consider the job done, because data is the continuous fuel feeding the engine, and privacy is a constant, major overhead.

Crucially, privacy concerns do not stop at deployment. Every time you use the engine, you feed it more data, and these models can memorise that personal information. When Copilot or Claude recalls a question that you asked three days ago to answer you today, you get a clear insight into why privacy is a completely different beast now — especially if someone else has access to your device.

Since UK privacy law guarantees the right to be forgotten, how can an individual realistically withdraw consent and demand their data back when that information is already woven into the very weights, nodes, and parameters of the trained AI model itself?

These are the questions AI designers must answer. They cannot embed your personal data into a system without permission, but how often do we actually read the small print when accepting cookies or installing retail apps? We must decide whether reduced functionality is a price worth paying for privacy, whilst remaining mindful of the value of what we are giving up.

Furthermore, major large language models are only being developed in the US and China. Even if a UK organisation uses its own customer data, if it is leveraging an overseas model, it faces entirely new challenges. How do you keep the model up to date whilst protecting your proprietary data, and how do you guarantee a user's right to exit? Navigating these jurisdictional and privacy issues represents a massive shift from traditional IT, and as professionals, we must have the confidence to answer them.

Given that AI models are probabilistic and constantly learning, should we move away from day-one certifications and instead adopt a model of continuous assurance to ensure they behave responsibly and safely over time?

Let us be clear: I am a big proponent of AI, but it cannot be left unpoliced — especially by the people who make it. We cannot have the creators marking their own homework. Consumers must have confidence that the organisations introducing these tools are actively managing and policing them.

This is why AI assurance is key, and it must address three areas. First, AI professionals must be accredited and willing to sign their names to declare they are responsible practitioners. Second, we need continuous checks and assurance to ensure training data is thoughtfully collated and monitored as the system learns. Finally, the core engines themselves must be accredited. Thinking about assurance across these three pillars of people, data, and engines is absolutely vital.

Listen to this episode of BCS Insight Exchange and more