Andrew Leigh MBCS CITP, Software Development Group Head, and Dr Amie Vella, Head of Research with MASS, explore a framework for deploying mission critical AI.

Summary:

  • Research suggests the hallucination rate of leading AI models falls between 26% and 93%, a challenge secure industries in particular must address
  • MASS' Double Diamond design methodology helps build appropriate, secure systems by following the four stages of 'discover, define, develop, deliver'
  • The Double Diamond emphasises that businesses must start with the problem, not the technology, and approach AI technologies as a socio-technical discipline as the risks are not purely technical
  • Building an experimental system which enables trialling different models allows clear evaluation and guardrail development
  • Confidence can be built by experimenting with real data and robust governance is key for successful delivery

Since the emergence of LLMs, the world has witnessed significant hype about the potential for this kind of AI. Chatbots powered by LLMs are increasingly ubiquitous as they become embedded in everyday systems such as search engines and word processors. In the defence and national security domains, it’s no different; as general awareness of LLMs increases, customers, users and investors frequently ask how AI can be used in products and services. The added complexity, of course, is the highly secure, business- and mission-critical systems the industry requires.

It’s paramount that the output of the systems they work with can be trusted, given the unique challenges the defence industry faces — operating with sensitive data in highly secure environments being the main one. This presents a problem for the application of LLMs, because they always say yes confidently, which can be misleading when model confidence is not always a measure of information accuracy — independent research has found that the hallucination rate of some leading LLMs falls within a range of between 26% and 93% — and in business- and mission-critical systems, confident wrong answers are highly problematic. So, how can highly secure industries like defence apply this promising yet tricky technology in a trustworthy and safe way?

Mission-critical AI methodology

At MASS, we started by adopting an existing robust design methodology, such as the Design Council’s Double Diamond. This methodology includes four principle stages:

  • ‘Discover’ explores the problem space to ensure all dimensions of the problem have been considered
  • ‘Define’ converges on a clear definition of the problem to be solved and considers what success looks like
  • ‘Develop’ explores the design and solution options that could be used to solve the problem
  • ‘Deliver’ implements the solution that best solves the well-defined problem within the known constraints

Discover

The discovery phase recognises that we should start with the problem, not the solution. The MASS user experience (UX) team conducted an intensive, AI-agnostic study to identify pain points and opportunities, using techniques such as contextual enquiry (observation), desk research, workshops, surveys and interviews. We found user research to be an essential method for validating our initial ideas about the kinds of problems that AI might help to solve and the constraints that need to be taken into account when designing solutions. For example, the need for locally hosted, secure LLMs rather than cloud-based solutions stems from the highly sensitive nature of our customers’ data.

Define

During this phase, we invited our customers to a hackathon to further explore and validate our initial ideas about the problems that AI could help solve. Whilst we began with a technology-agnostic exploration of user needs, we found that exposing users to working prototypes helped reveal additional requirements that would otherwise remain unarticulated. This is because people tend to think in terms of the constraints they already know due to confirmation and belief biases. When new technology comes along offering new possibilities, they need to see it to overcome those biases and understand the art of the possible.

For you

Be part of something bigger, join BCS, The Chartered Institute for IT.

To facilitate the hackathon, we built a pluggable architecture with a simple chat interface that allowed different models and inference engineers to be easily swapped in and out. This meant we could experiment with different combinations to see which worked best for each problem. Having a secure local LLM enabled us to experiment with real customer data, which helped participants engage more easily with the hackathon because they could validate results using recognised data.

More significantly, allowing users to be hands-on with the technology revealed that much of the risk is not just technical. System and software engineering are socio-technical disciplines because their success depends on the joint optimisation of systems and people. The non-deterministic and hallucinatory nature of LLMs amplifies this risk, a fact clearly on our customers' minds. AI changes not only system behaviour but also human behaviour; users may begin to over-rely on automated outputs or discount their own expertise, creating new operational risks that cannot be mitigated solely through technical controls.

Develop

Drawing on insights from the discover and define stages, we gained the confidence to ideate on how LLMs could solve several problems our customers are facing. During the development stage, we returned to the discipline of user experience/user-centred design to explore options for LLM-enabled features through wireframes and prototypes. Users took part in usability testing of the candidate wireframes and prototypes. This offered further opportunity to surface the kinds of guardrails and affordances that users felt would be needed for AI-enabled features to be trusted, such as AI explainability and where to put humans in the loop.

Delivery

Following the double diamond approach and conducting disciplined discovery, define, and develop stages has given us a clear vision for leveraging this promising technology. It has done so in ways that will solve customer problems whilst fostering trust in what is also a risky technology, owing to its non-deterministic and hallucinating nature, which can amplify the socio-technical risks. It has also revealed the skill sets, governance structures, human oversight, validation methodologies and quality procedures needed to deliver trustworthy, safe implementations in mission- and business-critical environments.

Conclusion

In summary, there are seven things to remember when looking at using AI for business and mission-critical systems:

  1. Start with the problem, not with the technology – focus on user research
  2. Confidence is a hidden danger of AI technology
  3. Double down on approaching AI technologies as a socio-technical discipline — because the non-deterministic and hallucinating nature of AI amplifies socio-technical risk
  4. Discipline unlocks efficiency — stick to long-proven design methodologies such as the double diamond whilst practising user experience/user centred design, system/software architecture, and agile development in combination with data science and machine learning
  5. Build an experimental system with a pluggable architecture so that different models can be evaluated for solving different problems and further elicit the necessary guardrails and affordances needed to foster trust in the AI-enabled system
  6. Build confidence by experimenting with real data — deploy the technology into a system that can host real data so stakeholders can more easily validate results
  7. Robust governance is needed to deliver AI solutions

Collectively, these principles help ensure that AI is not only technically capable but also delivers safe solutions that are trusted, governable, and fit for use in business and mission-critical environments.

Dr Andrew Leigh CITP MBCS is Software Development Group Head at MASS, leading software, product and research teams. A software engineering leader with over 25 years' experience, he is an award-winning innovator, BCS member and Open University Associate Lecturer. Dr Amie Vella holds a PhD in Social Science and combines evidence-based research with industry expertise to understand how people, organisations and technology interact across defence and security.