Martin Cooper MBCS speaks with key members of the IT Leaders’ Forum Availability Working Group about engineering the resilience of IT systems within an organisation.

Summary

  • Shift focus to recovery: true IT resilience requires shifting beyond basic defensive perimeters to engineering systems that accept inevitable failures, contain localised disruptions and prioritise the rapid recovery of core operations
  • Prioritise impact over prevention: organisations must classify ‘important business services’ — where prolonged outages cause intolerable harm — and measure resilience through end-user impact metrics (such as lost user hours) rather than just technical uptime or prevention
  • Foster psychological safety: building operational resilience requires leadership to eliminate a blame culture, encouraging teams to openly report internal system vulnerabilities and conduct realistic simulation tests without fear of reprisal

Modern society is increasingly dependent on IT systems that are complex and prone to failure, creating risks so large they can push national GDP growth into negative territory. These are the findings of extensive research published by the BCS IT Leaders Forum (ITLF).

To maintain control and protect core services, organisations must look beyond basic cyber security perimeters and adopt resilience engineering — as the draft Cyber Security and Resilience Bill emphasises.

True resilience means accepting that systems will fail, identifying vulnerabilities early, and designing processes that enable essential operations to recover smoothly when disruptions occur.

Why is resilience critical right now?

Digital infrastructure has grown increasingly fragile as systems become more interconnected. When a complex IT system collapses, the consequences extend far beyond technical teams. Outages disrupt daily life, cause severe economic loss, and can threaten the stability of entire industries.

The UK Government's Cyber Security and Resilience Bill addresses this issue by expanding requirements beyond traditional defensive perimeters. Building walls around software is no longer enough to stop every incident. Organisations must design systems with graceful degradation, localised failure containment and reliable data protection so that operations can continue even when individual components fail.
‘The complexity of the systems that we've been building means they're very fragile’, says Gill Ringland, Secretary of the ITLF. ‘It's having an impact on people's lives, and it's significantly affecting GDP.’

What is the difference between security and resilience?

Cybersecurity primarily focuses on preventing external attacks from breaching internal networks. Resilience engineering focuses on what happens during and after a failure, regardless of whether the cause was an external threat or an internal glitch.

Regulatory approaches, such as the NIS framework, measure resilience by tracking direct user impact, such as lost user hours. While internal technical metrics such as mean time to repair remain useful, organisations must prioritise the services that matter most to end users.

‘It's cyber security’, explains Gill. ‘And it’s thinking about methods of failure, methods of recovery, methods of protection of data, thinking about graceful degradation so that if there is a failure, it's localised.’

What distinguishes a resilient organisation?

A resilient organisation accepts that failure will happen and actively plans for it. Many organisations operate on the assumption that proper prevention will stop all disruptions.

When minor incidents occur, resilient organisations handle them using tested fallback procedures. For example, when retail IT infrastructure experiences issues, an organisation that can quickly revert to manual processing can continue trading, even if its online channels are temporarily disrupted. Conversely, many organisations risk catastrophic outages where core services remain offline for months or years.

Why do organisations invest more in prevention than recovery?

Many leadership teams hold a mindset that preventing disruption removes the need to invest in recovery capabilities. Investing in recovery requires testing that can uncover uncomfortable vulnerabilities within existing processes.

Simulation exercises help uncover hidden gaps between documented procedures and practical reality. In one instance, Gill explained, a support department capped its phone log capacity at 10 messages, even though the emergency procedure required 500 employees to call in during an incident. Without running realistic simulations, these operational mismatches go unnoticed until a real emergency occurs.

‘This really reflects a mindset: if disruptions are prevented, why invest in recovery capability?’ says Paul Reason. ‘However, you cannot cover every eventuality, and when [disaster] happens, you do not have the wherewithal in place to understand what the problem is and deal with it.’

What are the early warning signs of system fragility?

System fragility often develops gradually before causing a total outage. Leaders can spot creeping instability by tracking operational indicators, including:

  • Software deployments that regularly require rollbacks
  • A lack of routine, structured testing
  • Spikes in unsuccessful login attempts that could indicate ransomware or brute-force activity

Many businesses operate relying on extended supply chains without realising their risk exposure. Comprehensive logging and centralised data analysis enable teams to detect changes in activity before systems fail.

How can leadership identify modern system vulnerabilities?

Executives can assess system fragility by asking targeted questions about operational processes and observability. A structured assessment framework helps identify which departments hold critical data and who is responsible for system health.

Building resilience requires a systematic method based on four key steps:

  1. Planning: define critical workflows and potential failure points across all operations
  2. Observability: maintain visibility across software supply chains to detect real-time changes
  3. Response procedures: Establish clear protocols to contain localised failure
  4. Repair: implement tested recovery mechanisms to restore full service safely

‘There's not a silver bullet on this; it's messy’, emphasises Gill. ‘It needs a systematic approach: you have to start with planning, you have to start with good observability of your system so that you can understand what's happening, and processes for mending it when it breaks.’

How should leaders prioritise critical services?

Organisations cannot protect every application equally when budgets and resources are limited. The financial services sector introduced the concept of important business services to tackle this challenge. An important business service is defined as any service where a prolonged failure would cause intolerable harm to the organisation or its customers.

For you

Be part of something bigger, join BCS, The Chartered Institute for IT.

‘Setting the level of what constitutes intolerable sets the priority for making decisions about which services deserve the greatest protection and activity about resilience’, explains Professor Ed Steinmueller. ‘We would put important business services at the centre of efforts to decide which services need to be considered first.’

Defining intolerable harm establishes priorities for resilience investments. Frameworks like the Digital Operational Resilience Act (DORA) apply this principle across multiple sectors. Categorising core services creates productive alignment between business executives and IT teams, ensuring resources go to the systems that matter most.

Can an organisation become resilient without changing its culture?

Technical fixes alone cannot build resilience if an organisation maintains a culture of blame or overconfidence. Leadership must foster psychological safety so employees can report flaws, log errors, and discuss system failures without fear of reprisal.

‘If the leadership takes the view that systems will fail and we have to be prepared for their failure, then that can begin the process of cultural change that's necessary’, says Professor Steinmueller. ‘In that sense, it very much is a people problem.’

In blame-heavy environments, teams often attribute outages to external cyberattacks because such events carry less personal accountability. However, operational data show that external cyberattacks account for only about 10% of system failures. More than half of all outages stem from internal issues, such as failed software updates deployed across complex, unstable architecture. Addressing these root causes requires leadership to model openness and encourage honest discussion around internal risks.

What is the next step in engineering resilience for IT systems?

The ITLF has recently published a detailed resource: Availability: Engineering Resilience for IT Systems (https://tinyurl.com/ydycsyrz). This summarises and codifies much of their research over the past four years, as published on the ITLF web page.

Professor Ed Steinmueller FBCS is a Professor Emeritus of both IT and Economics at the University of Sussex. He is co-chair of the Availability Working Group of the ITLF. Paul Reason FBCS CITP has extensive experience as an Enterprise Architect and Programme Manager, and is the Secretary of the Availability Working Group of the ITLF. Gill Ringland is a Life FBCS, and author or co-author of 13 books on strategy and on IT, most recently Resilience of Services, LPP, 2024 with Ed. She is Secretary of the ITLF and co-chair of the Availability Working Group of the ITLF.