SCADA and BMS in the SOC: Which Events Truly Matter

Integrating SCADA and BMS systems into a Security Operations Center (SOC) often leads to information overload. For critical infrastructure leaders, it's crucial to identify which events are truly relevant for cybersecurity and operational resilience to avoid 'alert fatigue' and ensure effective incident response.

The convergence of Operational Technology (OT) and Information Technology (IT) creates new opportunities for monitoring and automation, but simultaneously expands the attack surface and complicates cybersecurity management. For CISOs and operational teams, integrating events from SCADA (Supervisory Control and Data Acquisition) and BMS (Building Management Systems) into a unified SOC is a strategic move. However, without clear selection criteria and processing mechanisms, this process can become a source of noise, threatening the effectiveness of responding to genuine threats.

Challenges of OT/IT event integration: Why 'everything' is too much

Traditional IT systems and OT environments have fundamental differences. While IT prioritizes confidentiality, OT prioritizes availability, safety, and reliability. SCADA and BMS systems often operate 24/7, contain legacy devices that are difficult to update or restart, and are critical for physical processes. Any downtime can lead to physical damage, environmental harm, or even human casualties.

Attempting to integrate all events from OT systems into a SOC leads to 'alert fatigue.' SOC analysts are overwhelmed by a massive volume of alerts, many of which are false positives or low-priority. This results in slowed response times, missed threats, and staff burnout. Studies show that out of thousands of daily alerts, about 83% turn out to be false positives, with only a small fraction being real threats. Consequently, the quality of investigations declines, and critical incidents may be overlooked.

Furthermore, the lack of uniform protocols and data formats across different SCADA/BMS vendors complicates integration. Legacy RTUs/PLCs may use proprietary dialects and store data in incompatible formats, requiring complex transformations and gateways.

Criteria for selecting critical events: What truly matters for security

For effective integration, it's essential to focus on events directly relevant to cybersecurity and operational resilience. CISA (Cybersecurity and Infrastructure Security Agency) and NIST (National Institute of Standards and Technology) provide recommendations for monitoring OT systems that help prioritize. Specifically, CISA Cybersecurity Performance Goals (CPG 2.0) offer a unified set of goals for critical infrastructure, covering both IT and OT environments.

Critical events for a SOC include:

  • Unauthorized access and configuration changes: Attempts at unauthorized access to controllers, HMI (Human-Machine Interface), changes to PLC logic, or SCADA configurations.
  • Abnormal device/network behavior: Unusual sensor readings, sensor disconnections, anomalous network traffic between OT segments, unauthorized connections.
  • Firmware or software modification attempts: Any attempts to update or modify OT device firmware without proper authorization.
  • Indicators of Compromise (IoC): Events matching known adversary tactics, techniques, and procedures (TTPs) outlined in frameworks like MITRE ATT&CK for ICS.
  • Events impacting safety and the environment: Any events that could lead to physical harm, environmental disasters, or the shutdown of critical production processes.
  • Authentication and authorization failures: Multiple failed login attempts, use of default accounts, privilege escalation attempts.

NIST SP 800-82 Revision 3, “Guide to Operational Technology (OT) Security,” recommends a consequence-driven risk management approach, where the highest priority is given to scenarios that could cause physical harm or failure of safety systems.

Normalization and contextual enrichment: From raw data to actionable intelligence

Raw data from SCADA and BMS systems often lack sufficient context for rapid analysis in a SOC. Normalizing events to a common format (e.g., Syslog, CEF, LEEF) and enriching them with additional information is critically important.

Contextual enrichment means adding data to an event such as:

  • Asset criticality: Is the device generating the event critical to the production process or safety?
  • Location: Where is the device located (workshop, building, line)?
  • Owner/Responsible department: Which team or individual is responsible for this asset?
  • System type: SCADA, BMS, fire safety system, access control, etc.
  • Related vulnerabilities: Are there known vulnerabilities for this type of device or software?
  • User identifiers: Who performed the action that led to the event (integration with IAM/Active Directory).

Integration with a CMDB (Configuration Management Database), Identity and Access Management (IAM) systems, and inventory systems allows for automatic event enrichment, providing SOC analysts with a complete picture of an incident. This accelerates investigations and enables quicker determination of an event's potential impact.

Defining responsibilities and response procedures: Who acts and how

The convergence of IT and OT requires clear definition of roles and responsibilities for incident response. A lack of such protocols can lead to delays and conflicts between IT and OT teams.

Key aspects:

  • RACI Matrix: Develop a RACI (Responsible, Accountable, Consulted, Informed) matrix for cybersecurity incidents involving OT/BMS.
  • Joint response plans: Create unified incident response plans that account for the specifics of OT systems, including prioritizing safety and availability. CISA recommends annual training and updating of response plans.
  • Communication channels: Ensure clear communication channels between SOC, engineering, operations, and security teams.
  • SLAs for OT incidents: Define Service Level Agreements (SLAs) for responding to incidents in OT environments, considering their criticality.
  • Training and cross-functional teams: Train both IT and OT teams on convergence risks and response specifics.

Initiatives such as ICS4ICS (Incident Command System for Industrial Control Systems), developed by ISAGCA in collaboration with CISA, offer a structured approach to incident management in ICS, based on proven emergency response systems.

Practical checklist for CISOs: Which SCADA/BMS events to integrate into the SOC

For CISOs and operational leaders, having a clear mechanism for deciding on event integration is critical. This checklist will help systematize the selection and prioritization process:

Criterion Yes/No Comments
Does the event indicate a potential cybersecurity threat (e.g., unauthorized access, configuration change, anomalous behavior)? Example: Attempted PLC logic change, unauthorized HMI login.
Does the event impact personnel safety or the environment? Example: Critical deviation in temperature, pressure, level.
Could the event lead to the shutdown of critical production processes or significant financial losses? Example: Shutdown of main production equipment, cooling system malfunction.
Is the event an Indicator of Compromise (IoC) according to known frameworks (e.g., MITRE ATT&CK for ICS)? Example: Detection of known malware, atypical network connections.
Does the event have clear context allowing a SOC analyst to quickly understand its essence and potential impact? Is it possible to enrich the event with data about the asset, its criticality, and location?
Is it technically feasible to automatically normalize and enrich this event for SIEM integration? Assessment of technical feasibility and integration effort.
Is a responsible department and response procedure defined for this event? Is there a clear incident owner and action plan?
Is it possible to filter out 'noisy' events, leaving only critical ones? Using correlation rules and thresholds.
Is the event unique and not duplicated by other monitoring sources? Avoiding data redundancy.
What is the cost of integrating and monitoring this event compared to the potential risk? Cost-benefit analysis.

How AZIOT implements this

The AZIOT platform is designed to address OT/IT integration challenges by providing centralized collection and processing of data from various protocols, such as MQTT, Modbus, BACnet, KNX, SCADA, and BMS. Through edge processing capabilities, AZIOT allows for on-site filtering and pre-processing of telemetry, reducing the volume of data transmitted to the central SOC and minimizing 'noise.' The use of rules and scenarios enables automatic event normalization and contextual enrichment, which is critical for rapid and relevant response. Unity Base's audit and access control systems ensure that only authorized users have access to critical data and functions, supporting the principles of least privilege and zero trust. Enterprise system architects, such as Serhii Boiko, can leverage these principles to design integration solutions that consolidate disparate OT/BMS systems into a manageable whole, ensuring effective monitoring and response to cybersecurity incidents in the SOC.

Effective integration of SCADA and BMS events into a SOC is not merely a technical task but a strategic decision requiring a deep understanding of both IT and OT environments. By focusing on critical events, ensuring their normalization and contextualization, and clearly defining responsibilities, organizations can build a resilient cybersecurity system that protects physical assets and ensures operational continuity. This allows for a shift from reactive to proactive defense, minimizing risks and optimizing SOC resources.

Learn more about Intecracy solutions at Intecracy solutions and inbase.com.ua solutions.

Source list

  1. ucertify.comICS Incident Response and Risk Management for Industrial Security
  2. sans.orgICS/OT Incident Response: A Guide to Industrial Cyber Resilience | SANS Institute
  3. sans.orgA Guide to OT Security Best Practices | SANS Institute
  4. opsiocloud.com
  5. rapid7.comWhat Is Alert Fatigue in Cybersecurity? | Rapid7
  6. lrqa.comWhat is alarm fatigue in cyber security? | LRQALogoCloseSearch open
  7. corelight.comThe key to alleviating alert fatigue in cybersecurity - Corelight
  8. prophetsecurity.aiAlert Fatigue in Cybersecurity: Why Tuning Isn’t Enough Anymore | Prophet Security