Simon Kupjetz, M. Sc., Fraunhofer Institute for Structural Durability and System Reliability LBF, Darmstadt
Maritime Autonomous Surface Ships (MASS) rely on AI‑based perception to support collision avoidance and COLREGs‑compliant navigation in complex environments. For classification societies, flag states, and operators, the central question is not which specific neural network is used but whether the overall architecture demonstrably keeps risk within acceptable limits over the vessel’s life cycle. This works presents a framework for quantitative risk assessment that links AI perception performance to system‑level safety in a way that is compatible with established safety standards (MIL‑STD‑882E) and emerging regulatory expectations (EU AI Act, DNV MASS guidelines, …). Starting from a functional architecture of an autonomous ship (camera, radar, sensor fusion, charts, GNSS), the authors derive perception‑related failure modes and embed it into a Bayesian Network (BN) that describes how these failures might propagate to hazardous system states and severity categories. The key idea is that the BN is parameterized using empirical performance data from the perception function (i.e. the confusion matrix) and is influenced by external factors such as visibility, sea state and sensor impairments. In a case study, a computer vision model is trained on a publicly available maritime dataset. Its classification statistics are then used to populate the BN’s conditional probability tables for different object classes and environmental conditions. Rather than focusing on the details of the vision model, the framework treats it as a component that provides measurable detection and misclassification rates for relevant maritime objects. These rates are combined with a scenario‑based severity model that assigns MIL‑STD‑882E categories (I–IV) to perception failures, reflecting the worst credible consequences under the current design and mitigations. In this way, one can quantify, for example, what share of operational situations falls into “catastrophic”, “critical”, or “acceptable” regions, considering both AI misclassification errors and environmental factors. The BN can be used interactively to perform “what‑if” analyses. Analysts can vary perception performance, operating conditions, or mitigation assumptions and immediately see the effect on the distribution of hazard severity. This allows them to answer questions that are central for approval and class: Which improvements in the perception chain (e.g. better training data, additional sensors) reduce system‑level risk? In what circumstances do current safety goals no longer apply? Which failure modes contribute most to unacceptable risk and therefore require design changes or operational limits? Because the model structure remains relatively compact and easy to manipulate, these analyses can be communicated to stakeholders without requiring deep knowledge of the underlying AI model. Compared with common practice, where AI components are often judged on accuracy metrics alone, the proposed framework explicitly links perception performance to potential navigational consequences. This enables more defensible safety cases for MASS: evidence from testing and trials can be translated into quantitative statements about residual risk and the effect of additional mitigation measures. Future work will extend the framework to multi‑sensor architectures, additional AI functions (e.g. route planning), hardware failures, and real sea‑trial data, with the goal of providing a decision‑transparent basis for design optimization and regulatory approval of AI‑enabled maritime systems.