Evaluating Maritime AI

Windward’s Maritime AI shows that smarter data can sharpen maritime risk detection—but without transparent, independent validation, bold AI performance claims remain signals, not proof.

Sanchez P.

9/18/202652 min read

Abstract

The increasing use of artificial intelligence (AI) in maritime intelligence has created new opportunities for detecting anomalous vessel behaviour, identifying potential sanctions evasion and supporting maritime compliance. Windward’s Maritime AI™ represents a commercial application of this approach, combining AIS, satellite imagery, synthetic aperture radar (SAR), radio-frequency data and behavioural analytics to generate predictive maritime risk intelligence. This paper critically examines the technological and evidential basis of Windward’s claims by comparing its stated capabilities with findings from peer-reviewed research on maritime anomaly detection, data fusion, AIS and GNSS manipulation, and machine-learning-based behavioural analysis.

The analysis finds that Windward’s underlying technological proposition is broadly consistent with established academic research. Peer-reviewed studies demonstrate that machine-learning methods can identify unusual maritime behaviour, that AIS data are vulnerable to missing observations, inconsistencies and manipulation, and that combining AIS with complementary observation sources such as SAR can improve the detection of suspected abnormal activity. However, the evidence provides substantially weaker support for Windward’s specific commercial performance claims. The publicly available information reviewed does not provide sufficient methodological detail to independently validate claims of 94% accuracy in illicit ship-to-ship transfer detection, approximately 75% fewer false positives in sanctions-related AIS spoofing detection, or 99% pre-designation identification of vessels subsequently sanctioned.

The paper argues that the principal analytical distinction is between anomaly detection, risk assessment and proof of wrongdoing. AI can identify deviations from expected behaviour and prioritise potentially significant risks, but these outputs do not independently establish causation, intent or unlawful conduct. Consequently, maritime AI should be understood primarily as an evidence-generating and decision-support technology rather than an autonomous adjudication mechanism. The study concludes that greater methodological transparency, independent validation and evaluation across multiple performance measures and operational environments are necessary for the reliable assessment of commercial maritime AI in high-stakes compliance applications.

Keywords: Maritime AI; AIS; anomaly detection; sanctions screening; ship-to-ship transfers; maritime intelligence; machine learning; GNSS spoofing; deceptive shipping practices; risk analytics

1. Introduction

Maritime transport is fundamental to global commerce, yet the structural characteristics that make shipping economically efficient also create significant challenges for maritime surveillance and risk management. International mobility, complex ownership and corporate structures, multiple jurisdictions, and operations across vast and often poorly observable ocean spaces can provide opportunities for sanctions evasion, smuggling and other forms of illicit activity. These challenges are compounded by the increasing dependence of maritime monitoring on digital traces, particularly Automatic Identification System (AIS) data, satellite imagery and other remote-sensing technologies. Importantly, the literature demonstrates that these data should not be treated as inherently reliable representations of vessel activity: AIS transmissions may be incomplete, invalid, anomalous or deliberately manipulated, making the quality and interpretation of maritime data central to effective risk detection (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Campbell, J et al., 2022).

This creates a distinctive analytical problem. The absence or manipulation of a vessel's digital signal is not necessarily a failure of surveillance; in some circumstances, it may itself constitute a meaningful behavioural indicator. Research on maritime anomaly detection has consequently moved beyond simple vessel tracking towards the identification of deviations from expected patterns of movement and behaviour. Höhle et al. (2022) and Ribeiro, Paes and de Oliveira (2023), for example, demonstrate the growing use of machine-learning and data-driven techniques to identify anomalous AIS trajectories and distinguish unusual vessel behaviour from normal operational patterns. Related research has also examined the identification of suspicious behaviour in specific maritime domains, including fishing activity, demonstrating how tracking-data anomalies can provide indicators for subsequent investigation rather than conclusive evidence of illicit conduct (Identification of suspicious behavior through anomalies in the tracking data of fishing vessels, 2024).

Windward positions its Maritime AI™ platform within this emerging technological paradigm. The company describes its platform as providing a unified maritime intelligence picture through the integration of AIS, so-called ‘dark vessel’ signals, electro-optical and synthetic aperture radar (SAR) imagery, radio-frequency information and other data sources (Windward, 2026). Its proposition therefore extends beyond conventional AIS monitoring: rather than relying on a single stream of positional information, the platform seeks to combine heterogeneous observations and apply behavioural analytics to identify patterns associated with operational and compliance risk. This approach is broadly consistent with the direction of academic research, where the combination of AIS with satellite and other observation technologies is increasingly investigated as a means of improving the detection of maritime activity that cannot be reliably identified from AIS alone. In particular, recent research on ship-to-ship transfer detection demonstrates the analytical value of combining AIS with synthetic aperture radar and other observations to identify activity that may be obscured in conventional tracking data (Cai et al., 2026).

The academic literature also provides support for the technological plausibility of detecting manipulated maritime signals. Research into invalid AIS messages shows that machine-learning techniques can be applied to distinguish potentially erroneous or abnormal transmissions from valid AIS data (Campbell, J et al. 2022). Similarly, research examining GNSS interference and spoofing demonstrates that satellite-navigation signals can be disrupted or manipulated, reinforcing the need to treat positional data as an observable signal that requires validation rather than as an unquestionable representation of vessel location (Appel, M., at al. 2019; Gattis, B. et al., 2026). The significance of this finding for maritime intelligence is substantial: if AIS or GNSS data can be manipulated, then detecting inconsistencies between different data sources may itself become an important component of risk assessment.

Against this academic background, Windward makes substantially stronger commercial claims. The company states that its models can identify illicit ship-to-ship transfers with 94% accuracy and detect intentional sanctions-related AIS spoofing with 75% fewer false positives (Windward, 2026). Its current materials further state that its models flagged 99% of vessels that were subsequently designated for sanctions in 2024 before their official designation (Windward, 2026). These claims are potentially significant because they suggest that behavioural and predictive analytics may identify indicators of sanctions-related risk before conventional screening mechanisms produce a formal designation. However, the significance of such claims depends critically on how the reported performance measures are defined, measured and independently validated.

In particular, the term ‘accuracy’ is insufficient on its own to establish the effectiveness of a risk-detection system. In anomaly-detection and classification problems, performance depends on the underlying prevalence of the behaviour being detected, the quality of the ground-truth labels, the composition of the test dataset, and the relative consequences of false positives and false negatives (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). A model may achieve a high overall accuracy while performing poorly on a relatively rare illicit activity. Conversely, a system designed to identify high-risk events may deliberately generate more alerts in order to reduce the probability of missing genuinely significant cases. Consequently, measures such as precision, recall, F1 score, false-positive and false-negative rates, calibration, and performance on independent datasets are necessary to interpret claims of predictive performance meaningfully.

This distinction is particularly important when considering the concept of ‘predictive intelligence’. Academic research supports the proposition that behavioural anomalies and multi-source observations can generate indicators of unusual or suspicious maritime activity (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Cai et al., 2026). However, identifying an anomaly is analytically different from establishing that a vessel is engaged in sanctions evasion or other illicit conduct. An anomalous trajectory, unexplained AIS gap, unusual ship-to-ship encounter or discrepancy between AIS and satellite observations may constitute a risk signal, but it does not by itself establish intent, illegality or sanctions liability. This distinction is fundamental to the responsible deployment of maritime AI because the output of an AI system is best understood as an evidential input into an investigative or compliance process rather than an autonomous determination of wrongdoing.

The same issue applies to Windward's claim that its models identified 99% of vessels subsequently sanctioned in 2024 before their designation (Windward, 2026b). The claim may indicate potentially valuable anticipatory detection, but its evidential significance cannot be assessed fully without information about the definition of ‘flagged’, the timing and duration of the alerts, the number of vessels assessed, the characteristics of the underlying population, and how many non-sanctioned vessels were simultaneously identified. A high proportion of subsequently sanctioned vessels being flagged does not, by itself, establish a high predictive value: the interpretation depends on the number of vessels flagged overall and the rate at which alerts correctly distinguish genuinely significant cases from benign or explainable behaviour. This is a general methodological problem in predictive risk systems rather than a criticism specific to Windward.

The purpose of this paper is therefore not to accept or reject Windward's commercial proposition in its entirety, but to critically examine the relationship between its claims and the peer-reviewed evidence supporting the underlying technological approach. Three questions structure the analysis. First, to what extent is Windward's multi-source, AI-driven approach consistent with established research on maritime anomaly detection, satellite-based observation, AIS validation and behavioural analysis? Second, are the company's specific performance claims sufficiently transparent and independently substantiated to support the strength of the conclusions implied by them? Third, what methodological, operational and governance limitations arise when AI-generated maritime signals are incorporated into sanctions screening, compliance and broader maritime risk-management processes?

The central issue is therefore not whether artificial intelligence can detect maritime anomalies—it demonstrably can—but whether the transition from detecting anomalous behaviour to producing reliable, predictive and decision-relevant maritime intelligence is supported by sufficiently transparent evidence. The distinction between technical capability, predictive performance and evidential reliability provides the critical framework for evaluating Windward's Maritime AI™ proposition.

2. Windward's Maritime AI Proposition

Maritime transport is fundamental to global commerce, yet the structural characteristics that make shipping economically efficient also create significant challenges for maritime surveillance and risk management. International mobility, complex ownership and corporate structures, multiple jurisdictions, and operations across vast and often poorly observable ocean spaces can provide opportunities for sanctions evasion, smuggling and other forms of illicit activity. These challenges are compounded by the increasing dependence of maritime monitoring on digital traces, particularly Automatic Identification System (AIS) data, satellite imagery and other remote-sensing technologies. Importantly, the literature demonstrates that these data should not be treated as inherently reliable representations of vessel activity: AIS transmissions may be incomplete, invalid, anomalous or deliberately manipulated, making the quality and interpretation of maritime data central to effective risk detection (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Campbell, J et al., 2022).

This creates a distinctive analytical problem. The absence or manipulation of a vessel's digital signal is not necessarily a failure of surveillance; in some circumstances, it may itself constitute a meaningful behavioural indicator. Research on maritime anomaly detection has consequently moved beyond simple vessel tracking towards the identification of deviations from expected patterns of movement and behaviour. Höhle et al. (2022) and Ribeiro, Paes and de Oliveira (2023), for example, demonstrate the growing use of machine-learning and data-driven techniques to identify anomalous AIS trajectories and distinguish unusual vessel behaviour from normal operational patterns. Related research has also examined the identification of suspicious behaviour in specific maritime domains, including fishing activity, demonstrating how tracking-data anomalies can provide indicators for subsequent investigation rather than conclusive evidence of illicit conduct (Identification of suspicious behavior through anomalies in the tracking data of fishing vessels, 2024).

Windward positions its Maritime AI™ platform within this emerging technological paradigm. The company describes its platform as providing a unified maritime intelligence picture through the integration of AIS, so-called ‘dark vessel’ signals, electro-optical and synthetic aperture radar (SAR) imagery, radio-frequency information and other data sources (Windward, 2026a). Its proposition therefore extends beyond conventional AIS monitoring: rather than relying on a single stream of positional information, the platform seeks to combine heterogeneous observations and apply behavioural analytics to identify patterns associated with operational and compliance risk. This approach is broadly consistent with the direction of academic research, where the combination of AIS with satellite and other observation technologies is increasingly investigated as a means of improving the detection of maritime activity that cannot be reliably identified from AIS alone. In particular, recent research on ship-to-ship transfer detection demonstrates the analytical value of combining AIS with synthetic aperture radar and other observations to identify activity that may be obscured in conventional tracking data (Cai et al., 2026).

The academic literature also provides support for the technological plausibility of detecting manipulated maritime signals. Research into invalid AIS messages shows that machine-learning techniques can be applied to distinguish potentially erroneous or abnormal transmissions from valid AIS data (Campbell, J et al., 2022). Similarly, research examining GNSS interference and spoofing demonstrates that satellite-navigation signals can be disrupted or manipulated, reinforcing the need to treat positional data as an observable signal that requires validation rather than as an unquestionable representation of vessel location (Appel, M., at al. 2019; Gattis, B. et al., 2026). The significance of this finding for maritime intelligence is substantial: if AIS or GNSS data can be manipulated, then detecting inconsistencies between different data sources may itself become an important component of risk assessment.

Against this academic background, Windward makes substantially stronger commercial claims. The company states that its models can identify illicit ship-to-ship transfers with 94% accuracy and detect intentional sanctions-related AIS spoofing with 75% fewer false positives (Windward, 2026). Its current materials further state that its models flagged 99% of vessels that were subsequently designated for sanctions in 2024 before their official designation (Windward, 2026). These claims are potentially significant because they suggest that behavioural and predictive analytics may identify indicators of sanctions-related risk before conventional screening mechanisms produce a formal designation. However, the significance of such claims depends critically on how the reported performance measures are defined, measured and independently validated.

In particular, the term ‘accuracy’ is insufficient on its own to establish the effectiveness of a risk-detection system. In anomaly-detection and classification problems, performance depends on the underlying prevalence of the behaviour being detected, the quality of the ground-truth labels, the composition of the test dataset, and the relative consequences of false positives and false negatives (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). A model may achieve a high overall accuracy while performing poorly on a relatively rare illicit activity. Conversely, a system designed to identify high-risk events may deliberately generate more alerts in order to reduce the probability of missing genuinely significant cases. Consequently, measures such as precision, recall, F1 score, false-positive and false-negative rates, calibration, and performance on independent datasets are necessary to interpret claims of predictive performance meaningfully.

This distinction is particularly important when considering the concept of ‘predictive intelligence’. Academic research supports the proposition that behavioural anomalies and multi-source observations can generate indicators of unusual or suspicious maritime activity (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Cai et al., 2026). However, identifying an anomaly is analytically different from establishing that a vessel is engaged in sanctions evasion or other illicit conduct. An anomalous trajectory, unexplained AIS gap, unusual ship-to-ship encounter or discrepancy between AIS and satellite observations may constitute a risk signal, but it does not by itself establish intent, illegality or sanctions liability. This distinction is fundamental to the responsible deployment of maritime AI because the output of an AI system is best understood as an evidential input into an investigative or compliance process rather than an autonomous determination of wrongdoing.

The same issue applies to Windward's claim that its models identified 99% of vessels subsequently sanctioned in 2024 before their designation (Windward, 2026). The claim may indicate potentially valuable anticipatory detection, but its evidential significance cannot be assessed fully without information about the definition of ‘flagged’, the timing and duration of the alerts, the number of vessels assessed, the characteristics of the underlying population, and how many non-sanctioned vessels were simultaneously identified. A high proportion of subsequently sanctioned vessels being flagged does not, by itself, establish a high predictive value: the interpretation depends on the number of vessels flagged overall and the rate at which alerts correctly distinguish genuinely significant cases from benign or explainable behaviour. This is a general methodological problem in predictive risk systems rather than a criticism specific to Windward.

The purpose of this paper is therefore not to accept or reject Windward's commercial proposition in its entirety, but to critically examine the relationship between its claims and the peer-reviewed evidence supporting the underlying technological approach. Three questions structure the analysis. First, to what extent is Windward's multi-source, AI-driven approach consistent with established research on maritime anomaly detection, satellite-based observation, AIS validation and behavioural analysis? Second, are the company's specific performance claims sufficiently transparent and independently substantiated to support the strength of the conclusions implied by them? Third, what methodological, operational and governance limitations arise when AI-generated maritime signals are incorporated into sanctions screening, compliance and broader maritime risk-management processes?

The central issue is therefore not whether artificial intelligence can detect maritime anomalies—it demonstrably can—but whether the transition from detecting anomalous behaviour to producing reliable, predictive and decision-relevant maritime intelligence is supported by sufficiently transparent evidence. The distinction between technical capability, predictive performance and evidential reliability provides the critical framework for evaluating Windward's Maritime AI™ proposition.

3. Claim One: Multi-Source Intelligence Improves Maritime Risk Detection

Windward presents multi-source data integration as a central source of its Maritime AI™ platform's value. The company states that its system incorporates more than 30 data sources and argues that this redundancy enables information to be cross-checked while maintaining analytical resilience when individual sources are incomplete, unavailable or manipulated (Windward, 2026). The proposition is therefore not simply that more data produces better intelligence, but that combining different forms of maritime observation can reduce dependence on any single signal and improve the ability to identify inconsistencies in vessel behaviour.

The academic literature provides substantive support for this underlying proposition, particularly where different sensing technologies have complementary strengths and weaknesses. Cai et al. (2026), for example, developed a multi-source observation framework combining synthetic aperture radar (SAR) and AIS to detect suspected abnormal ship-to-ship (STS) activity. Their research is particularly relevant to Windward's proposition because it demonstrates why the combination of physical and digital observations can be analytically valuable. AIS can provide information about a vessel's reported identity, position and movement, but is vulnerable to transmission failures, intentional shutdown and manipulation. SAR, by contrast, can provide physical observation of vessels independently of their AIS transmissions, although imagery alone does not necessarily establish the digital identity of the vessel being observed (Cai et al., 2026).

The analytical value of combining these sources therefore lies partly in their complementarity. Where AIS and SAR observations are consistent, confidence in the observed activity may be strengthened; where they diverge, the discrepancy itself may constitute a signal requiring further investigation. This is particularly relevant to maritime environments in which vessels may attempt to obscure their movements or identities. Cai et al. (2026) demonstrate this principle through experiments conducted in the Panama, Singapore and Gibraltar regions, finding that SAR-AIS integration can assist in identifying suspected abnormal STS activity and potential inconsistencies between reported and physically observed vessel activity.

The findings provide meaningful academic support for Windward's broader architectural proposition. They demonstrate that combining heterogeneous maritime observations can provide analytical information that may not be available from either source independently. This is consistent with the wider literature on maritime anomaly detection, which identifies the limitations of AIS-only approaches and increasingly considers the use of complementary data and sensing technologies to improve the observation of anomalous behaviour (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). Multi-source analysis can consequently be understood as a means of addressing an information problem: when no individual data source provides a complete or consistently reliable representation of maritime activity, combining observations can potentially improve situational awareness and anomaly detection.

However, an important evidential distinction must be maintained between demonstrating the usefulness of data fusion and demonstrating the superiority of a particular commercial implementation. Cai et al. (2026) evaluate a defined research framework using specified datasets, locations and experimental procedures. The resulting findings therefore provide evidence about the performance and feasibility of that particular methodological approach under those conditions. They do not establish that all multi-source maritime AI systems will perform similarly, nor do they demonstrate the effectiveness of Windward's proprietary platform.

This distinction is especially important because Windward's claim of using more than 30 sources describes the scale of its data architecture but does not, by itself, establish the incremental analytical value of those sources. The number of inputs does not necessarily correspond to the quality of the resulting intelligence. For additional sources to improve risk assessment, they must contribute information that is sufficiently accurate, timely, relevant and, where possible, independently informative. Otherwise, data proliferation may increase analytical complexity without producing a proportional improvement in predictive performance. The academic literature's continuing concerns regarding data quality, inconsistent datasets, limited labelled anomalies and evaluation methodology reinforce the importance of this qualification (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

There is also a further methodological issue concerning the interpretation of cross-source discrepancies. An inconsistency between AIS and SAR, for example, may indicate manipulation, but it may also arise from differences in observation time, positional accuracy, vessel identification, data latency or other benign causes. Consequently, the identification of a discrepancy should be interpreted as an investigative signal rather than definitive evidence of misconduct. This distinction is consistent with the wider anomaly-detection literature, in which the detection of statistically or behaviourally unusual observations does not necessarily establish the underlying cause of the anomaly (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

Windward's multi-source architecture can therefore be regarded as scientifically credible in principle, particularly because independent research demonstrates the value of combining AIS with other observation technologies such as SAR (Cai et al., 2026). Nevertheless, the available evidence does not justify extending this conclusion into an assertion that Windward's particular implementation has demonstrated superior performance. The appropriate distinction is between evidence supporting the technological architecture and evidence validating the commercial product. The former is supported by emerging peer-reviewed research; the latter requires independent evidence concerning Windward's datasets, model development, validation procedures, benchmark comparisons and performance under operational conditions.

The strongest conclusion is therefore deliberately qualified: multi-source maritime intelligence has a credible scientific basis because complementary observations can expose inconsistencies that individual data streams may miss; however, the existence of a technically plausible architecture does not in itself demonstrate that Windward's specific platform is more accurate, resilient or predictive than alternative approaches. Establishing that stronger claim requires transparent and independently reproducible performance evidence.

4. Claim Two: AI Can Detect Anomalous or Deceptive Vessel Behaviour

Windward places considerable emphasis on behavioural analytics as a means of identifying deceptive maritime practices. Its Maritime AI™ platform claims to detect patterns associated with AIS manipulation, GNSS manipulation, ‘dark’ vessel activity and unusual ship-to-ship (STS) behaviour (Windward, 2026). The underlying proposition is that vessel behaviour can be modelled sufficiently to identify deviations from expected patterns and, from those deviations, generate indicators of operational or compliance risk. This represents an important development beyond conventional rule-based monitoring because the analytical focus shifts from identifying predefined events to detecting behavioural patterns that may not conform to an established baseline.

There is substantial peer-reviewed support for the technical feasibility of this approach. Machine-learning techniques have been applied to the detection of anomalous maritime trajectories, invalid AIS transmissions and other forms of abnormal vessel behaviour (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). In particular, Campbell, J et al. (2022) demonstrates that classification algorithms, including decision trees, random forests and neural networks, can identify fabricated or invalid AIS observations with high F1 scores under experimental conditions. The findings provide evidence that machine-learning models can extract behavioural or technical patterns from AIS data that are difficult to identify reliably through simple rule-based approaches.

Research on suspicious fishing activity provides further support for behavioural anomaly detection. Identification of suspicious behavior through anomalies in the tracking data of fishing vessels (2024) demonstrates how deviations in AIS-derived movement patterns can provide useful indicators of potentially suspicious activity. Significantly, however, the research does not equate an anomalous observation with intentional wrongdoing. An unusual trajectory may be associated with deliberate manipulation, but it may also arise from legitimate operational circumstances or limitations in the underlying tracking data. This qualification is central to interpreting the wider literature on maritime AI.

The distinction can be expressed as a hierarchy of analytical claims:

anomaly detection → risk indication → attribution of intent.

The first involves identifying behaviour that differs from an established or learned pattern. The second involves interpreting that deviation as potentially relevant to a particular risk category. The third involves determining why the deviation occurred and, in particular, whether it reflects deliberate conduct. The evidential requirements become progressively stronger at each stage. The academic literature provides considerable evidence for the feasibility of anomaly detection, but substantially less basis for treating anomalous behaviour as direct evidence of intent (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

This distinction is crucial because maritime behaviour is inherently context-dependent. A vessel deviating from its expected trajectory may be responding to weather, traffic conditions, navigational constraints, port instructions, operational requirements, mechanical problems or other legitimate circumstances. Similarly, an interruption in AIS transmission may reflect technical failure or coverage limitations rather than deliberate concealment. An anomaly-detection model can identify that an observation is inconsistent with expected behaviour; it cannot necessarily determine the causal explanation for that inconsistency from the observation alone.

The problem is therefore partly one of causal inference. Machine-learning models are generally effective at identifying statistical relationships and patterns in data, but the presence of a correlation between a particular behavioural pattern and previously identified suspicious activity does not necessarily establish that the same underlying cause is present in a new case. A model trained on historical examples may recognise a combination of features associated with known deceptive behaviour without establishing that those features were deliberately produced for the same purpose in every subsequent observation. This limitation is particularly important when model outputs are used in sanctions screening or other high-consequence compliance decisions.

Windward's proposition becomes more ambitious when behavioural anomalies are translated into concepts such as ‘deceptive practices’, ‘risk’, ‘intentional’ manipulation or sanctions-related activity (Windward, 2026). These terms imply more than the identification of statistical abnormality. For example, identifying that a vessel's AIS behaviour is inconsistent with expected patterns is analytically different from concluding that the vessel intentionally manipulated its AIS signal. Establishing the latter requires additional evidence capable of distinguishing deliberate conduct from technical error, environmental conditions or legitimate operational behaviour.

The literature on AIS and GNSS manipulation reinforces this distinction. Research demonstrates that maritime positioning and identification signals can be invalid, manipulated or otherwise unreliable, establishing the technical possibility of deceptive digital behaviour (Campbell, J et al., 2022). Research concerning GNSS interference and spoofing similarly demonstrates that navigation signals can be disrupted or manipulated, highlighting the vulnerability of digital positioning systems as sources of maritime evidence (Appel et al., 2019; Gattis, B et al,, 2026). However, demonstrating that manipulation is technically possible does not establish that a particular anomaly was intentionally produced. Detection of the signal abnormality and attribution of its cause remain analytically distinct tasks.

This creates an important limitation for the interpretation of AI-generated maritime risk scores. A model may be highly effective at identifying vessels whose observed behaviour resembles previously identified suspicious cases while still being unable to establish whether the behaviour is intentional. The resulting output should therefore be understood as a risk signal or investigative lead, rather than as a definitive finding of misconduct. Such an interpretation is consistent with the wider maritime anomaly-detection literature, which treats anomalous behaviour as a basis for further analysis rather than as an unequivocal classification of illegality (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

The distinction also has implications for explainability. If an AI system identifies a vessel as presenting elevated risk, users need to understand which observable features contributed to that assessment and what alternative explanations remain plausible. An explanation that merely identifies an anomaly does not necessarily explain its cause. Consequently, ‘explainable intelligence’ should not be interpreted simply as the ability to display the factors associated with a model output; meaningful explanation also requires sufficient contextual information for analysts to evaluate competing interpretations of the observed behaviour.

Windward's behavioural-analytics proposition is therefore supported at the level of pattern recognition and anomaly detection, where peer-reviewed research demonstrates that machine learning can identify unusual maritime trajectories, invalid AIS observations and potentially suspicious behavioural patterns (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Campbell, J et al., 2022). The evidence is considerably less conclusive when the analytical claim moves from detecting an anomaly to establishing intent, deception or sanctions-related misconduct. The critical issue is not whether AI can recognise unusual behaviour—it can—but whether the available evidence is sufficient to determine why the behaviour occurred.

Accordingly, the strongest interpretation of Windward's behavioural AI claims is that the technology can potentially improve the identification and prioritisation of maritime anomalies for investigation. It should not, without additional contextual and corroborating evidence, be treated as an autonomous mechanism for determining intent or illegality. The critical boundary is between recognising that behaviour is unusual and establishing why it is unusual. That boundary represents one of the most important methodological limitations when AI-based behavioural analytics are applied to maritime compliance and risk intelligence.

5. Claim Three: Detecting Ship-to-Ship Transfers

One of Windward's more specific and consequential performance claims is that its technology can identify illicit ship-to-ship (STS) activity with 94% accuracy (Windward, 2026). The claim is potentially significant because STS transfers occupy an analytically ambiguous position in maritime risk assessment. An STS transfer is not inherently illicit: vessels may legitimately transfer cargo at sea for commercial, logistical and operational reasons. At the same time, STS activity can be used to obscure cargo movements, complicate the identification of cargo provenance, or reduce the transparency of vessel-to-vessel transactions. The analytical challenge is therefore not simply to detect whether an STS event has occurred, but to distinguish ordinary maritime operations from activity warranting further investigation.

Recent peer-reviewed research supports the feasibility of detecting and analysing STS activity through maritime data. Cai et al. (2026), for example, develop a multi-source observation framework combining synthetic aperture radar (SAR) and AIS to identify suspected abnormal STS activity. Their findings are particularly relevant to Windward's proposition because they demonstrate how complementary observations can improve the identification of potentially unusual transfers. AIS can provide information concerning reported vessel identity and movement, while SAR can provide physical observations that are less dependent on the vessels' own transmissions (Cai et al., 2026).

However, the terminology used by Cai et al. (2026) is methodologically important. Their research identifies suspected abnormal STS activity, rather than establishing that a detected transfer is necessarily illicit. This distinction mirrors the broader problem identified throughout the maritime anomaly-detection literature: identifying an unusual or suspicious pattern is analytically different from determining its underlying cause or legal status (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). The same distinction is therefore necessary when interpreting Windward's claim of 94% accuracy. Detecting STS activity, detecting abnormal STS activity and determining that an STS event is illicit are three different classification problems, each requiring different forms of ground truth.

The interpretation of the reported 94% accuracy is consequently dependent on information that Windward's publicly available material does not fully disclose. Accuracy is calculated as the proportion of correctly classified observations across the total dataset. While this can provide a useful summary measure, it may be misleading when the classes being detected are highly imbalanced. In maritime surveillance, legitimate vessel movements and legitimate STS activity are likely to substantially outnumber genuinely illicit events. Under such circumstances, a classifier could achieve a high overall accuracy while still performing poorly in identifying the relatively rare illicit cases.

The problem is particularly important because the costs of classification errors are asymmetric. A false negative occurs when an illicit event is not detected, potentially allowing significant activity to escape further scrutiny. A false positive, by contrast, occurs when legitimate activity is incorrectly classified as suspicious, potentially generating unnecessary investigations, operational costs or reputational consequences for vessels and counterparties. Consequently, the practical value of an STS detection system cannot be established from overall accuracy alone. The relative frequency and consequences of these two types of error must also be considered.

This issue is consistent with the methodological challenges identified in the broader maritime anomaly-detection literature. Höhle et al. (2022) identify limited labelled anomalies, inconsistent datasets and the absence of a universally established ground-truth dataset as important limitations in evaluating maritime anomaly-detection approaches. Ribeiro, Paes and de Oliveira (2023) similarly identify missing labels, heterogeneous data and difficulties in evaluating detection performance as continuing challenges. These limitations are directly relevant to any commercial claim involving the classification of rare and potentially consequential maritime events.

A rigorous assessment of Windward's 94% figure would therefore require considerably more information than the headline percentage provides. At minimum, evaluation would need to establish:

  • the population and number of STS events used in model development and testing;

  • the geographic distribution of those events, since performance in one maritime region may not generalise to another;

  • the operational definition of ‘illicit’ STS activity;

  • the method used to establish ground truth, including whether classifications were based on regulatory findings, confirmed enforcement cases, expert annotation or another source;

  • the proportion of legitimate and illicit events in the evaluation dataset;

  • precision and recall, which indicate respectively how frequently positive classifications are correct and how effectively relevant events are identified;

  • F1 score, which provides a combined measure of precision and recall;

  • false-positive and false-negative rates;

  • the baseline or benchmark against which the 94% performance is being compared; and

  • performance on an independent test dataset not used for model development or calibration.

These requirements are not merely technical preferences. They determine what the reported 94% actually means. For example, a model achieving 94% accuracy on a balanced experimental dataset may have very different operational value from a model achieving the same figure in a dataset reflecting the much lower prevalence of genuinely illicit STS activity. Similarly, performance obtained from a geographically concentrated dataset may not generalise to different shipping routes, vessel populations or operational environments. The absence of transparent information about these dimensions therefore limits the extent to which the headline figure can be interpreted as evidence of real-world predictive performance.

There is also a fundamental distinction between benchmark performance and operational effectiveness. Even a model that performs strongly on an independently constructed test dataset may encounter different conditions when deployed operationally. Maritime environments change over time, vessel behaviour evolves, new evasion techniques emerge and the availability and quality of observation data can vary by region. The academic literature's emphasis on dynamic maritime data and changing anomaly patterns reinforces the importance of evaluating models beyond a single static performance measure (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

Windward's 94% claim should therefore be interpreted cautiously. The peer-reviewed literature provides credible evidence that STS activity can be detected using machine-learning and multi-source observation techniques, particularly through the combination of AIS and SAR data (Cai et al., 2026). However, this evidence does not independently validate Windward's specific accuracy figure. Nor does it establish that the company's definition of ‘illicit STS’ corresponds directly to the ‘suspected abnormal STS’ classification investigated in academic research.

The most defensible conclusion is therefore that the academic literature supports the feasibility of automated STS anomaly detection, but does not provide sufficient evidence to independently substantiate Windward's reported 94% accuracy for identifying illicit STS activity. Until the underlying dataset, ground-truth methodology, class distribution, evaluation metrics and independent test performance are disclosed, the 94% figure should be treated as a company-reported performance claim rather than independently established evidence of 94% real-world accuracy. The distinction is important because the scientific question is not simply whether a model can classify observations correctly under specified conditions, but whether its performance remains reliable when applied to the heterogeneous, dynamic and highly consequential environment of operational maritime risk assessment.

6. Claim Four: AIS and GNSS Spoofing Detection

Windward also claims that its proprietary models can detect intentional manipulation of AIS and GNSS signals and that its approach produces approximately 75% fewer false positives than an unspecified benchmark (Windward, 2026). The claim addresses an important problem in maritime intelligence. Digital positioning and identification systems provide much of the information used to monitor vessel movements, yet these signals are not inherently immune to error, disruption or deliberate manipulation. The academic literature therefore provides a strong basis for treating signal integrity as a substantive maritime-security and compliance concern.

Research on AIS demonstrates that maritime identification data can contain invalid or fabricated observations and that machine-learning techniques can assist in identifying anomalous messages (Campbell, J et al., 2022). This is consistent with the broader literature on AIS anomaly detection, which identifies missing observations, inconsistent transmissions and abnormal trajectories as persistent challenges for automated maritime monitoring (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). These findings support the general proposition that analysing the internal consistency and behavioural characteristics of AIS transmissions can provide information that is relevant to identifying potentially manipulated or unreliable data.

The security of maritime positioning systems presents a related challenge. Experimental research demonstrates the feasibility of detecting GNSS repeater attacks in maritime applications, highlighting the vulnerability of satellite-navigation signals to deliberate interference and manipulation (Appel et al., 2019). More recent research in the Baltic Sea has further demonstrated the detection and localisation of GNSS jamming and spoofing sources in real time, providing evidence that interference with positioning signals represents a practical maritime security concern (Gattis, Cydejko and Akos, 2026). Taken together, these studies indicate that GNSS manipulation and interference are not merely theoretical possibilities but represent genuine challenges to the reliability and integrity of the digital positioning signals used to monitor maritime activity.

The academic evidence therefore provides support for two elements of Windward's proposition. First, AIS and GNSS signals can be unreliable or deliberately manipulated. Second, analytical techniques, including machine learning and the comparison of multiple observations, can potentially assist in identifying such abnormalities (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Campbell, J et al., 2022). However, these findings do not independently substantiate Windward's specific claim of approximately 75% fewer false positives.

The methodological problem begins with the meaning of ‘false positive’. A false positive occurs when a system classifies an observation as suspicious or manipulated when the underlying behaviour is subsequently established to be benign. Measuring a reduction in false positives therefore requires a reliable ground truth against which the model's classifications can be assessed. This is particularly challenging in maritime environments because an apparent AIS or GNSS anomaly may have several possible explanations. A discrepancy between reported and observed vessel position could result from deliberate spoofing, technical malfunction, transmission problems, data latency, sensor limitations or other operational circumstances. Without sufficiently reliable ground-truth labels, it may be difficult to determine whether an alert represents a genuine false positive or an unresolved case.

Windward's public materials also do not appear to provide sufficient methodological detail to determine what the 75% reduction is measured against (Windward, 2026). In particular, the public claim does not clearly establish whether ‘industry standard’ refers to a named commercial system, a conventional rule-based approach, a specific benchmark algorithm or another defined baseline. Nor is the relevant test population, observation period, geographical coverage or prevalence of genuine AIS/GNSS manipulation clearly specified. Without these details, the percentage cannot readily be reproduced or compared with results reported in the peer-reviewed literature.

The interpretation of false-positive rates is particularly sensitive to the composition of the evaluation dataset. Maritime anomaly-detection research identifies substantial variation in datasets, limited labelled examples and the absence of a universally accepted ground-truth dataset, all of which can affect reported model performance (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). A system evaluated on a population containing a high proportion of benign or unexplained anomalies may therefore produce a different false-positive rate from the same system tested in an environment where genuinely manipulated signals are more prevalent. Similarly, changes in how a positive case is defined or labelled can materially alter the resulting performance measures. Consequently, a reported percentage reduction in false positives has limited interpretive value unless the underlying evaluation population, denominator, classification criteria and ground-truth methodology are disclosed.

This issue is particularly important when considering Windward’s claim of approximately 75% fewer false positives in sanctions-related AIS spoofing detection. A reduction in false positives does not necessarily demonstrate an overall improvement in detection performance. Classification systems inherently involve a trade-off between identifying relevant cases and avoiding unnecessary alerts, and changes in the threshold for generating an alert can affect both outcomes. A system may produce fewer false alarms simply by becoming more conservative in flagging suspicious activity; however, the same change may increase the number of genuine cases that are missed. In a maritime compliance environment, these false negatives may be particularly consequential because an undetected instance of significant manipulation could carry greater operational implications than an additional investigative alert.

The underlying technical problem is therefore not simply how many false alerts a system generates, but how effectively it distinguishes relevant cases from benign or ambiguous behaviour. Research on maritime anomaly detection reinforces this point because unusual vessel behaviour does not necessarily indicate deliberate manipulation or wrongdoing. AIS anomalies may result from data-quality problems, operational circumstances or other legitimate causes, while studies of suspicious vessel behaviour similarly demonstrate the value of anomalies as indicators for further investigation rather than definitive evidence of intentional misconduct (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Identification of suspicious behavior through anomalies in the tracking data of fishing vessels, 2024). Likewise, research demonstrates that GNSS interference and spoofing are technically feasible in maritime environments, but detecting anomalous positioning signals remains distinct from establishing the cause or intent behind them (Appel et al., 2019; Gattis, Cydejko and Akos, 2026).

A rigorous evaluation must therefore consider both sides of the classification problem. Precision measures the proportion of positive alerts that are genuinely positive, while recall measures the proportion of genuine positive cases that the system successfully identifies. F1 score provides a combined measure of precision and recall, while false-positive and false-negative rates provide additional information about the nature of the system’s errors. These measures are particularly important for maritime compliance applications because genuinely manipulated AIS or GNSS signals may represent a relatively rare subset of a much larger population of benign, ambiguous or otherwise anomalous observations. In such settings, headline reductions in false positives can be misleading if they are not considered alongside recall and false-negative performance.

Accordingly, Windward’s claim of fewer false positives should be interpreted as an incomplete performance indicator rather than evidence of superior detection capability in isolation. Establishing that the reduction represents a genuine improvement would require disclosure of the evaluation population, ground-truth methodology, baseline against which the reduction was calculated, classification threshold and corresponding changes in recall or false-negative rates. Without these details, the reported percentage indicates a claimed change in one dimension of model performance, but does not establish whether the system provides more reliable overall detection of maritime manipulation.

The issue can be represented conceptually as a four-part classification problem:

  • True positive: genuine manipulation correctly identified;

  • False positive: benign behaviour incorrectly classified as manipulation;

  • True negative: benign behaviour correctly recognised as benign; and

  • False negative: genuine manipulation not detected.

Windward's claim focuses on the second category, but the operational value of a maritime risk-detection system depends on the distribution across all four categories. A 75% reduction in false positives could represent a substantial improvement if recall remains stable or improves. Conversely, the same reduction could be achieved at the expense of materially lower recall. Without disclosure of both dimensions, the direction and magnitude of the overall performance improvement cannot be established.

There is a further distinction between detecting manipulation and establishing intentional manipulation. The academic literature provides evidence that AIS messages can be fabricated or invalid and that GNSS signals can be spoofed or disrupted. However, identifying an anomalous signal does not necessarily establish that the anomaly was intentionally produced. As with other forms of maritime behavioural analytics, technical anomaly detection and attribution of intent should therefore be treated as separate analytical tasks (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). Additional contextual evidence may be required before an anomaly can reasonably be characterised as deliberate deception.

Windward's multi-source approach could potentially assist with this problem by allowing discrepancies between independent observations to be investigated. For example, an inconsistency between AIS-reported position and another observation source may increase the evidential basis for investigating possible manipulation. Nevertheless, the presence of multiple conflicting signals does not itself establish intent. Their interpretation depends on the reliability, timing, resolution and independence of the underlying observations.

Overall, the peer-reviewed literature provides substantial support for Windward's underlying technological proposition: AIS and GNSS manipulation represent genuine maritime-security challenges, and analytical methods can assist in identifying anomalous or potentially manipulated signals (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Campbell, J et al., 2022). However, the available evidence does not independently establish the company's reported 75% reduction in false positives. To substantiate that claim, Windward would need to disclose the benchmark against which the reduction is measured, the composition and size of the test population, the definition and ground-truthing of manipulation events, the evaluation period and geographic scope, and—critically—the corresponding effect on recall and false-negative rates.

The appropriate conclusion is therefore qualified: the academic literature supports the problem Windward is attempting to solve, but the reported reduction in false positives cannot be interpreted as evidence of superior detection performance without transparency about the underlying benchmark and the corresponding trade-off between false positives and false negatives. In a maritime compliance environment, fewer alerts are not necessarily better intelligence; the relevant question is whether unnecessary alerts are reduced without sacrificing the detection of genuinely significant manipulation.

7. Claim Five: Predicting Sanctions Before Designation

Perhaps the most consequential of Windward's performance claims is that its models flagged 99% of vessels that were subsequently designated for sanctions before their official designation (Windward, 2026). Windward presents this capability as evidence of predictive risk intelligence, suggesting that behavioural and other maritime indicators can identify elevated risk before a vessel appears on an official sanctions list. The claim is potentially significant because it positions maritime AI not simply as a tool for screening against known sanctions subjects, but as a mechanism for identifying emerging risk before formal government action occurs.

The conceptual distinction between conventional sanctions screening and predictive risk intelligence is important. Conventional list-based screening is, by definition, dependent on an existing designation: once a vessel or associated entity appears on an applicable sanctions list, screening systems can identify the relevant match. Predictive analytics attempts to operate earlier in the process by identifying behavioural, transactional or relational characteristics associated with elevated risk. Windward's claim therefore concerns a more difficult analytical task than simply identifying an already designated vessel. It implies that patterns observable in maritime data may provide an early indication of risk before a formal designation has occurred (Windward, 2026).

The broader academic literature provides a plausible foundation for this proposition. Research on maritime anomaly detection demonstrates that machine-learning techniques can identify deviations in vessel trajectories and other behavioural patterns (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). Research combining AIS and SAR further demonstrates the potential for multi-source observations to identify suspicious or abnormal maritime activity that may not be apparent from a single data stream (Cai et al., 2026). These findings support the general possibility that behavioural indicators can provide information about maritime risk before an external authority formally classifies a vessel as problematic.

However, the interpretation of Windward’s 99% statistic requires particular methodological caution. At most, the statistic indicates that a very high proportion of vessels subsequently designated for sanctions had previously generated a qualifying Windward risk signal. It does not establish that Windward’s system independently determined that those vessels were engaged in sanctionable conduct, nor does it demonstrate that the system’s signals were the principal or decisive basis for the subsequent government designations. The statistic therefore establishes a temporal association between an earlier model-generated signal and a later sanctions designation; it does not, without further evidence, establish predictive causality or independent predictive validity.

This distinction can be understood through the difference between temporal precedence and predictive discrimination. A model can identify a vessel before a subsequent event occurs and therefore satisfy a temporal definition of prediction, while simultaneously generating a large number of alerts for vessels that are never subsequently sanctioned. For a predictive system to demonstrate meaningful discrimination, evaluation must establish not only how many eventual positive cases were identified, but also how many negative cases were incorrectly or unnecessarily flagged. This is consistent with the broader methodological concerns identified in maritime anomaly-detection research, where incomplete labels, inconsistent datasets and the absence of universally accepted ground-truth populations complicate the interpretation and comparison of model performance (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

The first major limitation of the 99% claim is therefore the denominator. If the evaluation population consists only of vessels that were eventually sanctioned, the statistic describes the proportion of positive cases that had previously generated a qualifying signal. Conceptually, this resembles a measure of recall or sensitivity rather than a complete assessment of predictive performance. It does not indicate how many vessels that were never sanctioned also generated comparable alerts. Without that broader population, it is not possible to assess the system’s false-positive rate or positive predictive value, nor to determine how selectively the model distinguishes vessels that will subsequently be sanctioned from the much larger population of vessels that will not.

This limitation is particularly important because the size of the non-sanctioned population is likely to be substantially larger than the population of vessels subsequently designated for sanctions. A system could therefore identify 99% of eventually sanctioned vessels while simultaneously producing very different numbers of alerts among vessels that are never sanctioned. For example, two hypothetical systems could both identify 99% of subsequently sanctioned vessels, while one generates relatively few additional alerts and the other flags a very large number of non-sanctioned vessels. The headline statistic would be identical, but the systems would have materially different precision, false-positive profiles and operational implications. The 99% figure alone therefore cannot establish the discriminative value of the underlying model.

There is also a further selection problem in evaluating predictive performance solely from vessels that were eventually sanctioned. Starting the evaluation with the eventual positive cases makes it possible to demonstrate how many of those cases were previously identified, but it does not test whether the model can reliably distinguish those cases prospectively from comparable vessels that never become sanctions subjects. A stronger evaluation would therefore require a clearly defined population containing both subsequently sanctioned and non-sanctioned vessels, together with a specified observation period and a consistent definition of what constitutes a qualifying pre-designation signal.

The timing of the signal is also relevant. A vessel identified months before designation represents a different form of predictive value from one flagged shortly before sanctions are announced, particularly where the purpose of the system is to support proactive compliance decisions. Meaningful evaluation should therefore report not simply whether a vessel was flagged before designation, but when it was first flagged, how consistently the risk signal persisted, and whether the lead time was sufficient to support an actionable intervention. Temporal separation between model evaluation and subsequent outcomes would further strengthen the evidence by reducing the possibility that performance is being assessed retrospectively.

Finally, the subsequent government designation should not automatically be treated as an independent confirmation of the model’s underlying inference. Sanctions decisions may incorporate information, investigative processes and evidential sources that are unavailable to commercial systems. A correlation between a Windward alert and a later designation may therefore demonstrate that the alert contained information associated with subsequent sanctions action, but it does not establish that the model independently reproduced the evidential reasoning underlying the government decision.

Accordingly, Windward’s 99% statistic may provide evidence that its risk signals contain potentially useful information about vessels that are subsequently sanctioned, but the statistic alone does not establish strong predictive discrimination. Demonstrating such capability would require evaluation against the broader population of vessels, including those that are never sanctioned, together with measures of precision, recall, false-positive and false-negative performance, calibration and the temporal lead time of alerts. The distinction is fundamental: identifying most future positive cases is not the same as reliably distinguishing future positive cases from the far larger population of negative cases.

This issue is particularly important because sanctions-related events are likely to constitute a relatively small proportion of the overall maritime population. In such settings, even a model with strong sensitivity can generate substantial numbers of false positives if its decision threshold is sufficiently broad. Conversely, a model may generate a relatively small number of alerts while missing a significant proportion of genuinely risky vessels. The appropriate assessment must therefore consider sensitivity or recall alongside specificity, precision, false-positive rates and the underlying prevalence of the event being predicted (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

A second limitation concerns the ground truth represented by sanctions designation. A government sanctions designation is not necessarily a direct measurement of maritime behavioural risk. Government investigations may incorporate intelligence from sources unavailable to commercial systems, including confidential or classified information, financial intelligence, human intelligence, law-enforcement investigations and other evidence. Consequently, a commercial AI system may identify behavioural characteristics that correlate with a later designation without independently reproducing the evidentiary basis on which the government reached its decision.

This distinction is particularly important because a subsequent sanctions designation should not automatically be treated as proof that the earlier AI signal correctly identified the underlying cause of the vessel's behaviour. The designation represents an external institutional decision made under a particular legal and evidential framework. Windward's earlier alert may have identified genuinely relevant behavioural indicators, but it may also have captured broader patterns associated with the vessel or its network that were not themselves sufficient to establish sanctionable conduct. The two forms of classification should therefore not be treated as interchangeable.

A third issue concerns selection effects and retrospective evaluation. If a model's performance is assessed by beginning with vessels that were ultimately sanctioned and then asking whether those vessels had previously been flagged, the evaluation starts from the positive cases. Such an approach can provide useful information about detection coverage, but it does not reproduce the conditions under which the model would have operated prospectively. A genuine prospective evaluation would need to consider the full population of vessels assessed at the relevant time, including those that were flagged but never sanctioned and those that were not flagged but were subsequently designated.

The appropriate research question is therefore broader than ‘What percentage of sanctioned vessels were previously flagged?’ A more informative evaluation would ask:

Of all vessels assessed by the system, what proportion of subsequently sanctioned vessels were identified in advance, what proportion of non-sanctioned vessels were also flagged, and how does this performance compare with an appropriate baseline?

Answering this question would permit assessment of the model's sensitivity, precision and false-positive rate, while also establishing whether its performance represents a meaningful improvement over simpler alternatives. An appropriate baseline might include conventional sanctions screening, rule-based risk indicators or another independently defined maritime-risk model, depending on the purpose of the evaluation.

The timing of the alert would also be relevant. A vessel flagged one day before designation and a vessel flagged twelve months before designation both satisfy the basic condition of being identified ‘before’ designation, but they provide different forms of predictive value. Earlier identification may provide substantially more operational opportunity for investigation or intervention. Consequently, a rigorous evaluation would ideally report not only whether a vessel was flagged before designation but also the distribution of lead times between the initial risk signal and the subsequent designation.

The definition of ‘flagged’ is similarly important. A broad risk signal generated at some point during a vessel's history is analytically different from a high-confidence alert that explicitly identified sanctions-related risk before designation. Without clarity concerning the threshold, persistence, type and interpretation of the alert, the 99% statistic remains difficult to translate into operational predictive performance.

These considerations do not invalidate Windward's claim. Rather, they place it within the appropriate evidential framework. The peer-reviewed literature supports the broader proposition that behavioural and multi-source maritime analytics can identify patterns associated with anomalous or suspicious activity (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Cai et al., 2026). Windward's reported result may therefore constitute evidence that its signals contain information correlated with subsequent sanctions designations. What it does not establish, without additional information, is the model's ability to distinguish future sanctions subjects from the wider population of vessels that will not subsequently be designated.

The most defensible interpretation is consequently that the 99% figure provides evidence of predictive association rather than, by itself, proof of predictive accuracy. To establish the stronger claim, Windward would need to disclose the population evaluated, the number and proportion of non-sanctioned vessels flagged, the definition of a qualifying alert, the distribution of lead times, the relevant performance metrics and the baseline against which its results are compared. Independent prospective or temporally separated validation would further strengthen the evidential basis.

The distinction is fundamental to evaluating AI-based sanctions intelligence. Identifying most vessels that are later sanctioned demonstrates that the system's signals contain potentially relevant information; demonstrating that those signals reliably discriminate future sanctions subjects from the much larger population of vessels that are never sanctioned would provide substantially stronger evidence of predictive capability. On the evidence publicly described, the former proposition is more supportable than the latter.

8. Claim Six: Explainability and Decision Support

Windward describes its maritime intelligence as both predictive and explainable, stating that users can trace alerts and recommendations back to underlying data, behavioural patterns and ownership information (Windward, 2026). Explainability is particularly important in maritime compliance because AI-generated risk assessments can influence commercially and operationally significant decisions, including whether to engage with a vessel, provide financial services, enter into a transaction or escalate a counterparty for further investigation. In such contexts, the ability to understand why a system generated an alert is not merely a usability feature; it is an important component of responsible human oversight.

However, a critical distinction must be made between explaining an AI output and establishing the substantive validity of that output. An explanation may identify the variables, observations or behavioural characteristics that contributed to a model's classification without demonstrating that those characteristics actually establish the underlying conduct being investigated. This distinction is particularly important for complex machine-learning systems, where a model may identify statistical relationships that are useful for classification without providing a causal explanation of the observed behaviour.

The distinction is especially relevant to maritime behavioural analytics. A vessel's unusual movement pattern, ownership structure, AIS anomaly or STS activity may provide legitimate grounds for further investigation, but none of these indicators necessarily establishes sanctions evasion, criminal conduct or intentional deception. As discussed above, anomalous behaviour can arise from multiple causes, including legitimate operational circumstances, technical limitations and environmental conditions (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). An explanation that accurately identifies the features associated with an alert does not therefore remove the need to consider alternative explanations.

This creates an important distinction between model transparency and evidential sufficiency. Transparency concerns whether users can understand the basis on which a system generated a particular output. Evidential sufficiency concerns whether the available evidence is adequate to support the substantive conclusion that a user might draw from that output. The two are related but not equivalent. A system can be highly transparent about the factors associated with a risk score while those factors remain insufficient to establish the underlying cause or legal significance of the observed behaviour.

The appropriate role of explainable AI in this context is therefore primarily one of decision support rather than autonomous adjudication. An AI system can identify, combine and prioritise potentially relevant signals, allowing investigators to focus attention on cases that warrant further examination. Human analysts can then assess the wider evidential context, consider alternative explanations, seek additional information and determine the appropriate response. This interpretation is consistent with the limitations identified in the maritime anomaly-detection literature, where anomalous observations are generally better understood as indicators for further investigation than as definitive classifications of misconduct (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

The distinction is also important when considering Windward's use of terms such as ‘predictive’ and ‘explainable intelligence’. Predictive performance requires evidence that the model can reliably distinguish relevant future outcomes from non-events, while explainability concerns the ability to understand the basis of individual model outputs. Neither property, independently, establishes that the model has correctly identified the cause of a particular maritime event. A system can therefore be predictive without being causally explanatory, or explainable without being substantively accurate.

For consequential maritime compliance decisions, the combination of predictive performance, transparent reasoning and human review is consequently more important than explainability in isolation. The practical objective should not be to replace human judgement with an apparently interpretable model, but to make the model's contribution sufficiently transparent that investigators can critically assess both the supporting evidence and its limitations.

9. Data Quality as a Fundamental Limitation

Windward describes its maritime intelligence as both predictive and explainable, stating that users can trace alerts and recommendations back to underlying data, behavioural patterns and ownership information (Windward, 2026). Explainability is particularly important in maritime compliance because AI-generated risk assessments can influence commercially and operationally significant decisions, including whether to engage with a vessel, provide financial services, enter into a transaction or escalate a counterparty for further investigation. In such contexts, the ability to understand why a system generated an alert is not merely a usability feature; it is an important component of responsible human oversight.

However, a critical distinction must be made between explaining an AI output and establishing the substantive validity of that output. An explanation may identify the variables, observations or behavioural characteristics that contributed to a model's classification without demonstrating that those characteristics actually establish the underlying conduct being investigated. This distinction is particularly important for complex machine-learning systems, where a model may identify statistical relationships that are useful for classification without providing a causal explanation of the observed behaviour.

The distinction is especially relevant to maritime behavioural analytics. A vessel's unusual movement pattern, ownership structure, AIS anomaly or STS activity may provide legitimate grounds for further investigation, but none of these indicators necessarily establishes sanctions evasion, criminal conduct or intentional deception. As discussed above, anomalous behaviour can arise from multiple causes, including legitimate operational circumstances, technical limitations and environmental conditions (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). An explanation that accurately identifies the features associated with an alert does not therefore remove the need to consider alternative explanations.

This creates an important distinction between model transparency and evidential sufficiency. Transparency concerns whether users can understand the basis on which a system generated a particular output. Evidential sufficiency concerns whether the available evidence is adequate to support the substantive conclusion that a user might draw from that output. The two are related but not equivalent. A system can be highly transparent about the factors associated with a risk score while those factors remain insufficient to establish the underlying cause or legal significance of the observed behaviour.

The appropriate role of explainable AI in this context is therefore primarily one of decision support rather than autonomous adjudication. An AI system can identify, combine and prioritise potentially relevant signals, allowing investigators to focus attention on cases that warrant further examination. Human analysts can then assess the wider evidential context, consider alternative explanations, seek additional information and determine the appropriate response. This interpretation is consistent with the limitations identified in the maritime anomaly-detection literature, where anomalous observations are generally better understood as indicators for further investigation than as definitive classifications of misconduct (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

The distinction is also important when considering Windward's use of terms such as ‘predictive’ and ‘explainable intelligence’. Predictive performance requires evidence that the model can reliably distinguish relevant future outcomes from non-events, while explainability concerns the ability to understand the basis of individual model outputs. Neither property, independently, establishes that the model has correctly identified the cause of a particular maritime event. A system can therefore be predictive without being causally explanatory, or explainable without being substantively accurate.

For consequential maritime compliance decisions, the combination of predictive performance, transparent reasoning and human review is consequently more important than explainability in isolation. The practical objective should not be to replace human judgement with an apparently interpretable model, but to make the model's contribution sufficiently transparent that investigators can critically assess both the supporting evidence and its limitations.

One of the strongest conclusions emerging from the academic literature is that the effectiveness of maritime AI is fundamentally constrained by the quality, completeness and representativeness of the data on which it operates. AIS data should not be treated as a neutral or complete observation of maritime reality. It can contain missing observations, transmission errors, inconsistencies, delays and potentially manipulated information (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). These limitations are not peripheral technical issues: they directly affect the reliability of any model trained on or dependent upon such data.

The systematic literature reviewed by Höhle et al. (2022) identifies inconsistent datasets, limited labelled anomalies and the absence of a universally accepted ground-truth dataset as important barriers to the development and evaluation of maritime anomaly-detection systems. Ribeiro, Paes and de Oliveira (2023) similarly identify large data volumes, dynamic and inconsistent observations, missing labels and difficulties in evaluating detection performance as persistent challenges. Together, these studies suggest that the principal constraint on maritime AI is not necessarily a lack of sophisticated algorithms, but the difficulty of obtaining sufficiently reliable and representative evidence against which those algorithms can learn and be evaluated.

This makes Windward's emphasis on multi-source data logically attractive. If AIS represents only one imperfect observation of maritime activity, combining it with SAR, radio-frequency information, satellite observations, ownership information and other sources can provide opportunities for cross-validation and contextual interpretation (Windward, 2026). Research combining AIS and SAR provides empirical support for this general proposition, demonstrating how complementary observations can assist in identifying abnormal STS activity and potential inconsistencies between reported and physically observed vessel behaviour (Cai et al., 2026).

Nevertheless, more data do not automatically produce better intelligence. Multi-source systems introduce additional methodological challenges, including differences in temporal resolution, spatial accuracy, data latency, coverage and reliability. Observations may also conflict with one another, requiring the system to determine which source should receive greater evidential weight. The existence of several signals pointing in different directions does not necessarily resolve uncertainty; it can instead create a more complex inference problem.

This is particularly important because the sources incorporated into a maritime intelligence platform may not be equally independent. Several apparently distinct observations may ultimately reflect related underlying information or common reporting infrastructure. Consequently, the number of data sources incorporated into a system should not be treated as a direct proxy for evidential strength. What matters is the quality, relevance, independence and reliability of the information contributed by each source.

Data fusion can therefore be understood as an opportunity for cross-validation rather than an automatic solution to data-quality problems. For example, disagreement between AIS and SAR observations may provide a valuable investigative signal, but the discrepancy could also result from differences in observation time, positional accuracy or vessel identification. Similarly, an AIS gap may be significant in one operational context but relatively benign in another. The analytical system must therefore account not only for the observations themselves but also for the uncertainty associated with those observations.

This suggests that data provenance and uncertainty estimation should be treated as core components of maritime AI governance rather than secondary technical considerations. Data provenance concerns the origin, transformation and reliability of the information entering the analytical system. Uncertainty estimation concerns the extent to which confidence should be placed in the resulting inference. Both are particularly important when AI outputs are used to support consequential compliance decisions.

The importance of provenance is also connected to explainability. An investigator cannot fully evaluate why an alert was generated merely by knowing which variables contributed to a model output. They also need to understand where those variables originated, how current the underlying information is, whether the source is independently corroborated and what limitations affect its reliability. In this sense, meaningful explainability should extend beyond the model itself to encompass the evidential chain from observation to inference.

The resulting analytical principle is therefore more demanding than simply asking whether a maritime AI system uses many data sources. The relevant question is whether the system can establish a sufficiently reliable chain between data provenance, observation, inference and decision. A sophisticated model operating on poorly characterised or systematically biased data may generate highly sophisticated outputs without producing correspondingly reliable intelligence.

Overall, the academic literature supports Windward's emphasis on multi-source intelligence and explainable analysis, but it also highlights the conditions under which those approaches can be reliable. Multi-source data can improve observability and provide opportunities for cross-validation, while explainability can enable users to scrutinise the basis of AI-generated alerts (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Cai et al., 2026). Neither, however, eliminates uncertainty arising from imperfect observations or establishes the truth of the model's interpretation. Trustworthy maritime AI therefore depends not simply on more data or more explainable models, but on the integrity of the evidential chain connecting data, inference and human decision-making.

10. From Detection to Risk: The Interpretation Problem

The most important conceptual issue in evaluating maritime AI is the distinction between anomaly detection, risk assessment and proof of wrongdoing. These represent different analytical tasks and require progressively stronger forms of evidence. Anomaly detection asks whether observed behaviour deviates from an established or expected pattern. Risk assessment goes further by considering whether that deviation is associated with an increased likelihood of an undesirable outcome. Determining wrongdoing, however, requires substantially more than statistical or behavioural deviation: it requires contextual evidence capable of supporting a substantive conclusion about what occurred, why it occurred and, where relevant, whether it was intentional or unlawful.

Peer-reviewed research provides relatively strong support for the first of these functions. Studies of maritime AIS data demonstrate that machine-learning and statistical techniques can identify unusual trajectories, invalid messages and behavioural deviations (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Campbell, J et al., 2022). Research also indicates that combining different observation modalities can strengthen the identification of suspected abnormal activity. For example, Cai et al. (2026) demonstrate how the combination of SAR and AIS can be used to identify suspected abnormal ship-to-ship transfer activity, illustrating the analytical value of integrating physical observations with vessel-tracking data. Similarly, research on suspicious fishing behaviour shows that anomalous tracking patterns can provide indicators for further investigation (Identification of suspicious behavior through anomalies in the tracking data of fishing vessels, 2024).

However, identifying an anomaly does not establish its cause. A deviation from expected behaviour may arise from legitimate operational circumstances, environmental conditions, technical problems, data-quality issues or deliberate manipulation. This creates an important inferential gap between recognising that behaviour is unusual and concluding that it represents sanctions evasion or other unlawful conduct. The literature on maritime anomaly detection therefore supports the use of AI to identify and prioritise potentially significant behaviour, but it does not establish that machine-learning models can independently determine criminal intent or legal responsibility (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

This distinction is particularly important when interpreting Windward's commercial language. Identifying apparent AIS manipulation can constitute an anomaly-detection output; identifying a combination of behavioural indicators associated with sanctions evasion represents a risk-assessment inference; but determining that a vessel is actually evading sanctions is a substantially stronger substantive conclusion requiring corroborating evidence and contextual assessment. Treating these three outputs as interchangeable risks overstating what the underlying analytical methods can establish.

The distinction also has direct implications for governance and responsible use. Maritime AI should therefore be understood primarily as an evidence-generating and prioritising technology, rather than an autonomous adjudication mechanism. Its appropriate function is to identify patterns, integrate heterogeneous evidence and direct human attention towards vessels or behaviours warranting further investigation. The subsequent determination of whether an anomaly reflects legitimate activity, deliberate deception, sanctions evasion or another form of wrongdoing remains dependent on contextual evidence and human assessment. In this respect, the strongest defensible interpretation of maritime AI is not that it replaces investigative judgement, but that it can make that judgement more targeted by identifying where further scrutiny is warranted.

11. Discussion

The examination of Windward’s claims produces a mixed but substantively important conclusion. The company’s underlying technological proposition is broadly consistent with the peer-reviewed literature. AIS-based anomaly detection is an established area of maritime research; machine-learning techniques can identify unusual vessel trajectories and behavioural patterns; AIS data are subject to missing observations, inconsistencies and potential manipulation; and the integration of AIS with other observation modalities, including SAR, can provide complementary information that is unavailable from either source independently (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Cai et al., 2026; Campbell, J et al., 2022). The academic literature therefore provides a credible technical basis for the general architecture described by Windward.

The evidential position becomes considerably weaker, however, when moving from technological plausibility to specific commercial performance claims. The publicly available evidence reviewed in this paper does not provide sufficient methodological detail to independently reproduce or validate Windward’s claims of 94% accuracy in illicit STS detection, approximately 75% fewer false positives in sanctions-related AIS spoofing detection, or 99% pre-designation identification of vessels subsequently sanctioned. These figures may represent genuine performance within Windward’s internal evaluation framework, but their academic verifiability is constrained by the absence of publicly reported information on the relevant datasets, ground-truth construction, benchmark conditions, class distributions, independent test populations and complete performance metrics. This limitation is particularly significant given that maritime anomaly-detection research itself identifies limited labelled data, inconsistent datasets and the absence of universally accepted ground-truth benchmarks as persistent methodological challenges (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023).

This does not, in itself, invalidate Windward’s product or demonstrate that its commercial claims are incorrect. Rather, it illustrates a broader methodological challenge in the evaluation of proprietary AI systems. Commercial systems may incorporate datasets, models and validation procedures that cannot be fully disclosed because of intellectual-property or commercial-confidentiality constraints. Consequently, external researchers and customers may be unable to determine precisely how models were trained, which observations constituted ground truth, how cases were sampled, or whether reported performance remains stable across different geographical regions, vessel populations and operational conditions. The distinction between not independently verifiable and demonstrably false is therefore important: the former describes the evidential limitation identified in this analysis without implying the latter.

The issue becomes particularly consequential in high-stakes maritime compliance environments, where the costs of different classification errors are asymmetric. A reduction in false positives may generate substantial operational value by reducing unnecessary investigations and allowing compliance resources to be concentrated on higher-risk cases. However, such a reduction is meaningful only if it is achieved without a disproportionate increase in false negatives. Similarly, identifying vessels before they are subsequently sanctioned may provide valuable early-warning intelligence, but a high proportion of eventually sanctioned vessels being identified does not by itself establish strong predictive discrimination. Evaluation must also establish how frequently comparable warnings are generated for vessels that are never sanctioned and how far in advance meaningful alerts are produced. These considerations reinforce the distinction between detecting relevant signals and demonstrating reliable predictive performance.

A mature evaluation framework should therefore assess maritime AI using a portfolio of complementary performance measures rather than a single headline accuracy statistic. Precision and recall are important for understanding the balance between correctly identifying relevant cases and generating unnecessary alerts, while F1 score can provide a combined measure where appropriate. False-negative rates are particularly important in compliance applications because missed high-risk cases may carry greater consequences than additional investigative workload. Calibration is relevant where outputs are interpreted as risk probabilities, while temporal stability and geographic generalisation are necessary to determine whether performance persists as maritime conditions and operating environments change. Independent test datasets are especially important because evaluation on data closely related to model development can produce an overly optimistic estimate of real-world performance.

The need for such validation is reinforced by the characteristics of maritime activity itself. Vessel behaviour varies according to geography, vessel type, trade route, port environment and operational circumstances, while the availability and quality of AIS and other observational data can also vary across contexts (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). A model that performs strongly within one population or evaluation environment cannot automatically be assumed to generalise to another. Consequently, the strongest evidence for commercial maritime AI would not be a single high accuracy figure, but transparent, independently reproducible evidence demonstrating robust performance across relevant populations and conditions.

Overall, the academic literature supports the technical plausibility of Windward’s approach more strongly than it supports the independent verification of its headline performance claims. The distinction is fundamental. Windward’s use of multi-source data, anomaly detection and behavioural analytics is compatible with established research directions, but the magnitude and reliability of its claimed operational benefits remain difficult to assess without greater methodological transparency and external validation. For high-stakes maritime intelligence, therefore, the appropriate standard is not simply whether an AI system can identify unusual behaviour, but whether its outputs remain accurate, calibrated, generalisable and evidentially interpretable when applied to real-world decisions.

12. Conclusion

This paper has critically examined the technological proposition and evidential basis of Windward’s Maritime AI™ claims against the findings of peer-reviewed research on maritime anomaly detection, data fusion, AIS and GNSS manipulation, and behavioural analytics. The analysis produces a nuanced conclusion. The technology underlying Windward’s approach is broadly supported by the academic literature, but the evidence available in the public domain is insufficient to independently establish the magnitude of its specific commercial performance claims.

The peer-reviewed literature provides a credible foundation for many of the technologies incorporated into Windward’s platform. Research demonstrates that AIS data can be used to identify anomalous maritime behaviour, while machine-learning techniques can detect unusual trajectories, invalid messages and other deviations from expected vessel activity (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023; Campbell, J et al., 2022). Research also demonstrates the value of combining complementary sources. Cai et al. (2026), for example, shows how SAR and AIS can be integrated to identify suspected abnormal ship-to-ship transfer activity. These findings support the general proposition that multi-source maritime intelligence can improve situational awareness and provide information that may not be available from a single data source.

However, technical plausibility should not be confused with demonstrated commercial performance. The central limitation identified in this paper concerns the methodological transparency surrounding Windward’s headline claims. The publicly available evidence reviewed does not provide sufficient information to independently reproduce or validate the reported 94% STS detection accuracy, approximately 75% reduction in false positives or 99% pre-designation identification of vessels subsequently sanctioned. The absence of detailed information concerning datasets, sampling procedures, ground-truth construction, class distributions, benchmark definitions, independent test populations and complete performance measures makes it difficult to determine how these figures should be interpreted or how well they would generalise beyond Windward’s internal evaluation environment.

This limitation is particularly significant because the academic literature identifies methodological challenges that directly affect the evaluation of maritime AI. Maritime datasets can contain missing observations, inconsistent measurements and limited labelled examples, while there is no universally accepted ground-truth dataset for maritime anomaly detection (Höhle et al., 2022; Ribeiro, Paes and de Oliveira, 2023). Maritime behaviour also varies across vessel types, geographical regions, trade routes and operational environments. Consequently, performance established within one dataset or population cannot automatically be assumed to represent performance across the wider maritime domain.

The analysis also identifies an important conceptual boundary between anomaly detection, risk assessment and proof of wrongdoing. An algorithm may identify behaviour that differs from an expected pattern and may reasonably associate that behaviour with elevated risk. Neither output, however, establishes why the behaviour occurred or whether it constitutes unlawful conduct. AIS manipulation, unusual STS activity or anomalous vessel movements may warrant investigation, but they do not independently establish sanctions evasion or criminal intent. This distinction is particularly important where AI-generated intelligence influences decisions concerning financial transactions, vessel engagement, compliance escalation or investigative prioritisation. The appropriate role of such systems is therefore to generate and prioritise evidence for human assessment rather than to function as autonomous adjudication mechanisms.

The implications extend beyond Windward to the broader development and procurement of commercial maritime AI. Headline performance statistics provide limited information when considered without the methodological conditions under which they were generated. A more rigorous evaluation should consider precision, recall, F1 score, false-negative rates, calibration and temporal stability, alongside performance on independent datasets and across different geographical and vessel populations. Particular attention should be given to the trade-off between false positives and false negatives: reducing unnecessary alerts can create operational value, but only if this reduction does not result in a corresponding deterioration in the detection of genuinely significant risks. Similarly, identifying vessels before subsequent sanctions designation may indicate useful predictive signals, but meaningful evaluation requires consideration of the wider population of vessels that are flagged but never sanctioned.

The overall finding is therefore neither that commercial maritime AI claims should be accepted uncritically nor that the underlying technology lacks validity. Rather, the academic evidence supports the technological foundations of Windward’s approach more strongly than it supports the independent verification of its claimed operational performance. This distinction provides a more appropriate basis for evaluating AI in maritime compliance: the relevant question is not simply whether a system can detect unusual behaviour, but whether its outputs are demonstrably accurate, generalisable, calibrated and sufficiently transparent to support consequential decisions.

Ultimately, trustworthy maritime AI depends on the integrity of the evidential chain connecting data, algorithmic inference and human decision-making. Multi-source data and sophisticated machine-learning models can increase the ability to observe and prioritise potentially significant maritime activity, but they cannot eliminate uncertainty or substitute for contextual evidence. Greater transparency in evaluation methodologies, independent validation and clear differentiation between risk signals and determinations of wrongdoing are therefore essential if commercial maritime AI is to move from technologically plausible innovation towards reliably validated intelligence for high-stakes compliance environments.

References

Appel, M., Iliopoulos, A., Fohlmeister, F., Pérez Marcos, E., Cuntz, M., Konovaltsev, A., Antreich, F. and Meurer, M. (2019) ‘Experimental validation of GNSS repeater detection based on antenna arrays for maritime applications’, CEAS Space Journal, 11, pp. 7–19. doi: 10.1007/s12567-018-0232-6.

Cai, P., Liu, B., Li, X., Li, X., Wang, S., Liu, P., Chen, P. and Li, Y. (2026) ‘Detecting ship-to-ship transfer by MOSA: Multi-source Observation framework with SAR and AIS’, Remote Sensing, 18(3), 473. doi: 10.3390/rs18030473.

Campbell, J et al. (2022) ‘Detection of invalid AIS messages using machine learning techniques’, Procedia Computer Science, 205, pp. 229–238. doi: 10.1016/j.procs.2022.09.024.

Gattis, B., Cydejko, J. and Akos, D. (2026) ‘Baltic sea GNSS jamming and spoofing emitter detection and localization in real-time using a time difference of arrival (TDOA) system’, GPS Solutions, 30, 96. doi: 10.1007/s10291-026-02061-5.

Höhle, M., et al. (2022) ‘Anomaly Detection in Maritime AIS Tracks: A Review of Recent Approaches’, Journal of Marine Science and Engineering, 10(1), 112. doi: 10.3390/jmse10010112.

Rodríguez, J et al. (2024) ‘Identification of suspicious behavior through anomalies in the tracking data of fishing vessels’, EPJ Data Science, 13, 23.

Ribeiro, C.V., Paes, A. and de Oliveira, D. (2023) ‘AIS-based maritime anomaly traffic detection: A review’, Expert Systems with Applications, 231, 120561. doi: 10.1016/j.eswa.2023.120561.

Windward (2026) Windward Maritime AI™

Contact

Reach out via email for inquiries.

Email

Subscribe to newsletter

info@grcadvisory.ch

© 2025. All rights reserved.