Next Article in Journal
Artificial Intelligence Adoption and the Need for Artificial Intelligence Literacy in Veterinary Education: A Cross-Sectional Survey
Previous Article in Journal
Context-Aware Path Planning: A Unified Approach for Navigation in Heterogeneous Environments
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Large Language Models in the Analysis of Radar Plotting Images in Accordance with COLREGs

Faculty of Navigation, Maritime University of Szczecin, Wały Chrobrego 1-2, 72-500 Szczecin, Poland
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(18), 9294; https://doi.org/10.3390/app16189294 (registering DOI)
Submission received: 1 July 2026 / Revised: 15 September 2026 / Accepted: 16 September 2026 / Published: 19 September 2026
(This article belongs to the Section Marine Science and Engineering)

Abstract

The development of artificial intelligence (AI) opens new possibilities for the automatic interpretation of radar images in maritime navigation. Traditionally, the interpretation of radar indications requires the experience of the officer of the watch, who assesses the risk of collision and makes navigational decisions based on echo position, motion vectors, course, speed, Closest Point of Approach (CPA) and Time to Closest Point of Approach (TCPA). Systems using artificial intelligence can support this process through automatic object detection, tracking of their movement, classification of encounter situations, and indication of possible maneuvers in accordance with the Convention on the International Regulations for Preventing Collisions at Sea, 1972 (COLREGs). Maritime safety depends largely on the correct interpretation and application of the relevant COLREGs. The aim of this article is to present the capabilities, limitations and significance of generally available artificial intelligence tools like Large Language Models (LLMs) in the field of human–machine interaction. Attention is paid to the application of LLMs in radar plotting image analysis, the interpretation of navigational situations, and decision-making support from the perspective of marine traffic safety.

1. Introduction

The fundamental and wide-ranging responsibility of a ship’s crew is to ensure the safe passage of the vessel from the point of departure to the point of destination. Proper navigation of a vessel is a task that depends primarily on the qualifications and competence of the officer of the watch. The usage of equipment available on the navigation bridge requires knowledge of both its capabilities and its limitations.
Marine radar is one of the basic devices supporting the officer’s decision-making process. It facilitates observation of the vessel’s surroundings regardless of prevailing weather conditions, such as sea state or restricted visibility caused by various types of precipitation. By using the radar image and Automatic Radar Plotting Aids (ARPAs) [1], the officer of the watch can determine the position of other vessels, their course, speed, and the predicted CPA and TCPA in relation to the position of the own ship.
The importance of the human factor and communication in collision situations has also been emphasized by Misztal and Hatlas-Sowinska [2]. The authors indicate that communication problems between navigators may significantly affect maritime safety and propose an automatic communication model based on natural language processing and ontology-based methods. This perspective is relevant to the present study because LLM-based radar plotting analysis should be considered not only as an image interpretation task but also as a human–machine communication process in which technical navigational data are translated into clear and understandable information for the navigator.
In recent years, systems using artificial intelligence have gained increasing importance. Recent review studies indicate that LLMs may contribute to maritime safety by improving communication, supporting decision-making, facilitating compliance, enabling interactive training, supporting automated reporting, and assisting in real-time risk assessment. Miller et al. [3] emphasize that these capabilities may be particularly useful in complex maritime operations, although issues such as data privacy, system integration and ethical considerations must be addressed before operational implementation.
LLMs are increasingly investigated in other expert and safety-critical domains, including medicine, aviation, robotics, construction safety and disaster management [4,5,6,7,8,9,10]. These studies show that LLMs may support clinical decision-making, air traffic safety analysis, robotic task planning, safety-related assessments and crisis information processing. However, their practical use requires validation, human supervision, constrained outputs and clearly defined responsibility, especially when AI-generated outputs may influence safety-related decisions. This broader context is relevant to maritime navigation, where LLM-generated radar plotting interpretations and maneuvering suggestions should be verified against navigational data, COLREGs-based rules, collision-risk assessment and expert supervision.
In the transport sector, the use of LLMs is also not limited to maritime navigation. Xu et al. [11] proposed DriveGPT4, an interpretable autonomous driving system based on a large language model, capable of processing visual data, answering user questions, justifying the vehicle’s actions, and predicting control signals. This points to a growing interest in the use of multimodal models as tools that integrate perception, situation interpretation, and action planning. Research on Maritime Autonomous Surface Ships (MASS) indicates that AI can support route planning, collision risk assessment, and decision-making in accordance with COLREGs [12]. Literature reviews concerning autonomous vessels emphasize that one of the key challenges is not only hazard detection but also the correct interpretation of navigational situations and the compliance of decisions with the rules of the road at sea. A recent example of this research direction is Navigation-GPT, proposed by Ma et al. [13]. The authors present a dual-core LLM agent designed for intelligent marine navigation, combining high-level reasoning with navigation-related recommendations. The framework is intended to improve adaptability in unknown or non-predefined navigation scenarios and to support the generation of navigation hints consistent with COLREGs and other maritime rules.
A radar image shows the position of objects located within radar range in relation to the own vessel. The most important information that can be read from radar includes the bearing and distance from the own ship to the echo. A marine radar system equipped with ARPA facilitates the analysis of the navigational situation by enabling rapid determination of the object’s course and speed, CPA and TCPA.
Technological progress aims to use AI systems by introducing them into fully automated and unmanned vessels. At present, such systems are not used in merchant shipping. Easy access to advanced AI-based LLMs, such as ChatGPT-4o or Microsoft 365 Copilot, raises questions about their capabilities and limitations in supporting the decisions of navigators. Trained on data of unprecedented scale, large language models (LLMs), such as ChatGPT and GPT-4o, demonstrate the emergence of significant reasoning abilities resulting from model scaling [14]. In 2024, research from the Swedish COLREG3 project conducted at the Maritime Competence Centre was published, showing the weak spatial awareness of LLMs, which significantly limits their use for maritime traffic analysis and often leads to misinterpretations [15]. LLMs such as ChatGPT or Copilot operate based on advanced AI language models. Based on neural networks, they use deep learning to generate text. GPT not only supports automation but also opens the door to new possibilities in the field of human–machine interaction. Since LLMs have no consciousness and no access to current information other than that used during their training process, they may provide incorrect or imprecise answers. An important point is that COLREGs are written regulations and are, to some extent, ambiguous from the perspective of algorithms [16]. Kristić and Žuškin [17] discussed the ambiguity by analyzing linguistic variables used in COLREGs and quantified expert knowledge concerning the term “Very Large Ship” from Rule 7. Their study shows that some COLREGs expressions are fuzzy by nature and can be interpreted differently unless they are formally described, for example, by fuzzy sets. This problem underscores the essence of the research presented in this article, namely the correct interpretation of radar images, the analysis of encounters between vessels, and the proper implementation of COLREGs, which may contain linguistic and interpretive uncertainties.
Despite the growing number of studies on LLMs in maritime navigation and autonomous vessel decision-making, there remains a research gap concerning the ability of generally available LLMs to interpret radar plotting images and translate them into COLREGs-compliant maneuvering recommendations without additional navigational data input. Previous studies mainly focused on theoretical knowledge of ship handling, autonomous navigation frameworks, or LLM-based decision-making architectures supported by structured data. This indicates the gap of practical visual-spatial interpretation of radar plotting images by widely accessible tools such as ChatGPT-4o and Copilot. The contribution of this study is therefore threefold: first, it evaluates the ability of selected commercially available LLMs to analyze radar plotting images and to adapt their responses to previous conversational context within a single interaction; second, it identifies typical errors related to spatial interpretation, vector reading, distance estimation and COLREGs application; and third, it discusses the role of LLMs as explanatory and training-support tools rather than independent navigational decision-making systems.

2. Materials and Methods

Advanced multimodal LLMs can process image input and may therefore be used to analyze radar plotting images. In such an analysis, the model attempts to recognize the position of the own ship, echoes of other vessels, motion vectors, and radar range scale. It must also determine the position of the target vessel, the distance from the center of the screen, the bearing, the direction and length of the motion vector, the approximate speed, and the predicted CPA and TCPA.
The research demonstrated a pattern of radar plotting image analysis performed by LLMs. The models carry out the analysis independently and in stages:
  • Reading the Settings
At the initial stage, a general analysis of the image is performed; for example:
observation range, e.g., 6 Nm,
range rings scale,
position of the own ship.
Without the aforementioned data, the analysis may be incorrect. For example, an incorrectly read radar scale leads to an erroneous determination of the target’s distance and speed.
2.
Target Detection
Next, AI identifies the echo of the target vessel. Depending on the manufacturer, the target on the radar image may appear as a point, a spot or a rectangle.
3.
Determining the Target’s Position
After detecting the target, the LLMs can determine:
on which side of the own ship the target vessel is located,
what the bearing is,
what the distance is,
whether the target is ahead, astern, to port or to starboard.
4.
Reading Course and Speed
If a motion vector is visible on the radar, AI can estimate the course and speed of the target vessel. The direction of the vector indicates the direction of movement, while its length may correspond to speed.
5.
Calculating CPA and TCPA and Suggesting a Maneuver
The most important element of the analysis is the assessment of whether there is a risk of collision. The following parameters are used for this purpose:
CPA—Closest Point of Approach,
TCPA—Time to Closest Point of Approach.
Assuming the user’s minimum safe setting, for example, a CPA of 1 Nm, the user expects guidance on selecting an appropriate maneuver, i.e., one compliant with COLREGs. For this purpose, after reading the position and motion of the target vessel, the system should classify the encounter and determine the type of navigational situation.
The analysis was carried out for 24 different simulations of situations presented in the form of radar image plotting, performed on the POLARIS [18] multifunctional navigation bridge simulator manufactured by KONGSBERG.
The 24 radar plotting scenarios were selected as a representative set of simulator-based navigational training exercises. The selection was made to include different types of ship-to-ship encounters and different levels of interpretative difficulty, including crossing, overtaking, nearly head-on and no-action-required situations. The scenario set included 16 crossing situations, 3 head-on, or nearly head-on situations, and 3 overtaking situations. Each scenario was assessed in terms of risk of collision. If the verified CPA exceeded the predefined minimum safe passing distance, the situation was classified as a no-action-required encounter, because the target vessel was expected to pass the own ship at a safe distance. Four scenarios were classified as no-action-required cases on this basis. The scenarios also differed in the target vessel’s position relative to the own ship, the distance from the own ship, the direction of the motion vector and the expected COLREGs-based assessment. For each scenario, the ground truth was established before the LLM analysis. The correct encounter classification, applicable COLREGs rule or rules, give-way and stand-on roles, and expected own-ship maneuver were determined by an expert navigator on the basis of the radar plotting image, simulator scenario assumptions, relative position and motion of the vessels, the assumed minimum CPA criterion of 1 Nm and the relevant COLREGs rules. This expert-defined ground truth was then used as the reference for evaluating the correctness of the LLM-generated responses.
The purpose of Table 1 is to clarify the scope and diversity of the analyzed scenarios before presenting the model outputs and accuracy results.
The initial assumption was to examine the recognition capabilities and effectiveness of three selected commercially available LLMs in maritime accident analysis.
DeepSeek-R1 was initially considered as an additional LLM-based tool because of its growing availability. However, under the applied research conditions, it did not provide sufficient radar image recognition capability; therefore, it was excluded from the final analysis.
The study was designed as a preliminary exploratory assessment rather than a comprehensive benchmark of all available LLMs. Therefore, the analysis was limited to two generally available multimodal tools: ChatGPT-4o and Microsoft 365 Copilot. These tools were selected because they were accessible to users and allowed image-based interaction with radar plotting screenshots.
Since the ChatGPT-4o multimodal model is capable of processing text and image inputs [19], it was selected to analyze images from navigation radars. By incorporating the Copilot tool, which is integrated with the Microsoft 365 suite, both solutions were used to conduct preliminary research on the recognition process and the proper application of COLREGs.
The experiment was conducted in November 2025. The radar plotting images were exported from the POLARIS navigation bridge simulator and used as image inputs for the tested LLM-based tools. The images were provided to the models in JPG format, with a resolution of approximately 1101 × 1055 px (78.9 KB). No additional structured navigational data, such as numerical CPA/TCPA tables, AIS data or ARPA target data, were provided to the models apart from the information visible in the radar plotting image and the short text prompt.
The tested tools were OpenAI’s GPT-4o multimodal model accessed through the ChatGPT Business environment with image input enabled, and Microsoft 365 Copilot, version 19.2608.54041.0, accessed through the Microsoft 365 environment. The same radar plotting images and the same prompting procedure were used for both tools. For the first analyzed image, the following prompt was used: “Situation 01, radar range: 6 nautical miles. Minimum CPA: 1 nautical mile. According to COLREGs, what should I do as a vessel?” For the subsequent images, the prompts contained only the sequence number of the situation and a request for recommendation, for example: “Situation 3, what do you suggest?” or simply “Sit 10”. All analyses for each model were conducted within a single conversation history in order to assess whether the model could maintain and use contextual information during the interaction.
In this study, a correct maneuvering suggestion was defined as a response in which the model correctly interpreted the encounter situation sufficiently to select the relevant COLREGs rule or rules, correctly assigned the give-way and stand-on roles where applicable, and proposed an own-ship maneuver consistent with COLREGs and the assumed minimum CPA criterion of 1 Nm. Minor inaccuracies in the estimation of descriptive parameters, such as distance, were accepted if they did not affect the final rule application or maneuvering recommendation. For example, a response with an approximate distance error was still classified as correct if the model correctly identified the crossing situation, assigned the own ship as the give-way vessel, and recommended an appropriate maneuver. Conversely, a response was classified as incorrect when an error in spatial interpretation, vector reading, distance assessment, COLREGs application, or maneuvering recommendation led to an incorrect or unsafe suggestion.
To ensure consistency in the evaluation, the following scoring rubric was applied (see Table 2). The assessment focused primarily on the correctness of the final maneuvering suggestion; however, the reasoning process was also considered when an error in reasoning affected the encounter classification, COLREGs application, assignment of vessel obligations or safety of the recommended action.
If the final maneuvering recommendation was safe but the reasoning contained a substantive error that affected the legal or navigational justification, such as incorrect encounter classification or wrong assignment of give-way and stand-on roles, the response was classified as incorrect. This approach was adopted because, in navigational decision-making, a safe-looking maneuver based on incorrect reasoning may lead to unsafe decisions in similar or slightly modified situations.
Incorrect answers were then corrected by an expert. Expert validation was performed by a senior officer of the watch with 15 years of navigational experience on international voyages on various types of merchant vessels. The expert is also a lecturer at the Faculty of Navigation at the Maritime University of Szczecin, teaching, among other subjects, maritime law. However, expert correction did not always lead the model to verify and apply the COLREGs rules correctly.
All analyses for each model were conducted within one conversation history. This procedure was intentionally used to examine whether the model could maintain and use previous conversational context, including earlier prompts, previous radar plotting analyses and expert corrections. The observed effect should not be interpreted as self-learning in the training sense, because the model parameters were not updated during the interaction. Instead, it reflects context-dependent response adaptation, or in-context adaptation, in which later responses may be influenced by information provided earlier in the same conversation. This design was therefore used to assess the practical behavior of generally available chatbot tools during a continuous analytical session. At the same time, it is acknowledged as a methodological limitation because previous conversational context may have influenced later responses and reduced the independence of individual scenario assessments.
Two selected situations are presented below as illustrative examples. Figure 1 shows a case in which ChatGPT-4o provided a correct analysis, whereas Figure 2 shows a case in which the initial analysis was incorrect.
The COLREGs cover various signals, such as sound signals, day shapes, and navigation lights, which play a key role in preventing collisions and maneuvering vessels. Without an operator involved in the decision-making process to take into account additional information from visual and auditory signals, important information may be overlooked [20]. The analyses performed by LLMs were based solely on the radar plotting image, i.e., information about the relative positions of the vessels; without light/daytime signals and/or sound signals, the correct application of COLREGs is practically impossible. For research purposes of this article, it was assumed that vessels maneuver without any weather restrictions, restrictions from the nature of their work, and are power-driven vessels underway.
The two examples presented above illustrate two typical outcomes of the analysis: a correct interpretation of a crossing situation and an incorrect interpretation caused by a wrong assessment of the target vessel’s side. Full chatbot responses are provided in the Supplementary Material. Table 3 presents only the elements relevant to the evaluation: situation interpretation, suggested COLREGs rule, maneuvering recommendation, and expert assessment.

3. Results

The analysis of Situation 1, presented in Figure 1, showed that ChatGPT-4o correctly interpreted the general encounter situation and selected the appropriate COLREGs rule. The model identified the target vessel as approaching from the starboard side and classified the situation as a crossing encounter. Consequently, the own vessel was correctly indicated as the give-way vessel under Rule 15 of COLREGs.
However, a minor inaccuracy was observed in the estimation of distance. ChatGPT-4o assessed the distance to the target vessel as approximately 5.0–5.5 Nm, whereas the actual distance was approximately 5.8 Nm. In this case, the error did not affect the classification of the encounter or the recommended maneuver. Nevertheless, similar inaccuracies in distance estimation were observed in other analyzed cases and may influence the assessment of collision risk, especially when CPA and TCPA values are close to the assumed safety limits.
The analysis of the situation presented in Figure 2 revealed several irregularities in the response generated by ChatGPT-4o. The model incorrectly interpreted the relative position of the target vessel and initially identified the vessel as being located on the port side of its own ship. This error affected the identification of the give-way and stand-on vessels and consequently led to an incorrect maneuvering recommendation. Although the model referred to the correct COLREGs rule for a crossing situation, the rule was applied to an incorrectly interpreted spatial configuration.
After an expert correction indicating that the target vessel was located on the starboard side of its own ship, ChatGPT-4o revised its response and generated a more appropriate recommendation in accordance with COLREGs. This suggests that expert input may improve the quality of the model’s response. However, it also indicates that the model may not reliably detect its own errors in the interpretation of relative vessel positions.
Across the 24 scenarios, the most frequent causes of incorrect responses were related to spatial interpretation and radar reading rather than to the inability to mention COLREGs rules. In several cases, the models suggested a plausible rule but applied it to an incorrectly interpreted situation. This was especially evident when the model misidentified the side of the target vessel, misread the motion vector, or incorrectly assessed whether its own ship should act as the give-way or stand-on vessel.
Based on the analysis of all 24 radar plotting situations (see Table 4, Table 5 and Table 6), the following results were obtained:
  • ChatGPT-4o provided correct maneuvering suggestions in 7 out of 24 situations analyzed.
  • Copilot provided correct maneuvering suggestions in 11 out of 24 analyzed situations.
  • In the repeated analysis performed with ChatGPT-4o, correct suggestions were obtained in 4 out of 24 situations.
The accuracy of the models was calculated as the ratio of the number of correct maneuvering suggestions to the total number of analyzed scenarios:
Accuracy (%) = (Number of correct maneuvering suggestions/Total number of analyzed scenarios) × 100.
The corresponding accuracy values were:
  • 29.2% for ChatGPT-4o in the first analysis,
  • 45.8% for Copilot,
  • 16.7% for the repeated ChatGPT-4o analysis.
A comparative summary of LLM performance in radar plotting analysis is presented in Table 7.
Due to the limited number of analyzed scenarios, the reported accuracy values should be interpreted as descriptive results obtained under the applied experimental conditions. The study was exploratory in nature and was not designed to provide a statistical comparison of model performance. Therefore, the observed difference between ChatGPT-4o and Microsoft 365 Copilot should not be interpreted as evidence of a statistically significant or generally meaningful performance advantage of one model over the other. A larger dataset and multiple independent repetitions for each model would be required to quantify uncertainty and compare model performance statistically.
The repeated ChatGPT-4o analysis was included as an exploratory check of response stability rather than as a full repeatability study. In this repeated run, correct maneuvering suggestions were obtained in 4 out of 24 situations, compared with 7 out of 24 in the first ChatGPT-4o analysis. This indicates variability between the two analyses of the same scenario set. However, because only one repeated run was performed, the results should not be interpreted as a complete statistical assessment of repeatability or reproducibility. Instead, they should be treated as a preliminary indication that LLM-generated radar plotting assessments may vary between repeated interactions and that more systematic repetition is required in future studies.
The most frequently observed errors in the responses generated by the analyzed LLMs included:
  • incorrect reading of the own ship’s course;
  • incorrect reading of the target vessel’s course;
  • incorrect interpretation of the target vessel’s position relative to the own ship;
  • difficulty in determining whether the target vessel was located on the port or starboard side;
  • inaccuracies in estimating the distance from the own ship;
  • incorrect selection or application of the relevant COLREGs rule;
  • difficulty in applying expert corrections consistently.
To identify the source of incorrect responses more precisely, the observed errors were additionally grouped into visual-spatial interpretation errors, navigational reasoning errors and mixed errors. Visual-spatial interpretation errors included incorrect identification of the target vessel’s side, incorrect reading of the own ship or target vessel’s course, incorrect interpretation of relative geometry, and distance or CPA-related interpretation errors. Navigational reasoning errors included incorrect COLREGs rule application, wrong assignment of give-way and stand-on roles, or incorrect maneuver recommendations despite sufficient interpretation of the radar situation. Mixed errors were defined as cases in which an initial visual-spatial misinterpretation led directly to incorrect COLREGs application or maneuver recommendation. The classification criteria are presented in Table 8, while the quantitative summary of these errors is presented in Table 9.
Overall, the obtained results indicate that the tested LLMs were able to generate structured and convincing descriptions of radar plotting situations. However, their correctness was limited, particularly in tasks requiring visual–spatial reasoning, the interpretation of motion vectors, and the application of COLREGs to specific encounter situations.

4. Discussion

The results obtained in this study show that generally available LLMs may support the interpretation of radar plotting images, but their reliability remains limited, especially in tasks requiring visual-spatial reasoning and variable radar interpretation. This conclusion is directly supported by the quantitative results: ChatGPT-4o provided correct maneuvering suggestions in 7 out of 24 situations (29.2%), Copilot in 11 out of 24 situations (45.8%), and the repeated ChatGPT-4o analysis produced correct suggestions in only 4 out of 24 situations (16.7%).
These values indicate limited accuracy under the applied experimental conditions. However, because the study included only 24 scenarios, the differences between ChatGPT-4o and Microsoft 365 Copilot should be interpreted descriptively and with caution. The results do not provide sufficient statistical evidence to claim a meaningful performance advantage of one model over the other.
The main sources of error were not only incorrect application of COLREGs but also incorrect interpretation of radar-derived spatial information. The most frequent errors included incorrect reading of the own ship’s course, incorrect interpretation of the target vessel’s course, wrong assessment of whether the target vessel was located on the port or starboard side, inaccurate distance estimation, and incorrect application of the relevant COLREGs rule. These findings are consistent with the observation by Xie et al. [21] that LLMs may fail in tasks requiring numerical, physical, and spatial reasoning. They are also consistent with the conclusions of Silwal and Dubey [22], who reported limited LLM performance in COLREGs-compliant autonomous maritime navigation, and with the Swedish COLREG3 project [15], which indicated that poor spatial reasoning significantly limits the use of LLMs in maritime traffic analysis.
The comparison with previous studies should also be interpreted in the context of methodological differences. Pei et al. [23] reported a higher accuracy for GPT-4o in ship-handling-related tasks; however, their study focused mainly on ship-handling knowledge and skills, whereas the present study required direct visual interpretation of radar plotting images, including relative position, motion vectors, CPA-related assessment and COLREGs-based maneuvering suggestions. Similarly, LLM-based maritime decision-making frameworks, such as the system proposed by Agyei, Sarhadi and Naeem [16], and the CORALL concept [24], use structured collision-risk indicators, COLREGs guidance or additional decision-support modules. In contrast, the present study evaluated generally available commercial LLMs used as multimodal chatbots, without deterministic verification modules or structured navigational input. Therefore, the lower accuracy obtained in this study reflects the specific difficulty of direct radar image interpretation rather than only the general ability of LLMs to recall or discuss COLREGs.
A key problem observed during the analyses was the incorrect interpretation of spatial relationships between vessels. In several cases, the model incorrectly determined whether the target vessel was located on the port or starboard side of the own ship. This error is critical because the relative position of the target is one of the main elements used to classify an encounter situation under COLREGs. In a crossing situation, for example, the vessel that has the other vessel on her starboard side is generally required to keep out of the way. Therefore, an incorrect reading of the side on which the target is located may directly lead to the wrong identification of the give-way and stand-on vessels.
From the communication perspective, the role of LLMs is particularly sensitive. Misztal and Hatlas-Sowinska [2] show that communication during collision situations is strongly affected by the human factor and that unclear or incorrectly interpreted information may reduce the effectiveness of collision avoidance. In the present study, the analyzed LLMs generated fluent and persuasive explanations of radar plotting situations; however, some of these explanations were based on incorrect spatial interpretation. This means that chatbot-generated communication may support the navigator, but it may also increase risk if the underlying analysis is wrong.
Another important limitation was the incorrect reading of motion vectors. In radar plotting analysis, the vector of the target vessel is essential for determining its course, speed, and future movement relative to the own ship. The study showed that LLMs may confuse the direction of movement, misread the course of the target vessel, or incorrectly interpret the own ship’s course. Such errors have a direct impact on the assessment of CPA and TCPA and may result in an inappropriate recommendation of action required to navigate safely. In practice, even a small error in vector interpretation may change the classification of the encounter from crossing to overtaking or from crossing to a nearly head-on situation.
The results also indicate that the models may have difficulties with the correct estimation of distance from the radar display. Errors in reading the radar range scale or interpolating the distance rings may lead to incorrect assumptions about the proximity of the target vessel. In collision avoidance, distance estimation is closely related to the time available for action and to the assessment of whether a risk of collision exists. If the distance is overestimated, the model may underestimate the urgency of the situation. If it is underestimated, the model may recommend unnecessary or excessive maneuvers.
The repeated analysis performed with ChatGPT-4o should be interpreted with caution. It was conducted as an exploratory check of response stability, not as a complete repeatability or reproducibility assessment. Since only one repeated analysis was performed, the study cannot provide statistically robust conclusions regarding consistency. Nevertheless, the difference between the first ChatGPT-4o analysis and the repeated analysis suggests that model responses may vary when the same or comparable radar plotting scenarios are analyzed again. Future research should include multiple independent repetitions for each model, preferably in separate conversation histories and under controlled prompting conditions, in order to assess repeatability and reproducibility more rigorously.
The study also showed that expert correction may improve the quality of the model’s response. In Situation 2 (see Figure 2), after the expert indicated that the target vessel was located on the starboard side of the own ship, the model was able to revise the COLREGs assessment and provide a more appropriate recommendation. This demonstrates that LLMs may be useful in an interactive human–machine process, where the navigator verifies the model’s interpretation and provides corrections when necessary. However, it also confirms that the model itself may not reliably detect its own error. The need for expert intervention limits the possibility of using such tools as autonomous decision-making systems, but it supports their potential role as auxiliary training or explanatory tools.
The persuasive language generated by LLMs represents both an advantage and a risk. On the one hand, the ability to explain a radar situation in simple language is valuable, especially in training, simulation exercises, and decision-support environments. The model can structure the analysis, recall the relevant COLREGs rule, and present the reasoning in an accessible form. On the other hand, an incorrect answer may be presented with the same confidence and clarity as a correct one. This may be dangerous for inexperienced users who may not be able to identify the underlying error in image interpretation or rule application.
It should also be emphasized that the analyses in this study were based only on radar plotting images. In real navigation, the officer of the watch uses many additional sources of information, including visual observation, AIS, sound signals, navigation lights, day shapes, VHF communication, environmental conditions, traffic density, maneuvering characteristics of the vessel, and the ordinary practice of seamen. COLREGs are not applied solely on the basis of a static radar image. Therefore, even a correct interpretation of the radar plotting picture may not be sufficient to determine the safest maneuver in an operational situation.
Consistent with the COLREG3 project [15], this confirms that the main weakness of LLMs in this context is not only knowledge of COLREGs but also the ability to correctly connect visual information with navigational reasoning. The models often know the content of the rules and can quote or paraphrase them correctly, but they may apply them to a wrongly interpreted situation. The problem is therefore not only legal or procedural knowledge but also perception, geometry, and spatial reasoning.
Future research should include a larger number of radar situations, different radar ranges, different vector lengths, more than one target vessel, and varying levels of traffic complexity. It would also be useful to compare model performance under different prompting strategies. Future research should also include a broader comparison of multimodal LLMs, including additional commercially available and open-source models, with more independent repetitions for each model in order to assess not only accuracy but also repeatability and model-specific differences in radar plotting interpretation. The present study used relatively simple prompts in order to observe the models’ spontaneous analytical capabilities. However, more structured prompts containing information about radar range, own ship course, vector time, and minimum acceptable CPA may improve the quality of the responses. Such research would help determine whether LLM performance depends mainly on visual interpretation limitations or on insufficient task formulation.
In summary, the discussion of the results shows the dual nature of LLMs in maritime navigation. They have clear educational and explanatory potential, especially in reading radar plotting data into understandable language and recalling the structure of COLREGs reasoning. At the same time, their errors in spatial interpretation, vector reading, distance estimation, and rule application prevent their use as independent decision-making systems. The results of this study support the conclusion that LLMs may become valuable elements of future human–machine interaction on the bridge, but only as carefully validated decision-support tools operating under expert supervision.

5. Conclusions

The results of this study indicate that, under the tested experimental conditions, generally available LLMs such as ChatGPT-4o and Microsoft 365 Copilot showed limited reliability in radar plotting image analysis and COLREGs-based maneuvering suggestions. Although the models were able to generate structured and convincing explanations, their accuracy was relatively low, and their responses included fundamental errors in visual-spatial interpretation, motion vector reading, distance assessment and COLREGs application. Therefore, the tested systems cannot currently be considered reliable tools for independent collision-risk assessment or operational navigational decision-making.
Their potential use should be limited to supervised training, explanation, post-encounter review and human–machine communication. In such contexts, LLMs may help structure information, explain possible encounter situations and recall relevant COLREGs considerations, but all outputs must be verified by a qualified officer. The final navigational assessment and maneuvering decision must remain under the responsibility of the officer of the watch and/or the master.
The study also indicates a need for further development of hybrid systems. A more reliable solution may involve combining LLMs with deterministic navigational algorithms, ARPA data, sensor fusion, and rule-based COLREGs modules. Research on COLREGs-compliant deep reinforcement learning for multi-vessel collision avoidance also supports the need for hybrid approaches. Xie et al. [25] proposed a DRL-based decision-making method for autonomous surface vessels that addresses multi-vessel collision avoidance under COLREGs constraints. In this context, LLMs should not replace validated collision-avoidance algorithms, but may complement them by explaining verified outputs, supporting human–machine interaction and presenting the navigational reasoning in natural language. A similar direction is consistent with the CORALL concept [24], where the LLM is guided by COLREGs and risk-awareness mechanisms rather than being used as an unrestricted conversational model. Such an approach may reduce the risk of incorrect interpretation by separating the explanatory and high-level decision-support role of the LLM from validated risk assessment and maneuver execution modules. In such a system, the LLM would not be responsible for independently reading the radar image or calculating CPA and TCPA. Instead, it could explain verified data generated by certified or validated systems. This would allow the model’s strength—natural language explanation and human–machine interaction—to be used while reducing the risk resulting from incorrect visual interpretation.
The potential value of LLMs lies mainly in their ability to convert verified technical radar data into a comprehensible explanatory message. In a supervised context, such systems may help describe:
where the target vessel appears to be located;
how it appears to be moving;
whether a potential risk of collision may require further assessment;
which COLREGs rule may be relevant;
what maneuvering options may require consideration by the navigator.
The analyses of the situations presented are delivered in a highly convincing manner; however, they are unfortunately often incorrect. Similar conclusions were also drawn in the Swedish COLREG3 project, which emphasized that “poor spatial reasoning completely hinders their use for the analysis of marine traffic situations” and that careful validation and verification frameworks will be needed to operationalize LLMs as decision-support systems for maritime navigation [15].
Preliminary analysis concerning maneuvering suggestions provided by LLMs may differ significantly from the actions taken by an operator. For example, according to COLREGs, the interpretation of a situation involving a stand-on vessel is that it should maintain its course and speed. However, depending on the circumstances, the stand-on vessel may decide to take very early action to avoid a close-quarters situation by carrying out a maneuver [26].
LLMs may support situational awareness and more structured navigational reasoning only in supervised training or advisory contexts. It is therefore useful to distinguish their potential explanatory advantages from their operational limitations. It is worth mentioning not only the disadvantages but also the advantages of using LLMs. Possible advantages of LLMs in supervised training or advisory contexts include:
explaining radar plotting information in simple language;
recalling relevant COLREGs rules;
supporting the interpretation of encounter situations;
helping the user consider selected aspects of the situation;
providing an additional explanatory layer for human–machine interaction.
These potential advantages correspond with the review by Miller et al. [3], who indicate that LLMs may enhance maritime safety through improved communication, training, decision support, compliance support and real-time risk assessment. In the context of radar plotting analysis, these functions may support the navigator’s situational awareness, provided that the model’s output is verified by a qualified officer.
Tested LLMs are not sufficiently advanced models to make navigational decisions independently. The results of this study confirm the general thesis presented by Guo et al. [27] that the evaluation of LLMs should encompass not only their task-specific capabilities but also aspects of safety, reliability, and responsible use. In the context of maritime navigation, this means that every maneuvering suggestion generated by a chatbot must be verified by a qualified expert. Such tools may support human interpretation, but they do not replace qualified officers of the watch or certified navigational systems. The preliminary research conducted in this study showed several limitations that require continuous expert supervision.
The most important limitations include:
possible incorrect reading of the image,
incorrect interpretation of vectors,
difficulties in assessing the intentions of the other vessel,
the necessity of applying good seamanship,
legal responsibility remaining with the human operator.
Considering these limitations, the safest near-term application of LLMs in radar plotting analysis may be found in simulator training, post-encounter review, and advisory bridge displays rather than in direct maneuver control. In such applications, the LLM may explain the encounter situation, indicate missing information, summarize the relevant COLREGs context, and support the navigator’s situational awareness, while maneuvering authority remains with the officer of the watch and/or the master. Therefore, responses provided by algorithms used in LLMs should be treated as decision-making support, not as an automatic maneuvering order. Future use of LLMs in autonomous or semi-autonomous navigation systems will require integration with validated navigational systems, reliable sensor data, collision-risk assessment modules and expert supervision.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/app16189294/s1.

Author Contributions

Conceptualization, K.D.; methodology, K.D. and W.J.; validation, K.D., P.Z. and D.J.; formal analysis, K.D.; investigation, K.D.; resources, K.D.; data curation, K.D.; writing—original draft preparation, K.D.; writing—review and editing, K.D., P.Z. and D.J.; visualization, K.D.; supervision, K.D. and P.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors on request.

Acknowledgments

The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
AISAutomatic Identification System
ARPAAutomatic Radar Plotting Aid
COLREGsConvention on the International Regulations for Preventing Collisions at Sea, 1972.
CPAClosest Point of Approach
LLMLarge Language Model
MASSMaritime Autonomous Surface Ships
NmNautical Mile
TCPATime to Closest Point of Approach
VHFVery High Frequency

References

  1. International Maritime Organization. Resolution, A.823(19): Performance Standards for Automatic Radar Plotting Aids (ARPAs); IMO: London, UK, 1995. [Google Scholar]
  2. Misztal, L.; Hatlas-Sowinska, P. The Impact of the Human Factor on Communication During a Collision Situation in Maritime Navigation. Appl. Sci. 2025, 15, 2797. [Google Scholar] [CrossRef] [Scilit]
  3. Miller, T.; Durlik, I.; Kostecka, E.; Łobodzińska, A.; Łazuga, K.; Kozlovska, P. Leveraging Large Language Models for Enhancing Safety in Maritime Operations. Appl. Sci. 2025, 15, 1666. [Google Scholar] [CrossRef] [Scilit]
  4. Omar, M.; Nadkarni, G.N.; Klang, E.; Glicksberg, B.S. Large Language Models in Medicine: A Review of Current Clinical Trials across Healthcare Applications. PLoS Digit. Health 2024, 3, e0000662. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Vrdoljak, J.; Boban, Z.; Vilović, M.; Kumrić, M.; Božić, J. A Review of Large Language Models in Medical Education, Clinical Decision Support, and Healthcare Administration. Healthcare 2025, 13, 603. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Deng, Q.; Zhang, M.; Yang, Y.; Gao, Z. SCOPE: A Lightweight-Training LLM Framework for Air Traffic Control Readback Monitoring. arXiv 2026, arXiv:2605.29543. [Google Scholar] [CrossRef] [Scilit]
  7. Darrell, T.; Ghazanfari, M.; Kam, J.; Bayen, A.; Tabrizian, A.; Wei, P. Towards Automated Air Traffic Safety Assessment Around Non-Towered Airports Using Large Language Models. arXiv 2026, arXiv:2605.12332. [Google Scholar] [CrossRef] [Scilit]
  8. Ahn, M.; Brohan, A.; Brown, N.; Chebotar, Y.; Cortes, O.; David, B.; Finn, C.; Fu, C.; Gopalakrishnan, K.; Hausman, K.; et al. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. arXiv 2022, arXiv:2204.01691v2. [Google Scholar] [CrossRef] [Scilit]
  9. Sammour, F.; Xu, J.; Wang, X.; Hu, M.; Zhang, Z. Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering. arXiv 2024, arXiv:2411.08320. [Google Scholar] [CrossRef] [Scilit]
  10. Lei, Z.; Dong, Y.; Li, W.; Ding, R.; Wang, Q.; Li, J. Harnessing Large Language Models for Disaster Management: A Survey. arXiv 2025, arXiv:2501.06932. [Google Scholar] [CrossRef] [Scilit]
  11. Xu, Z.; Zhang, Y.; Xie, E.; Zhao, Z.; Guo, Y.; Wong, K.-Y.K.; Li, Z.; Zhao, H. DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model. arXiv 2023, arXiv:2310.01412. [Google Scholar] [CrossRef] [Scilit]
  12. International Maritime Organization. Convention on the International Regulations for Preventing Collisions at Sea, 1972 (COLREGs); as amended; IMO: London, UK, 1972. [Google Scholar]
  13. Ma, F.; Wang, X.-M.; Chen, C.; Xu, X.-B.; Yan, X.-P. Navigation-GPT: A Robust and Adaptive Framework Utilizing Large Language Models for Navigation Applications. IET Intell. Transp. Syst. 2026, 20, e70233. [Google Scholar] [CrossRef] [Scilit]
  14. Zhou, G.; Hong, Y.; Wu, Q. Explicit Reasoning in Vision-and-Language Navigation with Large Language Models. Proc. AAAI Conf. Artif. Intell. 2024, 38, 7641–7649. [Google Scholar] [CrossRef] [Scilit]
  15. Sanchez-Heres, L.; Weber, R.; Ahlgren, F.; Olsson, F.; Lundstrom, O. COLREG3—Exploring the Potential of Large Language Models in Marine Navigation Systems; Lighthouse—Swedish Marine Competence Center: Gothenburg, Sweden, 2024; Available online: https://lighthouse.nu/sv/publikationer/lighthouse-rapporter/colreg3 (accessed on 10 January 2025).
  16. Agyei, K.; Sarhadi, P.; Naeem, W. Large Language Model-based Decision-making for COLREGs and the Control of Autonomous Surface Vehicles. arXiv 2024, arXiv:2411.16587. [Google Scholar] [CrossRef] [Scilit]
  17. Kristić, M.; Žuškin, S. Quantification of Expert Knowledge in Describing COLREGs Linguistic Variables. J. Mar. Sci. Eng. 2024, 12, 849. [Google Scholar] [CrossRef] [Scilit]
  18. SO-0612-Z; Polaris Technical Manual Section 5a Instructors Manual, 1997–2016. Polaris: Medina, MN, USA, 2016.
  19. OpenAI. GPT-4o System Card; OpenAI: San Francisco, CA, USA, 2024; Available online: https://openai.com/index/gpt-4o-system-card/ (accessed on 1 May 2025).
  20. Weber, R.; Sanchez-Heres, L.; Sjoblom, T. COLREG2–Potential Consequences of Varying Algorithms in Traffic Situations; Lighthouse—Swedish Marine Competence Center: Gothenburg, Sweden, 2023; Available online: https://lighthouse.nu/en/publications/lighthouse-reports/colreg-2-potential-consequences-of-varying-algorithms-in-traffic-situations (accessed on 10 January 2025).
  21. Xie, Y.; Yu, C.; Zhu, T.; Bai, J.; Gong, Z.; Soh, H. Translating Natural Language to Planning Goals with Large-Language Models. arXiv 2023, arXiv:2302.05128. [Google Scholar] [CrossRef] [Scilit]
  22. Silwal, A.; Dubey, R. Are LLMs Ready for COLREGs? A Study in Autonomous Maritime Vessel Navigation. In Proceedings of the OCEANS 2025—Great Lakes; IEEE: New York, NY, USA, 2025. [Google Scholar] [CrossRef] [Scilit]
  23. Pei, D.; He, J.; Liu, K.; Chen, M.; Zhang, S. Application of Large Language Models and Assessment of Their Ship-Handling Theory Knowledge and Skills for Connected Maritime Autonomous Surface Ships. Mathematics 2024, 12, 2381. [Google Scholar] [CrossRef] [Scilit]
  24. Agyei, K.; Sarhadi, P.; Naeem, W. CORALL: A COLREGs—Guided Risk-Aware LLM for Decision-Making in Maritime Autonomous Surface Ships. IEEE J. Ocean. Eng. 2026, in press. [Google Scholar] [CrossRef] [Scilit]
  25. Xie, W.; Gang, L.; Zhang, M.; Liu, T.; Lan, Z. Optimizing Multi-Vessel Collision Avoidance Decision Making for Autonomous Surface Vessels: A COLREGs—Compliant Deep Reinforcement Learning Approach. J. Mar. Sci. Eng. 2024, 12, 372. [Google Scholar] [CrossRef] [Scilit]
  26. Weber, R.; Aylward, K.; MacKinnon, S.; Lundh, M. Operationalizing COLREGs in SMART Ship Navigation; Lighthouse—Swedish Marine Competence Center: Gothenburg, Sweden, 2022; Available online: https://lighthouse.nu/en/publications/lighthouse-reports/operationalizing-colregs-in-smart-ship-navigation (accessed on 2 February 2025).
  27. Guo, Z.; Jin, R.; Liu, C.; Huang, Y.; Shi, D.; Supryadi; Yu, L.; Liu, Y.; Li, J.; Xiong, B.; et al. Evaluating Large Language Models: A Comprehensive Survey. arXiv 2023, arXiv:2310.19736. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Radar plotting image with the prompt (correct analysis). Source: own study.
Figure 1. Radar plotting image with the prompt (correct analysis). Source: own study.
Applsci 16 09294 g001
Figure 2. Radar plotting image with the prompt (incorrect analysis). Source: own study.
Figure 2. Radar plotting image with the prompt (incorrect analysis). Source: own study.
Applsci 16 09294 g002
Table 1. General characteristics of the 24 radar plotting scenarios used in the study.
Table 1. General characteristics of the 24 radar plotting scenarios used in the study.
NoSituation IDEncounter TypeExpected COLREGs Rule(s) for Own ShipTarget Position Relative to Own Ship/Distance [Nm]
1. 01Crossing 15, 16Starboard side/5.8
2. 02Crossing 15,16Starboard side/5.5
3.03Head-on 14Reciprocal/5.7
4.04Head-on 14Starboard side + Nearly reciprocal courses/5.8
5.05Head-on 14Port side + Nearly reciprocal courses/5.8
6.06Crossing 15, 16Starboard side/5.9
7.07Crossing,
stand-on
15, 17Port side/5.8
8.08Overtaking,
stand-on
13, 17Starboard side, astern/1.8
9.09Overtaking,
stand-on
13, 17Port side, astern/1.8
10.10Overtaking,
stand-on
13,17Starboard side, astern/1.5
11.21Crossing 15, 16Starboard side/5.0
12.22N/ANo action requiredStarboard side/4.7
13.23Crossing15, 17Port side on a beam/2.6
14.24Crossing 15, 17Port side/3.3
15.25Crossing 15, 17Port side/5.2
16.26Crossing 15, 17Port side/5.8
17.27Crossing 15, 17Port side/5.9
18.30Crossing15, 16Starboard side/2.9
19.31Crossing15, 16Starboard side/2.8
20.32Crossing15, 16Starboard side/3.3
21.33CrossingNo action requiredStarboard side/3.3
22.34Crossing15, 17Port side/3.5
23.35N/ANo action requiredPort side/3.3
24.36CrossingNo action requiredPort side/3.9
Source: own study. N/A-not applicable.
Table 2. Scoring rubric used to classify LLM-generated maneuvering suggestions.
Table 2. Scoring rubric used to classify LLM-generated maneuvering suggestions.
ClassificationCriteriaExample
CorrectThe model correctly identifies the encounter situation, selects the applicable COLREGs rule or rules, correctly assigns give-way and stand-on roles where applicable, and recommends a maneuver consistent with COLREGs and the assumed minimum CPA criterion of 1 Nm. Minor descriptive inaccuracies are accepted if they do not affect the final rule application or maneuvering recommendation.Slightly inaccurate distance estimation, but correct crossing classification, correct give-way role and safe maneuver to starboard.
IncorrectThe model misinterprets the radar situation in a way that affects encounter classification, COLREGs rule application, assignment of give-way/stand-on roles, or maneuvering recommendations.Target vessel incorrectly identified on the port side instead of the starboard side, leading to the wrong stand-on/give-way assessment.
Incorrect despite safe final maneuverThe final maneuver may appear safe or conservative, but the response is classified as incorrect if it is based on wrong encounter classification, wrong COLREGs rule, wrong assignment of obligations, or incorrect explanation of why the maneuver should be taken.The model recommends reducing speed, which may be safe, but it incorrectly classifies an overtaking situation as crossing and assigns the wrong COLREGs obligations.
Correct despite minor parameter errorThe response is classified as correct if a minor error in a descriptive parameter, such as approximate distance, does not affect the encounter classification, applicable rule, give-way/stand-on roles, or recommended maneuver.The model estimates the distance as 5.0–5.5 Nm instead of approximately 5.8 Nm but correctly identifies Rule 15 and recommends that the own ship give way.
Incorrect after expert correctionThe model remains incorrect if, after expert correction, it still fails to apply the correct COLREGs rule, assign the correct vessel obligations, or recommend an appropriate own-ship maneuver.The expert corrects the target’s side, but the model still recommends maintaining course and speed when the own ship should give way.
Source: own study.
Table 3. Summary of selected ChatGPT-4o responses used as illustrative examples.
Table 3. Summary of selected ChatGPT-4o responses used as illustrative examples.
ExampleModelSituation InterpretationSuggested RuleSuggested
Maneuver
Expert
Assessment
Figure 1, Situation 01ChatGPT-4oCrossing situation; target on starboard sideRule 15Alter course to starboard/reduce speedCorrect
Figure 2,
Situation 20
ChatGPT-4oTarget incorrectly interpreted as being on port sideRule 15Own ship should maintain course and speedIncorrect
Figure 2, Situation 20, after expert correctionChatGPT-4oTarget vessel correctly reconsidered as being on the starboard sideRule 15Own ship should give way, preferably by altering course to starboardCorrect after expert correction
Source: own study.
Table 4. Analysis of radar plotting image generated by ChatGPT-4o.
Table 4. Analysis of radar plotting image generated by ChatGPT-4o.
Analysis of Radar Plotting Image Performed by ChatGPT-4o
Situation NoCorrectness of Situation InterpretationSuggested COLREGs RuleCorrectness of Rule IndicationCorrectness of Maneuver SuggestionExpert Comment AreaCorrectness after Expert Correction
101YRule 15 YY
202N Rule 13 NNSideY, Rule 16 + 8 (missed rule 15)
303N Rule 8 + 15NYRule 14Y, Rule 14 + 8
404NRule 15NYRule 14N, Rule 15
505N Rule 15 + 8NYRule 14N, Rule 17
606YRule 15 + 8YY
707NRule 15 + 17YYCourse of
another vessel
Y, Rule 15 + 17
808NRule 15 + 17NYCourse of
another vessel
Y, Rule 13
909YRule 13YY
1010NRule 13YNSideY, Rule 13 + 17b
1121NRule 15YNSideY, Rule 15
1222NRule 15NNCPAY, No rule
1323NNo Rule NNCPAY, Rule 15 + 17
1424YRule 15 + 17YY
1525YRule 15 + 17YY
1626YRule 15 + 17YY
1727YRule 15 + 17YY
1830YRule 15YNCorrection of the maneuverY, Rule 15
1931NRule 15YY
2032NRule 15 + 8YY
2133NR 15NNCPAN, Rule 15
2234N Rule 15YNSideY, Rule 15
2335NRule 15NNSide + CPAY
No Rule + Rule 7, 8
2436NRule 15NNSideY, No Rule
Source: own study. Y—correct answer; N—incorrect answer; Rule—COLREGs rule; CPA—Closest Point of Approach; Side—incorrect interpretation of the position of the other vessel relative to the side of our vessel. Marked in RED color—situations deemed incorrect.
Table 5. Analysis of radar plotting generated by Copilot.
Table 5. Analysis of radar plotting generated by Copilot.
Analysis of Radar Plotting Image Performed by Copilot
NoSituation NoCorrectness of Situation InterpretationSuggested COLREGs RuleCorrectness of Rule IndicationCorrectness of
Maneuver Suggestion
Expert Comment AreaCorrectness After Expert Correction
101YRule 15 YNCorrection of the maneuverN
202YRule 15 + 13 YNCorrection of the maneuverY, Rule 15 + 8
303YRule 14YY
404YRule 15 + 8NYRule 14N Rule 15
505YRule 15 + 8NYRule 14N Rule 17
606YRule 15 + 8YY
707NRule 15 + 17YNSideY, Rule 15 + 17
808N Rule 13YNSideY, Rule 13 + 17
909YRule 13 + 17YY
1010YRule 15NNAnother vessel is overtakingY, Rule 13 + 17b
1121NRule 15 + 17YNSideY, Rule 15 + 8
1222NRule 15NNCPAY, Rule 17
1323NRule 15YNCourse of own vesselY, Rule 15 + 17
1424YRule 15 + 17YY
1525YRule 15 + 17YY
1626YRule 15 + 17YY
1727YRule 15 + 17YY
1830YRule 15 + 16 + 17YY
1931YRule 15 + 16 + 8YY
2032YRule 15 + 16 + 8YY
2133N Rule 15NNCPA + Course of another vesselN, Rule 15
2234Y Rule 8YY
2335NRule 15NNCPA + SideY, No Rule + Rule 7, 8
2436NRule 15NNCPA + SideY, Rule 8
Source: own study. Y—correct answer; N—incorrect answer; Rule—COLREGs rule; CPA—Closest Point of Approach; Side—incorrect interpretation of the position of the other vessel relative to the side of our vessel. Marked in RED color—situations deemed incorrect.
Table 6. Repeated analysis of radar plotting generated by ChatGPT-4o.
Table 6. Repeated analysis of radar plotting generated by ChatGPT-4o.
Analysis of Radar Plotting Image Performed by ChatGPT-4o—Repeated Analysis
NoSituation NoCorrectness of Situation InterpretationSuggested COLREGs RuleCorrectness of Rule IndicationCorrectness of Maneuver SuggestionExpert Comment AreaCorrectness after Expert Correction
101’NRule 15 YNCourse of own vesselY, Rule 15
202’NRule 15 YYCourse of another vesselY, Rule 15 + 16
303’YRule 14YY
404’NRule 15NY Repeated twice Rule 14N, Rule 15 > Y, Rule 14
505’N Rule 15 + 8NN1. Side
2. nearly reciprocal
1.N Rule 15
2. Y Rule 14
606’YRule 15 + 8YNCorrection of the maneuverY, Rule 15
707’NRule 14NNCourse of another vesselY, Rule 15 + 17
808’N Rule 15 + 17NNCourse of another vesselY, Rule 13
909’YRule 13YY
1010’NRule 15NN1. speed of another vessel
2. Side
1.N, Rule 13 + 17 (side)
2. N, Rule 15
1121’NRule 15YNSideY, Rule 15
1222’NRule 15 + 8NNCPAY, Rule 7
1323’YRule 15 + 17YYYY, Rule 15
1424’NRule 13NYSpeed of both vesselsY, Rule 15 + 17
1525’NRule 15 + 17YY
1626’YRule 15 + 17YY
1727’NRule 14NNCourse of own vessel + SideY, Rule 15 + 17
1830’N Rule 15YNCourse of own vessel + SideY, Rule 15
1931’NRule 15YY
2032’N Rule 13NNCourse of another vesselY, Rule 15 + 8
2133’NRule 15 + 8NNCourse of another vessel + CPAN, Rule 15
2234N Rule 15YNSide + position of own vessel Y, Rule 15
2335’NRule 15NNSide Y, Rule 15 + 17
2436’NRule 15NNSideY, Rule 15
Source: own study. Y—correct answer; N—incorrect answer; Rule—COLREGs rule; CPA—Closest Point of Approach; Side—incorrect interpretation of the position of the other vessel relative to the side of our vessel. Marked in RED color—situations deemed incorrect.
Table 7. Comparative summary of LLM performance in radar plotting analysis.
Table 7. Comparative summary of LLM performance in radar plotting analysis.
ModelCorrect
Suggestions
Total
Situations
Accuracy [%]Main Error Types
ChatGPT-4o
(first analysis)
72429.2Spatial interpretation, port/starboard side, vector reading, COLREGs application
Copilot112445.8Maneuver suggestion, encounter classification, CPA/side interpretation
ChatGPT-4o
(repeated analysis)
42416.7 Repeatability, own/target course interpretation, side identification, COLREGs application
Source: own study.
Table 8. Classification of error types in LLM-generated radar plotting analyses.
Table 8. Classification of error types in LLM-generated radar plotting analyses.
Error GroupError
Category
Classification Criterion
Visual-Spatial Interpretation ErrorPort/starboard
identification
error
The model incorrectly identified whether the target vessel was located on the port or starboard side of the own ship
Own-ship or
target-vessel course
error
The model incorrectly interpreted the course or direction of movement of the own ship or the target vessel from the radar plotting image.
Relative
Geometry error
The model incorrectly interpreted the relative position, sector, bearing, or movement geometry of the vessels.
Distance/CPA
Interpretation error
The model incorrectly estimated distance or interpreted CPA-related risk in a way that affected the assessment.
Navigational Reasoning ErrorCOLREGs rule
application error
The model interpreted the visual situation sufficiently correctly but selected or applied an incorrect COLREGs rule.
Give-way/stand onThe model incorrectly assigned the obligations of the own ship or the target vessel despite sufficient interpretation of the encounter.
Maneuver
recommendation error
The model identified the situation or rule sufficiently correctly but proposed an incorrect, unclear, or unsafe maneuver.
Mixed
Error
Visual-spatial
error leading to
navigational error
The model first misinterpreted the radar geometry, and this error then caused incorrect COLREGs application or maneuver recommendation.
Source: own study.
Table 9. Quantitative summary of visual–spatial and navigational reasoning errors.
Table 9. Quantitative summary of visual–spatial and navigational reasoning errors.
Model/AnalysisVisual-Spatial
Interpretation Errors
Navigational Reasoning
Errors
Mixed
Errors
Correct Final
Maneuvering Suggestions
ChatGPT-4o (first analysis)161137/24
Microsoft 365 Copilot76311/24
ChatGPT-4o (repeated analysis)154124/24
Note: One response could contain more than one error type; therefore, the number of error categories may exceed the number of incorrect responses
Source: own study.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Drwięga, K.; Zalewski, P.; Juszkiewicz, W.; Jarząbek, D. Large Language Models in the Analysis of Radar Plotting Images in Accordance with COLREGs. Appl. Sci. 2026, 16, 9294. https://doi.org/10.3390/app16189294

AMA Style

Drwięga K, Zalewski P, Juszkiewicz W, Jarząbek D. Large Language Models in the Analysis of Radar Plotting Images in Accordance with COLREGs. Applied Sciences. 2026; 16(18):9294. https://doi.org/10.3390/app16189294

Chicago/Turabian Style

Drwięga, Kinga, Paweł Zalewski, Wiesław Juszkiewicz, and Dorota Jarząbek. 2026. "Large Language Models in the Analysis of Radar Plotting Images in Accordance with COLREGs" Applied Sciences 16, no. 18: 9294. https://doi.org/10.3390/app16189294

APA Style

Drwięga, K., Zalewski, P., Juszkiewicz, W., & Jarząbek, D. (2026). Large Language Models in the Analysis of Radar Plotting Images in Accordance with COLREGs. Applied Sciences, 16(18), 9294. https://doi.org/10.3390/app16189294

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop