1. Introduction
The fundamental and wide-ranging responsibility of a ship’s crew is to ensure the safe passage of the vessel from the point of departure to the point of destination. Proper navigation of a vessel is a task that depends primarily on the qualifications and competence of the officer of the watch. The usage of equipment available on the navigation bridge requires knowledge of both its capabilities and its limitations.
Marine radar is one of the basic devices supporting the officer’s decision-making process. It facilitates observation of the vessel’s surroundings regardless of prevailing weather conditions, such as sea state or restricted visibility caused by various types of precipitation. By using the radar image and Automatic Radar Plotting Aids (ARPAs) [
1], the officer of the watch can determine the position of other vessels, their course, speed, and the predicted CPA and TCPA in relation to the position of the own ship.
The importance of the human factor and communication in collision situations has also been emphasized by Misztal and Hatlas-Sowinska [
2]. The authors indicate that communication problems between navigators may significantly affect maritime safety and propose an automatic communication model based on natural language processing and ontology-based methods. This perspective is relevant to the present study because LLM-based radar plotting analysis should be considered not only as an image interpretation task but also as a human–machine communication process in which technical navigational data are translated into clear and understandable information for the navigator.
In recent years, systems using artificial intelligence have gained increasing importance. Recent review studies indicate that LLMs may contribute to maritime safety by improving communication, supporting decision-making, facilitating compliance, enabling interactive training, supporting automated reporting, and assisting in real-time risk assessment. Miller et al. [
3] emphasize that these capabilities may be particularly useful in complex maritime operations, although issues such as data privacy, system integration and ethical considerations must be addressed before operational implementation.
LLMs are increasingly investigated in other expert and safety-critical domains, including medicine, aviation, robotics, construction safety and disaster management [
4,
5,
6,
7,
8,
9,
10]. These studies show that LLMs may support clinical decision-making, air traffic safety analysis, robotic task planning, safety-related assessments and crisis information processing. However, their practical use requires validation, human supervision, constrained outputs and clearly defined responsibility, especially when AI-generated outputs may influence safety-related decisions. This broader context is relevant to maritime navigation, where LLM-generated radar plotting interpretations and maneuvering suggestions should be verified against navigational data, COLREGs-based rules, collision-risk assessment and expert supervision.
In the transport sector, the use of LLMs is also not limited to maritime navigation. Xu et al. [
11] proposed DriveGPT4, an interpretable autonomous driving system based on a large language model, capable of processing visual data, answering user questions, justifying the vehicle’s actions, and predicting control signals. This points to a growing interest in the use of multimodal models as tools that integrate perception, situation interpretation, and action planning. Research on Maritime Autonomous Surface Ships (MASS) indicates that AI can support route planning, collision risk assessment, and decision-making in accordance with COLREGs [
12]. Literature reviews concerning autonomous vessels emphasize that one of the key challenges is not only hazard detection but also the correct interpretation of navigational situations and the compliance of decisions with the rules of the road at sea. A recent example of this research direction is Navigation-GPT, proposed by Ma et al. [
13]. The authors present a dual-core LLM agent designed for intelligent marine navigation, combining high-level reasoning with navigation-related recommendations. The framework is intended to improve adaptability in unknown or non-predefined navigation scenarios and to support the generation of navigation hints consistent with COLREGs and other maritime rules.
A radar image shows the position of objects located within radar range in relation to the own vessel. The most important information that can be read from radar includes the bearing and distance from the own ship to the echo. A marine radar system equipped with ARPA facilitates the analysis of the navigational situation by enabling rapid determination of the object’s course and speed, CPA and TCPA.
Technological progress aims to use AI systems by introducing them into fully automated and unmanned vessels. At present, such systems are not used in merchant shipping. Easy access to advanced AI-based LLMs, such as ChatGPT-4o or Microsoft 365 Copilot, raises questions about their capabilities and limitations in supporting the decisions of navigators. Trained on data of unprecedented scale, large language models (LLMs), such as ChatGPT and GPT-4o, demonstrate the emergence of significant reasoning abilities resulting from model scaling [
14]. In 2024, research from the Swedish COLREG3 project conducted at the Maritime Competence Centre was published, showing the weak spatial awareness of LLMs, which significantly limits their use for maritime traffic analysis and often leads to misinterpretations [
15]. LLMs such as ChatGPT or Copilot operate based on advanced AI language models. Based on neural networks, they use deep learning to generate text. GPT not only supports automation but also opens the door to new possibilities in the field of human–machine interaction. Since LLMs have no consciousness and no access to current information other than that used during their training process, they may provide incorrect or imprecise answers. An important point is that COLREGs are written regulations and are, to some extent, ambiguous from the perspective of algorithms [
16]. Kristić and Žuškin [
17] discussed the ambiguity by analyzing linguistic variables used in COLREGs and quantified expert knowledge concerning the term “Very Large Ship” from Rule 7. Their study shows that some COLREGs expressions are fuzzy by nature and can be interpreted differently unless they are formally described, for example, by fuzzy sets. This problem underscores the essence of the research presented in this article, namely the correct interpretation of radar images, the analysis of encounters between vessels, and the proper implementation of COLREGs, which may contain linguistic and interpretive uncertainties.
Despite the growing number of studies on LLMs in maritime navigation and autonomous vessel decision-making, there remains a research gap concerning the ability of generally available LLMs to interpret radar plotting images and translate them into COLREGs-compliant maneuvering recommendations without additional navigational data input. Previous studies mainly focused on theoretical knowledge of ship handling, autonomous navigation frameworks, or LLM-based decision-making architectures supported by structured data. This indicates the gap of practical visual-spatial interpretation of radar plotting images by widely accessible tools such as ChatGPT-4o and Copilot. The contribution of this study is therefore threefold: first, it evaluates the ability of selected commercially available LLMs to analyze radar plotting images and to adapt their responses to previous conversational context within a single interaction; second, it identifies typical errors related to spatial interpretation, vector reading, distance estimation and COLREGs application; and third, it discusses the role of LLMs as explanatory and training-support tools rather than independent navigational decision-making systems.
2. Materials and Methods
Advanced multimodal LLMs can process image input and may therefore be used to analyze radar plotting images. In such an analysis, the model attempts to recognize the position of the own ship, echoes of other vessels, motion vectors, and radar range scale. It must also determine the position of the target vessel, the distance from the center of the screen, the bearing, the direction and length of the motion vector, the approximate speed, and the predicted CPA and TCPA.
The research demonstrated a pattern of radar plotting image analysis performed by LLMs. The models carry out the analysis independently and in stages:
At the initial stage, a general analysis of the image is performed; for example:
- ▪
observation range, e.g., 6 Nm,
- ▪
range rings scale,
- ▪
position of the own ship.
Without the aforementioned data, the analysis may be incorrect. For example, an incorrectly read radar scale leads to an erroneous determination of the target’s distance and speed.
- 2.
Target Detection
Next, AI identifies the echo of the target vessel. Depending on the manufacturer, the target on the radar image may appear as a point, a spot or a rectangle.
- 3.
Determining the Target’s Position
After detecting the target, the LLMs can determine:
- ▪
on which side of the own ship the target vessel is located,
- ▪
what the bearing is,
- ▪
what the distance is,
- ▪
whether the target is ahead, astern, to port or to starboard.
- 4.
Reading Course and Speed
If a motion vector is visible on the radar, AI can estimate the course and speed of the target vessel. The direction of the vector indicates the direction of movement, while its length may correspond to speed.
- 5.
Calculating CPA and TCPA and Suggesting a Maneuver
The most important element of the analysis is the assessment of whether there is a risk of collision. The following parameters are used for this purpose:
- ▪
CPA—Closest Point of Approach,
- ▪
TCPA—Time to Closest Point of Approach.
Assuming the user’s minimum safe setting, for example, a CPA of 1 Nm, the user expects guidance on selecting an appropriate maneuver, i.e., one compliant with COLREGs. For this purpose, after reading the position and motion of the target vessel, the system should classify the encounter and determine the type of navigational situation.
The analysis was carried out for 24 different simulations of situations presented in the form of radar image plotting, performed on the POLARIS [
18] multifunctional navigation bridge simulator manufactured by KONGSBERG.
The 24 radar plotting scenarios were selected as a representative set of simulator-based navigational training exercises. The selection was made to include different types of ship-to-ship encounters and different levels of interpretative difficulty, including crossing, overtaking, nearly head-on and no-action-required situations. The scenario set included 16 crossing situations, 3 head-on, or nearly head-on situations, and 3 overtaking situations. Each scenario was assessed in terms of risk of collision. If the verified CPA exceeded the predefined minimum safe passing distance, the situation was classified as a no-action-required encounter, because the target vessel was expected to pass the own ship at a safe distance. Four scenarios were classified as no-action-required cases on this basis. The scenarios also differed in the target vessel’s position relative to the own ship, the distance from the own ship, the direction of the motion vector and the expected COLREGs-based assessment. For each scenario, the ground truth was established before the LLM analysis. The correct encounter classification, applicable COLREGs rule or rules, give-way and stand-on roles, and expected own-ship maneuver were determined by an expert navigator on the basis of the radar plotting image, simulator scenario assumptions, relative position and motion of the vessels, the assumed minimum CPA criterion of 1 Nm and the relevant COLREGs rules. This expert-defined ground truth was then used as the reference for evaluating the correctness of the LLM-generated responses.
The purpose of
Table 1 is to clarify the scope and diversity of the analyzed scenarios before presenting the model outputs and accuracy results.
The initial assumption was to examine the recognition capabilities and effectiveness of three selected commercially available LLMs in maritime accident analysis.
DeepSeek-R1 was initially considered as an additional LLM-based tool because of its growing availability. However, under the applied research conditions, it did not provide sufficient radar image recognition capability; therefore, it was excluded from the final analysis.
The study was designed as a preliminary exploratory assessment rather than a comprehensive benchmark of all available LLMs. Therefore, the analysis was limited to two generally available multimodal tools: ChatGPT-4o and Microsoft 365 Copilot. These tools were selected because they were accessible to users and allowed image-based interaction with radar plotting screenshots.
Since the ChatGPT-4o multimodal model is capable of processing text and image inputs [
19], it was selected to analyze images from navigation radars. By incorporating the Copilot tool, which is integrated with the Microsoft 365 suite, both solutions were used to conduct preliminary research on the recognition process and the proper application of COLREGs.
The experiment was conducted in November 2025. The radar plotting images were exported from the POLARIS navigation bridge simulator and used as image inputs for the tested LLM-based tools. The images were provided to the models in JPG format, with a resolution of approximately 1101 × 1055 px (78.9 KB). No additional structured navigational data, such as numerical CPA/TCPA tables, AIS data or ARPA target data, were provided to the models apart from the information visible in the radar plotting image and the short text prompt.
The tested tools were OpenAI’s GPT-4o multimodal model accessed through the ChatGPT Business environment with image input enabled, and Microsoft 365 Copilot, version 19.2608.54041.0, accessed through the Microsoft 365 environment. The same radar plotting images and the same prompting procedure were used for both tools. For the first analyzed image, the following prompt was used: “Situation 01, radar range: 6 nautical miles. Minimum CPA: 1 nautical mile. According to COLREGs, what should I do as a vessel?” For the subsequent images, the prompts contained only the sequence number of the situation and a request for recommendation, for example: “Situation 3, what do you suggest?” or simply “Sit 10”. All analyses for each model were conducted within a single conversation history in order to assess whether the model could maintain and use contextual information during the interaction.
In this study, a correct maneuvering suggestion was defined as a response in which the model correctly interpreted the encounter situation sufficiently to select the relevant COLREGs rule or rules, correctly assigned the give-way and stand-on roles where applicable, and proposed an own-ship maneuver consistent with COLREGs and the assumed minimum CPA criterion of 1 Nm. Minor inaccuracies in the estimation of descriptive parameters, such as distance, were accepted if they did not affect the final rule application or maneuvering recommendation. For example, a response with an approximate distance error was still classified as correct if the model correctly identified the crossing situation, assigned the own ship as the give-way vessel, and recommended an appropriate maneuver. Conversely, a response was classified as incorrect when an error in spatial interpretation, vector reading, distance assessment, COLREGs application, or maneuvering recommendation led to an incorrect or unsafe suggestion.
To ensure consistency in the evaluation, the following scoring rubric was applied (see
Table 2). The assessment focused primarily on the correctness of the final maneuvering suggestion; however, the reasoning process was also considered when an error in reasoning affected the encounter classification, COLREGs application, assignment of vessel obligations or safety of the recommended action.
If the final maneuvering recommendation was safe but the reasoning contained a substantive error that affected the legal or navigational justification, such as incorrect encounter classification or wrong assignment of give-way and stand-on roles, the response was classified as incorrect. This approach was adopted because, in navigational decision-making, a safe-looking maneuver based on incorrect reasoning may lead to unsafe decisions in similar or slightly modified situations.
Incorrect answers were then corrected by an expert. Expert validation was performed by a senior officer of the watch with 15 years of navigational experience on international voyages on various types of merchant vessels. The expert is also a lecturer at the Faculty of Navigation at the Maritime University of Szczecin, teaching, among other subjects, maritime law. However, expert correction did not always lead the model to verify and apply the COLREGs rules correctly.
All analyses for each model were conducted within one conversation history. This procedure was intentionally used to examine whether the model could maintain and use previous conversational context, including earlier prompts, previous radar plotting analyses and expert corrections. The observed effect should not be interpreted as self-learning in the training sense, because the model parameters were not updated during the interaction. Instead, it reflects context-dependent response adaptation, or in-context adaptation, in which later responses may be influenced by information provided earlier in the same conversation. This design was therefore used to assess the practical behavior of generally available chatbot tools during a continuous analytical session. At the same time, it is acknowledged as a methodological limitation because previous conversational context may have influenced later responses and reduced the independence of individual scenario assessments.
Two selected situations are presented below as illustrative examples.
Figure 1 shows a case in which ChatGPT-4o provided a correct analysis, whereas
Figure 2 shows a case in which the initial analysis was incorrect.
The COLREGs cover various signals, such as sound signals, day shapes, and navigation lights, which play a key role in preventing collisions and maneuvering vessels. Without an operator involved in the decision-making process to take into account additional information from visual and auditory signals, important information may be overlooked [
20]. The analyses performed by LLMs were based solely on the radar plotting image, i.e., information about the relative positions of the vessels; without light/daytime signals and/or sound signals, the correct application of COLREGs is practically impossible. For research purposes of this article, it was assumed that vessels maneuver without any weather restrictions, restrictions from the nature of their work, and are power-driven vessels underway.
The two examples presented above illustrate two typical outcomes of the analysis: a correct interpretation of a crossing situation and an incorrect interpretation caused by a wrong assessment of the target vessel’s side. Full chatbot responses are provided in the
Supplementary Material.
Table 3 presents only the elements relevant to the evaluation: situation interpretation, suggested COLREGs rule, maneuvering recommendation, and expert assessment.
3. Results
The analysis of Situation 1, presented in
Figure 1, showed that ChatGPT-4o correctly interpreted the general encounter situation and selected the appropriate COLREGs rule. The model identified the target vessel as approaching from the starboard side and classified the situation as a crossing encounter. Consequently, the own vessel was correctly indicated as the give-way vessel under Rule 15 of COLREGs.
However, a minor inaccuracy was observed in the estimation of distance. ChatGPT-4o assessed the distance to the target vessel as approximately 5.0–5.5 Nm, whereas the actual distance was approximately 5.8 Nm. In this case, the error did not affect the classification of the encounter or the recommended maneuver. Nevertheless, similar inaccuracies in distance estimation were observed in other analyzed cases and may influence the assessment of collision risk, especially when CPA and TCPA values are close to the assumed safety limits.
The analysis of the situation presented in
Figure 2 revealed several irregularities in the response generated by ChatGPT-4o. The model incorrectly interpreted the relative position of the target vessel and initially identified the vessel as being located on the port side of its own ship. This error affected the identification of the give-way and stand-on vessels and consequently led to an incorrect maneuvering recommendation. Although the model referred to the correct COLREGs rule for a crossing situation, the rule was applied to an incorrectly interpreted spatial configuration.
After an expert correction indicating that the target vessel was located on the starboard side of its own ship, ChatGPT-4o revised its response and generated a more appropriate recommendation in accordance with COLREGs. This suggests that expert input may improve the quality of the model’s response. However, it also indicates that the model may not reliably detect its own errors in the interpretation of relative vessel positions.
Across the 24 scenarios, the most frequent causes of incorrect responses were related to spatial interpretation and radar reading rather than to the inability to mention COLREGs rules. In several cases, the models suggested a plausible rule but applied it to an incorrectly interpreted situation. This was especially evident when the model misidentified the side of the target vessel, misread the motion vector, or incorrectly assessed whether its own ship should act as the give-way or stand-on vessel.
Based on the analysis of all 24 radar plotting situations (see
Table 4,
Table 5 and
Table 6), the following results were obtained:
ChatGPT-4o provided correct maneuvering suggestions in 7 out of 24 situations analyzed.
Copilot provided correct maneuvering suggestions in 11 out of 24 analyzed situations.
In the repeated analysis performed with ChatGPT-4o, correct suggestions were obtained in 4 out of 24 situations.
The accuracy of the models was calculated as the ratio of the number of correct maneuvering suggestions to the total number of analyzed scenarios:
The corresponding accuracy values were:
A comparative summary of LLM performance in radar plotting analysis is presented in
Table 7.
Due to the limited number of analyzed scenarios, the reported accuracy values should be interpreted as descriptive results obtained under the applied experimental conditions. The study was exploratory in nature and was not designed to provide a statistical comparison of model performance. Therefore, the observed difference between ChatGPT-4o and Microsoft 365 Copilot should not be interpreted as evidence of a statistically significant or generally meaningful performance advantage of one model over the other. A larger dataset and multiple independent repetitions for each model would be required to quantify uncertainty and compare model performance statistically.
The repeated ChatGPT-4o analysis was included as an exploratory check of response stability rather than as a full repeatability study. In this repeated run, correct maneuvering suggestions were obtained in 4 out of 24 situations, compared with 7 out of 24 in the first ChatGPT-4o analysis. This indicates variability between the two analyses of the same scenario set. However, because only one repeated run was performed, the results should not be interpreted as a complete statistical assessment of repeatability or reproducibility. Instead, they should be treated as a preliminary indication that LLM-generated radar plotting assessments may vary between repeated interactions and that more systematic repetition is required in future studies.
The most frequently observed errors in the responses generated by the analyzed LLMs included:
incorrect reading of the own ship’s course;
incorrect reading of the target vessel’s course;
incorrect interpretation of the target vessel’s position relative to the own ship;
difficulty in determining whether the target vessel was located on the port or starboard side;
inaccuracies in estimating the distance from the own ship;
incorrect selection or application of the relevant COLREGs rule;
difficulty in applying expert corrections consistently.
To identify the source of incorrect responses more precisely, the observed errors were additionally grouped into visual-spatial interpretation errors, navigational reasoning errors and mixed errors. Visual-spatial interpretation errors included incorrect identification of the target vessel’s side, incorrect reading of the own ship or target vessel’s course, incorrect interpretation of relative geometry, and distance or CPA-related interpretation errors. Navigational reasoning errors included incorrect COLREGs rule application, wrong assignment of give-way and stand-on roles, or incorrect maneuver recommendations despite sufficient interpretation of the radar situation. Mixed errors were defined as cases in which an initial visual-spatial misinterpretation led directly to incorrect COLREGs application or maneuver recommendation. The classification criteria are presented in
Table 8, while the quantitative summary of these errors is presented in
Table 9.
Overall, the obtained results indicate that the tested LLMs were able to generate structured and convincing descriptions of radar plotting situations. However, their correctness was limited, particularly in tasks requiring visual–spatial reasoning, the interpretation of motion vectors, and the application of COLREGs to specific encounter situations.
4. Discussion
The results obtained in this study show that generally available LLMs may support the interpretation of radar plotting images, but their reliability remains limited, especially in tasks requiring visual-spatial reasoning and variable radar interpretation. This conclusion is directly supported by the quantitative results: ChatGPT-4o provided correct maneuvering suggestions in 7 out of 24 situations (29.2%), Copilot in 11 out of 24 situations (45.8%), and the repeated ChatGPT-4o analysis produced correct suggestions in only 4 out of 24 situations (16.7%).
These values indicate limited accuracy under the applied experimental conditions. However, because the study included only 24 scenarios, the differences between ChatGPT-4o and Microsoft 365 Copilot should be interpreted descriptively and with caution. The results do not provide sufficient statistical evidence to claim a meaningful performance advantage of one model over the other.
The main sources of error were not only incorrect application of COLREGs but also incorrect interpretation of radar-derived spatial information. The most frequent errors included incorrect reading of the own ship’s course, incorrect interpretation of the target vessel’s course, wrong assessment of whether the target vessel was located on the port or starboard side, inaccurate distance estimation, and incorrect application of the relevant COLREGs rule. These findings are consistent with the observation by Xie et al. [
21] that LLMs may fail in tasks requiring numerical, physical, and spatial reasoning. They are also consistent with the conclusions of Silwal and Dubey [
22], who reported limited LLM performance in COLREGs-compliant autonomous maritime navigation, and with the Swedish COLREG3 project [
15], which indicated that poor spatial reasoning significantly limits the use of LLMs in maritime traffic analysis.
The comparison with previous studies should also be interpreted in the context of methodological differences. Pei et al. [
23] reported a higher accuracy for GPT-4o in ship-handling-related tasks; however, their study focused mainly on ship-handling knowledge and skills, whereas the present study required direct visual interpretation of radar plotting images, including relative position, motion vectors, CPA-related assessment and COLREGs-based maneuvering suggestions. Similarly, LLM-based maritime decision-making frameworks, such as the system proposed by Agyei, Sarhadi and Naeem [
16], and the CORALL concept [
24], use structured collision-risk indicators, COLREGs guidance or additional decision-support modules. In contrast, the present study evaluated generally available commercial LLMs used as multimodal chatbots, without deterministic verification modules or structured navigational input. Therefore, the lower accuracy obtained in this study reflects the specific difficulty of direct radar image interpretation rather than only the general ability of LLMs to recall or discuss COLREGs.
A key problem observed during the analyses was the incorrect interpretation of spatial relationships between vessels. In several cases, the model incorrectly determined whether the target vessel was located on the port or starboard side of the own ship. This error is critical because the relative position of the target is one of the main elements used to classify an encounter situation under COLREGs. In a crossing situation, for example, the vessel that has the other vessel on her starboard side is generally required to keep out of the way. Therefore, an incorrect reading of the side on which the target is located may directly lead to the wrong identification of the give-way and stand-on vessels.
From the communication perspective, the role of LLMs is particularly sensitive. Misztal and Hatlas-Sowinska [
2] show that communication during collision situations is strongly affected by the human factor and that unclear or incorrectly interpreted information may reduce the effectiveness of collision avoidance. In the present study, the analyzed LLMs generated fluent and persuasive explanations of radar plotting situations; however, some of these explanations were based on incorrect spatial interpretation. This means that chatbot-generated communication may support the navigator, but it may also increase risk if the underlying analysis is wrong.
Another important limitation was the incorrect reading of motion vectors. In radar plotting analysis, the vector of the target vessel is essential for determining its course, speed, and future movement relative to the own ship. The study showed that LLMs may confuse the direction of movement, misread the course of the target vessel, or incorrectly interpret the own ship’s course. Such errors have a direct impact on the assessment of CPA and TCPA and may result in an inappropriate recommendation of action required to navigate safely. In practice, even a small error in vector interpretation may change the classification of the encounter from crossing to overtaking or from crossing to a nearly head-on situation.
The results also indicate that the models may have difficulties with the correct estimation of distance from the radar display. Errors in reading the radar range scale or interpolating the distance rings may lead to incorrect assumptions about the proximity of the target vessel. In collision avoidance, distance estimation is closely related to the time available for action and to the assessment of whether a risk of collision exists. If the distance is overestimated, the model may underestimate the urgency of the situation. If it is underestimated, the model may recommend unnecessary or excessive maneuvers.
The repeated analysis performed with ChatGPT-4o should be interpreted with caution. It was conducted as an exploratory check of response stability, not as a complete repeatability or reproducibility assessment. Since only one repeated analysis was performed, the study cannot provide statistically robust conclusions regarding consistency. Nevertheless, the difference between the first ChatGPT-4o analysis and the repeated analysis suggests that model responses may vary when the same or comparable radar plotting scenarios are analyzed again. Future research should include multiple independent repetitions for each model, preferably in separate conversation histories and under controlled prompting conditions, in order to assess repeatability and reproducibility more rigorously.
The study also showed that expert correction may improve the quality of the model’s response. In Situation 2 (see
Figure 2), after the expert indicated that the target vessel was located on the starboard side of the own ship, the model was able to revise the COLREGs assessment and provide a more appropriate recommendation. This demonstrates that LLMs may be useful in an interactive human–machine process, where the navigator verifies the model’s interpretation and provides corrections when necessary. However, it also confirms that the model itself may not reliably detect its own error. The need for expert intervention limits the possibility of using such tools as autonomous decision-making systems, but it supports their potential role as auxiliary training or explanatory tools.
The persuasive language generated by LLMs represents both an advantage and a risk. On the one hand, the ability to explain a radar situation in simple language is valuable, especially in training, simulation exercises, and decision-support environments. The model can structure the analysis, recall the relevant COLREGs rule, and present the reasoning in an accessible form. On the other hand, an incorrect answer may be presented with the same confidence and clarity as a correct one. This may be dangerous for inexperienced users who may not be able to identify the underlying error in image interpretation or rule application.
It should also be emphasized that the analyses in this study were based only on radar plotting images. In real navigation, the officer of the watch uses many additional sources of information, including visual observation, AIS, sound signals, navigation lights, day shapes, VHF communication, environmental conditions, traffic density, maneuvering characteristics of the vessel, and the ordinary practice of seamen. COLREGs are not applied solely on the basis of a static radar image. Therefore, even a correct interpretation of the radar plotting picture may not be sufficient to determine the safest maneuver in an operational situation.
Consistent with the COLREG3 project [
15], this confirms that the main weakness of LLMs in this context is not only knowledge of COLREGs but also the ability to correctly connect visual information with navigational reasoning. The models often know the content of the rules and can quote or paraphrase them correctly, but they may apply them to a wrongly interpreted situation. The problem is therefore not only legal or procedural knowledge but also perception, geometry, and spatial reasoning.
Future research should include a larger number of radar situations, different radar ranges, different vector lengths, more than one target vessel, and varying levels of traffic complexity. It would also be useful to compare model performance under different prompting strategies. Future research should also include a broader comparison of multimodal LLMs, including additional commercially available and open-source models, with more independent repetitions for each model in order to assess not only accuracy but also repeatability and model-specific differences in radar plotting interpretation. The present study used relatively simple prompts in order to observe the models’ spontaneous analytical capabilities. However, more structured prompts containing information about radar range, own ship course, vector time, and minimum acceptable CPA may improve the quality of the responses. Such research would help determine whether LLM performance depends mainly on visual interpretation limitations or on insufficient task formulation.
In summary, the discussion of the results shows the dual nature of LLMs in maritime navigation. They have clear educational and explanatory potential, especially in reading radar plotting data into understandable language and recalling the structure of COLREGs reasoning. At the same time, their errors in spatial interpretation, vector reading, distance estimation, and rule application prevent their use as independent decision-making systems. The results of this study support the conclusion that LLMs may become valuable elements of future human–machine interaction on the bridge, but only as carefully validated decision-support tools operating under expert supervision.
5. Conclusions
The results of this study indicate that, under the tested experimental conditions, generally available LLMs such as ChatGPT-4o and Microsoft 365 Copilot showed limited reliability in radar plotting image analysis and COLREGs-based maneuvering suggestions. Although the models were able to generate structured and convincing explanations, their accuracy was relatively low, and their responses included fundamental errors in visual-spatial interpretation, motion vector reading, distance assessment and COLREGs application. Therefore, the tested systems cannot currently be considered reliable tools for independent collision-risk assessment or operational navigational decision-making.
Their potential use should be limited to supervised training, explanation, post-encounter review and human–machine communication. In such contexts, LLMs may help structure information, explain possible encounter situations and recall relevant COLREGs considerations, but all outputs must be verified by a qualified officer. The final navigational assessment and maneuvering decision must remain under the responsibility of the officer of the watch and/or the master.
The study also indicates a need for further development of hybrid systems. A more reliable solution may involve combining LLMs with deterministic navigational algorithms, ARPA data, sensor fusion, and rule-based COLREGs modules. Research on COLREGs-compliant deep reinforcement learning for multi-vessel collision avoidance also supports the need for hybrid approaches. Xie et al. [
25] proposed a DRL-based decision-making method for autonomous surface vessels that addresses multi-vessel collision avoidance under COLREGs constraints. In this context, LLMs should not replace validated collision-avoidance algorithms, but may complement them by explaining verified outputs, supporting human–machine interaction and presenting the navigational reasoning in natural language. A similar direction is consistent with the CORALL concept [
24], where the LLM is guided by COLREGs and risk-awareness mechanisms rather than being used as an unrestricted conversational model. Such an approach may reduce the risk of incorrect interpretation by separating the explanatory and high-level decision-support role of the LLM from validated risk assessment and maneuver execution modules. In such a system, the LLM would not be responsible for independently reading the radar image or calculating CPA and TCPA. Instead, it could explain verified data generated by certified or validated systems. This would allow the model’s strength—natural language explanation and human–machine interaction—to be used while reducing the risk resulting from incorrect visual interpretation.
The potential value of LLMs lies mainly in their ability to convert verified technical radar data into a comprehensible explanatory message. In a supervised context, such systems may help describe:
- ▪
where the target vessel appears to be located;
- ▪
how it appears to be moving;
- ▪
whether a potential risk of collision may require further assessment;
- ▪
which COLREGs rule may be relevant;
- ▪
what maneuvering options may require consideration by the navigator.
The analyses of the situations presented are delivered in a highly convincing manner; however, they are unfortunately often incorrect. Similar conclusions were also drawn in the Swedish COLREG3 project, which emphasized that “poor spatial reasoning completely hinders their use for the analysis of marine traffic situations” and that careful validation and verification frameworks will be needed to operationalize LLMs as decision-support systems for maritime navigation [
15].
Preliminary analysis concerning maneuvering suggestions provided by LLMs may differ significantly from the actions taken by an operator. For example, according to COLREGs, the interpretation of a situation involving a stand-on vessel is that it should maintain its course and speed. However, depending on the circumstances, the stand-on vessel may decide to take very early action to avoid a close-quarters situation by carrying out a maneuver [
26].
LLMs may support situational awareness and more structured navigational reasoning only in supervised training or advisory contexts. It is therefore useful to distinguish their potential explanatory advantages from their operational limitations. It is worth mentioning not only the disadvantages but also the advantages of using LLMs. Possible advantages of LLMs in supervised training or advisory contexts include:
- ▪
explaining radar plotting information in simple language;
- ▪
recalling relevant COLREGs rules;
- ▪
supporting the interpretation of encounter situations;
- ▪
helping the user consider selected aspects of the situation;
- ▪
providing an additional explanatory layer for human–machine interaction.
These potential advantages correspond with the review by Miller et al. [
3], who indicate that LLMs may enhance maritime safety through improved communication, training, decision support, compliance support and real-time risk assessment. In the context of radar plotting analysis, these functions may support the navigator’s situational awareness, provided that the model’s output is verified by a qualified officer.
Tested LLMs are not sufficiently advanced models to make navigational decisions independently. The results of this study confirm the general thesis presented by Guo et al. [
27] that the evaluation of LLMs should encompass not only their task-specific capabilities but also aspects of safety, reliability, and responsible use. In the context of maritime navigation, this means that every maneuvering suggestion generated by a chatbot must be verified by a qualified expert. Such tools may support human interpretation, but they do not replace qualified officers of the watch or certified navigational systems. The preliminary research conducted in this study showed several limitations that require continuous expert supervision.
The most important limitations include:
- ▪
possible incorrect reading of the image,
- ▪
incorrect interpretation of vectors,
- ▪
difficulties in assessing the intentions of the other vessel,
- ▪
the necessity of applying good seamanship,
- ▪
legal responsibility remaining with the human operator.
Considering these limitations, the safest near-term application of LLMs in radar plotting analysis may be found in simulator training, post-encounter review, and advisory bridge displays rather than in direct maneuver control. In such applications, the LLM may explain the encounter situation, indicate missing information, summarize the relevant COLREGs context, and support the navigator’s situational awareness, while maneuvering authority remains with the officer of the watch and/or the master. Therefore, responses provided by algorithms used in LLMs should be treated as decision-making support, not as an automatic maneuvering order. Future use of LLMs in autonomous or semi-autonomous navigation systems will require integration with validated navigational systems, reliable sensor data, collision-risk assessment modules and expert supervision.