Next Article in Journal
A Preliminary Study on the Role of Orexin A in Leydig Cell Steroidogenesis and Its Implications for Fertility in Alpacas (Vicugna pacos)
Previous Article in Journal
Swine Influenza Virus Introduction in Pig Farms: A Semi-Quantitative Risk Assessment in Northern Italy
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters

1
Department of Agricultural, Forest and Food Sciences, University of Turin, Largo Paolo Braccini 2, 10095 Grugliasco, TO, Italy
2
Department of Prevention, Asl TO3, Veterinary Service, Animal Health Unit, Via Torino 62, 10045 Piossasco, TO, Italy
3
Department of Economics, Quantitative Methods and Business Strategies, University of Milan-Bicocca, Piazza dell’Ateneo Nuovo 1, 20126 Milan, MI, Italy
4
Department of Agricultural and Environmental Sciences—Production, Landscape, Agroenergy, University of Milan, Via Celoria 2, 20133 Milan, MI, Italy
5
Department of Veterinary Sciences, University of Turin, Largo Paolo Braccini 2, 10095 Grugliasco, TO, Italy
*
Author to whom correspondence should be addressed.
Animals 2026, 16(4), 546; https://doi.org/10.3390/ani16040546
Submission received: 14 January 2026 / Revised: 6 February 2026 / Accepted: 7 February 2026 / Published: 10 February 2026
(This article belongs to the Section Animal Welfare)

Simple Summary

Nowadays, the evaluation of inter-observer reliability is of outmost importance for ensuring the introduction of individual animal-based welfare indicators within animal welfare protocols. The present study focuses on the evaluation of inter-observer reliability of dichotomous and trichotomous individual animal-based welfare indicators (assessed through two/three levels scoring system), which is guaranteed calculating the concordance among three raters during the evaluation process through some statistical indices proposed in the current literature, defined as agreement indices. In this regard, the performance of the most popular agreement indices is compared to understand which ones are the most suitable to assess the inter-observer reliability. The most exploited agreement indices (e.g., the indices belonging to the Kappa statistic) are shown to be inappropriate to evaluate the inter-observer reliability in the presence of three raters. On the contrary, some less known agreement indices, such as Gwet’s γ(AC1), Gwet’s γ(AC2), Quatto’s S, Quatto’s weighted S, Brennan and Prediger’s BP coefficient and Brennan and Prediger’s weighted BP, were able to confer more reliable agreement results.

Abstract

This study deals with the evaluation of inter-observer reliability (IOR) among three raters in the case of dichotomous and trichotomous individual animal-based welfare indicators. The performance of the most documented agreement indices proposed in the literature was compared, using udder asymmetry (UA) as a dichotomous indicator and body condition score (BCS) as a trichotomous indicator, both obtained from the AWIN Goat protocol. Nine dairy goat farms, exploiting three alpine pastures (AP1 to AP3), were used for data collection. Krippendorff’s α, the agreement indices belonging to the Kappa statistic and their weighted forms were in some cases affected by the paradox behaviour. This phenomenon was observed for both UA and BCS [e.g., P0(BCS-AP2) = 80%; Fleiss’ K = 0.22]. In the case of UA, Gwet’s γ(AC1), followed by BP coefficient and Quatto’s S, gave the best agreement results [e.g., P0(UA-AP1) = 86%; γ(AC1) = 0.84]. In the case of BCS, the best agreement results were obtained with Gwet’s γ(AC2), followed by the weighted forms of BP and S. When the evaluation is performed by three raters, γ(AC1), BP and S are suggested to evaluate IOR in the case of both dichotomous and trichotomous indicators, while the related weighted forms are suitable for trichotomous indicators only.

1. Introduction

Animal-based welfare indicators confer accurate information regarding the real welfare status of a subject [1], measuring the reactions of the animal to both the resources present inside the environment where it lives (resource-based welfare indicators) and its manipulation by humans (management-based welfare indicators) [2].
Together with validity and feasibility, reliability is one of the most important features for an animal-based welfare indicator to be included into animal welfare protocols [3]. A relevant type of reliability is the inter-observer reliability (IOR) [4], which is linked to the level of agreement between two or more raters, when they classify independent sample units inside a predetermined category [5] at the same time and without influencing each other [6].
A proper IOR evaluation is guaranteed developing some statistical indices proposed in the literature, defined as agreement indices, which confer a value of the concordance among the raters during the evaluation process [7]. Subsequently, the obtained concordance value is compared to the concordance rate (P0), which is given by the ratio between the number of times that the raters agree out of the total number of observations [8]. Thus, the lower is the concordance among the raters, the lower will be the IOR of an indicator.
Most of the animal-based welfare indicators contained inside currently available animal welfare protocols are dichotomous and trichotomous variables which are evaluated using two- and three-level scoring, respectively. Giammarino et al. [9] and Torsiello et al. [6] identified the most suitable agreement indices to assess the IOR for dichotomous and trichotomous variables in the presence of two raters. However, welfare assessment can be performed by different assessors for different purposes (e.g., self-assessment, official controls, and certification procedures); therefore, it is also fundamental to assess which are the most suitable indices to evaluate the agreement among multiple raters for the above-mentioned variables.
In this regard, as already reported for two raters [6,9], also in the presence of multiple raters, a critical point is represented by the implementation of the rate of agreement that occurs by chance (Pe), which must be removed from the P0 [10]. Gwet [11] stated that adjusting the concordance rate for the chance agreement becomes crucial in the presence of multiple raters, due to the constraints imposed by the experimental design. Specifically, Gwet [11] made the example of three raters, who can classify a subject into two categories only. In this case, the three raters will not have the possibility to completely disagree, due to the low presence of categories. Consequently, two out of three raters will necessarily agree on this classification, with the possibility that some of the agreements will be due to chance.
The Pe is calculated in different ways, depending on the formula of the implemented agreement index [7]. For example, Fleiss [12] defined the chance agreement as the probability to assign a subject into the same category by a single pair of raters. Krippendorff [13] followed the same approach, even if, differently from Fleiss [12], he proposed a statistic which is applicable in the presence of both two and multiple raters [6]. Conger [14] criticised this assumption, stating that the Pe can be implemented averaging the probabilities of assignments of a subject into the same category not only by a single pair but by all the involved pairs of raters.
However, as already pointed out for two raters [6,9], also in the presence of three raters, the agreement indices belonging to the Kappa statistic (Fleiss’ K [12], Light’s K [15], Hubert’s K [16], and Conger’s K [14]) can be affected by the paradox behaviour [17]. Indeed, when the Pe assumes high values, the agreement values obtained implementing the Kappa statistic can sometimes result very low if compared to the P0 [17]. On the contrary, some agreement indices which calculate the Pe taking into consideration the number of categories that characterises the variable under analysis are free from the paradox and are suitable to assess the IOR properly [6].
Based on the above-reported considerations, the aim of this study is to identify the most suitable agreement indices for properly assessing the IOR of dichotomous and trichotomous animal-based welfare indicators in the presence of three raters. For this purpose, we selected two indicators from a modified version of the original AWIN welfare assessment protocol for goats [18], namely the udder asymmetry (UA) and body condition score (BCS), which are evaluated using two- and three-level scoring, respectively.

2. Materials and Methods

2.1. Dichotomous and Trichotomous Animal-Based Welfare Indicators

A modified version of the AWIN protocol developed for the welfare assessment of dairy goats kept under semi-extensive farming systems [18] was applied by three raters in nine dairy goat farms, exploiting three alpine pastures (APs) which were breeding a total of 160 goats (AP1: n = 44; AP2: n = 70; and AP3: n = 46), in north-west Italy, between June and August 2021. Two raters were enrolled in the second year of the MSc in Animal Science, while the third observer was enrolled in the first year of the MSc in Forestry and Environmental Sciences, at the University of Turin (Italy). The raters had no previous experience with dairy goats. Before data collection, the raters received both a theoretical and a practical training by one of the authors of the original AWIN protocol developed for the welfare assessment of dairy goats kept under intensive and semi-intensive farming systems [19]. In addition, as training material, they received both the original AWIN protocol [19] and the adapted AWIN protocol to be applied in semi-extensive farming conditions [18]. The theoretical training was based on these protocols and on additional training material developed within the AWIN project and made available to the raters, consisting of Power Point presentations divided into the following different sections: (1) Definition of the indicator; (2) How to assess it; (3) How to score it; (4) Examples (four photos for UA and 15 for BCS, plus three detailed drawings for BCS); and (5) Self-assessment (six questions for each indicator, where raters can test their knowledge and ability to correctly assess the indicator with immediate feedback on the correct/incorrect answer). After all the raters correctly assessed the indicators using the theoretical training material, a practical training was carried out with an expert trainer (one of the authors of the AWIN protocols) on a farm raising about 80 lactating goats. The indicators were further described, including methods and scoring systems, and discussed with practical examples. Raters were then asked to simultaneously assess UA and BCS and then their assessment was discussed to agree on the scoring and to clarify uncertainties. Only when all the raters assessed goats giving the same scores, the training was considered satisfactory (one full day was required).
The UA and BCS were chosen as dichotomous and trichotomous categorical animal-based welfare indicators, respectively. The UA was confirmed when one half of the udder was at least 25% longer than the other, excluding the teats [19]. According to the original AWIN protocol [19], for UA each goat was assigned to one of two mutually exclusive and exhaustive categories (absence of asymmetry = 0; presence of asymmetry = 1). This binary classification had previously been confirmed as a good predictor of the somatic cell count and the microorganism present in the udder [20]. For BCS each goat was assigned to one of three mutually exclusive and exhaustive categories (very thin goat = −1; normal goat = 0; and very fat goat = 1). Both indicators were recorded for all the goats in the three APs.

2.2. Agreement Indices and Confidence Intervals

The simplest measure for assessing the reliability is the P0. Fleiss [12] defined P0 as the ratio between the number of times that the pairs of raters agree in assigning the scores to each subject involved in the evaluation and the total number of subjects. However, this assumption was already previously criticised by Cohen [21], as P0 does not consider the possibility of agreement occurring by chance [22]. For this reason, to estimate the IOR properly, it is fundamental to implement agreement indices which also consider the Pe.
The most documented agreement indices implemented to assess the IOR for dichotomous and trichotomous variables, in the presence of multiple raters, are reported in Table 1.
In particular, Krippendorff’s α [13], Fleiss’ K [12], Light’s K [15], Hubert’s K [16], Conger’s K [14], BP coefficient [24], Quatto’s S [25], Gwet’s γ(AC1) [27] and Andrès and Hernàndez’s multi-raters Δ [28] are the most documented in the current literature to evaluate the IOR among three raters for both dichotomous and trichotomous indicators. In the case of trichotomous indicators, in addition to the above-mentioned agreement indices, Krippendorff’s weighted α (α*; [29]), Gwet’s γ(AC2) [11], Fleiss’ weighted K (K*; [11]), Conger’s weighted K (K*; [11]), weighted BP coefficient (BP*; [11]), Quatto’s weighted S (S*; [26]) and Hubert’s weighted K (K*; [30]) are also available.
The closed formulas of the above-mentioned agreement indices are reported in the Supplementary Materials.
For each agreement index, the calculation of the confidence intervals is fundamental to gather information regarding the dispersion of the values assumed by the index itself inside the sample. For this purpose, the calculation of the variance estimates is a prerequisite to guarantee the implementation of the confidence intervals.

2.3. Statistical Analyses

Differently from the approach applied in the case of two raters by Giammarino et al. [9] and Torsiello et al. [6], in the current study the manual implementation of the agreement indices based on closed formulas was not performed, because the complexity of the calculation increases as the number of raters increases.
R Commander (version R × 64 4.2.2) was used to implement both the values and the confidence intervals of all the considered agreement indices. The Bootstrap t-method [31] was implemented to calculate values and confidence intervals of all the agreement indices, while some packages and R functions were specifically developed to calculate the values only (i.e., Krippendorff’s α, Fleiss’ K, Light’s K, and Andrès and Hernàndez’s multi-raters Δ) or both the values and the confidence intervals (i.e., Krippendorff’s α, Fleiss’ K, Conger’s K, BP coefficient, Quatto’s S, Krippendorff’s α*, Gwet’s γ(AC1), Gwet’s γ(AC2), Fleiss’s K*, Conger’s K*, BP* coefficient, and Quatto’s S*) of the agreement indices.
All the indices and confidence intervals were calculated separately for each AP.
All the exploited R packages (version R × 64 4.2.2) and functions are summarised in Table 1.

3. Results

3.1. Dichotomous Animal-Based Welfare Indicators

3.1.1. Agreement Indices for Udder Asymmetry

The values of the agreement indices obtained for UA in each AP are reported in Table 2.
Two different concordance rates (P01; P02) were obtained, based on the agreement index under analysis. The first one (P01) was the same for all the agreement indices, except for Hubert’s K and Andrès and Hernàndez’s multi-raters Δ, for which a different concordance rate (P02) was obtained.
In AP1, despite the fact that P01 and P02 were equal to 86% and 80%, respectively, Krippendorff’s α and all the indices belonging to the Kappa statistic (i.e., Fleiss’ K, Light’s K, Hubert’s K, and Conger’s K) resulted in negative values (−0.07 to −0.03). In both AP2 and AP3, the results obtained for the above-mentioned agreement indices were above zero; however, their values (AP2: 0.51 and 0.52; AP3: 0.68 and 0.69) were low if compared to their respective concordance rates (AP2: P01 = 92% and P02 = 89%; AP3: P01 = 94% and P02 = 91%).
The BP coefficient, Quatto’s S and Gwet’s γ(AC1) conferred agreement results close to P01 in all the considered cases. Gwet’s γ(AC1) values (0.84, 0.91 and 0.93 in AP1, AP2 and AP3, respectively) were closer to those of the concordance rate (P01) if compared to BP coefficient and Quatto’s S ones (0.73, 0.85 and 0.88, respectively). Andrès and Hernàndez’s multi-raters Δ gave agreement results in line with the respective concordance rate (P02) in two out of three APs (AP2: P02 = 89%, Δ = 0.79; AP3: P02 = 91%, Δ = 0.86); however, in AP1 this index exceeded the concordance rate (P02 = 80%, Δ = 0.86).

3.1.2. Confidence Intervals for Udder Asymmetry

The values of the confidence intervals obtained for UA in each AP are shown in Table 3.
The confidence intervals implemented using the Bootstrap t-method and, when available, using specific R functions, were very close to each other in all the APs, resulting identical in some cases.
Wide confidence intervals were obtained in AP2 and AP3 for Krippendorff’s α and for the indices belonging to the Kappa statistic. For the same agreement indices, the confidence intervals in AP1 were narrow, but characterised by negative values. Gwet’s γ(AC1), BP coefficient, Quatto’s S and Andrès and Hernàndez’s multi-raters Δ conferred tight confidence intervals in all the studied cases.

3.2. Trichotomous Animal-Based Welfare Indicators

3.2.1. Agreement Indices for Body Condition Score

The values of the agreement indices obtained for BCS are reported in Table 4.
As already observed for dichotomous indicators, also for trichotomous ones, the obtained concordance rates were not the same for all the implemented agreement indices (P01; P02; and P03). In this regard, P01 was the concordance rate obtained for Krippendorff’s α, almost all the agreement indices belonging to the Kappa statistic (except for Hubert’s K), BP coefficient, Quatto’s S and Gwet’s γ(AC1). P02 was the concordance rate obtained for Hubert’s K and Andrès and Hernàndez’s multi-raters Δ, while P03 was obtained for all the weighted indices (i.e., Krippendorff’s α*, Gwet’s γ(AC2), Fleiss’s K*, Conger’s K*, BP* coefficient, Quatto’s S* and Hubert’s K*).
For BCS the agreement results obtained for Krippendorff’s α, the indices belonging to the Kappa statistic and the related weighted forms were very low in all the APs if compared to their respective concordance rates. In AP3, despite high concordance rates (P01 = 80%; P02 = 70%; and P03 = 90%), such values were close to zero (0.02 to 0.07). On the other hand, BP coefficient, Quatto’s S and Gwet’s γ(AC1) showed agreement results in line with the related concordance rate (P01). Moreover, as already pointed out for dichotomous variables, Gwet’s γ(AC1) values were closer to the concordance rate if compared to those obtained for BP coefficient and Quatto’s S. The agreement values obtained for Andrès and Hernàndez’s multi-raters Δ were close to the related concordance rate in AP1 and AP2 (P02 = 77% and 70%; Δ = 0.75 and 0.65, respectively), while in AP3 the result for this index was quite low if compared to P02 (0.47 and 70%, respectively). Gwet’s γ(AC2), BP* coefficient and Quatto’s S* showed agreement values in line with the observed concordance rate (P03) in all the considered cases. The results given by Gwet’s γ(AC2) were closer to P03 if compared to those obtained for BP* coefficient and Quatto’s S*.

3.2.2. Confidence Intervals for Body Condition Score

As already observed for dichotomous indicators (Table 3), also for the trichotomous ones, the results of the confidence intervals obtained exploiting the Bootstrap t-method were very close to those obtained implementing the R functions (Table 5).
In particular, it is known that the R functions “concordance” and “wlin.conc”, used for the implementation of Quatto’s S and Quatto’s S*, respectively, were developed starting from the Bootstrap method.
Krippendorff’s α, the indices belonging to the Kappa statistic and their relative weighted forms conferred wide confidence intervals in AP1 and AP2, while in AP3 the confidence intervals obtained for the above-mentioned indices were tighter. On the other hand, BP coefficient, Quatto’s S and Gwet’s γ(AC1) were characterised by narrow confidence intervals in all the APs. Such trends resemble what was already observed for dichotomous indicators (Table 3).
The confidence intervals for Andrès and Hernàndez’s multi-raters Δ were tight in AP1 and AP2, while they were wide in AP3. Gwet’s γ(AC2), BP* coefficient and Quatto’s S* showed narrow confidence intervals in all the considered APs.

4. Discussion

The UA is treated as a categorical variable, identifying the absence or the presence of the welfare problem only. On the other hand, the BCS can be considered both as a categorical and ordinal variable, as it is possible to identify a pre-ordered scale, based on the accumulation of body fat reserves [6].
Most of the agreement indices implemented in the current study are suitable to assess the IOR in the presence of both two and multiple raters (i.e., Krippendorff’s α; BP coefficient; Quatto’s S; Krippendorff’s α*; Gwet’s γ(AC1); Gwet’s γ(AC2); BP* coefficient; and Quatto’s S*).
Based on the type of concordance matrix (i.e., agreement matrix) developed for the calculation of the agreement indices considered in the current study, three different P0 are obtained (Table 2 and Table 4; Supplementary Materials). Fleiss [12] proposed a concordant matrix where the sum of the squared probabilities through which each subject is attributed to a specific category allows to obtain the P0. The same approach is used when implementing Krippendorff’s α, Light’s K, Conger’s K, BP coefficient, Quatto’s S and Gwet’s γ(AC1). Andrès and Hernàndez [28] proposed an alternative type of concordance matrix, valid for the implementation of both Andrès and Hernàndez’s multi-raters Δ and Hubert’s K, based on the contingencies table implemented by Dillon and Mulani [32]. Finally, in the case of weighted agreement indices, which are specifically developed to assess the reliability of variables evaluated through an ordered scale, Krippendorff [29], Gwet [11], Marasini et al. [26] and Andrès and Hernàndez [30] proposed a concordance matrix for the calculation of Krippendorff’s α*, Gwet’s γ(AC2), Fleiss’ K*, Conger’s K*, BP* coefficient, Quatto’s S* and Hubert’s K*, where the level of disagreement among the raters is also considered.
In this regard, the unweighted indices only allow to quantify and to verify the presence or absence of the agreement among the raters when they classify a subject within a predetermined category. On the contrary, the weighted forms of these indices also measure the degree of the disagreement present among the raters during this classification [33]. For example, concerning the ordinal variable analysed in the current study (BCS), a higher disagreement among the raters when they classify the subjects within the categories −1 or 1 (very thin goat; very fat goat) could be considered more serious if compared to disagreement observed for the categories 0 and 1 (normal goat; very fat goat), as the difference present between these two latter categories is lower if compared to the previous ones, and the possibility of error during this last classification is higher and more acceptable. Consequently, the disagreement among the raters for the categories −1 and 1 “weighs more” during the evaluation if compared to the disagreement obtained for the categories 0 and 1. The seriousness of this kind of disagreement is evaluated implementing a matrix where specific weights, proposed in the current literature [11], are exploited.
Our results show that the agreement indices belonging to the Kappa statistic can be affected by the paradox behaviour for both dichotomous (Table 2) and trichotomous (Table 4) indicators, conferring, in some cases, very low agreement values despite high P0 [17]. The paradox has been widely analysed in the published literature [17,34,35], especially when discussing the evaluation of agreement among two raters using Cohen’s K [21]. However, it is known that the paradox behaviour also affects the Kappa indices in the presence of multiple raters. In this regard, Falotico and Quatto [36] stated that the statistic proposed by Fleiss [12] should be invariant to the permutation, that is, the different combinations of assignments when couples of raters identify a subject within a specific category. However, this statement is not verified for Fleiss’ K, which gets worse when the marginal distributions within the concordance matrix are not fixed. This means that the raters are unaware regarding the exact number of times each subject is attributed to each of the categories characterising the variable. Thus, each subject involved in the evaluation process can be attributed with no limits to a specific category, producing constant assignments, but resulting in higher variations and unbalanced marginal distributions [37]. Consequently, this factor conduces to the obtainment of lower agreement values if compared to P0 (Table 2 and Table 4). Furthermore, Fleiss’ K is an extension of Scott’s π [38,39], which is sometimes affected by the paradox behaviour when assessing the IOR in the presence of two raters [6,9]. Conger [14] tried to improve the Kappa statistic proposed by Fleiss [12], but our results clearly show that Conger’s K can also sometimes confer low agreement values if compared to P0, therefore being susceptible to paradox behaviours (Table 2 and Table 4). Similarly, Light’s K [15] and Hubert’s K [16], being generalisations of Cohen’s K in the presence of multiple raters [39], are also affected by the above-mentioned problem (Table 2 and Table 4). Krippendorff’s α suffers from the paradox behaviour too, when the evaluation is performed by two [6,9] or multiple raters (current study). Furthermore, the weighted forms of the above-mentioned agreement indices are affected by the same problem, even if they were able to confer slightly higher agreement results if compared to the respective unweighted forms (Table 4). The same trend was observed for Cohen’s weighted K [40] when assessing the IOR of trichotomous indicators in the presence of two raters [6].
For all these indices, the occurrence of paradox behaviour is also highlighted when observing the obtained confidence intervals. The complexity of the manual implementation of closed formulas of variance estimates increases as the number of the raters and categories increase. While it can be quite easily applied in the presence of two raters [6,9], it becomes challenging already in the presence of three raters. For this reason, in the current study, the closed formulas of variance estimates were not implemented manually to calculate the confidence intervals. The Bootstrap t-method proposed by Efron [31] was already previously recommended to be used for confidence intervals calculation [6,9], as it is easier than the manual implementation of closed formulas. Moreover, through a resampling technique, bootstrapping also allows obtaining a more accurate implementation of the confidence intervals if compared to closed formulas, which are based on approximate calculation of the variance estimates, resulting worse in terms of accuracy and flexibility [41].
For the agreement indices affected by the paradox behaviour, the confidence intervals result is wide, showing a high dispersion of the values obtained within the sample (Table 3 and Table 5). Tight confidence intervals are sometimes found for agreement indices affected by the above-mentioned problem; in particular, this occurs when the indices, and their confidence intervals, result in negative values (Table 3 and Table 5). Such results confirm what was previously observed when assessing the IOR for dichotomous and trichotomous indicators in the presence of two raters [6,9].
Despite the occurrence of the paradox behaviour, the Kappa statistic, and especially Fleiss’ K, has been frequently exploited in the published literature to assess the IOR of animal-based welfare indicators in the presence of multiple raters. For example, Fleiss’ K was developed to assess the IOR of indicators used to evaluate the presence of respiratory diseases in pre-weaned dairy calves [42], the rumen fill and tail length in ewes kept outdoors [43], the presence of severe lameness in horses [44], the kneel fracture in laying hens [45], and ten different indicators of lamb welfare (i.e., demeanour, response to stimulation, shivering, standing ability, posture, abdominal fill, body condition, lameness, eye condition and salivation) [46]. Examples of the paradox behaviour are also highlighted in the literature when the IOR of animal-based welfare indicators was assessed for several species and in the presence of many raters (P0 = 86%; Fleiss’ K = 0.13 [42]; P0 = 73%; Fleiss’ K = 0.14 [43]; P0 = 61.9%; Fleiss’ K = 0.23 [44]; P0 = 94%; Fleiss’ K = 0.13 [47]; P0 = 81%; and Fleiss’ K = 0.43 [48]). Similarly, Torsiello et al. [6] recently highlighted examples of the paradox behaviour in the published literature when evaluating the IOR of trichotomous and four-level animal-based welfare indicators in the presence of two raters.
Alternative agreement indices have been proposed more recently to overcome the paradox that may affect older indices. Andrès and Marzo’s Δ [49] was developed to solve this problem and confers good agreement results when assessing the IOR for dichotomous indicators in the presence of two raters [9]. The same trend is visible in the current study for UA when implementing Andrès and Hernàndez’s multi-raters Δ [28] even if, in AP1, this index overestimates the P0 (Table 2). This phenomenon can be explained considering that multi-raters Δ is not developed starting from the implementation of the real agreement present among the raters but starting from the calculation of the estimated agreement when the raters classify the subjects involved within each predetermined category [28]. Consequently, based on estimations, the agreement values obtained for multi-raters Δ cannot be necessarily correct and could result in values which can exceed the P0. Another explanation of the overestimated value obtained for multi-raters Δ relies on the unbalance of the agreement present among the raters towards a specific category. Specifically, in the current study, the agreement was very high for the category 0 (Δ = 0.88) but resulted in a negative value for the category 1 (Δ = −0.02), producing an overall agreement which exceeded the P0 (Δ = 0.86) (Table 2). As reported by Andrès and Hernàndez [28], this means that the raters should try to improve the criteria used when they classify a subject within a specific category, in order to homogenise the degree of agreement for each of the considered categories (consistency), consequently obtaining more reliable agreement values.
The confidence intervals obtained for Δ are tighter in the presence of dichotomous indicators, resulting in a lower dispersion of the values within the sample, if compared to the indices belonging to the Kappa statistic (Table 3). Concerning the BCS, Andrès and Hernàndez’s multi-raters Δ [28] confers better agreement results in the presence of three raters (Table 4) if compared to the ones obtained with Andres and Marzo’s Δ in the presence of two raters [6]. Despite this, in AP3, the agreement value obtained for this index is worse when comparing it to P0 (Table 4), also resulting in wider confidence intervals (Table 5).
The agreement indices which implement the Pe considering the total number of categories characterising the variable are paradox-free and conferred in all the APs the best agreement results for both UA and BCS (Table 2 and Table 4). This occurs because these indices are not influenced by the unbalanced values assumed by the marginal distributions within the concordance matrix when the assignments of the subjects to a specific category are higher, as instead observed for the Kappa statistic [17]. For example, Quatto [25] calculated the Pe as the inverse of the total number of categories, and this principle is valid in the presence of both two and multiple raters. Furthermore, the S-statistic follows the same statistical approach of BP coefficient [24], which is a generalisation of Holley and Guilford’s G [50] and Bennett, Alpert and Goldstein’s S [51] in the presence of two raters for variables characterised by three categories or more [11]. Following the same statistical approach, Quatto’s S and BP coefficient showed identical agreement values in all considered cases (Table 2 and Table 4).
In the presence of multiple raters, Gwet [11] defined the Pe as the probability that pairs of raters, casually selected from a group of n raters, agree on assigning a subject into a predetermined category. Thus, after calculating the Pe for each pair of raters, the total chance agreement is given averaging the values of all the above-mentioned Pe.
Gwet’s γ(AC1), BP coefficient, and Quatto’s S gave the best results for both dichotomous and trichotomous categorical indicators. Moreover, their weighted forms (Gwet’s γ(AC2), BP* coefficient and Quatto’s S* (BP* coefficient and Quatto’s S* follow the same statistical approach)) conferred the best agreement values for trichotomous ordinal indicators. When using weighted agreement indices, a crucial point is always represented by the choice of the weights used for the implementation of the index itself. For the current study, as described in Torsiello et al. [6], the linear weights proposed by Cicchetti and Allison [52] were exploited (the same weights are also used for the development of the weighted forms of the Kappa indices, and for Krippendorff’s α*) as they are less sensitive to the total number of categories of the analysed variable [26]. Furthermore, Gwet’s γ(AC1), BP coefficient, and Quatto’s S, as well as their relative weighted forms, also conferred the narrowest confidence intervals (Table 3 and Table 5), pointing out a low dispersion of the values obtained within the sample. This trend has also been observed when assessing the IOR for both dichotomous and trichotomous indicators in the presence of two raters [6,9].

5. Conclusions

From the results obtained in this study, it is evident that Krippendorff’s α, the Kappa indices and their weighted forms can suffer from the paradox behaviour when evaluating the IOR of dichotomous and trichotomous animal-based welfare indicators in the presence of multiple raters. For this reason, Gwet’s γ(AC1), BP coefficient, Quatto’s S, and their weighted forms (Gwet’s γ(AC2), BP* coefficient, Quatto’s S*), being paradox-free, should be preferred when assessing the IOR of dichotomous and trichotomous animal-based welfare indicators in the presence of three raters. Gwet’s γ(AC1), BP coefficient and Quatto’s S consider the total number of categories characterising the variable when implementing the Pe; for this reason, they can be considered suitable to assess the IOR of categorical variables, characterised by any number of categories, and in the presence of both two and many raters. The same approach is followed by their weighted forms, which are recommended to assess the IOR in the presence of ordinal variables and in the presence of any number of raters.
The confidence intervals also give relevant information regarding the accuracy and the goodness of the implemented agreement indices. In this regard, the best agreement indices are those characterised by tighter confidence intervals. The Bootstrap t-method and R functions (when available) can be considered valid methods to calculate the agreement values and their relative confidence intervals for each considered agreement index, avoiding the use of cumbersome closed formulas, and demonstrating usefulness in the presence of both two [6,9] and many raters.
Finally, it has to be highlighted that UA and BCS, used in the present study, are only examples of dichotomous and trichotomous variables, and that the considerations drawn from our results can be generalised and applied to other dichotomous and trichotomous animal-based welfare indicators, particularly to those for which information on reliability is not available yet (which is a common issue, especially for pasture-based systems; [53]), also for different farm species. As stated above, reliability is one of the most important features for animal-based welfare indicators: the availability of appropriate tools to evaluate IOR is therefore of fundamental importance for selecting the most appropriate indicators, especially when different raters are called to assess welfare, above all for certification purposes, to ensure a fair assessment.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/ani16040546/s1, Supplementary Materials: Agreement indices.

Author Contributions

Conceptualisation, M.G., P.Q., M.B., S.M., L.B. and M.R.; Methodology, B.T., M.G. and P.Q.; Software, B.T., M.G. and P.Q.; Validation, B.T., M.G. and P.Q.; Formal Analysis, B.T. and M.G.; Investigation, B.T., M.B., S.M., L.B. and M.R.; Data Curation, B.T., M.G. and P.Q.; Writing—Original Draft Preparation, B.T.; Writing—Review and Editing, M.G., P.Q., M.B., S.M., L.B. and M.R.; Visualisation, B.T., M.G., P.Q., M.B., S.M., L.B. and M.R.; Supervision, M.G., P.Q. and M.R.; Project Administration, M.R.; Funding Acquisition, L.B. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Project “Increasing the resilience of livestock production” (grant number BIAD_RILO_23_01—2023), DISAFA, Funder: University of Turin (Italy).

Institutional Review Board Statement

The UA and BCS data used in this study were obtained performing a trial that was approved by the Bioethics Committee of the University of Turin (Italy) (protocol n. 0587791).

Informed Consent Statement

Informed consent was obtained from the owners of the animals involved in this study.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors on request.

Acknowledgments

We thank the farmers who allowed us to visit their farms. We also acknowledge the help of Gabriele My, Mauro Masino and Andrea Bonzanino for data collection.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
IORInter-observer reliability
UAUdder asymmetry
BCSBody condition score
APAlpine pasture

References

  1. Blokhuis, H.; Jones, B.; Veissier, I.; Miele, M. Introduction. In Improving Farm Animal Welfare; Blokhuis, H., Miele, M., Veissier, I., Jones, B., Eds.; Wageningen Academic Publishers: Wageningen, The Netherlands, 2013; pp. 1–13. [Google Scholar] [CrossRef][Green Version]
  2. EFSA Panel on Animal Health and Welfare (AHAW). Statement on the use of animal-based measures to assess the welfare of animals. EFSA J. 2012, 10, 2767. [Google Scholar] [CrossRef]
  3. Vieira, A.; Battini, M.; Can, E.; Mattiello, S.; Stilwell, G. Inter-observer reliability of animal-based welfare indicators included in the animal welfare indicators welfare assessment protocol for dairy goats. Animal 2018, 12, 1942–1949. [Google Scholar] [CrossRef] [PubMed]
  4. Martin, P.; Bateson, P. Measuring Behaviour: An Introductory Guide, 3rd ed.; Cambridge University Press: Cambridge, UK, 2007. [Google Scholar]
  5. Popping, R. Interrater agreement. In Introduction to Interrater Agreement for Nominal Data; Springer: Cham, Switzerland, 2019; pp. 21–78. [Google Scholar] [CrossRef]
  6. Torsiello, B.; Giammarino, M.; Quatto, P.; Battini, M.; Mattiello, S.; Battaglini, L.; Renna, M. Evaluation of inter-observer reliability in the case of trichotomous and four-level animal-based welfare indicators with two observers. Ital. J. Anim. Sci. 2024, 23, 938–960. [Google Scholar] [CrossRef]
  7. Taylor, J.; Watkinson, D. Indexing reliability for condition survey data. Conservator 2007, 30, 49–62. [Google Scholar] [CrossRef]
  8. Bajpai, S.; Bajpai, R.C.; Chaturvedi, H.K. Evaluation of inter-rater agreement and inter-rater reliability for observational data: An overview of concepts and methods. J. Indian Acad. Appl. Psychol. 2015, 41, 20–27. [Google Scholar]
  9. Giammarino, M.; Mattiello, S.; Battini, M.; Quatto, P.; Battaglini, L.M.; Vieria, A.C.L.; Stilwell, G.; Renna, M. Evaluation of inter-observer reliability of animal welfare indicators: Which is the best index to use? Animals 2021, 11, 1445. [Google Scholar] [CrossRef]
  10. Gwet, K.L. Handbook of Inter-Rater Reliability—How to Estimate the Level of Agreement Between Two or Multiple Raters; STATAXIS Publishing Company: Gaithersburg, MD, USA, 2001. [Google Scholar]
  11. Gwet, K.L. Handbook of Inter-Rater Reliability—The Definitive Guide to Measuring the Extent of Agreement Among Raters; Advanced Analytics, LLC: Gaithersburg, MD, USA, 2014. [Google Scholar]
  12. Fleiss, J.L. Measuring nominal scale agreement among many raters. Psychol. Bull. 1971, 76, 378–382. [Google Scholar] [CrossRef]
  13. Krippendorff, K. Estimating the reliability, systematic error and random error of interval data. Educ. Psychol. Meas. 1970, 30, 61–70. [Google Scholar] [CrossRef]
  14. Conger, A.J. Integration and generalization of kappas for multiple raters. Psychol. Bull. 1980, 88, 322–328. [Google Scholar] [CrossRef]
  15. Light, R.J. Measures of response agreement for qualitative data: Some generalizations and alternatives. Psychol. Bull. 1971, 76, 365–377. [Google Scholar] [CrossRef]
  16. Hubert, L. Kappa revisited. Psychol. Bull. 1977, 84, 289–297. [Google Scholar] [CrossRef]
  17. Feinstein, A.R.; Cicchetti, D.V. High agreement but low Kappa: I. the problems of two paradoxes. J. Clin. Epidemiol. 1990, 43, 543–549. [Google Scholar] [CrossRef] [PubMed]
  18. Battini, M.; Renna, M.; Giammarino, M.; Battaglini, L.; Mattiello, S. Feasibility and reliability of the AWIN welfare assessment protocol for dairy goats in semi-extensive farming conditions. Front. Vet. Sci. 2021, 8, 731927. [Google Scholar] [CrossRef] [PubMed]
  19. Mattiello, S.; Battini, M.; Vieira, A.; Stilwell, G. AWIN welfare assessment protocol for goats. AWIN 2015, 1–70. [Google Scholar] [CrossRef]
  20. Ajuda, I.; Vieira, A.; de Almeida, F.; Stilwell, G. Conformation of the Udder, Is That a Problem in Our Dairy Farms? Preliminary Results. In Proceedings of the Regional IGA Conference Goat Milk Quality, Tromsø, Norway, 4–6 June 2013; p. 1. Available online: https://www.iga-goatworld.com/2013-iga-regional-conference-norway.html (accessed on 3 February 2026).
  21. Cohen, J. A coefficient of agreement for nominal scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef]
  22. McHugh, M.L. Interrater reliability: The kappa statistic. Biochem. Med. 2012, 22, 276–282. [Google Scholar] [CrossRef]
  23. Andrès, A.M.; Hernàndez, A.M. Hubert’s multi-rater kappa revisited. Br. J. Math. Stat. Psychol. 2020, 73, 1–22. [Google Scholar] [CrossRef]
  24. Brennan, R.L.; Prediger, D.J. Coefficient Kappa: Some uses, misuses, and alternatives. Educ. Psycol. Meas. 1981, 41, 687–699. [Google Scholar] [CrossRef]
  25. Quatto, P. Un test di concordanza tra più esaminatori. [Testing agreement among multiple raters]. Statistica 2004, 1, 145–151. (In Italian) [Google Scholar] [CrossRef]
  26. Marasini, D.; Quatto, P.; Ripamonti, E. Assessing the inter-rater agreement for ordinal data through weighted indexes. Stat. Methods. Med. Res. 2016, 25, 2611–2633. [Google Scholar] [CrossRef]
  27. Gwet, K.L. Computing inter-rater reliability and its variance in the presence of high agreement. Br. J. Math. Stat. Psychol. 2008, 61, 29–48. [Google Scholar] [CrossRef] [PubMed]
  28. Andrès, A.M.; Hernàndez, M.A. Multi-rater delta: Extending the delta nominal measure of agreement between two raters to many raters. J. Stat. Comput. Simul. 2021, 92, 1877–1897. [Google Scholar] [CrossRef]
  29. Krippendorff, K. Reliability in content analysis: Some common misconceptions and recommendations. Hum. Commun. Res. 2004, 30, 411–433. [Google Scholar] [CrossRef]
  30. Andrès, A.M.; Hernàndez, M.A. Estimators of various kappa coefficients based on the unbiased estimator of the expected index of agreements. Adv. Data Anal. Classif. 2024, 19, 177–207. [Google Scholar] [CrossRef]
  31. Efron, B. Bootstrap methods: Another look at the jackknife. Ann. Stat. 1979, 7, 1–26. [Google Scholar] [CrossRef]
  32. Dillon, W.R.; Mulani, N. A probabilistic latent class model for assessing inter-judge reliability. Multivar. Behav. Res. 1984, 19, 438–458. [Google Scholar] [CrossRef]
  33. Warrens, M.J.; De Raadt, A.; Bosker, R.J.; Kiers, H.A.L. Weighted kappa for interobserver agreement and missing data. Mach. Learn. Knowl. Extr. 2025, 7, 18. [Google Scholar] [CrossRef]
  34. Lantz, C.A.; Nebenzahl, E. Behavior and interpretation of the k statistic: Resolution of the two paradoxes. J. Clin. Epidemiol. 1996, 49, 431–434. [Google Scholar] [CrossRef]
  35. Shoukri, M.M. Measures of Interobserver Agreement and Reliability, 1st ed.; CRC Press: Boca Raton, FL, USA, 2003. [Google Scholar] [CrossRef]
  36. Falotico, R.; Quatto, P. Fleiss’ kappa statistic without paradoxes. Qual. Quant. 2015, 49, 463–470. [Google Scholar] [CrossRef]
  37. Randolph, J.J. Free-marginal multirater kappa (multirater kfree): An alternative to Fleiss’ fixed-marginal multirater kappa. In Proceedings of the Oensuu University Learning and Instruction Symposium, Joensuu, Finland, 14–15 October 2005. [Google Scholar]
  38. Scott, W.A. Reliability of content analysis: The case of nominal scale coding. Public Opin. Q. 1955, 19, 321–325. [Google Scholar] [CrossRef]
  39. Warrens, M.J. Inequalities between multi-rater kappas. Adv. Data Anal. Classif. 2010, 4, 271–286. [Google Scholar] [CrossRef]
  40. Cohen, J. Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychol. Bull. 1968, 70, 213–220. [Google Scholar] [CrossRef] [PubMed]
  41. DiCiccio, T.J.; Efron, B. Bootstrap confidence intervals. Stat. Sci. 1996, 11, 189–228. [Google Scholar] [CrossRef]
  42. Buczinski, S.; Faure, C.; Jolivet, S.; Abdallah, A. Evaluation of inter-observer agreement when using a clinical respiratory scoring system in pre-weaned dairy calves. N. Z. Vet. J. 2016, 64, 243–247. [Google Scholar] [CrossRef]
  43. Munoz, C.; Campbell, A.; Hemsworth, P.; Doyle, R. Animal- based measures to assess the welfare of extensively managed ewes. Animals 2018, 8, 8. [Google Scholar] [CrossRef]
  44. Keegan, K.G.; Dent, E.V.; Wilason, D.A.; Janicek, J.; Kramer, J.; Lacarrubba, A.; Walsh, D.M.; Cassells, M.W.; Esther, T.M.; Schiltz, P.; et al. Repeatability of subjective evaluation of lameness in horses. Equine Vet. J. 2010, 42, 92–97. [Google Scholar] [CrossRef]
  45. Petrik, M.T.; Guerin, M.T.; Widowski, T.M. Keel fracture assessment of laying hens by palpation: Inter-observer reliability and accuracy. Vet. Rec. 2013, 173, 500. [Google Scholar] [CrossRef]
  46. Phythian, C.J.; Toft, N.; Cripps, P.J.; Michalopoulou, E.; Winter, A.C.; Jones, P.H.; Grove-White, D.; Duncan, J.S. Inter-observer agreement, diagnostic sensitivity and specificity of animal-based indicators of young lamb welfare. Animal 2013, 7, 1182–1190. [Google Scholar] [CrossRef]
  47. Contreras-Jodar, A.; Michel, V.; Vinco, L.J.; Vurvarò-Porter, A.; Velarde, A. Relevant indicators of consciousness after head-only electrical stunning in rabbits, stunning efficiency, and risk factors in commercial conditions. Animals 2025, 15, 587. [Google Scholar] [CrossRef]
  48. Croyle, S.L.; Nash, C.G.R.; Bauman, C.; LeBlanc, S.J.; Haley, D.B.; Khosa, D.K.; Kelton, D.F. Training method for animal-based measures in dairy cattle welfare assessments. J. Dairy Sci. 2018, 101, 9463–9471. [Google Scholar] [CrossRef]
  49. Andrès, A.M.; Marzo, P.F. Delta: A new measure of agreement between two raters. Br. J. Math. Stat. Psychol. 2004, 57, 1–19. [Google Scholar] [CrossRef]
  50. Holley, G.W.; Guilford, J.P. A note on the G-index of agreement. Educ. Psychol. Meas. 1964, 24, 749–753. [Google Scholar] [CrossRef]
  51. Bennet, E.M.; Alpert, R.; Goldstein, A.C. Communications through limited response questioning. Public Opin. Q. 1954, 18, 303–308. [Google Scholar] [CrossRef]
  52. Cicchetti, A.; Allison, T. A new procedure for assessing reliability of scoring EEG sleep recordings. Am. J. EEG Technol. 1971, 11, 101–109. [Google Scholar] [CrossRef]
  53. Spigarelli, C.; Zuliani, A.; Battini, M.; Mattiello, S.; Bovolenta, S. Welfare assessment on pasture: A review on animal-based measures for ruminants. Animals 2020, 10, 609. [Google Scholar] [CrossRef]
Table 1. Agreement indices implemented for each animal-based welfare indicator.
Table 1. Agreement indices implemented for each animal-based welfare indicator.
VARIABLESAGREEMENT
INDICES
REFERENCES
FOR THE AGREEMENT INDEX
CONFIDENCE INTERVALSREFERENCES FOR THE CONFIDENCE INTERVALSR PACKAGES (VERSION R × 64 4.2.2)R FUNCTIONS
UA
and
BCS
Krippendorff’s αKrippendorff [13]BootstrapKrippendorff [13]library (irr)
library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
kripp.alpha
krippen.alpha.raw
Fleiss’ KFleiss [12]BootstrapFleiss [12]library (irr)
library (raters)
library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
kappam.fleiss
fleiss.kappa.raw
concordance
Light’s KLight [15]BootstrapLight [15]library(irr)
library(boot)
boot_result
var (boot)
boot.ci
kappam.light
Hubert’s KHubert [16]BootstrapHubert [16]
Andrès and
Hernàndez [23]
library (boot)boot_result
var (boot)
boot.ci
Conger’s KConger [14]BootstrapConger [14]library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
conger.kappa.raw
BP coefficientBrennan and Prediger [24]BootstrapGwet [11]library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
bp.coeff.raw
Quatto’s SQuatto [25]BootstrapQuatto [25]
Marasini et al. [26]
library (raters)
library (boot)
boot_result
var (boot)
boot.ci
concordance
Gwet’s γ(AC1)Gwet [27]BootstrapGwet [27]library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
gwet.ac1.raw
Andrès and Hernàndez’s multi-raters ΔAndrès and Hernàndez [28]BootstrapAndrès and Hernàndez
[28]
library (DeltaMAN)
library (boot)
boot_result
var (boot)
boot.ci
multiDelta
BCSKrippendorff’s weighted αKrippendorff [29]BootstrapGwet [11]library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
krippen.alpha.raw
Gwet’s γ(AC2)Gwet [11]BootstrapGwet [11]library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
gwet.ac1.raw
Fleiss’ weighted KGwet [11]BootstrapGwet [11]library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
fleiss.kappa.raw
Conger’s weighted KGwet [11]BootstrapGwet [11]library(irrCAC)
library(boot)
boot_result
var (boot)
boot.ci
conger.kappa.raw
Weighted BP coefficientGwet [11]BootstrapGwet [11]library (irrCAC)
library (boot)
boot_result
var (boot)
boot.ci
bp.coeff.raw
Quatto’s weighted SMarasini et al. [26]BootstrapMarasini et al. [26]library (raters)
library (boot)
boot_result
var (boot)
boot.ci
wlin.conc
Hubert’s weighted KAndrès and Hernàndez [30]BootstrapAndrès and Hernàndez [30]library (boot)boot_result
var (boot)
boot.ci
Table 2. Values of the concordance rate and of the agreement indices obtained for udder asymmetry for the three alpine pastures.
Table 2. Values of the concordance rate and of the agreement indices obtained for udder asymmetry for the three alpine pastures.
Concordance RateAgreement IndicesAgreement Values
AP1
(n = 44)
P01 = 86%
P02 = 80%
α−0.07
Fleiss’ K−0.07
Light’s K−0.03
Hubert’s K−0.04
Conger’s K−0.04
BP0.73
S0.73
γ(AC1)0.84
Δ0.86
AP2
(n = 70)
P01 = 92%
P02 = 89%
α0.52
Fleiss’ K0.51
Light’s K0.51
Hubert’s K0.51
Conger’s K0.51
BP0.85
S0.85
γ(AC1)0.91
Δ0.79
AP3
(n = 46)
P01 = 94%
P02 = 91%
α0.68
Fleiss’ K0.68
Light’s K0.69
Hubert’s K0.68
Conger’s K0.68
BP0.88
S0.88
γ(AC1)0.93
Δ0.86
Abbreviations: n = sample size; P01 = concordance rate for Krippendorff’s α, Fleiss’ K, Light’s K, Conger’s K, Brennan and Prediger’s coefficient, Quatto’s S and Gwet’s γ(AC1); P02 = concordance rate for Hubert’s K and Andrès and Hernàndez’s multi-raters Δ; α = Krippendorff’s α; BP = Brennan and Prediger’s coefficient; S = Quatto’s S; γ(AC1) = Gwet’s γ(AC1); Δ = Andrés and Hernàndez’s multi-raters Δ; AP1 = alpine pasture 1; AP2 = alpine pasture 2; AP3 = alpine pasture 3.
Table 3. Values of the confidence intervals for the agreement indices obtained for udder asymmetry implemented using the Bootstrap t-method and R functions in the three alpine pastures.
Table 3. Values of the confidence intervals for the agreement indices obtained for udder asymmetry implemented using the Bootstrap t-method and R functions in the three alpine pastures.
Confidence Intervals
Agreement
Indices
By Bootstrap
t-Method
By R Functions
AP1
(n = 44)
α−0.11; −0.02krippen.alpha.raw: −0.11; −0.02
Fleiss’ K−0.12; −0.03fleiss.kappa.raw: −0.12; 0.03
Light’s K−0.05; 0.00N.A.
Hubert’s K−0.08; 0.00N.A.
Conger’s K−0.08; 0.00conger.kappa.raw: −0.08; 0.01
BP0.57; 0.89bp.coeff.raw: 0.56; 0.89
S0.57; 0.89concordance: 0.58; 0.88
γ(AC1)0.74; 0.95gwet.ac1.raw: 0.74; 0.95
Δ0.80; 0.94N.A.
AP2
(n = 70)
α0.18; 0.92krippen.alpha.raw: 0.20; 0.84
Fleiss’ K0.20; 0.88fleiss.kappa.raw: 0.19; 0.84
Light’s K0.19; 0.89N.A.
Hubert’s K0.20; 0.89N.A.
Conger’s K0.19; 0.90conger.kappa.raw: 0.19; 0.84
BP0.75; 0.94bp.coeff.raw: 0.75; 0.95
S0.75; 0.95concordance: 0.73; 0.94
γ(AC1)0.85; 0.98gwet.ac1.raw: 0.84; 0.98
Δ0.67; 0.92N.A.
AP3
(n = 46)
α0.37; 1.05krippen.alpha.raw: 0.38; 0.99
Fleiss’ K0.36; 1.10fleiss.kappa.raw: 0.37; 0.99
Light’s K0.35; 1.08N.A.
Hubert’s K0.34; 1.09N.A.
Conger’s K0.36; 1.07conger.kappa.raw: 0.38; 0.99
BP0.77; 1.00bp.coeff.raw: 0.77; 1.00
S0.78; 0.99concordance: 0.77; 0.97
γ(AC1)0.86; 1.00gwet.ac1.raw: 0.86; 1.00
Δ0.75; 0.96N.A.
Abbreviations: n = sample size; α = Krippendorff’s α; BP = Brennan and Prediger’s coefficient; S = Quatto’s S; γ(AC1) = Gwet’s γ(AC1); Δ = Andrés and Hernàndez’s multi-raters Δ; AP1 = alpine pasture 1; AP2 = alpine pasture 2; AP3 = alpine pasture 3; N.A. = not available (i.e., no R function available to calculate confidence intervals).
Table 4. Values of the concordance rate and of the agreement indices obtained for body condition score for the three alpine pastures.
Table 4. Values of the concordance rate and of the agreement indices obtained for body condition score for the three alpine pastures.
Concordance RateAgreement IndicesAgreement Values
AP1
(n = 44)
P01 = 85%
P02 = 77%
P03 = 92%
α0.35
Fleiss’ K0.35
Light’s K0.36
Hubert’s K0.33
Conger’s K0.35
BP0.77
S0.77
γ(AC1)0.83
Δ0.75
α*0.37
γ(AC2)0.91
Fleiss’ K*0.37
Conger’s K*0.37
BP*0.83
S*0.83
Hubert’s K*0.37
AP2
(n = 70)
P01 = 80%
P02 = 70%
P03 = 90%
α0.23
Fleiss’ K0.22
Light’s K0.24
Hubert’s K0.21
Conger’s K0.23
BP0.70
S0.70
γ(AC1)0.77
Δ0.65
α*0.24
γ(AC2)0.87
Fleiss’ K*0.24
Conger’s K*0.24
BP*0.78
S*0.78
Hubert’s K*0.24
AP3
(n = 46)
P01 = 80%
P02 = 70%
P03 = 90%
α0.04
Fleiss’ K0.04
Light’s K0.04
Hubert’s K0.02
Conger’s K0.04
BP0.70
S0.70
γ(AC1)0.77
Δ0.47
α*0.07
γ(AC2)0.88
Fleiss’ K*0.06
Conger’s K*0.07
BP*0.77
S*0.77
Hubert’s K*0.07
Abbreviations: n = sample size; P01 = concordance rate for Krippendorff’s α, Fleiss’ K, Light’s K, Conger’s K, Brennan and Prediger’s coefficient; Quatto’s S and Gwet’s γ(AC1); P02 = concordance rate for Hubert’s K and Andrès and Hernàndez’s multi-raters Δ; P03 = concordance rate for Krippendorff’s weighted α, Gwet’s γ(AC2), Fleiss’ weighted K, Conger’s weighted K, Brennan and Prediger’s weighted coefficient, Quatto’s weighted S and Hubert’s weighted K; α = Krippendorff’s α; BP = Brennan and Prediger’s coefficient; S = Quatto’s S; γ(AC1) = Gwet’s γ(AC1); Δ = Andrés and Hernàndez’s multi-raters Δ; α* = Krippendorff’s weighted α; γ(AC2) = Gwet’s γ(AC2); Fleiss’ K* = Fleiss’ weighted K; Conger’s K* = Conger’s weighted K; BP* = Brennan and Prediger’s weighted coefficient; S* = Quatto’s weighted S; Hubert’s K* = Hubert’s weighted K; AP1 = alpine pasture 1; AP2 = alpine pasture 2; AP3 = alpine pasture 3.
Table 5. Values of the confidence intervals for the agreement indices obtained for body condition score implemented using the Bootstrap t-method and R functions in three alpine pastures.
Table 5. Values of the confidence intervals for the agreement indices obtained for body condition score implemented using the Bootstrap t-method and R functions in three alpine pastures.
Confidence Intervals
Agreement
Indices
By Bootstrap
t-Method
By R Functions
AP1
(n = 44)
α0.11; 0.64krippen.alpha.raw: 0.09; 0.62
Fleiss’ K0.11; 0.62fleiss.kappa.raw: 0.08; 0.61
Light’s K0.13; 0.63N.A.
Hubert’s K0.07; 0.61N.A.
Conger’s K0.11; 0.62conger.kappa.raw: 0.09; 0.61
BP coefficient0.65; 0.90bp.coeff.raw: 0.64; 0.90
S0.65; 0.90concordance: 0.64; 0.89
γ(AC1)0.73; 0.93gwet.ac1.raw: 0.72; 0.94
Δ0.58; 0.89N.A.
α*0.13; 0.65krippen.alpha.raw: 0.11; 0.64
γ(AC2)0.85; 0.97gwet.ac1.raw: 0.84; 0.97
Fleiss’ K*0.12; 0.64fleiss.kappa.raw: 0.10; 0.63
Conger’s K*0.13; 0.65conger.kappa.raw: 0.11; 0.63
BP*0.74; 0.93bp.coeff.raw: 0.73; 0.93
S*0.73; 0.92wlin.conc: 0.73; 0.91
Hubert’s K*0.14; 0.65N.A.
AP2
(n = 70)
α0.06; 0.40krippen.alpha.raw: 0.06; 0.40
Fleiss’ K0.06; 0.40fleiss.kappa.raw: 0.05; 0.40
Light’s K0.07; 0.41N.A.
Hubert’s K0.05; 0.39N.A.
Conger’s K0.07; 0.40conger.kappa.raw: 0.06; 0.40
BP coefficient0.59; 0.80bp.coeff.raw: 0.59; 0.81
S0.59; 0.81concordance: 0.59; 0.80
γ(AC1)0.69; 0.86gwet.ac1.raw: 0.68; 0.86
Δ0.44; 0.88N.A.
α*0.09; 0.41krippen.alpha.raw: 0.08; 0.41
γ(AC2)0.82; 0.93gwet.ac1.raw: 0.82; 0.93
Fleiss’ K*0.08; 0.42fleiss.kappa.raw: 0.07; 0.41
Conger’s K*0.08; 0.41conger.kappa.raw: 0.09; 0.41
BP*0.69; 0.87bp.coeff.raw: 0.69; 0.86
S*0.69; 0.85wlin.conc: 0.70; 0.85
Hubert’s K*0.08; 0.41N.A.
AP3
(n = 46)
α−0.09; 0.20krippen.alpha.raw: −0.11; 0.20
Fleiss’ K−0.11; 0.20fleiss.kappa.raw: −0.12; 0.19
Light’s K−0.09; 0.17N.A.
Hubert’s K−0.12; 0.16N.A.
Conger’s K−0.10; 0.18conger.kappa.raw: −0.11; 0.19
BP coefficient0.56; 0.83bp.coeff.raw: 0.56; 0.83
S0.57; 0.82concordance: 0.57; 0.83
γ(AC1)0.66; 0.88gwet.ac1.raw: 0.66; 0.89
Δ0.09; 0.86N.A.
α*−0.07; 0.23krippen.alpha.raw: −0.09; 0.23
γ(AC2)0.81; 0.95gwet.ac1.raw: 0.81; 0.94
Fleiss’ K*−0.08; 0.22fleiss.kappa.raw: −0.09; 0.22
Conger’s K*−0.07; 0.22conger.kappa.raw: −0.08; 0.22
BP*0.67; 0.88bp.coeff.raw: 0.67; 0.88
S*0.68; 0.87wlin.conc: 0.66; 0.87
Hubert’s K*−0.07; 0.21N.A.
Abbreviations: n = sample size; α = Krippendorff’s α; S = Quatto’s S; γ(AC1) = Gwet’s γ(AC1); Δ = Andrés and Hernàndez’s multi-raters Δ; α* = Krippendorff’s weighted α; γ(AC2) = Gwet’s γ(AC2); Fleiss’ K* = Fleiss’ weighted K; Conger’s K* = Conger’s weighted K; BP* = Brennan and Prediger’s weighted BP coefficient; S* = Quatto’s weighted S; Hubert’s K* = Hubert’s weighted K; AP1 = alpine pasture 1; AP2 = alpine pasture 2; AP3 = alpine pasture 3; N.A. = not available (i.e., no R function available to calculate confidence intervals).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Torsiello, B.; Giammarino, M.; Quatto, P.; Battini, M.; Mattiello, S.; Battaglini, L.; Renna, M. Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters. Animals 2026, 16, 546. https://doi.org/10.3390/ani16040546

AMA Style

Torsiello B, Giammarino M, Quatto P, Battini M, Mattiello S, Battaglini L, Renna M. Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters. Animals. 2026; 16(4):546. https://doi.org/10.3390/ani16040546

Chicago/Turabian Style

Torsiello, Benedetta, Mauro Giammarino, Piero Quatto, Monica Battini, Silvana Mattiello, Luca Battaglini, and Manuela Renna. 2026. "Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters" Animals 16, no. 4: 546. https://doi.org/10.3390/ani16040546

APA Style

Torsiello, B., Giammarino, M., Quatto, P., Battini, M., Mattiello, S., Battaglini, L., & Renna, M. (2026). Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters. Animals, 16(4), 546. https://doi.org/10.3390/ani16040546

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop