Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters
Simple Summary
Abstract
1. Introduction
2. Materials and Methods
2.1. Dichotomous and Trichotomous Animal-Based Welfare Indicators
2.2. Agreement Indices and Confidence Intervals
2.3. Statistical Analyses
3. Results
3.1. Dichotomous Animal-Based Welfare Indicators
3.1.1. Agreement Indices for Udder Asymmetry
3.1.2. Confidence Intervals for Udder Asymmetry
3.2. Trichotomous Animal-Based Welfare Indicators
3.2.1. Agreement Indices for Body Condition Score
3.2.2. Confidence Intervals for Body Condition Score
4. Discussion
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| IOR | Inter-observer reliability |
| UA | Udder asymmetry |
| BCS | Body condition score |
| AP | Alpine pasture |
References
- Blokhuis, H.; Jones, B.; Veissier, I.; Miele, M. Introduction. In Improving Farm Animal Welfare; Blokhuis, H., Miele, M., Veissier, I., Jones, B., Eds.; Wageningen Academic Publishers: Wageningen, The Netherlands, 2013; pp. 1–13. [Google Scholar] [CrossRef][Green Version]
- EFSA Panel on Animal Health and Welfare (AHAW). Statement on the use of animal-based measures to assess the welfare of animals. EFSA J. 2012, 10, 2767. [Google Scholar] [CrossRef]
- Vieira, A.; Battini, M.; Can, E.; Mattiello, S.; Stilwell, G. Inter-observer reliability of animal-based welfare indicators included in the animal welfare indicators welfare assessment protocol for dairy goats. Animal 2018, 12, 1942–1949. [Google Scholar] [CrossRef] [PubMed]
- Martin, P.; Bateson, P. Measuring Behaviour: An Introductory Guide, 3rd ed.; Cambridge University Press: Cambridge, UK, 2007. [Google Scholar]
- Popping, R. Interrater agreement. In Introduction to Interrater Agreement for Nominal Data; Springer: Cham, Switzerland, 2019; pp. 21–78. [Google Scholar] [CrossRef]
- Torsiello, B.; Giammarino, M.; Quatto, P.; Battini, M.; Mattiello, S.; Battaglini, L.; Renna, M. Evaluation of inter-observer reliability in the case of trichotomous and four-level animal-based welfare indicators with two observers. Ital. J. Anim. Sci. 2024, 23, 938–960. [Google Scholar] [CrossRef]
- Taylor, J.; Watkinson, D. Indexing reliability for condition survey data. Conservator 2007, 30, 49–62. [Google Scholar] [CrossRef]
- Bajpai, S.; Bajpai, R.C.; Chaturvedi, H.K. Evaluation of inter-rater agreement and inter-rater reliability for observational data: An overview of concepts and methods. J. Indian Acad. Appl. Psychol. 2015, 41, 20–27. [Google Scholar]
- Giammarino, M.; Mattiello, S.; Battini, M.; Quatto, P.; Battaglini, L.M.; Vieria, A.C.L.; Stilwell, G.; Renna, M. Evaluation of inter-observer reliability of animal welfare indicators: Which is the best index to use? Animals 2021, 11, 1445. [Google Scholar] [CrossRef]
- Gwet, K.L. Handbook of Inter-Rater Reliability—How to Estimate the Level of Agreement Between Two or Multiple Raters; STATAXIS Publishing Company: Gaithersburg, MD, USA, 2001. [Google Scholar]
- Gwet, K.L. Handbook of Inter-Rater Reliability—The Definitive Guide to Measuring the Extent of Agreement Among Raters; Advanced Analytics, LLC: Gaithersburg, MD, USA, 2014. [Google Scholar]
- Fleiss, J.L. Measuring nominal scale agreement among many raters. Psychol. Bull. 1971, 76, 378–382. [Google Scholar] [CrossRef]
- Krippendorff, K. Estimating the reliability, systematic error and random error of interval data. Educ. Psychol. Meas. 1970, 30, 61–70. [Google Scholar] [CrossRef]
- Conger, A.J. Integration and generalization of kappas for multiple raters. Psychol. Bull. 1980, 88, 322–328. [Google Scholar] [CrossRef]
- Light, R.J. Measures of response agreement for qualitative data: Some generalizations and alternatives. Psychol. Bull. 1971, 76, 365–377. [Google Scholar] [CrossRef]
- Hubert, L. Kappa revisited. Psychol. Bull. 1977, 84, 289–297. [Google Scholar] [CrossRef]
- Feinstein, A.R.; Cicchetti, D.V. High agreement but low Kappa: I. the problems of two paradoxes. J. Clin. Epidemiol. 1990, 43, 543–549. [Google Scholar] [CrossRef] [PubMed]
- Battini, M.; Renna, M.; Giammarino, M.; Battaglini, L.; Mattiello, S. Feasibility and reliability of the AWIN welfare assessment protocol for dairy goats in semi-extensive farming conditions. Front. Vet. Sci. 2021, 8, 731927. [Google Scholar] [CrossRef] [PubMed]
- Mattiello, S.; Battini, M.; Vieira, A.; Stilwell, G. AWIN welfare assessment protocol for goats. AWIN 2015, 1–70. [Google Scholar] [CrossRef]
- Ajuda, I.; Vieira, A.; de Almeida, F.; Stilwell, G. Conformation of the Udder, Is That a Problem in Our Dairy Farms? Preliminary Results. In Proceedings of the Regional IGA Conference Goat Milk Quality, Tromsø, Norway, 4–6 June 2013; p. 1. Available online: https://www.iga-goatworld.com/2013-iga-regional-conference-norway.html (accessed on 3 February 2026).
- Cohen, J. A coefficient of agreement for nominal scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef]
- McHugh, M.L. Interrater reliability: The kappa statistic. Biochem. Med. 2012, 22, 276–282. [Google Scholar] [CrossRef]
- Andrès, A.M.; Hernàndez, A.M. Hubert’s multi-rater kappa revisited. Br. J. Math. Stat. Psychol. 2020, 73, 1–22. [Google Scholar] [CrossRef]
- Brennan, R.L.; Prediger, D.J. Coefficient Kappa: Some uses, misuses, and alternatives. Educ. Psycol. Meas. 1981, 41, 687–699. [Google Scholar] [CrossRef]
- Quatto, P. Un test di concordanza tra più esaminatori. [Testing agreement among multiple raters]. Statistica 2004, 1, 145–151. (In Italian) [Google Scholar] [CrossRef]
- Marasini, D.; Quatto, P.; Ripamonti, E. Assessing the inter-rater agreement for ordinal data through weighted indexes. Stat. Methods. Med. Res. 2016, 25, 2611–2633. [Google Scholar] [CrossRef]
- Gwet, K.L. Computing inter-rater reliability and its variance in the presence of high agreement. Br. J. Math. Stat. Psychol. 2008, 61, 29–48. [Google Scholar] [CrossRef] [PubMed]
- Andrès, A.M.; Hernàndez, M.A. Multi-rater delta: Extending the delta nominal measure of agreement between two raters to many raters. J. Stat. Comput. Simul. 2021, 92, 1877–1897. [Google Scholar] [CrossRef]
- Krippendorff, K. Reliability in content analysis: Some common misconceptions and recommendations. Hum. Commun. Res. 2004, 30, 411–433. [Google Scholar] [CrossRef]
- Andrès, A.M.; Hernàndez, M.A. Estimators of various kappa coefficients based on the unbiased estimator of the expected index of agreements. Adv. Data Anal. Classif. 2024, 19, 177–207. [Google Scholar] [CrossRef]
- Efron, B. Bootstrap methods: Another look at the jackknife. Ann. Stat. 1979, 7, 1–26. [Google Scholar] [CrossRef]
- Dillon, W.R.; Mulani, N. A probabilistic latent class model for assessing inter-judge reliability. Multivar. Behav. Res. 1984, 19, 438–458. [Google Scholar] [CrossRef]
- Warrens, M.J.; De Raadt, A.; Bosker, R.J.; Kiers, H.A.L. Weighted kappa for interobserver agreement and missing data. Mach. Learn. Knowl. Extr. 2025, 7, 18. [Google Scholar] [CrossRef]
- Lantz, C.A.; Nebenzahl, E. Behavior and interpretation of the k statistic: Resolution of the two paradoxes. J. Clin. Epidemiol. 1996, 49, 431–434. [Google Scholar] [CrossRef]
- Shoukri, M.M. Measures of Interobserver Agreement and Reliability, 1st ed.; CRC Press: Boca Raton, FL, USA, 2003. [Google Scholar] [CrossRef]
- Falotico, R.; Quatto, P. Fleiss’ kappa statistic without paradoxes. Qual. Quant. 2015, 49, 463–470. [Google Scholar] [CrossRef]
- Randolph, J.J. Free-marginal multirater kappa (multirater kfree): An alternative to Fleiss’ fixed-marginal multirater kappa. In Proceedings of the Oensuu University Learning and Instruction Symposium, Joensuu, Finland, 14–15 October 2005. [Google Scholar]
- Scott, W.A. Reliability of content analysis: The case of nominal scale coding. Public Opin. Q. 1955, 19, 321–325. [Google Scholar] [CrossRef]
- Warrens, M.J. Inequalities between multi-rater kappas. Adv. Data Anal. Classif. 2010, 4, 271–286. [Google Scholar] [CrossRef]
- Cohen, J. Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychol. Bull. 1968, 70, 213–220. [Google Scholar] [CrossRef] [PubMed]
- DiCiccio, T.J.; Efron, B. Bootstrap confidence intervals. Stat. Sci. 1996, 11, 189–228. [Google Scholar] [CrossRef]
- Buczinski, S.; Faure, C.; Jolivet, S.; Abdallah, A. Evaluation of inter-observer agreement when using a clinical respiratory scoring system in pre-weaned dairy calves. N. Z. Vet. J. 2016, 64, 243–247. [Google Scholar] [CrossRef]
- Munoz, C.; Campbell, A.; Hemsworth, P.; Doyle, R. Animal- based measures to assess the welfare of extensively managed ewes. Animals 2018, 8, 8. [Google Scholar] [CrossRef]
- Keegan, K.G.; Dent, E.V.; Wilason, D.A.; Janicek, J.; Kramer, J.; Lacarrubba, A.; Walsh, D.M.; Cassells, M.W.; Esther, T.M.; Schiltz, P.; et al. Repeatability of subjective evaluation of lameness in horses. Equine Vet. J. 2010, 42, 92–97. [Google Scholar] [CrossRef]
- Petrik, M.T.; Guerin, M.T.; Widowski, T.M. Keel fracture assessment of laying hens by palpation: Inter-observer reliability and accuracy. Vet. Rec. 2013, 173, 500. [Google Scholar] [CrossRef]
- Phythian, C.J.; Toft, N.; Cripps, P.J.; Michalopoulou, E.; Winter, A.C.; Jones, P.H.; Grove-White, D.; Duncan, J.S. Inter-observer agreement, diagnostic sensitivity and specificity of animal-based indicators of young lamb welfare. Animal 2013, 7, 1182–1190. [Google Scholar] [CrossRef]
- Contreras-Jodar, A.; Michel, V.; Vinco, L.J.; Vurvarò-Porter, A.; Velarde, A. Relevant indicators of consciousness after head-only electrical stunning in rabbits, stunning efficiency, and risk factors in commercial conditions. Animals 2025, 15, 587. [Google Scholar] [CrossRef]
- Croyle, S.L.; Nash, C.G.R.; Bauman, C.; LeBlanc, S.J.; Haley, D.B.; Khosa, D.K.; Kelton, D.F. Training method for animal-based measures in dairy cattle welfare assessments. J. Dairy Sci. 2018, 101, 9463–9471. [Google Scholar] [CrossRef]
- Andrès, A.M.; Marzo, P.F. Delta: A new measure of agreement between two raters. Br. J. Math. Stat. Psychol. 2004, 57, 1–19. [Google Scholar] [CrossRef]
- Holley, G.W.; Guilford, J.P. A note on the G-index of agreement. Educ. Psychol. Meas. 1964, 24, 749–753. [Google Scholar] [CrossRef]
- Bennet, E.M.; Alpert, R.; Goldstein, A.C. Communications through limited response questioning. Public Opin. Q. 1954, 18, 303–308. [Google Scholar] [CrossRef]
- Cicchetti, A.; Allison, T. A new procedure for assessing reliability of scoring EEG sleep recordings. Am. J. EEG Technol. 1971, 11, 101–109. [Google Scholar] [CrossRef]
- Spigarelli, C.; Zuliani, A.; Battini, M.; Mattiello, S.; Bovolenta, S. Welfare assessment on pasture: A review on animal-based measures for ruminants. Animals 2020, 10, 609. [Google Scholar] [CrossRef]
| VARIABLES | AGREEMENT INDICES | REFERENCES FOR THE AGREEMENT INDEX | CONFIDENCE INTERVALS | REFERENCES FOR THE CONFIDENCE INTERVALS | R PACKAGES (VERSION R × 64 4.2.2) | R FUNCTIONS |
|---|---|---|---|---|---|---|
| UA and BCS | Krippendorff’s α | Krippendorff [13] | Bootstrap | Krippendorff [13] | library (irr) library (irrCAC) library (boot) | boot_result var (boot) boot.ci kripp.alpha krippen.alpha.raw |
| Fleiss’ K | Fleiss [12] | Bootstrap | Fleiss [12] | library (irr) library (raters) library (irrCAC) library (boot) | boot_result var (boot) boot.ci kappam.fleiss fleiss.kappa.raw concordance | |
| Light’s K | Light [15] | Bootstrap | Light [15] | library(irr) library(boot) | boot_result var (boot) boot.ci kappam.light | |
| Hubert’s K | Hubert [16] | Bootstrap | Hubert [16] Andrès and Hernàndez [23] | library (boot) | boot_result var (boot) boot.ci | |
| Conger’s K | Conger [14] | Bootstrap | Conger [14] | library (irrCAC) library (boot) | boot_result var (boot) boot.ci conger.kappa.raw | |
| BP coefficient | Brennan and Prediger [24] | Bootstrap | Gwet [11] | library (irrCAC) library (boot) | boot_result var (boot) boot.ci bp.coeff.raw | |
| Quatto’s S | Quatto [25] | Bootstrap | Quatto [25] Marasini et al. [26] | library (raters) library (boot) | boot_result var (boot) boot.ci concordance | |
| Gwet’s γ(AC1) | Gwet [27] | Bootstrap | Gwet [27] | library (irrCAC) library (boot) | boot_result var (boot) boot.ci gwet.ac1.raw | |
| Andrès and Hernàndez’s multi-raters Δ | Andrès and Hernàndez [28] | Bootstrap | Andrès and Hernàndez [28] | library (DeltaMAN) library (boot) | boot_result var (boot) boot.ci multiDelta | |
| BCS | Krippendorff’s weighted α | Krippendorff [29] | Bootstrap | Gwet [11] | library (irrCAC) library (boot) | boot_result var (boot) boot.ci krippen.alpha.raw |
| Gwet’s γ(AC2) | Gwet [11] | Bootstrap | Gwet [11] | library (irrCAC) library (boot) | boot_result var (boot) boot.ci gwet.ac1.raw | |
| Fleiss’ weighted K | Gwet [11] | Bootstrap | Gwet [11] | library (irrCAC) library (boot) | boot_result var (boot) boot.ci fleiss.kappa.raw | |
| Conger’s weighted K | Gwet [11] | Bootstrap | Gwet [11] | library(irrCAC) library(boot) | boot_result var (boot) boot.ci conger.kappa.raw | |
| Weighted BP coefficient | Gwet [11] | Bootstrap | Gwet [11] | library (irrCAC) library (boot) | boot_result var (boot) boot.ci bp.coeff.raw | |
| Quatto’s weighted S | Marasini et al. [26] | Bootstrap | Marasini et al. [26] | library (raters) library (boot) | boot_result var (boot) boot.ci wlin.conc | |
| Hubert’s weighted K | Andrès and Hernàndez [30] | Bootstrap | Andrès and Hernàndez [30] | library (boot) | boot_result var (boot) boot.ci |
| Concordance Rate | Agreement Indices | Agreement Values | |
|---|---|---|---|
| AP1 (n = 44) | P01 = 86% P02 = 80% | α | −0.07 |
| Fleiss’ K | −0.07 | ||
| Light’s K | −0.03 | ||
| Hubert’s K | −0.04 | ||
| Conger’s K | −0.04 | ||
| BP | 0.73 | ||
| S | 0.73 | ||
| γ(AC1) | 0.84 | ||
| Δ | 0.86 | ||
| AP2 (n = 70) | P01 = 92% P02 = 89% | α | 0.52 |
| Fleiss’ K | 0.51 | ||
| Light’s K | 0.51 | ||
| Hubert’s K | 0.51 | ||
| Conger’s K | 0.51 | ||
| BP | 0.85 | ||
| S | 0.85 | ||
| γ(AC1) | 0.91 | ||
| Δ | 0.79 | ||
| AP3 (n = 46) | P01 = 94% P02 = 91% | α | 0.68 |
| Fleiss’ K | 0.68 | ||
| Light’s K | 0.69 | ||
| Hubert’s K | 0.68 | ||
| Conger’s K | 0.68 | ||
| BP | 0.88 | ||
| S | 0.88 | ||
| γ(AC1) | 0.93 | ||
| Δ | 0.86 |
| Confidence Intervals | |||
|---|---|---|---|
| Agreement Indices | By Bootstrap t-Method | By R Functions | |
| AP1 (n = 44) | α | −0.11; −0.02 | krippen.alpha.raw: −0.11; −0.02 |
| Fleiss’ K | −0.12; −0.03 | fleiss.kappa.raw: −0.12; 0.03 | |
| Light’s K | −0.05; 0.00 | N.A. | |
| Hubert’s K | −0.08; 0.00 | N.A. | |
| Conger’s K | −0.08; 0.00 | conger.kappa.raw: −0.08; 0.01 | |
| BP | 0.57; 0.89 | bp.coeff.raw: 0.56; 0.89 | |
| S | 0.57; 0.89 | concordance: 0.58; 0.88 | |
| γ(AC1) | 0.74; 0.95 | gwet.ac1.raw: 0.74; 0.95 | |
| Δ | 0.80; 0.94 | N.A. | |
| AP2 (n = 70) | α | 0.18; 0.92 | krippen.alpha.raw: 0.20; 0.84 |
| Fleiss’ K | 0.20; 0.88 | fleiss.kappa.raw: 0.19; 0.84 | |
| Light’s K | 0.19; 0.89 | N.A. | |
| Hubert’s K | 0.20; 0.89 | N.A. | |
| Conger’s K | 0.19; 0.90 | conger.kappa.raw: 0.19; 0.84 | |
| BP | 0.75; 0.94 | bp.coeff.raw: 0.75; 0.95 | |
| S | 0.75; 0.95 | concordance: 0.73; 0.94 | |
| γ(AC1) | 0.85; 0.98 | gwet.ac1.raw: 0.84; 0.98 | |
| Δ | 0.67; 0.92 | N.A. | |
| AP3 (n = 46) | α | 0.37; 1.05 | krippen.alpha.raw: 0.38; 0.99 |
| Fleiss’ K | 0.36; 1.10 | fleiss.kappa.raw: 0.37; 0.99 | |
| Light’s K | 0.35; 1.08 | N.A. | |
| Hubert’s K | 0.34; 1.09 | N.A. | |
| Conger’s K | 0.36; 1.07 | conger.kappa.raw: 0.38; 0.99 | |
| BP | 0.77; 1.00 | bp.coeff.raw: 0.77; 1.00 | |
| S | 0.78; 0.99 | concordance: 0.77; 0.97 | |
| γ(AC1) | 0.86; 1.00 | gwet.ac1.raw: 0.86; 1.00 | |
| Δ | 0.75; 0.96 | N.A. | |
| Concordance Rate | Agreement Indices | Agreement Values | |
|---|---|---|---|
| AP1 (n = 44) | P01 = 85% P02 = 77% P03 = 92% | α | 0.35 |
| Fleiss’ K | 0.35 | ||
| Light’s K | 0.36 | ||
| Hubert’s K | 0.33 | ||
| Conger’s K | 0.35 | ||
| BP | 0.77 | ||
| S | 0.77 | ||
| γ(AC1) | 0.83 | ||
| Δ | 0.75 | ||
| α* | 0.37 | ||
| γ(AC2) | 0.91 | ||
| Fleiss’ K* | 0.37 | ||
| Conger’s K* | 0.37 | ||
| BP* | 0.83 | ||
| S* | 0.83 | ||
| Hubert’s K* | 0.37 | ||
| AP2 (n = 70) | P01 = 80% P02 = 70% P03 = 90% | α | 0.23 |
| Fleiss’ K | 0.22 | ||
| Light’s K | 0.24 | ||
| Hubert’s K | 0.21 | ||
| Conger’s K | 0.23 | ||
| BP | 0.70 | ||
| S | 0.70 | ||
| γ(AC1) | 0.77 | ||
| Δ | 0.65 | ||
| α* | 0.24 | ||
| γ(AC2) | 0.87 | ||
| Fleiss’ K* | 0.24 | ||
| Conger’s K* | 0.24 | ||
| BP* | 0.78 | ||
| S* | 0.78 | ||
| Hubert’s K* | 0.24 | ||
| AP3 (n = 46) | P01 = 80% P02 = 70% P03 = 90% | α | 0.04 |
| Fleiss’ K | 0.04 | ||
| Light’s K | 0.04 | ||
| Hubert’s K | 0.02 | ||
| Conger’s K | 0.04 | ||
| BP | 0.70 | ||
| S | 0.70 | ||
| γ(AC1) | 0.77 | ||
| Δ | 0.47 | ||
| α* | 0.07 | ||
| γ(AC2) | 0.88 | ||
| Fleiss’ K* | 0.06 | ||
| Conger’s K* | 0.07 | ||
| BP* | 0.77 | ||
| S* | 0.77 | ||
| Hubert’s K* | 0.07 |
| Confidence Intervals | |||
|---|---|---|---|
| Agreement Indices | By Bootstrap t-Method | By R Functions | |
| AP1 (n = 44) | α | 0.11; 0.64 | krippen.alpha.raw: 0.09; 0.62 |
| Fleiss’ K | 0.11; 0.62 | fleiss.kappa.raw: 0.08; 0.61 | |
| Light’s K | 0.13; 0.63 | N.A. | |
| Hubert’s K | 0.07; 0.61 | N.A. | |
| Conger’s K | 0.11; 0.62 | conger.kappa.raw: 0.09; 0.61 | |
| BP coefficient | 0.65; 0.90 | bp.coeff.raw: 0.64; 0.90 | |
| S | 0.65; 0.90 | concordance: 0.64; 0.89 | |
| γ(AC1) | 0.73; 0.93 | gwet.ac1.raw: 0.72; 0.94 | |
| Δ | 0.58; 0.89 | N.A. | |
| α* | 0.13; 0.65 | krippen.alpha.raw: 0.11; 0.64 | |
| γ(AC2) | 0.85; 0.97 | gwet.ac1.raw: 0.84; 0.97 | |
| Fleiss’ K* | 0.12; 0.64 | fleiss.kappa.raw: 0.10; 0.63 | |
| Conger’s K* | 0.13; 0.65 | conger.kappa.raw: 0.11; 0.63 | |
| BP* | 0.74; 0.93 | bp.coeff.raw: 0.73; 0.93 | |
| S* | 0.73; 0.92 | wlin.conc: 0.73; 0.91 | |
| Hubert’s K* | 0.14; 0.65 | N.A. | |
| AP2 (n = 70) | α | 0.06; 0.40 | krippen.alpha.raw: 0.06; 0.40 |
| Fleiss’ K | 0.06; 0.40 | fleiss.kappa.raw: 0.05; 0.40 | |
| Light’s K | 0.07; 0.41 | N.A. | |
| Hubert’s K | 0.05; 0.39 | N.A. | |
| Conger’s K | 0.07; 0.40 | conger.kappa.raw: 0.06; 0.40 | |
| BP coefficient | 0.59; 0.80 | bp.coeff.raw: 0.59; 0.81 | |
| S | 0.59; 0.81 | concordance: 0.59; 0.80 | |
| γ(AC1) | 0.69; 0.86 | gwet.ac1.raw: 0.68; 0.86 | |
| Δ | 0.44; 0.88 | N.A. | |
| α* | 0.09; 0.41 | krippen.alpha.raw: 0.08; 0.41 | |
| γ(AC2) | 0.82; 0.93 | gwet.ac1.raw: 0.82; 0.93 | |
| Fleiss’ K* | 0.08; 0.42 | fleiss.kappa.raw: 0.07; 0.41 | |
| Conger’s K* | 0.08; 0.41 | conger.kappa.raw: 0.09; 0.41 | |
| BP* | 0.69; 0.87 | bp.coeff.raw: 0.69; 0.86 | |
| S* | 0.69; 0.85 | wlin.conc: 0.70; 0.85 | |
| Hubert’s K* | 0.08; 0.41 | N.A. | |
| AP3 (n = 46) | α | −0.09; 0.20 | krippen.alpha.raw: −0.11; 0.20 |
| Fleiss’ K | −0.11; 0.20 | fleiss.kappa.raw: −0.12; 0.19 | |
| Light’s K | −0.09; 0.17 | N.A. | |
| Hubert’s K | −0.12; 0.16 | N.A. | |
| Conger’s K | −0.10; 0.18 | conger.kappa.raw: −0.11; 0.19 | |
| BP coefficient | 0.56; 0.83 | bp.coeff.raw: 0.56; 0.83 | |
| S | 0.57; 0.82 | concordance: 0.57; 0.83 | |
| γ(AC1) | 0.66; 0.88 | gwet.ac1.raw: 0.66; 0.89 | |
| Δ | 0.09; 0.86 | N.A. | |
| α* | −0.07; 0.23 | krippen.alpha.raw: −0.09; 0.23 | |
| γ(AC2) | 0.81; 0.95 | gwet.ac1.raw: 0.81; 0.94 | |
| Fleiss’ K* | −0.08; 0.22 | fleiss.kappa.raw: −0.09; 0.22 | |
| Conger’s K* | −0.07; 0.22 | conger.kappa.raw: −0.08; 0.22 | |
| BP* | 0.67; 0.88 | bp.coeff.raw: 0.67; 0.88 | |
| S* | 0.68; 0.87 | wlin.conc: 0.66; 0.87 | |
| Hubert’s K* | −0.07; 0.21 | N.A. | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Torsiello, B.; Giammarino, M.; Quatto, P.; Battini, M.; Mattiello, S.; Battaglini, L.; Renna, M. Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters. Animals 2026, 16, 546. https://doi.org/10.3390/ani16040546
Torsiello B, Giammarino M, Quatto P, Battini M, Mattiello S, Battaglini L, Renna M. Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters. Animals. 2026; 16(4):546. https://doi.org/10.3390/ani16040546
Chicago/Turabian StyleTorsiello, Benedetta, Mauro Giammarino, Piero Quatto, Monica Battini, Silvana Mattiello, Luca Battaglini, and Manuela Renna. 2026. "Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters" Animals 16, no. 4: 546. https://doi.org/10.3390/ani16040546
APA StyleTorsiello, B., Giammarino, M., Quatto, P., Battini, M., Mattiello, S., Battaglini, L., & Renna, M. (2026). Comparing Agreement Indices to Assess Inter-Observer Reliability in the Case of Dichotomous and Trichotomous Animal-Based Welfare Indicators with Three Raters. Animals, 16(4), 546. https://doi.org/10.3390/ani16040546

