Next Article in Journal
DCANet: Diffusion-Coded Attention Network for Cross-Domain Semantic Noise Mitigation and Multi-Scale Context Fusion
Next Article in Special Issue
Towards Reliable Evaluation of Underwater Image Enhancement Using Subjective and Objective Analysis
Previous Article in Journal
Technical and Economic Feasibility Analysis of a Traction Substation-Based Microgrid
Previous Article in Special Issue
BAG-CLIP: Bifurcated Attention Graph-Enhanced CLIP for Zero-Shot Industrial Anomaly Detection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Impact of Environmental Control on Subjective Video Quality Assessment in Crowdsourced QoE Experiments

1
AGH University of Krakow, 30-059 Krakow, Poland
2
Ondokuz Mayıs University, 55139 Samsun, Türkiye
3
University of Belgrade, 11000 Belgrade, Serbia
*
Authors to whom correspondence should be addressed.
Electronics 2026, 15(8), 1666; https://doi.org/10.3390/electronics15081666
Submission received: 14 March 2026 / Revised: 10 April 2026 / Accepted: 13 April 2026 / Published: 16 April 2026

Abstract

This research investigates the influence of environmental regulation on subjective evaluations of video quality within the Quality of Experience (QoE) paradigm. This work presents a supplementary experiment conducted in a controlled laboratory setting, building on our previous crowdsourcing studies carried out in uncontrolled, web-based conditions using the Prolific platform. Both tests utilized the identical crowdsourcing platform and complied with the International Telecommunication Union Telecommunication (ITU-T) P.910 Recommendations, ensuring external validity and methodological consistency. Participants assessed a collection of processed video sequences (PVS) comprising 46 distinct video clips utilizing the 5-point Absolute Category Rating (ACR) scale, while their response times were documented in milliseconds as measures of cognitive exertion and decision delay. The comparison analysis employs nonparametric tests (Mann–Whitney U and Kolmogorov–Smirnov) and a hierarchical Linear Mixed-Effects Model (LMM) to examine disparities in reaction time distributions, rating consistency, and the incidence of outliers across both environments. The results indicate that controlled settings produce statistically significantly less response variability and enhanced data reliability, whereas uncontrolled settings encompass greater external diversity and real-world unpredictability. These findings offer significant insights into the balance between experimental control and external validity in crowdsourced video quality assessment, advancing the development of scalable approaches for Quality of Experience research.

1. Introduction

In recent years, the assessment of multimedia Quality of Experience (QoE) has gained significance within dynamic content delivery systems and streaming services. The precise evaluation of perceived video quality continues to be a significant concern, as it directly affects user satisfaction and service optimization tactics. Conventional laboratory investigations have always been the foundation for subjective video quality evaluation, owing to their regulated circumstances that provide uniform measurements and reduce extraneous disturbances. The emergence of crowdsourcing approaches has created opportunities for extensive and economical experimentation, but also prompting concerns over data reliability, user diversity, and environmental variability.
Our previous research [1,2] has extensively explored subjective video quality evaluations using a proprietary crowdsourcing platform designed in accordance with ITU-T P.910 Recommendations [3]. These studies, conducted in uncontrolled environments and using international participant sampling via the Prolific platform, demonstrated that while web-based assessments offer significant scalability and external validity, they also introduce several sources of potential variability. Factors such as participant device specifications, ambient lighting, network bandwidth, and display precision can influence rating behavior and response time dynamics, leading to heterogeneous measurement outcomes.
This study conducts a comparative analysis between two experimental settings: a controlled laboratory environment at AGH University of Krakow, utilizing undergraduate students on standardized laboratory computers, and an uncontrolled remote crowdsourcing environment previously implemented via the Prolific platform. The study seeks to quantify the impacts of environmental management on subjective judgments, rating consistency, and temporal response patterns by adhering to uniform experimental protocols: video stimuli, Absolute Category Rating (ACR) scale, and response time measurement in milliseconds.
The incorporation of response-time analytics provides a new viewpoint on QoE evaluation, connecting cognitive processing effort to subjective decision-making. Initial evidence [4] from prior research indicates that medium-quality video sequences (designated as “Fair”) provoke prolonged response times, signifying heightened cognitive load and perceptual ambiguity. This study expands the investigation by examining whether these response patterns consistently occur in controlled laboratory environments or are intensified in uncontrolled contexts.
This comparative analysis aims to enhance experimental approaches in video quality research. The study elucidates the trade-offs between external validity and data consistency across contexts, offering useful insights for the development of robust, scalable, and cognitively informed QoE evaluation frameworks suitable for both academic research and commercial applications.
The remainder of this paper is organized as follows. Section 2 presents a survey of related literature. Section 3 describes the methodology. Section 4 details the experimental procedures employed in both the controlled and uncontrolled settings. Section 5 presents the results. Section 6 discusses the key findings, including the limitations of crowdsourcing for temporally complex tasks and the threats to validity. Section 7 concludes the paper and outlines directions for future research.

2. Literature Survey

Subjective video quality assessment is generally regarded as the most reliable method from the end-user perspective, with core QoE concepts and technologies thoroughly discussed in [5,6,7]. Crowdsourcing has become a faster, more cost-effective alternative to traditional lab-based QoE evaluations, and has been increasingly adopted as a substitute for laboratory experiments, with several advances reported in task design, including paired-comparison approaches and integrated feedback mechanisms that trigger rating forms after playback or via interface buttons [8,9,10]. However, crowdsourcing still faces challenges such as dependence on stable internet connectivity, participant recruitment, limited guarantees of high-resolution display availability partially mitigated by higher resolution crowd testing methods and persistent concerns over the validity and reliability of results [11,12]. While some studies indicate that crowdsourced assessments can match lab-based outcomes under specific conditions, reliability, conceptual, technical, and motivational issues must be addressed through better test design tailored to crowdsourcing environments, and accuracy can further improve when workers evaluate others’ opinions rather than their own, with aggregation methods critically influencing final result quality [12,13].
Winther et al. showed that experts, due to their domain knowledge and motivation, generally outperform crowdworkers when selecting diverse video footage (e.g., shooting, tackling, running, walking, standing) to validate assumptions [14]. To enhance video quality assessment, an open-source crowdsourcing-based extension of the ITU-T P.910 test was proposed, but its 2021 implementation is limited by slow recruitment of suitable participants [15]. Although crowdsourcing lacks the accuracy of P.910 and remains only an alternative, controlled lab-based subjective tests are still the gold standard, highlighting the importance of rigorous video quality evaluation in engineering. To mitigate crowdsourcing’s weaknesses, Anegekuh et al. developed a screening algorithm using metadata such as workers’ devices and time-on-task, finding that about one-third of participants did not actually watch the video [16]. Seufert et al. proposed an adaptive crowdsourced QoE method that reallocates a fixed rating budget across conditions to boost reliability [17], while Nassar et al. discussed expert-identification techniques, including using rich social media data to exclude unreliable workers before task assignment [18]. Overall, quality control is central to crowdsourcing, yet Daniel et al.’s unified model of quality-related factors indicates that, despite extensive prior work and general agreement on its importance, current practices still fall short of consistently achieving quality levels that fully leverage human intelligence [19].
Crowdsourced mobile reporting faces challenges because many contributors lack the expertise and mobile editing skills needed to produce detailed, consistent bug reports, which undermines clear communication [20]. In response, research on crowdsourced subjective quality evaluation in mobile and video communications has developed and validated systems and methodologies against formal MPEG tests, proposed streaming approaches to cope with uncontrolled end-to-end network conditions and motivated Paired Comparison-based assessments, leveraged existing databases and web apps to exploit large observer pools, surveyed web-based frameworks compatible with major platforms and global workers, and even integrated eye-tracking to connect gaze behavior with users’ evolving understanding and potential gaps between actions and self-reports [21,22,23,24,25].
Crowdsourcing has been widely used for image and audio quality assessment, including HDR images and codecs, aesthetic evaluation, object segmentation, and audio MOS testing [26,27,28,29,30,31,32,33]. For HDR images, ref. [27] split participants into two groups and ensured reliability via repeated image presentations and comparison with a gold standard, while [28] compared aesthetic evaluations of Tone Mapping Operators across a lab study, an online replication, and a Prolific-based crowdsourced test. Implicit input for object segmentation was collected through a game in [29], revealing that crowd workers underperformed relative to computer vision experts, and [30] similarly used gamification for subjective image quality assessment. In audio, crowdMOS [31] offered a low-cost, web-based alternative to traditional lab MOS tests, suitable as a complement or early substitute for objective metrics. Image-related work in [32] found high agreement between lab and crowdsourced evaluations for content recognisability but not for aesthetic appeal, while a large-scale study in [33] showed that learning-based image codecs provide promising compression performance compared to anchor codecs.
Crowdsourcing has become an increasingly popular research method, particularly since the COVID-19 pandemic. For graphics-related QoE, it was proposed as an alternative to lab-based tests using multiple screening strategies [34]. Speech quality assessment, which is even more time-consuming and traditionally conducted in controlled labs following ITU-T P.800 [35], has similarly moved toward micro-task platforms, leading to standardized crowdsourcing methodologies in ITU-T P.808, with studies showing good agreement between crowdsourced and P.800 lab results [36,37]. Further work comparing lab and crowdsourced assessments of overall speech quality using ITU-T P.501 Annex D samples (each presented four times) found that inter-rater agreement increased with task exposure while intra-rater reliability remained stable [38]. Another web-based crowdsourcing study examined how the number of stimuli affects rating validity, noting that although 40 stimuli produced the strongest correlations, it also increased reports of listener fatigue [39].
Kurup et al. propose an aggregation method that evaluates submission quality by considering similarity among answers, worker reliability and expertise, and task difficulty, followed by a cost-minimization phase [40]. Similarly, Kazai et al. analyze how workers’ motivation, interest, domain familiarity, perceived task difficulty, and satisfaction with pay, together with task design, affect the accuracy of relevance labels [41]. In disaster management, Khajwal et al. introduce an uncertainty-aware framework that decomposes overall damage assessment into simple microtask questionnaires to reduce subjectivity and improve accuracy, while Weaver et al. examine the reliability of crowd reports and highlight the importance of enabling report modification and completion marking as key features [42,43]. More broadly, crowdsourcing in information retrieval demands careful task design and ongoing quality control during both design and execution [44], and in innovation and management, Ghezzi et al. systematically review the field to identify key themes such as open innovation, co-creation, information systems management, organizational theory and design, marketing, and strategy [45].
Narimanzadeh evaluates two labeling methods for handling task subjectivity, showing that comparison-based labeling either more effectively minimizes random error when the number of comparisons and votes is the same or scales more efficiently as task volume grows, whereas majority voting tends to plateau [46]. Gardlo et al. demonstrate that carefully designed incentive schemes can markedly enhance crowd performance and introduce real-time verification of participant reliability, which is applicable not only to video but also to images and audio [47]. Jin et al. propose the Subjectivity-and-Difficulty Response (SDR) model, which separates question subjectivity to derive worker-specific truths and captures how question difficulty influences the likelihood that a worker’s response aligns with their perceived subjective truth [48]. Focusing on task assignment, Boutsis and Kalogeraki design a crowdsourcing framework that effectively chooses appropriate worker groups for each task, meeting application requirements with minimal overhead while boosting the total number of tasks completed under fixed constraints [49]. Varshney presents a method that jointly guarantees privacy and reliability via stochastic perturbation of microtasks and associated fusion rules, offering a mathematical formulation with threshold conditions for privacy loss under collusion and examining the tradeoffs among privacy, reliability, and cost [50]. Blanco et al. further show, through a large-scale, multi-system study on ad hoc Web object retrieval, that crowd-based evaluation can be conducted repeatedly over long periods while maintaining reliability [51]. Complementing these findings, Behrend et al. investigate whether crowdsourcing can substitute for traditional university subject pools—often criticized for their homogeneity and limited work experience—by surveying targeted samples from both crowdsourcing platforms and university participant pools to compare data quality and better understand crowd workers; overall, the behavior of the crowdsourcing sample was largely comparable to that of participants from a standard psychology subject pool [52].
Jimenez et al. examined how effectively “trapping questions” and “outlier detection” contribute to reliable results, while noting that they may clutter tasks with stimuli irrelevant to the researcher or remove data points that still reflect genuine worker opinions. They conducted a speech quality assessment via a web-based crowdsourcing platform in accordance with ITU-T Recommendation P.800. Their results show that neither crowdsourcing nor laboratory testing alone improves accuracy; instead, combining both methods is required [53].
Hoßfeld et al. describe QoE as the overall acceptability or degree of delight or annoyance a user experiences with an application or service, framing it as a subjective evolution of the more objective QoS and emphasizing its multidimensional, factor-dependent nature [54]. A comprehensive review of QoE methodologies for video services underscores the scarcity of long-term studies and the need to observe user experiences over time, summarizing existing methods, subjective assessment techniques, and study durations, and identifying research gaps and future directions to better capture QoE variability and improve external validity [55]. To enhance the generalizability of QoE research, a unified theoretical model integrating Influencing Factors (IFs) and Perceptual Dimensions (PDs) from experiment design through data analysis enables additive measurement across comparable experiments [56], while crowdsourcing is noted as a less costly and time-consuming alternative to traditional lab-based QoE assessments. Additionally, Konaszyński et al. examine how memory and variability in video quality affect subjective ratings, introducing “measurement points” as critical moments shaping assessments, revealing the negative effect of quality fluctuations, the “last impression effect” whereby improving quality yields higher ratings, and the role of contextual factors, such as the quality of previously viewed videos [57].
A comprehensive review of video compression and optimization technologies and their impact on video streaming QoE is presented in [58], focusing on QoE measurement methodologies, major codecs (MPEG, Google, Apple), and future challenges, especially for high-resolution streaming under bandwidth and storage constraints. Crowdsourcing-based QoE evaluation tools and methods are explored in [1,2,59,60], including Kaleidoscope, a browser-extension platform enabling side-by-side webpage comparisons, and subjective experiments where human participants assess video summaries. These works highlight cost-effective large-scale testing, the role of demographics (e.g., gender) in rating reliability, and statistical validation of crowdsourced VQA outcomes. Objective no-reference quality metrics for consistent video evaluation are further addressed in [61].
Research on QoE in video streaming highlights the tension between internal and external validity and proposes theoretical models, especially the video QoE model, as tools to clarify assumptions, improve study comparability, and strengthen real-world relevance [62]. Recent work shifts from traditional subjective ratings toward behavior-based assessment and even direct brain activity measurements, with a focus on realistic user interaction, psychometric function-based analysis, and user-centered, real-time optimization of multimedia services [63,64]. Age-related differences in video quality perception appear minimal, as psychometric function fitting and ANOVA showed no significant variation in detection thresholds across three age groups [65]. In addition, the structure of video presentation—such as content variability, sequence order, and prior clip quality—has been shown to influence subjective ratings, with repeated viewings generally reducing perceived quality [66].
Recent work highlights the need for user-subjective criteria and externally valid conditions in QoE assessment. Aguilar et al. introduce a no-reference QoE model for video streaming that uses crowdsourcing and the TRIANGLE testbed to efficiently collect diverse user opinions, accounting for network conditions and human preferences [67]. Wielgus et al. propose an experimental protocol for YouTube QoE that better reflects real user behavior such as searching, pausing, changing resolution, and reading comments, arguing that traditional lab-based video quality tests lack external validity due to restricted interaction and content choice [68]. Similarly, Cieplinska et al. conduct a longitudinal study to explore how video quality is perceived over time and how it relates to behaviors like service abandonment in realistic usage scenarios [69]. Crowdsourcing-based studies using the Absolute Category Rating (ACR) scale show that short HD clips rated as “fair” lead to longer response times and reveal contextual biases, for instance due to monitor resolution, while confirming the value of crowdsourcing for collecting heterogeneous feedback [70]. Complementing these approaches, Wanat et al. propose behavioral experiments that emulate real-world conditions, such as the “Fix Your Netflix” experiment, where participants actively respond to quality degradations; psychometric analysis of these interactions uncovers individual differences in perceived quality [71].
A study on low-latency adaptive bitrate (LL-ABR) algorithms for HTTP Adaptive Streaming (HAS) evaluates whether reducing end-to-end delay compromises video quality, a key QoE factor in modern video applications [72]. Another work introduces an application for longitudinal QoE studies that simulates real-life video use and collects daily subjective ratings, offering customizable scheduling, user feedback, device orientation tracking, and instant results [73]. The COVID-19 pandemic rapidly expanded the use of video consultations between patients and general practitioners, compelling even inexperienced or initially reluctant users to adopt tele-consultations. This widespread uptake creates an opportunity to examine QoE and influencing factors in video consultations [74].
Crowdsourced subjective video and image quality assessments have been shown to be both effective and reliable under appropriate conditions. Naderi et al. compare three methods: Absolute Category Rating (ACR), ACR with Hidden Reference (ACR-HR), and Comparison Category Rating (CCR) and find that although ACR-HR is faster and cheaper, CCR is more sensitive and better at capturing quality improvements, making it preferable across diverse compression settings and content types [75]. Saupe et al. report a strong correlation (above 0.96) between crowdsourced DMOS and lab-based MOS, despite limitations in experimental control and preprocessing, indicating that crowdsourcing can provide trustworthy video quality evaluations [76]. A broader perspective on how crowdsourcing supports vendors, operators, and regulators in assessing QoE for emerging network applications and architectures is presented in [77], which also outlines key use cases and challenges of crowdsourced network and QoE measurements. Similarly, Ak et al. assess the use of crowdsourcing for evaluating tone mapping operators (TMOs) for HDR images, and show that results from the Prolific platform closely align with controlled lab experiments while offering faster and larger-scale data collection [27].
Existing outlier detection methods are largely evaluated on synthetic data, highlighting the need for more reliable comparative studies. To address this, the authors present a practical worst-case analysis based on adversarial attacks and propose two low-complexity, robust detection techniques that outperform many conventional methods by explicitly countering such attacks and improving the stability of quality labels [78]. In related work on digital image quality assessment, Gavrovska et al. propose a no-reference approach that avoids subjective ratings by using data augmentation over various distortions and combining Approximate Entropy (AppE), the Perception Image Quality Evaluator (PIQE), Weibull modeling, and color invariance factors to better align objective metrics with human perception and support more extensive experimental analysis [79,80].

3. Methodology

3.1. Experimental Design

The study employed a comparative experimental design to analyze differences in subjective video quality assessment between two distinct environments: (1) a controlled laboratory setting with undergraduate students from AGH University of Krakow, and (2) an uncontrolled online crowdsourcing environment using the Prolific platform.
Both experiments adhered to the ITU-T Recommendations P.910 standard for subjective multimedia evaluations, ensuring methodological consistency and comparability across sessions. The same video sequences, assessment procedures, and quality scales were used in both environments to isolate the influence of environmental factors on user responses and response time dynamics.
In the controlled environment, participants were 32 bachelor’s degree students aged 19–25, all with basic technical literacy and experience with video streaming services. They performed the assessment on pre-calibrated laboratory desktop computers with identical hardware configurations and Full HD (1920 × 1080) monitors. In the uncontrolled environment, participant recruitment was achieved through the Prolific crowdsourcing platform, ensuring demographic diversity across countries, age groups, and educational backgrounds. Only participants meeting the technical criteria—stable network connection (>40 Mbps) and Full HD display resolution—were admitted to the online test.
Both experiments utilized the CrowdQoE platform [81], a web-based system developed in-house using PHP, JavaScript, and HTML, designed for scalable subjective evaluation experiments. Before participation, the system automatically verified browser compatibility, screen resolution, and internet speed. Users failed verification if their resolution dropped below the Full HD threshold or their download speed was insufficient. All interactions, including video playback and rating submissions, were automatically logged with millisecond precision to capture response time after each video clip.
Database A contains 26 unique full HD short video clips and 20 unique full HD long videos, which are processed to create 60 short test videos and 20 long test videos, all exhibiting time-varying quality. In the experiment, each of the 80 videos is shown in a randomized order. In a similar manner, Database B is built from 26 full HD short clips taken from the same source material as Database A. The clips in Database B are encoded to preserve a constant quality level throughout playback. In total, 40 high-quality videos are generated for testing. During the test, each is shown 3 times, arranged in a pseudo-random sequence, yielding 120 video presentations to the participants. The quality levels are regulated by modifying the bitrate and compression settings. The derived clips represented five quality levels: Excellent, Good, Fair, Poor, and Bad, generated by adjusting parameters such as bitrate, quantization, and compression ratio. All videos shared identical content durations (20–60 s) and visual categories to maintain uniform cognitive demand. To eliminate any effect of auditory cues, all audio tracks are intentionally removed from the videos. The inclusion of both content types was a deliberate methodological decision. Database B (stable quality) serves as a cognitively minimal baseline, allowing participants to deliver immediate perceptual judgments without temporal integration, thereby isolating the effect of the evaluation environment from task cognitive complexity. Database A (time-varying quality), by contrast, requires continuous integration of fluctuating quality levels, a process referred to as temporal pooling. This juxtaposition allows the study to identify the boundary conditions of crowdsourcing validity as a function of cognitive demand, which constitutes the primary methodological finding of this work. The complete datasets, including all video sequences and associated metadata, are publicly available via the CrowdQoE-2025 repository [81] to which interested readers are directed for full dataset documentation and reproducibility purposes.
Participants were instructed to watch each video in full-screen mode (F11) and evaluate its perceived visual quality using a 5-point Absolute Category Rating (ACR) scale: Excellent, Good, Fair, Poor, or Bad. After each video, the evaluation interface appeared automatically, and participants had to select a score before proceeding to the next clip. The response time (in milliseconds) was recorded as the interval between the display of the evaluation page and the submission of the rating.

3.2. Environment Protocol

In the laboratory experiment, the controlled setting was standardized to remove potential confounding factors. All participants: Used the same AGH laboratory computers and network infrastructure, completed the tasks individually under supervision, worked in quiet rooms with regulated lighting conditions, and were instructed not to multitask or run other applications during the sessions. Each session lasted about 50 to 55 min to keep the cognitive load manageable.
In the Prolific-based experiment, participants accessed the platform on their personal computers in natural, uncontrolled environments. They followed on-screen instructions to ensure browser and display compliance, yet differences in ambient light, task focus, and system performance introduced inevitable variability. This phase captured the external reality of web-based subjective QoE testing.

3.3. Data Collection and Analysis

The main data collected comprised ACR rating scores (ordinal, on a 1–5 scale), response times (continuous, measured in milliseconds), and a set of demographic and contextual variables, including gender, age, education, country, and self-reported mood and fatigue.
In the post-processing phase, outlier responses were detected and removed using modified z-scores and Pearson correlation thresholds ( r < 0.75 ), in line with the reliability criteria specified in ITU-T P.910. Subsequent comparative statistical analysis considered the distributions of ratings and response times, applied two-sample Kolmogorov–Smirnov and Mann–Whitney U tests to examine differences between environments, and employed a Linear Mixed-Effects Model (LMM).
All statistical analyses were conducted using Python 3.14.1 (SciPy) and MATLAB R2025a.

3.4. Research Objectives

The methodological strategy aimed to:
  • Quantify differences in rating reliability and response time distributions between controlled and uncontrolled environments;
  • Determine whether increased control reduces variability without compromising external realism;
  • Identify potential cognitive load indicators manifested through response time behavior across both setups.
This approach enables a rigorously controlled yet externally valid comparison, advancing the understanding of environmental effects in crowdsourced Quality-of -Experience assessments.

3.5. Uncontrolled Environment: Prolific Study Deployment

The uncontrolled component of the experiment was conducted through the Prolific online crowdsourcing platform, which enabled the recruitment of geographically diverse participants and allowed data collection under naturalistic conditions. The study was configured as an external web-based experiment hosted on the AGH University of Krakow server and accessed via Prolific, ensuring institutional control over data storage while leveraging Prolific’s recruitment and screening infrastructure.

3.5.1. Study Configuration on Prolific

The study was developed using the Survey Builder feature of Prolific, facilitating integration with an external experimental platform instead of utilizing an internal Prolific questionnaire. The data collecting method was classified as a Survey, accompanied by a clear and detailed Study Label presented to participants before enrollment. The study mandated completion on a desktop computer, while mobile devices and tablets were excluded due to variability in display resolution and probable performance errors that could distort judgments of video quality is shown in Figure 1. No content notice was issued, as the study materials included neither explicit nor sensitive content. Participants were consequently provided with a neutral research description devoid of supplementary ethical or psychological guidance.

3.5.2. External Platform and URL Integration

Participants accessed the experiment through a dedicated university-hosted URL: https://s.agh.edu.pl/kI0GC (accessed on 8 March 2026).
This link guided participants to the internal CrowdQoE platform, where all experimental protocols, video presentations, and data logging were administered. The incorporation of dynamic Prolific characteristics facilitated traceability between Prolific demographic records and experimental submissions while safeguarding personally identifying information. The study was designed without limitations on concurrent access, permitting any number of individuals to join simultaneously. The decision was warranted by the resilience of the university server and the streamlined design of the web-based interface.

3.5.3. Prolific ID Recording

To ensure accurate correspondence between Prolific records and the experimental data, participant identifiers were transmitted via URL parameters. Specifically, PROLIFIC_PID recorded each participant’s distinct Prolific ID, STUDY_ID linked their responses to the relevant experiment, and SESSION_ID designated each unique participation instance. The identifiers were securely retained in the databases of both CrowdQoE and Prolific systems, utilized solely for validation, exclusion verification, and payment processing.

3.5.4. Screening and Eligibility Criteria

No customized screening completion pathway was established on Prolific. Eligibility was determined using Prolific’s normal screening procedures in conjunction with automated technical checks embedded in the CrowdQoE platform. Participants were required to be proficient in English, possess a desktop computer with a Full HD (1920 × 1080) monitor, maintain a reliable internet connection with a minimum download speed of 40 Mbps, and not have previously engaged in relevant QoE studies. The recruitment aimed at a substantial, internationally diversified cohort, encompassing participants from the United Kingdom, United States, Ireland, Germany, France, and various other nations. This method enhanced demographic diversity and bolstered the external validity of the findings.

3.5.5. Participant Recruitment and Sample

The Prolific cohort deliberately imposed no restrictions on age, educational background, or professional expertise, yielding a demographically heterogeneous sample across multiple countries. This is an intentional and valued feature of the uncontrolled crowdsourcing design, reflecting the natural diversity of a real-world participant pool and maximizing the external validity of findings from the uncontrolled environment. The intended sample size was 80 people. Out of these, 25 and 36 participants completed the survey and were included in the analysis, 5 were excluded for noncompliance with instructions, and 49 were automatically disqualified due to technical validation issues (insufficient internet speed or inadequate display resolution). The final authorized sample had both male and female participants in about equal numbers, indicating a balanced demographic distribution. It is noted that the admitted sample of 25 participants for Database A (uncontrolled) falls below the minimum of 35 specified in ITU-T P.910 Section 10.1 for uncontrolled subjective experiments. This shortfall was a direct consequence of the stringent automated technical validation criteria applied by the CrowdQoE platform. However, as demonstrated in Section 5.6.1, the subsequent Leave-One-Out (LOO) consistency screening produced a rejection rate of 92.0% for this cohort, resulting in only 2 valid subjects. This extreme attrition rendered the crowdsourced Database A dataset unsuitable for inferential statistical comparison, and it was therefore considered to be a pilot study and excluded from all further analysis. Consequently, the shortfall below the P.910 minimum had no bearing on the validity of the reported statistical findings. The core controlled-versus-uncontrolled comparison was conducted on Database B (stable content), where the uncontrolled environment yielded 36 initial participants and 32 valid subjects after LOO screening, comfortably exceeding the P.910 requirement.

3.5.6. Completion and Payment Procedure

After concluding the experiment, participants were instructed to return to Prolific via the standard completion route and input a confirmation code. Submissions underwent thorough inspection prior to approval to guarantee data integrity, technological compliance, and complete participation. The anticipated study time was 50 min, and participants received £ 9.00 per hour, which Prolific deemed a reasonable and competitive remuneration level.

3.5.7. Technical Validation and Data Integrity

Prior to commencing the video assessment, all participants underwent an automatic screening on the CrowdQoE platform to verify the use of a compatible web browser (preferably Google Chrome), a minimum screen resolution of 1920 × 1080, an internet connection with a download speed of no less than 40 Mbps, and access via a desktop computer, as mobile devices were prohibited. The internet connection speed was verified in real time by measuring the download time of a dedicated test payload file hosted on the AGH university server, ensuring that the measurement reflected actual network performance rather than self-reported values. Screen resolution compliance was verified using JavaScript properties that report the physical display resolution of the participant’s device. Participants failing either check were immediately and automatically redirected to a disqualification page and excluded from the study without consequences, prior to any exposure to the experimental stimuli. Participants who failed to meet any of the criteria were automatically excluded from the study without consequences. Response data, timestamps, and system logs were securely archived in a database, facilitating comprehensive temporal analysis and thorough post hoc reliability assessments, in accordance with ITU-T P.910. The Prolific-based, uncontrolled environment maintained essential experimental criteria while allowing for natural variability in user situations, rendering it appropriate for comparison with the controlled laboratory setup.

4. Experiments

4.1. Hardware Configuration

The experiments were conducted using two distinct hardware setups to represent the controlled and uncontrolled testing environments.
In the controlled laboratory environment, assessments were conducted in the multimedia laboratory at AGH University of Krakow. All participants used identically configured laboratory computers equipped with Intel Core i5 processors, 8 GB RAM, integrated Intel UHD Graphics, and 23-inch Full HD (1920 × 1080) monitors. Each station was connected via a wired 1 Gbps local network, ensuring consistent transmission quality. Room lighting and seating positions were standardized to maintain visual uniformity and reduce reflections or glare.
For comparison, the uncontrolled environment involved participants using their personal computers and network connections to access the same CrowdQoE platform via the Prolific website. The system automatically verified that only users with Full HD displays and a minimum download speed of 40 Mbps could proceed to the test. This verification step was essential for ensuring a comparable baseline quality of visual stimuli despite environmental variability.

4.2. Software Environment

Both experiments were conducted using the custom-built CrowdQoE platform, implemented in PHP, JavaScript, and HTML for internal data handling. The platform incorporated system validation scripts to verify browser compatibility, screen resolution, and internet connection speed; automated redirection mechanisms to exclude participants who did not satisfy the prerequisites; randomization procedures to determine the order in which videos were shown, thereby reducing order-related bias; and response logging modules that recorded user interactions with millisecond-accurate timestamps. Participants accessed the platform using the latest version of Google Chrome in Incognito mode, as recommended on the introduction page, which served a dual purpose: it helped reduce cache-related interference and limit the impact of prior browser performance differences, while simultaneously protecting participant anonymity by preventing the browser from retaining any session data, cookies, or locally stored activity, an aspect of increasing relevance in crowdsourcing studies involving human participants.

4.3. Experimental Procedure

Both groups completed the same testing procedure. First, during the Introduction & Verification phase, participants read the instructions, conducted browser and hardware checks, and provided informed consent. Next, they completed a brief demographic questionnaire covering age, gender, education, current mood (positive/neutral/negative), and fatigue level. They then proceeded to the viewing session, during which each participant watched a randomly ordered set of video clips to minimize learning effects and bias. After each clip, they performed a rating task evaluating perceived video quality on a 5-point Absolute Category Rating (ACR) scale from 1 (Bad) to 5 (Excellent). At the same time, the system automatically recorded the response time, defined as the interval in milliseconds between the appearance of the rating page and the submission of the participant’s response. At the end of the assessment, a completion message confirmed that the session had successfully finished.
In the laboratory environment, an experiment supervisor monitored participants to ensure that all instructions were followed and that they did not multitask, switch windows, or change browser tabs. Participants were offered short breaks roughly every 55 min to alleviate fatigue.
By contrast, the Prolific participants carried out the same procedure independently, without direct supervision, and under diverse, uncontrolled conditions.

4.4. Data Logging and Monitoring

All interaction-related information, including system diagnostics, demographic inputs, timestamps, and ACR scores, was securely retained in the platform’s database. To further ensure data reliability:
Millisecond-accurate response times for each clip were captured with a JavaScript event timer. Unique Session_ID codes preserved anonymity while allowing data tracking. Real-time synchronization automatically excluded incomplete or interrupted sessions.
Extensive system logging facilitated the identification of irregularities such as implausibly short response times (under 200 ms), thereby supporting subsequent outlier detection and analysis in line with ITU-T P.910 and established statistical quality assurance practices.

4.5. Environmental Control Summary

This setup ensured identical experimental logic and interface across both conditions while isolating environmental control as the primary variable of interest. The combination of rigorous software validation and procedural symmetry provided a robust framework for comparing laboratory reliability with real-world diversity in subjective video quality assessment (Table 1).

5. Results

5.1. Analysis of Raw Data

The experimental study is conducted under two distinct configurations: a controlled laboratory environment and an uncontrolled situation. The session showcases participants’ subjective evaluations of individual video sequences classified in Database A and Database B. Database A contains video content with dynamically fluctuating quality levels, while Database B includes videos with consistent, stable quality throughout their entirety. Table 2 summarizes the overall number of participants in the study and the average number of movies assessed per participant in each condition.
The distribution of the Mean Opinion Score (MOS) per film for database A under both controlled and uncontrolled conditions is shown in Figure 2. The computed confidence interval, which shows the likely range of the MOS values with a 95% confidence level, is represented by the shaded area. Additionally, for both the controlled and uncontrolled setups of database A, Figure 3 displays the statistically calculated standard deviation of the MOS for every video and frequency histograms that show how many videos fall within specific variability ranges. When taken as a whole, these numbers describe both the subjective scores’ central tendency and their variability in various experimental contexts.
In a similar manner, the distribution of MOS for videos in Database B, under both controlled and uncontrolled experimental conditions, is illustrated in Figure 4. Furthermore, the corresponding standard deviation values associated with these MOS scores are depicted in Figure 5, providing additional insight into the variability of subjective quality assessments.
The dataset must be pre-processed to remove replies that are thought to be unreliable before any further statistical studies can be carried out. In order to prevent ratings from being skewed by people who were confused, distracted, insufficiently engaged, or subjected to excessive cognitive demands during the experiment, this data-cleaning step is crucial to ensuring that the resulting analyses primarily capture the behavior of participants who provided coherent and stable judgments. By doing this, the empirical results’ interpretability and integrity are significantly enhanced.
A methodical, correlation-based process is used to identify and eliminate unreliable participants in accordance with the ITU-T’s standards. Specifically, the primary quantitative metric for evaluating the internal consistency and dependability of each participant’s MOS is Pearson’s correlation coefficient. A correlation value of 0.75 is frequently used as the bottom bound for acceptable reliability in this context. Individual correlation values below this cutoff are used to categorize participants as inconsistent raters, and their data is eliminated from all subsequent phases of research. This process reduces the impact of unpredictable or noise-dominated assessments on the total MOS and any inferred statistical findings.
All experimental data gathered under both controlled and uncontrolled testing conditions—corresponding to the A and B databases, respectively—are subjected to the same reliability filter in the same way. The methodology encourages comparability between circumstances and improves the overall findings’ robustness, repeatability, and external validity by applying a consistent exclusion criterion across these two experimental contexts.

5.2. Dataset Reliability and Subject Validation

A two-stage screening procedure that complied with ITU-R BT.500 Recommendations was used to strictly ensure data integrity. The final valid dataset included N = 6 subjects for the Time-Varying dataset (Database A) and N = 13 subjects for the Stable dataset (Database B) after incomplete sessions (95% completion threshold) and inconsistent raters (Leave-One-Out correlation screening with r t h 0.75 ) were eliminated. A thorough analysis of the participant retention data is shown in Table 3, which emphasizes the significant difference in rejection rates between the stable and time-varying settings in the controlled environment.
The dependability of the gathered data was verified by post hoc psychometric validation, even with the smaller sample size for the time-varying cohort. Internal consistency analysis produced Cronbach’s α values of 0.81 for Database A and 0.94 for Database B, both of which were higher above the suggested cutoff point of 0.70 for subjective QoE research.
The inter-subject correlation matrices shown in Figure 6 show consistent pairwise agreement between the individuals that were kept. Additionally, a Shapiro–Wilk test verified that Database A ( p = 0.089 ) and Database B ( p = 0.097 ) Mean Opinion Scores (MOS) have a normal distribution.
These metrics confirm that the retained subjects formed a cohesive measurement instrument capable of reliably assessing both stable and temporally complex video distortions.

5.3. Impact of Temporal Complexity on Quality Ratings

Investigating if the temporal variation in video quality causes a systematic bias in subjective judgments was the main goal. Subject ID was treated as a random intercept to account for individual rating baselines, and we used a Linear Mixed-Effects Model (LMM) to predict Score depending on Content Type (Time-Varying vs. Stable).
There was no statistically significant impact of content type on the final quality ratings, according to the analysis ( β = 0.114 , S E = 0.123 , p = 0.353 ). The observed difference in raw MOS was perceptually insignificant (Database A: 3.24 ± 0.09 ; Database B: 3.13 ± 0.05 ), as shown in Table 4. We computed the effect size, which was determined to be insignificant (Cohen’s d = 0.10 ) in order to confirm that this null result was not an artifact of low statistical power. Additionally, for both datasets, strong bootstrapping (1000 resamples) produced overlapping 95% confidence intervals. These findings show that people incorporate different quality into a final score in a controlled setting that is statistically comparable to steady information of similar average quality.
We conducted a bootstrapping validation (1000 resamples) to ascertain the robustness of this null conclusion. The ratings are statistically indistinguishable since the 95% bootstrapped confidence intervals for Database A [ 3.15 , 3.33 ] and Database B [ 3.07 , 3.18 ] overlap. Moreover, despite the constrained sample size diminishing statistical power, the effect size observed was negligible (Cohen’s d = 0.10 ), suggesting that genuine similarity, rather than a Type II error, accounts for the absence of statistical significance.

5.4. Cognitive Load and Reaction Time Analysis

The implicit behavioral measurements exhibited significant fluctuations, but the explicit quality ratings (MOS) remained unchanged. We utilized a physiologically validated filter to exclude anticipatory reflexes (less than 200 ms) and attentional lapses (more than 10,000 ms) to analyze Voting Reaction Time (ms) as a proxy for cognitive burden. A statistically significant difference in processing time ( p < 0.001 ) was identified using a Mann–Whitney U test. Compared to stable material (M = 2215 ms), respondents required an average of 506 ms longer to evaluate time-varying content (M = 2721 ms). Figure 7 depicts this distribution, with Database A’s violin plot exhibiting a significant tail that indicates frequent high-latency decision-making occurrences.
The notion of cognitive compensation is substantiated by the disparity between the markedly distinct Reaction Times and the associated Mean Opinion Scores (MOS). The quality of Stable content (Database B) remains constant, facilitating prompt evaluation. Conversely, Time-Varying material (Database A) requires participants to perform temporal pooling, which involves the continuous integration of fluctuating quality levels during the stimulus presentation. The cognitive expense of this pooling technique is indicated by the 506 ms delay seen in Database A. The statistically equivalent final MOS values ( p = 0.353 ) indicate that individuals effectively compensated for the heightened cognitive demand in a controlled environment, maintaining rating accuracy while sacrificing processing speed.

5.5. Influence of Human Factors

We conducted a further examination of the influence of subject-specific factors by employing a multivariable Linear Mixed-Effects Model (LMM) ( Score Database + Mood + Tiredness ). The findings demonstrate that psychological state greatly influenced QoE evaluations, often eclipsing technical aspects.
  • Positivity Bias: Subjects reporting a Positive mood rated videos significantly higher ( β = + 0.39 , p = 0.024 ) compared to the Negative baseline.
  • Fatigue Effect: A significant tiredness effect ( p = 0.002 ) was observed, whereby subjects reporting High tiredness provided higher ratings, consistent with a leniency effect or heuristic processing strategy aimed at expediting task completion.

5.6. Comprehensive Analysis: Controlled vs. Uncontrolled Environments

5.6.1. The Limits of Crowdsourcing for Temporally Complex Tasks

After establishing the baseline psychometric behavior and quantifying the cognitive costs linked to temporal complexity in a highly controlled laboratory setting, the final phase of this study assesses the feasibility of utilizing crowdsourcing in an uncontrolled environment as a scalable alternative methodology.
Before performing a direct statistical comparison, it is essential to rectify the significant disparity in participant retention rates noted in the crowdsourcing datasets after rigorous Leave-One-Out (LOO) consistency screening. When this cognitively taxing activity was administered to the uncontrolled crowd, the compensatory mechanisms effectively employed by laboratory subjects did not generalize. The crowdsourced cohort for Database A saw an extraordinary 92.0% participant rejection rate, retaining merely two valid subjects ( N = 2 ) who satisfied the reliability criterion ( r 0.75 ).
This study suggests that, in the absence of the environmental oversight and concentrated focus characteristic of a laboratory environment, crowd workers were either disinclined or unable of maintaining the cognitive exertion necessary to appropriately assimilate varying quality levels. The resultant attention impairment generated quasi-random or markedly inconsistent voting patterns. As a result, the crowdsourced Database A dataset was omitted from further inferential statistical studies. Table 5 presents a detailed summary of participant screening and retention statistics, emphasizing the significant difference in rejection rates between the Stable and Time-Varying conditions in the uncontrolled setting.

5.6.2. Objective of the Joint Comparison

In contrast, the crowdsourced Database B (Stable content) had a strong retention rate, maintaining a sufficient pool of very consistent individuals. The cognitive load of steady content is intrinsically smaller, enabling crowd workers to successfully accomplish the rating task. To thoroughly assess the Environmental Effect, the following analysis isolates Database B, contrasting the ground truth acquired in the controlled laboratory environment with the data gathered from the uncontrolled crowd. If the two distributions are statistically equivalent, this would substantiate crowdsourcing as a robust, scalable, and economical alternative to laboratory testing—assuming that video quality is temporally consistent.

5.6.3. Distributional Shift and Leniency Bias

A first distributional shift study was performed on the stable content (Database B) to build a baseline understanding of rating habits across contexts. The unregulated crowdsourced cohort ( N = 3840 ratings) produced a Mean Opinion Score (MOS) of 3.241 ( S D = 1.282 ), demonstrating a marginal positive increase relative to the controlled laboratory cohort ( N = 1560 ratings), which achieved a MOS of 3.126 ( S D = 1.156 ).
A Mann–Whitney U test revealed that the distributional change was statistically significant ( U = 2 , 808 , 719.0 , p < 0.001 ). This discovery was validated by a non-parametric bootstrapping study (1000 resamples), which demonstrated non-overlapping 95.
Although the statistical significance is strong, it must be contextualized within the limitations of Null Hypothesis Significance Testing (NHST) applied to big datasets. The absolute difference between the two settings is mathematically negligible ( Δ MOS = 0.115 ). In psychometric assessments employing a 5-point Absolute Category Rating (ACR) scale, a divergence of this magnitude is often regarded as perceptually insignificant. The modest increase in crowdsourced ratings corresponds with the established leniency bias, or positivity bias, commonly seen in unsupervised remote testing, when evaluators tend to exhibit somewhat greater leniency.
Due to the heightened sensitivity of classic non-parametric tests to large sample sizes, the rejection of the null hypothesis does not necessarily imply practical system equivalency. It requires a hierarchical modeling strategy to distinguish individual subject leniency from the actual environmental effect, followed by formal equivalence testing.

5.6.4. Hierarchical Confounding Control

A Linear Mixed-Effects Model (LMM) was utilized to meticulously separate the genuine impact of the evaluation setting from the confounding effect of individual rater leniency. Standard non-parametric tests, such as the Mann–Whitney U, presume sample independence and are hence excessively sensitive in extensive QoE datasets, frequently merging subject-level variance with ambient influences. The LMM resolves this by treating the Environment as a fixed effect and the Subject ID as a random intercept, thus accommodating the repeated-measures design of the experiment (120 trials per subject over 45 valid subjects).
The hierarchical analysis successfully mitigated the observable distributional change. The model output indicates that the fixed impact of moving from a controlled laboratory setting to an uncontrolled crowd scenario had a coefficient of β = 0.116 . Nonetheless, after adjusting for individual subject variability, this environmental effect was statistically negligible ( z = 1.131 , p = 0.258 , 95% CI: [−0.085, 0.316]). The variance was absorbed by the random group effect ( Group Var = 0.084 ), indicating that the inherent leniency of the selected workers, rather than the unsupervised characteristics of the crowdsourcing platform, was responsible for the small increase in raw scores.
Consequently, we ascertain that when the cognitive demands of the QoE task are limited (i.e., employing steady video content), the assessment environment does not induce a systematic bias in the quality ratings. The crowdsourced and laboratory cohorts essentially represent the same psychometric population.

5.6.5. Cross-Environment Rank Agreement and System-Level Equivalence

The Linear Mixed-Effects Model (LMM) validated the lack of systemic environmental bias at the population level, although practical video engineering necessitates accurate assessment of individual stimuli. To determine if the uncontrolled public maintains the same relative quality rankings of specific video sequences as the controlled laboratory, we performed a cross-environment correlation study on the per-video Mean Opinion Scores (MOS) for the 40 stable video sequences in Database B.
The inter-environment agreement was assessed utilizing three common benchmarking metrics: Pearson Linear Correlation Coefficient (PLCC) assesses linearity in predictions, Spearman Rank-Order Correlation Coefficient (SRCC) evaluates monotonic rank preservation, and Kendall’s Rank Correlation Coefficient (KRCC) measures resilience in the presence of linked ranks. Additionally, the Root Mean Square Error (RMSE) was computed to measure absolute scoring deviation.
The research demonstrated nearly flawless concordance between the laboratory and crowdsourcing groups. Figure 8 illustrates that both the linear and rank-order correlations surpassed the conventional ITU benchmarks for superior model concordance, resulting in a PLCC of 0.9829 ( p < 0.001 ) and an SRCC of 0.9809 ( p < 0.001 ). The KRCC was notably high at 0.8988 ( p < 0.001 ), affirming that the precise ordinal ranking of video quality was maintained. The absolute error margin was minimal, with an RMSE of merely 0.2548 MOS points on the 5-point ACR scale, well below the usual limits of human perceptual fluctuation.
These findings offer conclusive evidence of system-level equivalence. The nearly flawless SRCC ensures that a video engineer employing unsupervised crowdsourcing data will arrive at identical algorithmic or optimization judgments as one dependent on supervised laboratory ground truth. Thus, for consistent video content, stringent crowdsourcing methodologies (including LOO screening) provide a highly valid, economical, and scalable alternative to conventional laboratory testing.

5.7. Psychometric Alignment and Rating Certainty

We assessed whether crowdsourcing individuals employed the rating scale with equivalent psychometric reliability as supervised laboratory subjects by testing the Standard Deviation of Opinion Scores (SOS) hypothesis. The traditional SOS model asserts that rating variance adheres to a quadratic trajectory, reaching its zenith near the midway of the scale (MOS 3 ) owing to heightened user perplexity, while diminishing at the extremities.
We fitted the theoretical curve as in Equation (1)
SOS ( x ) = a · x ( 5 x )
to the collaboratively sourced per-video data. The study produced a parameter value of a = 0.318 , accompanied by a negative goodness-of-fit ( R 2 = 0.245 ), reflecting the psychometric behavior noted in the controlled laboratory group ( a = 0.291 , R 2 = 0.391 ).
The negative R 2 values suggest that neither group adhered to the traditional parabolic variance curve; instead, both displayed a flattened variance distribution. This indicates that crowdsourced individuals that passed the rigorous LOO screening process had consistently good inter-rater agreement across all quality levels, without displaying the increased uncertainty usually linked to distant mid-tier quality assessment. Thus, we affirm that the assessment setting did not modify the essential psychometric characteristics or rating confidence of the legitimate participants. Figure 9 elucidates this tendency, demonstrating that participants who passed LOO screening exhibited consistently elevated inter-rater agreement across all quality tiers.

5.8. Summary of System-Level Equivalence

Through a rigorous, multi-stage analytical pipeline, this study evaluated the viability of crowdsourcing as a substitute for controlled laboratory testing. Our findings indicate a clear dichotomy based on content complexity:
  • High Temporal Complexity (Time-Varying): Crowdsourcing is currently an invalid methodology. The lack of supervision leads to an inability to sustain the cognitive load required for temporal pooling, resulting in a near-total collapse of participant validity (92% rejection rate).
  • Low Temporal Complexity (Stable): When coupled with strict LOO consistency screening, crowdsourcing achieves system-level equivalence with controlled laboratory data. Hierarchical modeling (LMM) confirmed the absence of systemic environmental bias ( p = 0.258 ), rank-order correlation demonstrated near-perfect sequence alignment (SRCC = 0.9809 ), and SOS analysis verified identical psychometric certainty.
Therefore, for stable video content, remote crowdsourcing is not merely an approximation of laboratory testing, but a statistically equivalent, highly scalable alternative.

6. Discussion

6.1. The Hidden Cost of Temporal Pooling

Our findings contest the presumption that time-varying quality is intrinsically evaluated worse because of recency effects or instability penalties. In a controlled setting with vigilant participants, the MOS exhibited stability across various content categories ( p > 0.05 ). Nonetheless, the Reaction Time study uncovers the concealed expense of this stability: cognitive strain.
The observed approximately 500 ms delay indicates that respondents actively participate in a cognitive integration process, possibly via a weighted average mechanism, to consolidate varying quality into a singular scalar value. This indicates that although traditional MOS approach is applicable to time-varying content, it places a heightened cognitive load on viewers. In practical streaming situations, this heightened cognitive strain may lead to accelerated viewer weariness, despite the perceived quality being satisfactory.

6.2. Robustness of the Controlled Environment

The elevated internal consistency ( α > 0.8) and the congruence of ratings across databases validate the resilience of the controlled laboratory methods. In contrast to crowdsourcing settings, where fluctuating tasks frequently result in elevated rejection rates due to their complexity, the controlled laboratory environment guaranteed that participants stayed enough engaged to execute the required temporal integration.

6.3. Limitations and Threats to Validity

We recognize three fundamental constraints. The demographic homogeneity of the controlled laboratory cohort (95% male, 100% engineering students) was deliberate and methodologically essential. In controlled QoE experiments, participant homogeneity minimizes inter-subject variability arising from non-technical factors, ensuring that observed differences in ratings and response times are attributable to the experimental stimuli rather than to participant background. By contrast, the Prolific-based uncontrolled cohort imposed no restrictions on age, education, or expertise, producing a deliberately heterogeneous sample that maximizes external validity and reflects realistic crowdsourcing conditions. The contrast between these two cohort designs is therefore not a limitation but a defining and intentional feature of the comparative experimental framework. Secondly, although the sample size in the controlled setting ( N = 19 ) meets ITU-R BT.500 criteria for valid measurements, it limits the statistical power to identify nuanced interaction effects between content and human factors. Ultimately, crowdsourced investigations are intrinsically characterized by a lack of hardware consistency (e.g., monitor dimensions, luminosity, ambient illumination). The nearly perfect rank correlation between the controlled and uncontrolled environments in Database B (SRCC > 0.98 ) indicates that cognitive stress, rather than variability in display hardware, is the principal factor contributing to failures in temporally complex distant evaluations. Subsequent research will broaden this methodology to encompass a broader, gender-balanced population to validate these findings regarding cognitive strain.

7. Conclusions

This study presented a comprehensive, multi-environment analysis of QoE assessment methodologies, specifically investigating the viability of unsupervised crowdsourcing as an alternative to controlled laboratory testing. By integrating advanced statistical modeling (Linear Mixed-Effects Models), behavioral metrics (Reaction Time), and psychometric validation (Standard Deviation of Opinion Scores), we empirically demonstrated that the efficacy of crowdsourcing is strictly bounded by the temporal complexity of the video content and the cognitive load it imposes on viewers.
Our findings reveal a critical dichotomy in subjective QoE evaluation. For temporally complex, time-varying video content, viewers must execute temporal pooling, a continuous mental integration of fluctuating quality levels. While supervised laboratory subjects successfully engaged in this cognitive compensation (evidenced by a statistically significant ∼500 ms processing latency), this mechanism failed to generalize to the unsupervised crowd. The unprecedented 92.0% participant rejection rate for crowdsourced time-varying content empirically proves that, under current methodological paradigms, crowdsourcing is an invalid and highly unreliable approach for assessing complex, dynamic video distortions. Without laboratory supervision, viewers are generally unwilling or unable to sustain the requisite cognitive effort, leading to a collapse in inter-rater reliability.
Conversely, when evaluating stable video content, the cognitive burden is substantially reduced, allowing for immediate snapshot quality judgments. Under these constrained cognitive conditions, coupled with strict Leave-One-Out (LOO) consistency screening, the crowdsourced cohort achieved definitive system-level equivalence with the controlled laboratory. Hierarchical modeling confirmed the absence of systemic environmental bias ( p = 0.258 ), and cross-environment correlation demonstrated near-perfect sequence alignment ( SRCC = 0.9809 ). Furthermore, psychometric analysis confirmed that valid crowd workers exhibited the exact same rating certainty and flattened variance distribution as supervised subjects.
Overall, this study identifies an important limitation for using crowdsourcing in Quality of Experience (QoE) assessment methods. The findings indicate that crowdsourcing should neither be treated as a universally applicable solution nor rejected entirely because of the additional noise that may arise from uncontrolled testing environments. Rather, for video content that is relatively stable, well-structured crowdsourcing procedures can serve as a valid, scalable, and cost-effective alternative to conventional laboratory experiments. In addition, the results suggest that future QoE study designs should take into account not only the perceptual effects of visual impairments, but also the cognitive effort required from participants by the evaluation task itself. Future research may aim to broaden this framework to include more demographically varied subject groups and to explore the practicality of incorporating real-time cognitive load monitoring methods into remote testing setups.

Author Contributions

Conceptualization, A.D. and M.L.; methodology, A.D., M.L., D.J. and M.G.; software, A.D.; validation, A.D., D.J. and M.L.; formal analysis, S.A. and M.T.M.S.; resources, A.D., M.L. and D.J.; data curation, S.U. and A.S.; writing—original draft preparation, A.D. and A.S.; writing—review and editing, A.D., M.L., D.J. and M.G.; visualization, A.D.; supervision, M.L. and D.J.; project administration, M.L. and A.D.; funding acquisition, S.U. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Polish Ministry of Science and Higher Education with the subvention funds of the Faculty of Computer Science, Electronics and Telecommunications of AGH University.

Institutional Review Board Statement

The study was conducted in accordance with ITU-T P.910 Recommendations. All subjective user evaluations were performed in a crowdsourced environment, ensuring participant anonymity and voluntary consent.

Data Availability Statement

Supporting reported results, code, and datasets used are available publicly on Github.

Conflicts of Interest

The authors state that they have no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ITU-TInternational Telecommunication Union Telecommunication
ACRAbsolute Category Rating
MOSMean Opinion Score
HDRHigh Dynamic Range
TMOTone Mapping Operators
VQAVideo Quality Assessment
IFInfluencing Factors
QoSQuality of Service
PDPerceptual Dimensions
LL-ABRLow-Latency Adaptive Bitrate
PVSProcessed Video Sequences
HTTPHyperText Transfer Protocol
PHPHypertext Preprocessor
msmilliseconds
HASHTTP Adaptive Streaming
ACR-HRACR with a Hidden Reference
CCRComparison Category Rating
DMOSDifferential Mean Opinion Score
HDHigh Defination
HTMLHypertext Markup Language
JSJavaScript
QoEQuality of Experience
SESubjective Experiment
AppEApproximate Entropy
SDRSubjectivity-and-Difficulty Response
LMMLinear Mixed-Effects Model
NHSTNull Hypothesis Significance Testing
PIQEPerception Image Quality Evaluator
MPEGMotion Picture Expert Group
LOOLeave-One-Out
SOSStandard Deviation of Opinion Scores
RMSERoot Mean Square Error
PLCCPearson Linear Correlation Coefficient
SRCCSpearman Rank-Order Correlation Coefficient
KRCCKendall’s Rank Correlation Coefficient

References

  1. Dutta, A.; Juszka, D.; Leszczuk, M. Crowdsourcing Evaluation of Video Summarization Algorithm. Int. J. Electron. Telecommun. 2024, 70, 1063. [Google Scholar] [CrossRef] [Scilit]
  2. Dutta, A. Rejection outliers detection in video quality assessment a crowdsourcing study. Prz.-Telekomun.-Wiadomos´Ci Telekomun. 2025, 1, 481–484. [Google Scholar] [CrossRef] [Scilit]
  3. International Telecommunication Union. P.910: Subjective Video Quality Assessment Methods for Multimedia Applications, 2023. Available online: https://www.itu.int/rec/T-REC-P.910-202310-I/en (accessed on 7 November 2025).
  4. Dutta, A. Statistical Evaluation of Subjective Video Quality on a Crowdsourcing Platform. Prz.-Telekomun.-Wiadomos´Ci Telekomun. 2025, 1, 493–496. [Google Scholar] [CrossRef] [Scilit]
  5. Sumner, J.L.; Farris, E.M.; Holman, M.R. Crowdsourcing Reliable Local Data. Political Anal. 2020, 28, 244–262. [Google Scholar] [CrossRef] [Scilit]
  6. Haralabopoulos, G.; Tsikandilakis, M.; Torres Torres, M.; McAuley, D. Objective assessment of subjective tasks in crowdsourcing applications. In Proceedings of the LREC 2020 Workshop on “Citizen Linguistics in Language Resource Development”; European Language Resources Association: Marseille, France, 2020. [Google Scholar]
  7. Möller, S.; Raake, A. (Eds.) Quality of Experience: Advanced Concepts, Applications and Methods; T-Labs Series in Telecommunication Services; Springer International Publishing: Cham, Switzerland, 2014. [Google Scholar] [CrossRef] [Scilit]
  8. Shahid, M.; Sogaard, J.; Pokhrel, J.; Brunnstrom, K.; Wang, K.; Tavakoli, S.; Gracia, N. Crowdsourcing based subjective quality assessment of adaptive video streaming. In Proceedings of the 2014 Sixth International Workshop on Quality of Multimedia Experience (QoMEX), Singapore, 18–20 September 2014; pp. 53–54. [Google Scholar] [CrossRef] [Scilit]
  9. Egger-Lampl, S.; Redi, J.; Hoßfeld, T.; Hirth, M.; Möller, S.; Naderi, B.; Keimel, C.; Saupe, D. Crowdsourcing Quality of Experience Experiments. In Evaluation in the Crowd. Crowdsourcing and Human-Centered Experiments; Archambault, D., Purchase, H., Hoßfeld, T., Eds.; Springer International Publishing: Cham, Switzerland, 2017; Volume 10264, pp. 154–190. [Google Scholar] [CrossRef] [Scilit]
  10. Hossfeld, T.; Keimel, C.; Hirth, M.; Gardlo, B.; Habigt, J.; Diepold, K.; Tran-Gia, P. Best Practices for QoE Crowdtesting: QoE Assessment with Crowdsourcing. IEEE Trans. Multimed. 2014, 16, 541–558. [Google Scholar] [CrossRef] [Scilit]
  11. Rao, R.R.R.; Goring, S.; Raake, A. Towards High Resolution Video Quality Assessment in the Crowd. In Proceedings of the 2021 13th International Conference on Quality of Multimedia Experience (QoMEX), Montreal, QC, Canada, 14–17 June 2021; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  12. Keimel, C.; Habigt, J.; Diepold, K. Challenges in crowd-based video quality assessment. In Proceedings of the 2012 Fourth International Workshop on Quality of Multimedia Experience, Melbourne, VIC, Australia, 5–7 July 2012; pp. 13–18. [Google Scholar] [CrossRef] [Scilit]
  13. Erez, E.S.; Zhitomirsky-Geffet, M.; Bar-Ilan, J. Subjective vs. objective evaluation of ontological statements with crowdsourcing. Proc. Assoc. Inf. Sci. Technol. 2015, 52, 1–4. [Google Scholar] [CrossRef] [Scilit]
  14. Winther, B.; Riegler, M.; Calvet, L.; Griwodz, C.; Halvorsen, P. Why Design Matters: Crowdsourcing of Complex Tasks. In Proceedings of the Fourth International Workshop on Crowdsourcing for Multimedia, Brisbane, Australia, 30 October 2015; pp. 27–32. [Google Scholar] [CrossRef] [Scilit]
  15. Naderi, B.; Cutler, R. A Crowdsourcing Approach to Video Quality Assessment. In Proceedings of the ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Republic of Korea, 14–19 April 2024; pp. 2810–2814. [Google Scholar] [CrossRef] [Scilit]
  16. Anegekuh, L.; Sun, L.; Ifeachor, E. A screening methodology for crowdsourcing video QoE evaluation. In Proceedings of the 2014 IEEE Global Communications Conference, Austin, TX, USA, 8–12 December 2014; pp. 1152–1157. [Google Scholar] [CrossRef] [Scilit]
  17. Seufert, M.; Zach, O.; Hossfeld, T.; Slanina, M.; Tran-Gia, P. Impact of test condition selection in adaptive crowdsourcing studies on subjective quality. In Proceedings of the 2016 Eighth International Conference on Quality of Multimedia Experience (QoMEX), Lisbon, Portugal, 6–8 June 2016; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  18. Nassar, L.; Karray, F. Overview of the crowdsourcing process. Knowl. Inf. Syst. 2019, 60, 1–24. [Google Scholar] [CrossRef] [Scilit]
  19. Daniel, F.; Kucherbaev, P.; Cappiello, C.; Benatallah, B.; Allahbakhsh, M. Quality Control in Crowdsourcing: A Survey of Quality Attributes, Assessment Techniques, and Assurance Actions. Acm Comput. Surv. 2019, 51, 1–40. [Google Scholar] [CrossRef] [Scilit]
  20. Chen, X.; Jiang, H.; Li, X.; Nie, L.; Yu, D.; He, T.; Chen, Z. A systemic framework for crowdsourced test report quality assessment. Empir. Softw. Eng. 2020, 25, 1382–1418. [Google Scholar] [CrossRef] [Scilit]
  21. Figuerola Salas, Ó.; Adzic, V.; Kalva, H. Subjective quality evaluations using crowdsourcing. In Proceedings of the 2013 Picture Coding Symposium (PCS), San Jose, CA, USA, 8–11 December 2013; pp. 418–421. [Google Scholar] [CrossRef] [Scilit]
  22. Søgaard, J.; Shahid, M.; Pokhrel, J.; Brunnström, K. On subjective quality assessment of adaptive video streaming via crowdsourcing and laboratory based experiments. Multimed. Tools Appl. 2017, 76, 16727–16748. [Google Scholar] [CrossRef] [Scilit]
  23. Šakić, K.; Dumić, E.; Grgić, S. Crowdsourced subjective Video Quality Assessment. In Proceedings of the IWSSIP 2014 Proceedings, Dubrovnik, Croatia, 12–15 May 2014; pp. 223–226. [Google Scholar]
  24. Hoßfeld, T.; Hirth, M.; Korshunov, P.; Hanhart, P.; Gardlo, B.; Keimel, C.; Timmerer, C. Survey of web-based crowdsourcing frameworks for subjective quality assessment. In Proceedings of the 2014 IEEE 16th International Workshop on Multimedia Signal Processing (MMSP), Jakarta, Indonesia, 22–24 September 2014; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  25. Lebreton, P.; Hupont, I.; Mäki, T.; Skodras, E.; Hirth, M. Eye Tracker in the Wild: Studying the delta between what is said and measured in a crowdsourcing experiment. In Proceedings of the Fourth International Workshop on Crowdsourcing for Multimedia, Brisbane, Australia, 30 October 2015; pp. 3–8. [Google Scholar] [CrossRef] [Scilit]
  26. Xu, Q.; Huang, Q.; Yao, Y. Online crowdsourcing subjective image quality assessment. In Proceedings of the 20th ACM international conference on Multimedia, Nara, Japan, 29 October–2 November 2012; pp. 359–368. [Google Scholar] [CrossRef] [Scilit]
  27. Ak, A.; Goswami, A.; Callet, P.L.; Dufaux, F. A Comprehensive Analysis of Crowdsourcing for Subjective Evaluation of Tone Mapping Operators. Electron. Imaging 2021, 33, art00020. [Google Scholar] [CrossRef] [Scilit]
  28. Goswami, A.; Ak, A.; Hauser, W.; Callet, P.L.; Dufaux, F. Reliability of Crowdsourcing for Subjective Quality Evaluation of Tone Mapping Operators. In Proceedings of the 2021 IEEE 23rd International Workshop on Multimedia Signal Processing (MMSP), Tampere, Finland, 6–8 October 2021; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  29. Carlier, A.; Salvador, A.; Cabezas, F.; Giro-i Nieto, X.; Charvillat, V.; Marques, O. Assessment of crowdsourcing and gamification loss in user-assisted object segmentation. Multimed. Tools Appl. 2016, 75, 15901–15928. [Google Scholar] [CrossRef] [Scilit]
  30. Petscharnig, S. Subjective Quality Assessment with Human Computation, Gamification, and Crowdsourcing 2015. Available online: https://netlibrary.aau.at/obvuklhs/download/pdf/2415980 (accessed on 27 October 2025).
  31. Ribeiro, F.; Florêncio, D.; Zhang, C.; Seltzer, M. CROWDMOS: An approach for crowdsourcing mean opinion score studies. In Proceedings of the 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Prague, Czech Republic, 22–27 May 2011; pp. 2416–2419. [Google Scholar] [CrossRef] [Scilit]
  32. Redi, J.A.; Hoßfeld, T.; Korshunov, P.; Mazza, F.; Povoa, I.; Keimel, C. Crowdsourcing-based multimedia subjective evaluations: A case study on image recognizability and aesthetic appeal. In Proceedings of the 2nd ACM International Workshop on Crowdsourcing for Multimedia, Barcelona, Spain, 22 October 2013; pp. 29–34. [Google Scholar] [CrossRef] [Scilit]
  33. Upenik, E.; Testolina, M.; Ascenso, J.; Pereira, F.; Ebrahimi, T. Large-Scale Crowdsourcing Subjective Quality Evaluation of Learning-Based Image Coding. In Proceedings of the 2021 International Conference on Visual Communications and Image Processing (VCIP), Munich, Germany, 5–8 December 2021; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  34. Nehme, Y.; Callet, P.L.; Dupont, F.; Farrugia, J.P.; Lavoue, G. Exploring Crowdsourcing for Subjective Quality Assessment of 3D Graphics. In Proceedings of the 2021 IEEE 23rd International Workshop on Multimedia Signal Processing (MMSP), Tampere, Finland, 6–8 October 2021; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  35. Sato, Y.; Miyazawa, K. Statistical quality estimation for partially subjective classification tasks through crowdsourcing. Lang. Resour. Eval. 2023, 57, 31–56. [Google Scholar] [CrossRef] [Scilit]
  36. Naderi, B.; Hoßfeld, T.; Hirth, M.; Metzger, F.; Möller, S.; Jiménez, R.Z. Impact of the Number of Votes on the Reliability and Validity of Subjective Speech Quality Assessment in the Crowdsourcing Approach. In Proceedings of the 2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX), Athlone, Ireland, 26–28 May 2020; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  37. Naderi, B.; Zequeira Jiménez, R.; Hirth, M.; Möller, S.; Metzger, F.; Hoßfeld, T. Towards speech quality assessment using a crowdsourcing approach: Evaluation of standardized methods. Qual. User Exp. 2021, 6, 2. [Google Scholar] [CrossRef] [Scilit]
  38. Zequeira Jiménez, R.; Llagostera, A.; Naderi, B.; Möller, S.; Berger, J. Intra- and Inter-rater Agreement in a Subjective Speech Quality Assessment Task in Crowdsourcing. In Proceedings of the Companion Proceedings of The 2019 World Wide Web Conference, San Francisco, CA, USA, 13–17 May 2019; pp. 1138–1143. [Google Scholar] [CrossRef] [Scilit]
  39. Jimenez, R.Z.; Gallardo, L.F.; Moller, S. Influence of Number of Stimuli for Subjective Speech Quality Assessment in Crowdsourcing. In Proceedings of the 2018 Tenth International Conference on Quality of Multimedia Experience (QoMEX), Cagliari, Italy, 29 May–1 June 2018; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  40. Kurup, A.R.; Sajeev, G.P.; Swaminathan, J. Aggregating Reliable Submissions in Crowdsourcing Systems. IEEE Access 2021, 9, 153058–153071. [Google Scholar] [CrossRef] [Scilit]
  41. Kazai, G.; Kamps, J.; Milic-Frayling, N. An analysis of human factors and label accuracy in crowdsourcing relevance judgments. Inf. Retr. 2013, 16, 138–178. [Google Scholar] [CrossRef] [Scilit]
  42. Khajwal, A.B.; Noshadravan, A. An uncertainty-aware framework for reliable disaster damage assessment via crowdsourcing. Int. J. Disaster Risk Reduct. 2021, 55, 102110. [Google Scholar] [CrossRef] [Scilit]
  43. Weaver, A.C.; Boyle, J.P.; Besaleva, L.I. Applications and Trust Issues When Crowdsourcing a Crisis. In Proceedings of the 2012 21st International Conference on Computer Communications and Networks (ICCCN), Munich, Germany, 30 July–2 August 2012; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  44. Samimi, P.; Ravana, S.D. Creation of Reliable Relevance Judgments in Information Retrieval Systems Evaluation Experimentation through Crowdsourcing: A Review. Sci. World J. 2014, 2014, 1–13. [Google Scholar] [CrossRef] [Scilit]
  45. Ghezzi, A.; Gabelloni, D.; Martini, A.; Natalicchio, A. Crowdsourcing: A Review and Suggestions for Future Research. Int. J. Manag. Rev. 2018, 20, 343–363. [Google Scholar] [CrossRef] [Scilit]
  46. Narimanzadeh, H.; Badie-Modiri, A.; Smirnova, I.G.; Chen, T.H.Y. Crowdsourcing Subjective Annotations Using Pairwise Comparisons Reduces Bias and Error Compared to the Majority-vote Method. Proc. Acm.-Hum.-Comput. Interact. 2023, 7, 1–29. [Google Scholar] [CrossRef] [Scilit]
  47. Gardlo, B.; Egger, S.; Seufert, M.; Schatz, R. Crowdsourcing 2.0: Enhancing execution speed and reliability of web-based QoE testing. In Proceedings of the 2014 IEEE International Conference on Communications (ICC), Sydney, NSW, Australia, 10–14 June 2014; pp. 1070–1075. [Google Scholar] [CrossRef] [Scilit]
  48. Jin, Y.; Carman, M.; Zhu, Y.; Buntine, W. Distinguishing Question Subjectivity from Difficulty for Improved Crowdsourcing. In Proceedings of the 10th Asian Conference on Machine Learning, PMLR, Beijing, China, 14–16 November 2018; pp. 192–207. [Google Scholar]
  49. Boutsis, I.; Kalogeraki, V. On Task Assignment for Real-Time Reliable Crowdsourcing. In Proceedings of the 2014 IEEE 34th International Conference on Distributed Computing Systems, Madrid, Spain, 30 June–3 July 2014; pp. 1–10. [Google Scholar] [CrossRef] [Scilit]
  50. Varshney, L.R. Privacy and Reliability in Crowdsourcing Service Delivery. In Proceedings of the 2012 Annual SRII Global Conference, San Jose, CA, USA, 24–27 July 2012; pp. 55–60. [Google Scholar] [CrossRef] [Scilit]
  51. Blanco, R.; Halpin, H.; Herzig, D.M.; Mika, P.; Pound, J.; Thompson, H.S.; Tran Duc, T. Repeatable and reliable search system evaluation using crowdsourcing. In Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, Beijing, China, 24–28 July 2011; pp. 923–932. [Google Scholar] [CrossRef] [Scilit]
  52. Behrend, T.S.; Sharek, D.J.; Meade, A.W.; Wiebe, E.N. The viability of crowdsourcing for survey research. Behav. Res. Methods 2011, 43, 800–813. [Google Scholar] [CrossRef] [Scilit]
  53. Zequeira Jiménez, R.; Fernández Gallardo, L.; Möller, S. Outliers Detection vs. Control Questions to Ensure Reliable Results in Crowdsourcing.: A Speech Quality Assessment Case Study. In Proceedings of the Companion of the The Web Conference 2018 on The Web Conference 2018-WWW ’18, Lyon, France, 23–27 April 2018; pp. 1127–1130. [Google Scholar] [CrossRef] [Scilit]
  54. Hossfeld, T.; Keimel, C.; Timmerer, C. Crowdsourcing Quality-of-Experience Assessments. Computer 2014, 47, 98–102. [Google Scholar] [CrossRef] [Scilit]
  55. Cieplinska, N.; Janowski, L.; De Moor, K.; Wierzchon, M. Long-Term Video QoE Assessment Studies: A Systematic Review. IEEE Access 2022, 10, 133883–133897. [Google Scholar] [CrossRef] [Scilit]
  56. Koniuch, K. Factors Influencing Video Quality of Experience in Ecologically Valid Experiments: Measurements and a Theoretical Mode. In Proceedings of the 14th ACM Multimedia Systems Conference, Vancouver, BC, Canada, 7–10 June 2023; pp. 338–342. [Google Scholar] [CrossRef] [Scilit]
  57. Konaszyński, T.; Dutta, A.; Gizlice, B.; Juszka, D.; Leszczuk, M. Measuring points for video subjective assessment—Impact of memory and stimulus variability. Displays 2026, 92, 103283. [Google Scholar] [CrossRef] [Scilit]
  58. Uddin, S.; Leszczuk, M.; Grega, M. Video compression and optimization technologies-review. Int. J. Electron. Telecommun. 2024, 70, 733–742. [Google Scholar] [CrossRef] [Scilit]
  59. Wang, P.; Varvello, M.; Kuzmanovic, A. Kaleidoscope: A Crowdsourcing Testing Tool for Web Quality of Experience. In Proceedings of the 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS), Dallas, TX, USA, 7–10 July 2019; pp. 1971–1982. [Google Scholar] [CrossRef] [Scilit]
  60. Dutta, A. subjective quality assessment of video summarisation algorithms a crowdsourcing approach. Prz.-Telekomun.-Wiadomos´Ci Telekomun. 2023, 1, 335–338. [Google Scholar] [CrossRef] [Scilit]
  61. Korus, F. HD-dSEQUA Baza zróżnicowanych sekwencji HD do obiektywnej oceny jakości wraz z indykatorami. Prz.-Telekomun.-Wiadomos´Ci Telekomun. 2025, 1, 453–456. [Google Scholar] [CrossRef] [Scilit]
  62. Koniuch, K.; Janowski, L.; De Moor, K.; Wierzchoń, M.; Subramanian, S. The Role of Theoretical Models in Ecologically Valid Studies: The example of a video Quality of Experience model. In Proceedings of the 2023 15th International Conference on Quality of Multimedia Experience (QoMEX), Ghent, Belgium, 20–22 June 2023; pp. 67–72. [Google Scholar] [CrossRef] [Scilit]
  63. Wanat, D.; Janowski, L.; De Moor, K. Behavior as a Function of Video Quality in an Ecologically Valid Experiment. In Proceedings of the 2023 ACM International Conference on Interactive Media Experiences, Nantes, France, 12–15 June 2023; pp. 415–418. [Google Scholar] [CrossRef] [Scilit]
  64. Jakubiec, N.; Janowski, L. From Hemoglobin to MOS: Towards Neuro-Based QoE Assessment Using fNIRS. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025; pp. 12428–12436. [Google Scholar] [CrossRef] [Scilit]
  65. Wanat, D. subiektywna ocena qoe wideo na podstawie reakcji na degradację jakości w różnych grupach wiekowych. Prz.-Telekomun.-Wiadomos´Ci Telekomun. 2025, 1, 497–500. [Google Scholar] [CrossRef] [Scilit]
  66. Konaszyński, T.; Juszka, D.; Leszczuk, M. Impact of the Stimulus Presentation Structure on Subjective Video Quality Assessment. Electronics 2023, 12, 4593. [Google Scholar] [CrossRef] [Scilit]
  67. Aguilar, A.P.; Lecci, M.; Zayas, A.D.; Madueño, G.C.; Wang, H. An Empirical Study of QoE Estimation for Video Streaming Services Using Crowdsourcing. In Proceedings of the 2024 IEEE 35th International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), Valencia, Spain, 2–5 September 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  68. Wielgus, G.; Janowski, L.; Koniuch, K.; Leszczuk, M.; Figlus, R. Proposing more ecologically-valid experiment protocol using YouTube platform. Electron. Imaging 2023, 35, 261-1–261-6. [Google Scholar] [CrossRef] [Scilit]
  69. Cieplińska, N.; Janowski, L.; De Moor, K. Quality Assessment of Video Services in the Long Term. In Proceedings of the 2023 ACM International Conference on Interactive Media Experiences, Nantes, France, 12–15 June 2023; pp. 431–434. [Google Scholar] [CrossRef] [Scilit]
  70. Dutta, A. badanie crowdsourcingowe jakości sekwencji wizyjnych. Prz.-Telekomun.-Wiadomos´Ci Telekomun. 2024, 1, 323–326. [Google Scholar] [CrossRef] [Scilit]
  71. Wanat, D.; Juszka, D.; Leszczuk, M.; Janowski, L. Bridging the Lab and the Wild: Behavioral Experiments as a Pathway to QoE Research Closer to Realistic Environment. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025; pp. 7084–7092. [Google Scholar] [CrossRef] [Scilit]
  72. Uddin, S.; Grega, M.; Rahman, W.U.; Leszczuk, M. Crowd-Sourced Subjective Assessment of Adaptive Bitrate Algorithms in Low-Latency MPEG-DASH Streaming. Appl. Sci. 2025, 15, 13092. [Google Scholar] [CrossRef] [Scilit]
  73. Bułat, J.; Cieplińska, N.; Figlus, R.; Janowski, L. Daily Video: A tool for quality of experience (QoE) in long-term context research. SoftwareX 2024, 25, 101637. [Google Scholar] [CrossRef] [Scilit]
  74. Oie, E.B.; Koniuch, K.; Cieplinska, N.; De Moor, K. Factors influencing QoE of video consultations. In Proceedings of the 2021 13th International Conference on Quality of Multimedia Experience (QoMEX), Montreal, QC, Canada, 14–17 June 2021; pp. 137–140. [Google Scholar] [CrossRef] [Scilit]
  75. Naderi, B.; Cutler, R. Comparative Study of Subjective Video Quality Assessment Test Methods in Crowdsourcing for Varied Use Cases. arXiv 2025, arXiv:2509.20118. [Google Scholar] [CrossRef] [Scilit]
  76. Saupe, D.; Hahn, F.; Hosu, V.; Zingman, I.; Rana, M.; Li, S. Crowd workers proven useful: A comparative study of subjective video quality assessment 2016. In Proceedings of the QoMEX 2016: 8th International Conference on Quality of Multimedia Experience, Lisbon, Portugal, 6–8 June 2016. [Google Scholar]
  77. Hoßfeld, T.; Wunderer, S.; Beyer, A.; Hall, A.; Schwind, A.; Gassner, C.; Guillemin, F.; Wamser, F.; Wascinski, K.; Hirth, M.; et al. White Paper on Crowdsourced Network and QoE Measurements—Definitions, Use Cases and Challenges. arXiv 2020, arXiv:2006.16896. [Google Scholar] [CrossRef] [Scilit]
  78. Saupe, D.; Bleile, T. Robustness and accuracy of mean opinion scores with hard and soft outlier detection. arXiv 2025, arXiv:2509.06554. [Google Scholar] [CrossRef] [Scilit]
  79. Gavrovska, A.; Samčović, A.; Dujkovic, D.; Golub, Y.; Starovoitov, V. Weibull distribution based model behavior in color invariantspace for blind image quality evaluation. In Proceedings of the Conference: X International Conference on BIG DATA and Advanced Analytics, Minsk, Belarus, 3 March 2024; pp. 254–261. [Google Scholar]
  80. Gavrovska, A.; Samčović, A.; Dujković, D. No-Reference Image Quality Assessment Based on Machine Learning and Outlier Entropy Samples. Pattern Recognit. Image Anal. 2024, 34, 275–287. [Google Scholar] [CrossRef] [Scilit]
  81. Dutta, A. CrowdQoE-2025: Crowdsourcing Subjective Video Quality Assessment. 2025. Available online: https://github.com/dutta-agh/CrowdQoE-2025 (accessed on 8 March 2026).
Figure 1. Flowchart.
Figure 1. Flowchart.
Electronics 15 01666 g001
Figure 2. MOS distribution for database A in controlled and uncontrolled settings.
Figure 2. MOS distribution for database A in controlled and uncontrolled settings.
Electronics 15 01666 g002
Figure 3. (a) MOS Standard deviation distribution and (b) frequency distribution of videos for specific variability ranges for database A in controlled and uncontrolled settings.
Figure 3. (a) MOS Standard deviation distribution and (b) frequency distribution of videos for specific variability ranges for database A in controlled and uncontrolled settings.
Electronics 15 01666 g003
Figure 4. MOS distribution for database B in controlled and uncontrolled settings.
Figure 4. MOS distribution for database B in controlled and uncontrolled settings.
Electronics 15 01666 g004
Figure 5. (a) MOS Standard deviation distribution and (b) frequency distribution of videos for specific variability ranges for controlled and uncontrolled settings.
Figure 5. (a) MOS Standard deviation distribution and (b) frequency distribution of videos for specific variability ranges for controlled and uncontrolled settings.
Electronics 15 01666 g005
Figure 6. Inter-subject correlation matrices for the Time-Varying content (Database A) and the Stable content (Database B) are illustrated in Figure 6. The heatmaps display pairwise Pearson correlation coefficients (r) among the valid subjects ( N = 6 for Database A; N = 19 for Database B), confirming strong inter-rater agreement across both datasets despite the increased temporal complexity of the time-varying content.
Figure 6. Inter-subject correlation matrices for the Time-Varying content (Database A) and the Stable content (Database B) are illustrated in Figure 6. The heatmaps display pairwise Pearson correlation coefficients (r) among the valid subjects ( N = 6 for Database A; N = 19 for Database B), confirming strong inter-rater agreement across both datasets despite the increased temporal complexity of the time-varying content.
Electronics 15 01666 g006
Figure 7. Distribution of voting reaction times for Time-Varying (Database A) and Stable (Database B) content ( 200 ms t 10 s ). The pronounced upper tail in Database A indicates significantly higher cognitive load and decision latency ( p < 0.001 ) compared to Stable content.
Figure 7. Distribution of voting reaction times for Time-Varying (Database A) and Stable (Database B) content ( 200 ms t 10 s ). The pronounced upper tail in Database A indicates significantly higher cognitive load and decision latency ( p < 0.001 ) compared to Stable content.
Electronics 15 01666 g007
Figure 8. Per-video MOS correlation between the controlled laboratory and uncontrolled crowdsourced environments.
Figure 8. Per-video MOS correlation between the controlled laboratory and uncontrolled crowdsourced environments.
Electronics 15 01666 g008
Figure 9. Standard Deviation of Opinion Scores (SOS) versus MOS, illustrating flattened psychometric variance in both environments.
Figure 9. Standard Deviation of Opinion Scores (SOS) versus MOS, illustrating flattened psychometric variance in both environments.
Electronics 15 01666 g009
Table 1. Summary of participants in raw experimental data.
Table 1. Summary of participants in raw experimental data.
Environment TypeDevice StandardizationNetwork ControlSupervisionLighting & AmbienceValidation Scripts
Laboratory (Controlled)Identical AGH desktop PCs, Full HD1 Gbps LANDirect supervisionStandardized lightingYes
Prolific (Uncontrolled)Participant’s personal computers with Full HDVariable (>40 Mbps required)NoneNatural personal settingsYes
Table 2. Summary of participants’ raw experimental data.
Table 2. Summary of participants’ raw experimental data.
Database ADatabase B
Metric Controlled Uncontrolled Controlled Uncontrolled
Participants32252936
Total Videos8080120120
Average Videos/Participant57.0670.0498.55112.83
Table 3. Participant screening and retention statistics. Comparison of initial recruitment versus final valid subjects following strict Leave-One-Out (LOO) screening ( r 0.75 ). The significantly higher rejection rate for Database A (81.25%) reflects the increased cognitive complexity of the time-varying rating task.
Table 3. Participant screening and retention statistics. Comparison of initial recruitment versus final valid subjects following strict Leave-One-Out (LOO) screening ( r 0.75 ). The significantly higher rejection rate for Database A (81.25%) reflects the increased cognitive complexity of the time-varying rating task.
MetricTime-Varying Quality (Database A)Constant Quality (Database B)
Initial Recruits3229
Completed Session (>95%)2222
Valid Subjects ( r 0.75 )613
Overall Rejection Rate81.25%55.17%
Table 4. Results of the Linear Mixed-Effects Model (LMM) analysis ( N = 19 ), comparing the impact of content type on subjective quality ratings (MOS). The model includes subject-specific baseline variability as a random intercept. The results indicate no statistically significant difference ( β = 0.114 , p = 0.353 ) between constant (Database B) and time-varying (Database A) quality content.
Table 4. Results of the Linear Mixed-Effects Model (LMM) analysis ( N = 19 ), comparing the impact of content type on subjective quality ratings (MOS). The model includes subject-specific baseline variability as a random intercept. The results indicate no statistically significant difference ( β = 0.114 , p = 0.353 ) between constant (Database B) and time-varying (Database A) quality content.
MetricValue
ComparisonStable (B) vs. Time-Varying (A)
Fixed Effect ( β )0.114
Standard Error0.123
95% Confidence Interval[−0.127, 0.354]
p-value0.3531
SignificanceNot Significant ( p > 0.05 )
Table 5. Participant Screening Statistics for Controlled and Uncontrolled Environments. Initial cohort sizes, excluded participants, and final valid subjects after applying completion filters and Leave-One-Out (LOO) correlation screening ( r 0.75 ). The asymmetric rejection rate for Database A (Uncontrolled) illustrates the limitations of crowdsourcing for temporally complex assessment tasks.
Table 5. Participant Screening Statistics for Controlled and Uncontrolled Environments. Initial cohort sizes, excluded participants, and final valid subjects after applying completion filters and Leave-One-Out (LOO) correlation screening ( r 0.75 ). The asymmetric rejection rate for Database A (Uncontrolled) illustrates the limitations of crowdsourcing for temporally complex assessment tasks.
MetricTime-Varying Quality (Database A)Constant Quality (Database B)
Initial Recruits2536
Valid Subjects ( r 0.75 )232
Overall Rejection Rate92.00%11.10%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Dutta, A.; Saeed, M.T.M.; Arawade, S.; Samčović, A.; Uddin, S.; Juszka, D.; Grega, M.; Leszczuk, M. Impact of Environmental Control on Subjective Video Quality Assessment in Crowdsourced QoE Experiments. Electronics 2026, 15, 1666. https://doi.org/10.3390/electronics15081666

AMA Style

Dutta A, Saeed MTM, Arawade S, Samčović A, Uddin S, Juszka D, Grega M, Leszczuk M. Impact of Environmental Control on Subjective Video Quality Assessment in Crowdsourced QoE Experiments. Electronics. 2026; 15(8):1666. https://doi.org/10.3390/electronics15081666

Chicago/Turabian Style

Dutta, Avrajyoti, Mohamedalfateh T. M. Saeed, Swapnil Arawade, Andreja Samčović, Syed Uddin, Dawid Juszka, Michał Grega, and Mikołaj Leszczuk. 2026. "Impact of Environmental Control on Subjective Video Quality Assessment in Crowdsourced QoE Experiments" Electronics 15, no. 8: 1666. https://doi.org/10.3390/electronics15081666

APA Style

Dutta, A., Saeed, M. T. M., Arawade, S., Samčović, A., Uddin, S., Juszka, D., Grega, M., & Leszczuk, M. (2026). Impact of Environmental Control on Subjective Video Quality Assessment in Crowdsourced QoE Experiments. Electronics, 15(8), 1666. https://doi.org/10.3390/electronics15081666

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop