1. Introduction
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition diagnosed on the basis of two overarching behavioral characteristics: challenges with social interactions, and restricted, repetitive patterns of behavior, interests, or activities (
Regier et al., 2013). Importantly, heterogeneity extends beyond diagnostic features where individuals with ASD show distinct patterns of strengths and weaknesses across multiple domains (
Halder et al., 2024;
Mottron & Bzdok, 2020;
Uljarević et al., 2017). These domains include low-level perceptual processes (
Uljarević et al., 2017), such as visual and auditory signal detection and discrimination (
Chung & Son, 2020;
Ong et al., 2024); higher-order domains, including filtering stimuli (
Joseph et al., 2009;
Wasiuk et al., 2025), integrating information (
Bertone et al., 2003;
Harms et al., 2010;
Marco et al., 2011), working memory (
Habib et al., 2019); and other executive function domains (
Demetriou et al., 2019). As a result, such variability contributes to the contradictory findings in studies of perception (
Kaldy et al., 2016;
Simmons et al., 2009), emotion recognition (
Harms et al., 2010), and executive function (
St John et al., 2022). Together, this body of evidence highlights ASD as a multidimensional condition in which domain-specific cognitive differences may contribute to diverse behavioral outcomes (
Happé et al., 2006). Comprehensive assessments across perceptual and cognitive domains are therefore critical for characterizing individual variability and advancing our understanding of ASD heterogeneity.
Despite growing recognition that comprehensive sensory–cognitive assessments is essential for all individuals with ASD (
Lord et al., 2018), scalable and widely accessible implementations remain rare in practice. Traditional studies often rely on geographically restricted samples (
Masoomi et al., 2025) and narrowly focused tasks (
Kaldy et al., 2016;
Remington & Fairnie, 2017), constraining both generalizability and the ability to capture the breadth of sensory and cognitive variability observed in ASD (
Lenroot & Yeung, 2013;
Lombardo et al., 2019). Remote testing offers a promising alternative by reducing barriers related to geographic access, scheduling constraints, and anxiety associated with laboratory or clinical settings (
Katakis et al., 2025). However, remote assessment in autistic populations is frequently accompanied by high attrition rates (
Whitehouse et al., 2017), raising concerns about the feasibility of unsupervised, at-home assessments (
Germine et al., 2012). A key challenge is the absence of real-time scaffolding from trained researchers, which can hinder task comprehension and sustained engagement, particularly for individuals with attentional or sensory processing differences (
Calub et al., 2022;
Panesi et al., 2023). Additional factors such as environmental distractions (e.g., noise, lighting), inconsistent caregiver involvement (e.g., under- or over-assistance), and technical variability (e.g., device type, screen size, or audio fidelity) further threaten data usability and interpretability in remote contexts (
Reips, 2002).
To address these feasibility barriers and make remote assessment more valid, scalable, and replicable for the ASD population, we designed a digital assessment battery across multiple domains implicated in ASD heterogeneity (
Figure 1), including auditory and visual processing, cognitive control and working memory, and complex emotion recognition. Task and questionnaire selection was guided by the goal of capturing complementary aspects of sensory, cognitive, and behavioral functioning. Parent-report measures were included to provide standardized indices of social and developmental functioning, and to enable future investigation of links between perceptual performance and real-world outcomes. Visual and auditory tasks were chosen to probe central perceptual processing, with particular emphasis on lower- and mid-level mechanisms that have been relatively underexplored in ASD. Cognitive tasks targeted key executive functions, including working memory and inhibitory control, and were implemented in a gamified format to support engagement in a remote, child-friendly context. These tasks were implemented in a simple to use mobile application BGC Science (
Brain Game Center, n.d.) that participants could download on their own devices (phones or tablets). In addition, trained research assistants (RA) maintained continuous communication with participants throughout the assessment process, providing live support when needed to ensure task comprehension and sustained engagement.
In the present feasibility study, we examined acceptability of the PART/BGC Science (
Brain Game Center, n.d.) platform and its ability to collect usable data under remote administration. Feasibility was operationalized by examining patterns of task engagement and data usability across the assessment battery, with the goal of determining whether remote assessment could be implemented successfully and, if not, where feasibility breakdowns occurred. This early-stage work aims to set the ground for future studies that could use apps of these types to better characterize perceptual and cognitive profiles and how these may relate to higher-level social–emotional functioning.
2. Materials and Methods
2.1. Participant
Two hundred and four participants with a prior diagnosis of ASD, as reported by parents, were recruited through “SPARK for Autism”, “Children Helping Science”, or general lab outreach. Of the 204 participants screened, 153 met the inclusion criteria. Of these, 9 participants withdrew and 23 were lost to follow-up, resulting in a final analytic sample of 121 participants who completed at least part of the assessment protocol. Because the study was conducted across multiple remote sessions and tasks were introduced in a staggered manner, not all participants completed all assessments (see
Figure 2). Demographic characteristics of the final analytic sample are reported in
Table 1.
All procedures were approved by the University of California, Riverside Institutional Review Board (IRB-SB. NUMBER: #HS-20-177). Participants received a $10 per hour Amazon gift card as compensation for their time.
Inclusion criteria were designed to ensure both clinical appropriateness and feasibility of remote participation. Criteria included: (a) a parent-reported clinical diagnosis of ASD; (b) access to required technology for remote testing, including two internet-enabled devices (iOS or Android smartphone, tablet, or computer), headphones, at least 2.5 GB of available storage to download the BGC Science, an app made by Brain Game Center for Mental Fitness and Well-Being and distributed freely online, and an active email address for account activation; (c) basic English proficiency; (d) normal or corrected vision sufficient to see stimuli on a screen at a distance of approximately two feet; (e) no parent-reported diagnosis of any psychiatric disorder besides ASD (e.g., schizophrenia, bipolar disorder, ADHD, depression, anxiety), intellectual disability, or significant perinatal complications (i.e., birth weight under 4 lbs or gestational age under 36 weeks), (f) full-scale IQ of 70 or above. To ensure participants had a full-scale IQ of 70 or above, all participants completed the two-subtest version of Wechsler Abbreviated Scale of Intelligence–Second Edition (WASI-II).
2.2. Research Design
The study employed a supervised, remotely administered assessment design comprising three child testing sessions and parent-completed questionnaires. The three child sessions were structured to target distinct domains of sensory and cognitive function, whereas caregiver-report measures were completed independently by parents to provide contextual and symptom-level information.
Caregiver-report questionnaires (i.e., SRS-2, 5-15R) were completed online by parents outside of the child testing sessions.
Session 1 focused on emotion discrimination and general cognitive ability. Measures included Child Affective Measure—Child Version (CAM-C), which provided accuracy scores for facial and vocal emotion recognition, and the WASI-II, which yielded age-normed estimates of verbal and nonverbal intellectual functioning. Session 2 assessed visual perception and cognitive functioning. Visual tasks (orientation discrimination and contour integration) generated threshold estimates of fine-grained spatial sensitivity and global contour integration, respectively. Cognitive tasks (N-back, Corsi, and Cancellation) produced indices of working memory updating, visuospatial span, and inhibitory control. Session 3 evaluated auditory perception across multiple processing levels. Pure tone detection measured minimum audibility thresholds; gap-in-noise assessed temporal resolution; digits-in-noise and spatial release from masking indexed speech-in-competition ability; and spectrotemporal modulation tasks quantified sensitivity to dynamic spectral–temporal structure (see
Figure 1).
Feasibility of remote assessment was evaluated across sessions using structured data quality coding.
2.3. Procedure
Participants completed three remote child testing sessions, each lasting approximately 30–60 min, with a minimum interval of 24 h between sessions. Prior to the first session, parents received an email with instructions for downloading and activating the BGC Science application.
Caregiver-report questionnaires were distributed electronically and completed asynchronously by parents via Qualtrics. These measures did not require live supervision and could be completed at the parents’ convenience.
All child sessions were conducted via Zoom and facilitated by a trained RA. Training for RAs included detailed instruction on task procedures, standardized administration protocols, and participant interaction guidelines, as well as supervised practice sessions to ensure consistency across sessions. During each session, the RA read aloud on-screen task instructions, clarified questions as needed, and repeated instructions when necessary to ensure task comprehension. Throughout testing, the RA monitored participant behavior and documented contextual, behavioral, and technical factors that could compromise data quality, including excessive background noise, technical issues (e.g., app malfunction or faulty headphones), and signs of inattention or impatience. Behavioral indicators included both verbal expressions (e.g., complaints or off-task conversation) and nonverbal cues (e.g., slouching or leaving the screen).
2.4. Measures
2.4.1. Caregiver-Report Questionnaires
Social Responsiveness Scale–Second Edition (SRS-2). The SRS-2 is a 65-item questionnaire designed to assess social behavior impairments associated with ASD (
Bruni, 2014). Items are rated on a 4-point Likert scale ranging from 1 (Not True) to 4 (Almost Always True). A total T-score was computed based on five domains: Social Awareness, Social Cognition, Social Communication, Social Motivation, and Restricted Interests and Repetitive Behavior were computed. Higher T-score indicated greater social-communicative difficulties, with T-scores of 76 or higher indicating severe social challenges, 66–75 indicating moderate, 60–65 suggesting mild difficulties, and 59 and below suggesting minimal to no impairment. The scale demonstrated strong psychometric properties, with an internal consistency (α = 0.95) in clinical samples (
Bruni, 2014). There were 12 participants with SRS scores under 59 (i.e., scores do not indicate social-communicative challenges). A T-test for all questionnaires and tasks was run comparing participants with SRS scores above and below 59. Only the orientation discrimination task evidenced significant differences between groups. Therefore, participants with scores under 59 were included in the analyses.
Five to Fifteen re-standardized version Questionnaire (5-15R). 5-15R, a re-standardized version of Five to Fifteen Questionnaire, assessing neurodevelopmental functioning in children (
Korkman et al., 2004). In the present study, parents completed subscales assessing Learning Skills (reading, writing, and mathematics), Language, General Learning Difficulties, Exceptional Skills, and Artistic/Practical Skills. Items are rated on a 3-point scale (0 = does not apply, 1 = applies sometimes or to some extent, 2 = definitely applies), with higher scores indicating greater difficulty. This questionnaire has shown strong psychometric properties, including acceptable to good internal consistency, test–retest reliability, and inter-rater agreement (
Korkman et al., 2004).
2.4.2. Emotion and Cognitive Screening (Session 1)
Wechsler Abbreviated Scale of Intelligence—Second Edition (WASI-II). This is a standardized intelligence test used to estimate intelligence (
Irby & Floyd, 2013). In the current study, this was included as a brief screening measure to ensure the child’s cognitive ability was within the average range, and was not used for diagnostic purposes. A two-subtest version was administered remotely to assess participants’ general cognitive ability. The Vocabulary subtest assessed verbal knowledge and expressive language, and the Matrix Reasoning subtest assessed nonverbal fluid reasoning. Age-normed standard scores from these two subtests were combined to estimate full-scale IQ. The WASI-II has demonstrated strong reliability and validity for estimating cognitive functioning in both clinical and non-clinical populations.
Cambridge Mindreading Face-Voice Battery for Children (CAM-C). This task was used to assess complex emotion recognition in children with ASD (
Golan et al., 2015). The task consisted of two subtasks: (1) Face Subtask—participants watched silent video clips of actors displaying emotional expressions; and (2) Voice Subtask—participants listened to audio recordings of emotional speech without visual cues. Nine emotions were assessed (i.e., amused, bothered, disappointed, embarrassed, jealous, loving, nervous, undecided, and unfriendly). Each emotion was presented in six trials (three per subtask), yielding a total of 54 trials. On each trial, participants selected the emotion label that best matched the stimulus from four response options. Responses were scored as correct or incorrect (1 or 0), and total accuracy scores were calculated separately for the Face (maximum = 27) and Voice (maximum = 27) conditions, with higher scores indicating better emotion recognition performance. The Voice subtask was introduced later in the study, resulting in a smaller number of participants completing this subtask.
2.4.3. Visual and Cognitive Tasks (Session 2)
Orientation Discrimination (OD) Task. This task evaluated participants’ ability to discriminate fine spatial orientation differences (
Skottun et al., 1987). Performance was measured through a two-alternative forced-choice (2AFC) staircase procedure. Participants completed six practice trials prior to main task to familiarize themselves with the task format. Each trial started with a fixation cross, followed by a Gabor patch presenting 128 millisecond (ms), tilted either clockwise or counterclockwise relative to vertical (see
Figure 2a). Participants indicated the direction of tilt on each trial. Task difficulty was modulated by a 3-down-1-up (i.e., 3 consecutively correct responses make the task harder, with any incorrect responses the task becomes 1 step easier) adaptive staircase (converging at 79.4% accuracy). The main task started at an orientation difference of 45° and decreased in logarithmic steps to a minimum of 0.72°. The task ended after five reversals or 60 trials. The threshold was defined as the orientation delta (°) of the final trial, with lower values indicating higher orientation sensitivity.
Contour Integration (CI) Task. This task evaluated participants’ ability to perceive global structure by detecting shapes embedded in visual noise (
Field et al., 1993). Each trial displayed approximately 440 Gabors on the screen, with 16 aligned Gabors forming a circular contour (see
Figure 2b). Participants completed seven practice trials prior to the main task to ensure task comprehension. On each trial, participants had 1500 ms to tap anywhere within the perceived contour. Task difficulty was manipulated by the orientation jitter (i.e., angular deviation from perfect alignment), bounded between 0° and 90°. A hybrid staircase was used: a standard 3-down-1-up rule with a “streaking” feature that temporarily switched to a 1-up rule after four consecutive correct responses to accelerate difficulty. The task terminated after five reversals or 60 trials. The threshold was defined as the jitter level (°) of the last completed trial, with higher values indicating greater sensitivity to global visual structure.
Corsi Task. This task evaluated visuospatial working memory capacity with two subtasks: forward and backward span (
Kessels et al., 2000). Each subtask began with two practice trials. On each trial, cartoon gophers appeared sequentially at different screen locations (see
Figure 2d). In the forward span subtask, participants reproduced the sequence in the same order; in the backward span subtask, they reproduced the sequence in reverse order. The task began with a sequence length of two and increased by one following a correct response. If a response was incorrect, a second trial of the same sequence length was presented. Two consecutive incorrect responses resulted in a decrease of two sequence units and the loss of one life. Participants had two lives in total, and the task terminated after two sequence levels were failed twice consecutively. Sequence length never dropped below two. The primary outcome measure was the longest forward or backward sequence length correctly recalled at least once.
N-Back Task. This task evaluated working memory updating using a fixed 1-back paradigm (
Kirchner, 1958). Participants viewed a sequence of animal images and were prompted to tap the screen when the current image matched with the one shown immediately before (see
Figure 2c). Ten practice trials were administered prior to the task, followed by 61 test trials. Performance was quantified as accuracy, with higher scores indicating better working memory performance.
UCancellation Task. This task evaluated inhibitory control (
Pahor et al., 2022). Participants completed two practice trials prior to the main task. On each trial, participants were shown a row of eight cartoon monkeys that varied in orientation (upright or upside-down), facing direction (left or right), and color (light-brown or dark-brown). The target stimulus was a dark-brown monkey that was upside-down and facing left (see
Figure 2e). Participants were instructed to tap only target stimuli while scanning from left to right. Each row displayed presented for up to 6 s, and participants completed as many rows as possible within a 3 min time limit. Performance was quantified as concentrated performance, calculated by subtracting total false alarms from total correct hits. Higher scores indicated better inhibitory control.
2.4.4. Auditory Tasks (Session 3)
Pure Tone Task (PTT). This task evaluated minimum audibility under two subtasks: quiet and noisy (
American National Standards Institute, 2004;
Lelo de Larrea-Mancera et al., 2020). Participants were asked to detect a 2000 hertz (Hz) pure tone presented either in silence (quiet subtask) or embedded in continuous broadband white noise at 60 dB SPL (noise subtask). On each trial, stimulus was presented for 500 ms, followed by a 2AFC response screen on which participants indicated whether they heard the tone (see
Figure 2f). In the quiet subtask, tone intensity followed a fixed series (70, 50, 40, 30, 20, 10, 5, and 0 dB SPL). In the noise condition, tone intensity decreased in 5 dB steps every three trials, starting at 70 dB SPL and continuing to a minimum of 10 dB SPL. Catch trials (0–2 per block of six) were included to reduce bias. The task ended if the participant either produced four consecutive incorrect responses or reached the lowest tone level. The threshold was defined as the final correct intensity level, with lower thresholds reflecting better auditory sensitivity.
Gap in Noise (GIN) Task. This task assessed temporal sensitivity, defined as the ability to detect brief silent gaps in continuous noise (
Gallun et al., 2014;
Lelo de Larrea-Mancera et al., 2020). Participants were instructed to detect brief silent gaps in white noise bursts. Each trial consisted of four auditory intervals, represented on screen by four gray squares that changed color (blue) during playback (see
Figure 2g). Each interval contained two 4 ms white noise bursts (cropped Gaussian), presented diotically at 70 dB SPL. A silent gap was inserted into one of the two middle intervals. Participants selected the interval containing the gap, and no feedback was provided following responses. Task difficulty was adjusted using a two-stage adaptive staircase procedure with a 2-down–1-up rule. In Stage 1, gap duration began with a 20 ms gap and 5 ms steps, ending after three reversals. In stage 2 gap duration was adjusted in 2 ms steps and ended after six additional reversals. The final gap detection threshold was computed as the mean of the last six reversals in Stage 2. Gap durations ranged from 0 to 40 ms, with lower thresholds reflecting better temporal resolution.
Digits in Noise (DIN) Task. This task assessed general speech-in-noise recognition ability (
Lelo de Larrea-Mancera et al., 2020;
Smits et al., 2013). Participants needed to identify spoken digit triplets (0–9) embedded in continuous broadband white noise presented diotically at 65 dB SPL. On each trial, a digit triplet was presented with 500 ms interstimulus intervals between digits, as well as 500 ms of silence padding at stimulus onset and offset. Participants responded via an on-screen keypad and received visual feedback after each trial (see
Figure 2h). Task difficulty was adjusted using a 1-down-1-up staircase procedure based on target-to-masker ratio (TMR), starting at a 0 dB with 2 dB step sizes and bounding between −16 dB and +10 dB. Thresholds were calculated by averaging the last six reversals per run, and the final TMR threshold was the mean across two 25-trial runs. Lower TMR values indicated better speech-in-noise perception.
Spatial Release Masking (SRM) Task. This task assessed participants’ ability to use spatial cues to distinguish target speech from competing talkers (
Lelo de Larrea-Mancera et al., 2020;
Marrone et al., 2008). On each trial, participants listened to a target sentence (“Charlie goes to the [color] [number]”) drawn from the Coordinate Response Measure (CRM) corpus (
Bolia et al., 2000) and identified the perceived color-number pair using an on-screen grid (see
Figure 2i). The task comprised three subtasks. In the single talker subtask, only the target sentence was presented. In the co-located subtask, the target was presented simultaneously with two masking sentences from different talkers, all originating from 0° azimuth. In the separated subtask, the target remained at 0°, while the maskers were spatially positioned at ±45°, on the azimuth, using head-related transfer functions (HRTFs). Task difficulty in the single talker subtask was adjusted using a 1-up–1-down adaptive staircase procedure with 5 dB step sizes over 20 trials. Thresholds were calculated as the average level across the final six reversals. In the colocated and separated subtasks, masker levels increased linearly from 55 to 75 dB SPL in 2 dB steps every two trials, while the target level remained fixed. Masker thresholds were estimated at the level corresponding to approximately 50% accuracy. TMR thresholds were calculated by subtracting the masker level from the fixed target level, with lower TMRs indicating better speech-in-noise perception and greater spatial release from masking.
Spectrotemporal Modulation (STM) Task. This task assessed participants’ sensitivity to dynamic changes in both frequency and time, which are critical for decoding complex auditory signals such as speech (
Bernstein et al., 2013;
Stavropoulos et al., 2021). Participants were presented with four sequential noise carriers (300 ms each, centered at 2000 Hz) at 70 dB SPL (same format with
Figure 2g). The task included two subtasks. In the detection subtask, one interval contained a spectrotemporally modulated signal (2 cycles/octave; 3.33 Hz), whereas the remaining intervals were unmodulated. In the discrimination subtask, all intervals were spectrotemporally modulated, but one interval differed in modulation direction (upward vs. downward). Participants identified the target interval by selecting the corresponding option on an on-screen response grid. Task difficulty was controlled by adjusting modulation depth using a two-stage 2-up-1-down adaptive staircase: Stage 1 used 0.5 dB steps for three reversals, followed by 0.2 dB steps in Stage 2 until nine total reversals were reached. The final threshold was calculated as the average of the last six reversals in Stage 2, with lower threshold indicating better spectrotemporal sensitivity.
2.5. Data Quality Coding
2.5.1. Coding Procedures
To evaluate data quality and task feasibility, two examiners independently reviewed each participant’s trial-by-trial performance alongside session notes. Based on this review, a detailed coding scheme was developed for all tasks, outlining specific criteria for data inclusion and exclusion. Using the finalized code scheme, one examiner initially coded all completed tasks for each participant. For tasks that included multiple subtasks (e.g., Corsi Forward and Corsi Backwards are different subtasks within the Corsi Block task), each subtask was coded separately to account for cases in which a participant was unable to complete one subtask but successfully completed another. To assess coding consistency, a second rater independently coded a randomly selected 10% subset of the data. Agreement between raters was 100% for this subset. In addition, all cases identified by the primary rater as requiring exclusion were reviewed by the second rater. Data was excluded from a specific subtask only if both examiners agreed on their exclusion status.
2.5.2. Code Scheme
Participants’ performance in each task or subtask was classified into one of the following categories to determine data usability.
Inclusion Categories
This includes two categories.
Consistent Performance. Participants were classified as showing consistent performance if they completed the task as intended and demonstrated response patterns consistent with task instructions. Performance trajectories followed expected trends (i.e., thresholds converging appropriately where the participant was responding consistently enough for a reliable threshold to be estimated), and no technical disruptions were reported. Data meeting these criteria were included in the final analyses.
Inconsistent Performance. Participants were classified as showing poor compliance if they demonstrated reduced engagement or attentional difficulties during the task—such as a sudden and sustained increase in errors after a certain trial point. In these cases, only task segments that reflected expected performance were retained, and outcome measures (e.g., thresholds or accuracy scores) were recalculated based on the usable portion of the data. This partial-data retention approach was adopted to preserve valid information while minimizing the influence of disengagement on data quality.
Exclusion Categories
This includes three categories.
Unable to Complete. Participants were classified as unable to complete a task if they were unable to engage meaningfully, resulting in uninterpretable or invalid data. This included cases where participants consistently responded incorrectly or failed to stabilize at a meaningful threshold, suggesting a lack of perceptual sensitivity or task comprehension. For example, in the Pure Tone Detection task, some participants responded “Yes” on every trial, including catch trials, indicating a response bias rather than true detection. In other instances, RAs observed behaviors such as random clicking or clear signs of inattention despite reminders, reflecting disengagement. These patterns indicated that the task could not be completed in a valid manner.
Technical Issues. Participants were classified as having technical issues if their session was affected by technological problems that interfered with stimulus presentation or response recording. These issues included audio playback failure, broken headphones, frozen screens, or application crashes, all of which compromised data integrity.
Background Noise. Participants were classified under background noise if excessive environmental noise (e.g., conversation, television, traffic) was noted by the RA during the session. Such noise could mask auditory stimuli or distract participants, particularly in tasks requiring fine acoustic discrimination (e.g., Digits-in-Noise or Gap-in-Noise). Data from these sessions were flagged for potential exclusion depending on task sensitivity.
2.6. Data Analysis
Data analysis was conducted in two sequential phases. The first phase focused on characterizing feasibility and data usability across tasks using the predefined data quality coding scheme. For each task and subtask, the proportion of participants classified as producing usable data (i.e., Consistent Performance or Inconsistent Performance with usable segments retained) was calculated. Exclusion frequencies were summarized by category (Unable to Complete, Technical Issues, Background Noise) to identify common sources of data loss across the assessment battery. To examine whether exclusion frequencies differed across tasks or subtasks, chi-square tests of independence were conducted on exclusion classifications. Chi-square analyses were based on subtask-level coding outcomes and used two-tailed tests with an alpha level of 0.05.
The second phase of analysis focused on descriptive characterization of task performance among participants whose data met inclusion criteria. For each task and subtask, performance was summarized using descriptive statistics. To visualize the distribution, variability, and central tendency of performance across participants, violin plots were generated for each task.
4. Discussion
The current results support the feasibility of delivering a comprehensive, sensory–cognitive remote assessment battery in a sample of children with ASD. Our results indicate high inclusion rates across most tasks, minimal technical exclusions, and generally interpretable performance distributions. These findings are particularly encouraging given the length and complexity of the battery, and the challenges typically associated with remote testing in pediatric and neurodiverse populations.
Feasibility of remote assessment varied across domains, reflecting differences in task order, design feature, environment conditions, and stimuli algorithm. Cognitive tasks demonstrated the highest feasibility, with approximately 94% of participants producing usable data, followed by auditory tasks (87%) and visual tasks (75%). One likely contributor to this pattern is task order (
Charness et al., 2012). Visual tasks were administered first within the BGCScience, and lower usability may reflect participants’ initial unfamiliarity with both the software interface and psychophysical task demands. By contrast, cognitive tasks were administered after the visual tasks, at a point when participants were already accustomed to the interface and response structure, potentially reducing operational errors and misunderstandings.
These observations suggest actionable improvements for future implementations. While the cognitive tasks had elaborated visual tutorials the demonstrated how to conduct the tasks, the hearing and vision tasks only had written instructions, which some participants may skip or not fully process. The difference in the number of participants producing usable data may reflect a need to clarify the instructions. To this end, we are actively developing revised tutorials that minimize text and instead rely on images and animations to communicate task rules in a more accessible and engaging format. Further, increasing the number of practice trials and having RAs explicitly confirm rule com-prehension prior to initiating the main task may further improve usability (
Crump et al., 2013).
In addition to task order, design features may also have contributed to the observed differences. All cognitive tasks were gamified, which may have supported sustained engagement. Prior work suggests that game-based paradigms can enhance motivation in children with ASD (
Grynszpan et al., 2014), and this design element may have contributed to the particularly high usability observed for executive function measures.
Auditory tasks, while still demonstrating strong feasibility overall, showed lower usability relative to cognitive measures. This pattern is consistent with the inherent sensitivity of auditory testing to environmental conditions (
Woods et al., 2017). Despite explicit instructions to maintain a quiet environment, RA frequently documented background noise that caregivers were unable to fully control in home settings. Moreover, the adaptive algorithm used in all auditory tasks began with relatively high stimulus intensities as part of a descending staircase. Some participants reported that these initial stimuli were uncomfortably loud, which negatively impacted engagement. Although this issue had not emerged in prior implementations with typically developing samples, individuals with ASD may exhibit heightened sensory sensitivity (
Khalfa et al., 2004), such that intensities tolerable for neurotypical listeners may be distressing. These findings suggest that future remote auditory protocols for ASD populations may benefit from initiating tasks at intermediate stimulus levels rather than the highest intensities, thereby reducing discomfort while preserving measurement validity.
Beyond feasibility, descriptive performance patterns provide additional insight. Across most tasks, mean performance fell within expected ranges reported for comparable paradigms, suggesting that remotely administered adaptive procedures yielded valid and interpretable estimates of perceptual and cognitive functioning. An exception was observed in the minimum hearing measures, where thresholds were higher relative to reference values obtained using the same remote paradigms in prior samples. Importantly, methodological explanations are unlikely to fully account for this pattern. The comparison data were derived from identical task implementations administered remotely within the same laboratory framework, reducing concerns regarding paradigm or platform differences. In addition, participants who failed to demonstrate task understanding were excluded from analyses, making systematic misunderstanding an unlikely primary driver of elevated thresholds. One possibility is that the upward shift in mean thresholds—and the relatively large variability observed—reflects genuine heterogeneity in low-level auditory sensitivity within the ASD population (
O’Connor, 2012). That is, a subset of participants may exhibit subtle peripheral or low-level auditory processing differences. While the present study was not designed to diagnose hearing impairment, this distributional pattern is consistent with the broader premise motivating this work: that variability in basic perceptual processing may underlie meaningful differences in higher-level functioning. Identifying such low-level sensory variability may ultimately inform more individualized approaches to intervention, particularly if specific perceptual profiles are linked to downstream communication outcomes.
Several limitations should be considered when interpreting the findings of the present study. First, ASD diagnoses and certain participant characteristics (e.g., vision status) were based on parent report and were not independently verified through clinical evaluation. In addition, the sample was relatively high-functioning and selected based on specific inclusion criteria (e.g., IQ ≥ 70, absence of reported comorbidities, and access to compatible technology). This limits the availability of objective clinical information and may restrict generalizability to the broader ASD population. Second, the analyses conducted in this study were descriptive statistics, focusing on data usability and performance distributions rather than inferential comparisons. Therefore, conclusions are limited to feasibility and interpretability of task performance under remote testing conditions. Third, the protocol employed a fixed task order, which may have introduced learning or fatigue effects across sessions. As discussed above, task order may have influenced feasibility outcomes, particularly for early visual tasks and later auditory tasks. The fourth limitation is the relatively large number of tasks administered across sessions, which may have increased participant burden and introduced variability in engagement or fatigue, particularly in a remote testing context. Future implementations may benefit from counterbalancing task order or introducing adaptive sequencing to further optimize feasibility across domains. Finally, although the tasks were derived from well-established paradigms, their psychometric properties have not been formally validated in the present remote, app-based implementation. Factors inherent to remote testing, including variability in device characteristics, environmental conditions, and the use of gamified interfaces, may influence performance in ways that differ from traditional laboratory settings. Accordingly, the current findings should be interpreted as evidence of feasibility and interpretability rather than formal validation of these measures. Future work should directly evaluate the reliability and validity of these tasks under remote conditions, including comparisons with in-lab assessments and examination of test–retest reliability.
The present findings have several methodological implications for the design and implementation of remote sensory–cognitive assessments in ASD research. First, the results highlight the importance of structured supervision in remote settings. Although assessments were conducted outside the laboratory, real-time support from trained RA played a critical role in ensuring task comprehension, maintaining engagement, and documenting contextual factors. These findings suggest that supervised remote protocols may offer a practical balance between scalability and data integrity, particularly for pediatric ASD populations. Second, task design features appear to meaningfully influence feasibility under remote conditions. Gamified cognitive tasks demonstrated high data usability, consistent with prior evidence that game-based paradigms can support motivation and sustained engagement in children with ASD. In contrast, tasks administered early in the protocol or those requiring fine perceptual judgments were more sensitive to initial unfamiliarity with software and task demands. Together, these observations underscore the value of intuitive interface design, streamlined tutorials, and sufficient practice trials to support successful remote task execution. Third, auditory measures, while feasible, were particularly sensitive to environmental noise and individual differences in sensory sensitivity. These results suggest that remote auditory assessments may benefit from adaptive stimulus calibration, more flexible starting levels, or enhanced environmental screening procedures to improve participant comfort and data usability without compromising interpretability. Finally, the present study demonstrates the feasibility of integrating multiple sensory and cognitive measures within a single remote protocol. Such multi-domain batteries offer a promising avenue for capturing individual variability across sensory and cognitive processes in ASD, while reducing barriers associated with in-lab testing.