1. Introduction
Motor behavior and environmental parameter monitoring through smartphone devices and wearable sensors are currently some of the most promising frontiers in research on the well-being and cognitive health of older adults [
1]. In recent years, the adoption of wearable sensors (accelerometers, gyroscopes, step sensors) and smartphones for detecting daily mobility has become a practical solution for in situ monitoring of individuals over 65 [
2]. Moreover, the monitoring of physical activity through smartphone devices has seen significant development, with the integration of advanced sensors and access to real-time environmental data. Modern smartphone applications use sensors integrated into smartphones, such as GPS, accelerometers, and motion sensors, to collect data on patient physical activity [
3,
4,
5]. The use of GPS devices (dedicated trackers or smartphone GPS functionalities) has been extensively studied in the literature [
6] for managing individuals with dementia: recent literature has evaluated the feasibility, acceptability, and practical impact on care (e.g., reducing emergency service involvement), but ethical and operational issues remain to be addressed. There is growing evidence suggesting that exposure to air pollutants (particularly PM2.5 and PM10) may have a potential impact on cognitive health. Therefore, integrating geolocated environmental measures with mobility data provides a more comprehensive framework to explore possible relationships between environmental factors and cognitive functions, as well as to contextualize behavioral variations [
7]. Assistive technologies have undergone a profound evolution in recent years, shifting from a reactive supervision model to a proactive risk management approach focused on prevention and continuous monitoring [
8]. Modern smartphone devices and dedicated applications are no longer conceived merely as emergency tools, but as intelligent systems capable of detecting anomalous behaviors and sending timely alerts to caregivers, thereby preventing hazardous situations [
9].
The analysis of currently available smartphone applications highlights three main categories of applications:
Medication reminder applications (e.g.,
MyTherapy, Medisafe), which enhance treatment adherence and support patients in the autonomous management of their therapy [
10,
11];
Emergency management and GPS localization applications (e.g.,
Recall Connect, Pulsante Panico), which reduce response times in cases of falls or medical emergencies [
12,
13];
Digital health profile management applications (e.g., Medical ID, TeraPiù), which ensure immediate access to essential clinical information even in critical situations.
Despite recent advances in telemedicine and digital health platforms, many existing applications still exhibit significant limitations in terms of data integration. These applications are often fragmented and fail to effectively combine physiological, behavioral, and environmental information within a unified and context-aware monitoring system. In this scenario, there is a clear need to develop advanced digital solutions capable of synergistically integrating data on physical activity and environmental conditions, providing contextual analyses and intelligent functionalities to support informed health and wellness decisions. However, currently available Service-Oriented Architecture (SOA) applications do not allow for the extraction of behavioral data suitable for predicting human activity, nor do they enable the integration of environmental pollution data factors that are strongly associated with health outcomes.
The Care-MOVE application was developed to enable continuous patient monitoring, integrated with cognitive assessment derived from clinical analysis. Building on these integrated datasets, the present study aimed to investigate the potential correlation between mobility patterns and cognitive status, specifically assessing whether dementia status could be inferred in a binary classification framework based on movement-related features. This experimental challenge was addressed through the extraction and analysis of data collected via the Care-MOVE application. In this context, the Care-MOVE application has emerged as an innovative application developed by a multidisciplinary team of clinicians, psychologists, sociologists, and computer scientists. The Care-MOVE application is designed to deliver a personalized and continuous monitoring system, linking patients’ physical parameters with the surrounding air quality.
Based on these capabilities, the main objectives of the present study are threefold:
Design and implementation of the Care-MOVE application, ensuring a continuous and integrated monitoring system;
Creation of a comprehensive Care-MOVE_dataset containing behavioral, motor, and environmental information and a cognitive target suitable for scientific analysis;
Comparison of shallow and Deep Learning model performance using the Care-MOVE dataset, aimed at evaluating the potential to predict cognitive status, based on movement patterns and collected environmental data.
Each participant recruited for the testing phase contributed an average of over 30,000 longitudinal samples, collected over five days of data extraction. However, we are aware that this wealth of intra-participant information does not replace the need for a larger sample of subjects for learning robust temporal representations. For this reason, this study was designed and presented as an exploratory proof-of-concept study rather than a definitive clinical validation.
This integrated approach not only enables the collection of detailed and contextualized data but also opens new perspectives for proactive digital healthcare, aimed at risk prevention, improving quality of life, and promoting autonomy in older patients.
2. State-of-the-Art
To contextualize the present study, the state-of-the-art analysis was conducted considering three main aspects. First, the existence of SOA applications like Care-MOVE was investigated, assessing whether commercial or experimental tools offered comparable functionalities for continuous and integrated patient monitoring. Second, the availability of SOA datasets containing behavioral information, combined with environmental data such as air quality or pollution levels, was examined, as these elements are crucial for analyzing potential correlations between environment, behavior, and cognitive status. Building on the availability of these applications and datasets, the scientific literature was examined to evaluate how researchers have combined behavioral and environmental data, and which analytical methods have been applied to predict cognitive status and assess health outcomes. Finally, the scientific literature was reviewed to identify studies addressing similar topics, with particular attention to the methods used for analyzing behavioral and environmental data and for predicting cognitive status. Between these two last aspects, it is important to highlight the role of recent technological advancements.
In particular, the rapid development of smartphone technologies and wearable devices has greatly expanded the potential for monitoring both physical activity and environmental conditions in real-time. This evolution has enabled the collection of detailed behavioral and contextual data, providing a foundation for the creation of integrated patient datasets and opening new avenues for personalized health and wellness interventions. In recent years, the rapid development of smartphone technologies and wearable devices has significantly expanded the possibilities for monitoring both physical activity and environmental conditions in real-time. This technological evolution has enabled researchers and practitioners to collect detailed behavioral and contextual data, opening new avenues for personalized health and wellness interventions. Within this landscape, several advanced applications have emerged, aiming to provide users with actionable information about their activity patterns and the surrounding environment. Some advanced applications also leverage external environmental sensors to monitor real-time air quality. For example, the “
Air Quality Monitor” application provides information on air quality and suggests physical activities based on pollution levels. Similarly, the Airly application offers real-time air quality data, along with educational and customizable features for users [
14]. Despite the advancements in smartphone applications for monitoring physical activity and environmental conditions, most existing solutions do not fully integrate data from motion sensors with environmental measurements, often providing separate information without contextualized analysis. Moreover, many applications lack advanced features, such as dynamic estimation of caloric expenditure based on individual parameters or the integration of historical data for in-depth analyses in the health field. Concurrently, several studies have highlighted the importance of considering physical activity and air quality in an integrated manner. For instance, some research has shown that exposure to high levels of air pollution can reduce individuals’ likelihood of engaging in physical exercise [
14,
15,
16,
17,
18,
19]. Other studies, such as those based on the Chinese General Social Survey, have demonstrated that the effects of pollution and physical activity on health are interactive processes: the benefits of exercise vary depending on the severity of pollution exposure [
20]. These findings underscore the need for solutions that integrate physiological and environmental data in a contextualized way, providing a more comprehensive understanding of user behaviors and their health implications. Nevertheless, most of these studies rely on aggregated data or theoretical models, with limited real-time practical applications. Recently, innovative approaches have emerged that combine smartphone sensors and machine learning techniques to enhance real-time air quality monitoring and estimate localized, adaptive exposure levels [
21]. Recent studies emphasize a significant association between exposure to fine particulate matter and markers of brain aging [
22,
23,
24,
25]. Advances in activity recognition algorithms (including deep learning-based approaches) have improved the ability to classify motor states from accelerometry and GPS streams on smartphones, even in older populations [
26,
27,
28]. Subsequently, clustering techniques, anomaly detection, and supervised predictive models are applied to identify significant behavioral patterns (e.g., variations in routine that may precede cognitive decline). However, clinical validation and generalizability remain open challenges [
29,
30]. The practical effectiveness of these systems heavily depends on ergonomic and adherence aspects: recent studies show that acceptance by older adults, continuous device usage duration, and limitations imposed by smartphone operating systems (
background process management, energy consumption) influence the quality and continuity of the collected data [
31]. Usability designs tailored for populations with cognitive deficits and cross-device testing protocols are necessary to ensure robustness. Continuous collection of geolocation and behavioral data raises important ethical and legal issues (informed consent, pseudonymization/anonymization, management of sensitive data). Adherence to GDPR and the adoption of data minimization and pseudonymization practices are essential for projects involving vulnerable populations [
32,
33]. Several publicly available datasets collect information on physical activity, calorie consumption, and observation timestamps; however, few also include explicit cognitive measures or categorizations and features related to air quality. For example, the LifeSnaps dataset provides a multimodal collection of smartwatch-derived data, including step count, distance traveled, heart rate, and calories burned over time, but it does not include structured cognitive measures [
34]. The DAMMI dataset (Daily Activities in a Psychologically Annotated Multi-sensor Dataset), on the other hand, integrates data on daily activities with psychological and contextual annotations, offering a potential link between observed behavior and mental states, although it lacks direct energetic measures such as caloric expenditure [
35]. Another relevant resource is AI4FoodDB, which includes data on physical activity and caloric intake within a personalized nutritional monitoring context, yet does not provide cognitive or psychological information [
36]. Overall, the availability of datasets that simultaneously integrate physical activity, energy expenditure, temporal timestamps, cognitive categorizations, and air quality data remains fragmented. This highlights the need for a unified data collection approach capable of correlating activity, cognitive state, and environmental conditions simultaneously paving the way for new interdisciplinary analyses on the impact of environmental factors on human health and behavior. In this regard, the application developed introduces a new dataset, named the “Care-MOVE Dataset”. The Care-MOVE application introduces an integrated digital ecosystem designed to continuously and automatically collect, process, and analyze physiological, behavioral, environmental data and cognitive targets. This work aimed to develop tools that enable personal assistance, i.e., tailored to the characteristics, habits, and individual needs of each patient, and evidence-based, i.e., grounded on concrete and measurable data systematically collected using the Care-MOVE application and standardized clinical assessments. This approach aims to support clinical evaluation by providing objective information on motor behavior, environmental exposure, and cognitive functions, and to prevent or slow the progression of neurodegenerative conditions, thanks to timely and personalized interventions derived from data analysis. The Care-MOVE application stands out for its scalable and modular client-server architecture, combining an Android smartphone application developed in Dart/Flutter with a Python v3.9.21 microservices backend. The application collects geolocated data related to the user’s daily mobility, including position, activity type (
still, walking, running, on_bicycle, in_vehicle), and caloric expenditure, integrating them with environmental information on air quality obtained in real-time from public sources (
ARPA Puglia). All data are managed securely and pseudonymized, in full compliance with current personal data protection regulations (
GDPR).
Thanks to the use of the
flutter_background_geolocation library, the application allows continuous background monitoring, maintaining low energy consumption and ensuring data acquisition even in the case of forced process closure. Regarding the use of the
flutter_background_geolocation plugin, we acknowledge that this is a proprietary solution, distributed under a commercial license included in the application code. This choice was driven by requirements of operational stability, long-term reliability, and functional completeness, particularly in the context of continuous monitoring of elderly subjects. The plugin provides advanced functionalities that are difficult to replicate natively using open-source solutions, including headless background task execution, automatic activity recognition, automatic HTTP data synchronization, and advanced energy management policies. However, we recognize that the use of a proprietary component limits the full reproducibility of the system by third parties who do not possess the required license, representing a genuine concern in a scientific context. Several open-source alternatives were evaluated, including
geolocator,
location, and
background_locator_2. Nonetheless, these solutions present relevant limitations for the specific use case, such as the lack of native support for persistent background execution, the absence of integrated activity recognition modules, and reduced stability across different devices and operating system versions. The adoption of the proprietary plugin therefore represents a trade-off between reproducibility and operational robustness in a real-world deployment scenario. With respect to sensor fusion, the app relies on the native mechanisms provided by the Android operating system, which combine data from GPS, Wi-Fi, cellular networks, and motion sensors to estimate user location and activity. No custom sensor fusion strategy was implemented at the application level. This choice is consistent with mobile development best practices, as system-level frameworks are generally more energy-efficient, better optimized, and more consistently maintained than custom implementations. That mobile sensor-based monitoring systems can be affected by various factors, including device hardware heterogeneity, sensor drift, operating system updates, and environmental influences, which may compromise the quality and comparability of data over time. Such issues generally require specific strategies for quality control, signal normalization, and calibration. From an implementation perspective, activity detection relies on the native Android activity recognition services, which integrate internal mechanisms for sensor fusion, normalization, and cross-device variability management. Leveraging activity recognition services through the
flutter_background_geolocation plugin allows delegating the management of sensor calibration, hardware heterogeneity, and system updates to Google’s proprietary frameworks, which are designed to ensure the consistency and stability of estimates across different devices. Prior to data collection, an extensive testing phase was conducted on smartphones from different manufacturers and Android versions to identify potential systematic anomalies and verify the robustness of the acquisition process. Given the short monitoring window adopted in this study, we therefore consider the effects of sensor drift and long-term inconsistency to be limited [
37].
Concerning real-time processing, our application supports continuous collection and transmission of location and activity data but does not implement advanced on-device techniques such as smoothing, adaptive buffering, or predictive filtering. Data processing is intentionally kept lightweight to reduce computational load and preserve battery life, delegating more sophisticated analyses to the offline post-processing phase. Regarding energy optimization, the system employs consolidated strategies, including dynamic adjustment of sampling frequency based on the user’s movement state, the use of low-power modes when the user is stationary, and efficient management of background services. The system can tolerate brief interruptions in GPS signal or network connectivity; however, it does not include advanced mechanisms such as temporal interpolation, intelligent recovery of missing packets, or adaptive data quality management.
An internal module dynamically calculates the caloric expenditure based on the MET (Metabolic Equivalent of Task) value, personalized according to weight and type of activity performed. Simultaneously, the environmental module associates the user’s movement data with atmospheric pollutant values (PM10, PM2.5, NO2, CO, C6H6, O3) from the nearest ARPA station, enabling the construction of an integrated overview of physical and contextual conditions in which daily activity takes place. This integration of mobility, physiological parameters, and environmental data provides a high-resolution informational basis for studying correlations between cognitive functioning and movement characteristics. The approach adopted by the Care-MOVE application, multidimensional and data-driven, allows not only for the characterization of individual motor activity profiles but also for the identification of behavioral patterns and variations potentially indicative of cognitive decline or changes in health status. In this perspective, the analysis of data collected through the Care-MOVE application represents a fundamental step towards the development of personalized predictive models, based on artificial intelligence, capable of correlating cognitive performance indices with features derived from movement sensors and environmental indicators. This approach fits within the broader vision of proactive digital healthcare, integrating smartphone technology, machine learning, and neuropsychological assessment to promote continuous, ecological, and non-invasive monitoring of cognitive well-being.
3. Study Design, Participant Screening and Enrollment Procedures
The study was structured into three main phases. The first phase focused on the development of a companion smartphone application (Care-MOVE application). The second phase involved the assessment of the neuropsychological and neurocognitive characteristics of the recruited sample. Finally, the third phase consisted of the experimental evaluation of the system in real-world social contexts.
The study was conducted in 2025 and involved the recruitment of 168 patients, all of whom underwent a comprehensive screening procedure. The main inclusion criteria were a minimum age of 65 years, preserved functional autonomy, absence of clinically relevant diagnosed neurological or psychiatric disorders, and ownership of an Android-based smartphone. Following enrollment, participants were administered a battery of psychological tests, a sociological questionnaire, and a medical–clinical interview focusing on medication use.
For each enrolled patient, the Care-MOVE application was installed free of charge on their personal Android smartphone. Given the collection and processing of large volumes of sensitive health-related data, high security standards and strict data protection protocols were implemented in full compliance with the General Data Protection Regulation (GDPR). All study procedures received approval from the competent Ethics Committee. Before taking part in the study, participants provided their informed consent. To ensure data anonymity, each participant was assigned a unique anonymized numerical identifier (UID), allowing for secure linkage of personal and clinical data without direct traceability to individual identities.
4. Smartphone Application: Care-MOVE Dataset and Architecture
The Care-MOVE smartphone application was designed following a comprehensive requirements elicitation phase and an iterative, user-centered development methodology. High-fidelity wireframes and functional prototypes were produced to support the design of an intuitive user interface, incorporating secure authentication mechanisms, multimodal sensor data acquisition (including GPS, accelerometer, and pedometer data), and automated background data transmission at 15-minute intervals, in accordance with Android operating system limitations. All data collection procedures were conducted in full compliance with the GDPR and were approved by the relevant Ethics Committee, which oversaw the entire study to ensure the protection of participants. All subjects provided explicit informed consent, and the data were pseudonymized using unique identifiers (UIDs), with no possibility of directly tracing the data back to individual participants. The monitoring framework was designed to be passive and non-intrusive, minimizing the level of required user interaction and, consequently, the cognitive burden on participants, thereby improving study adherence. GPS data were used exclusively to characterize the type of motor activity, rather than to track personal locations.
A dedicated module based on the
Metabolic Equivalent of Task (MET) framework was integrated to estimate caloric expenditure dynamically, accounting for individual user characteristics. Usability evaluations conducted with elderly participants informed successive refinement cycles, leading to improvements in interface accessibility, system responsiveness, and overall user experience. All development phases were systematically documented to ensure traceability and reproducibility. In parallel, a web-based dashboard was implemented to collect, integrate, and visualize patient activity data in conjunction with contextual environmental indicators, such as air quality metrics. The dashboard was built upon a secure, modular cloud infrastructure, featuring robust API-based communication and near real-time data updates. Interactive visualization tools allow users to explore and filter historical data through dynamic graphical representations, while synchronization mechanisms and validation routines preserve data integrity and support downstream analytical workflows. The Care-MOVE system adopts a client–server architecture. The client component consists of an Android mobile application developed using Dart and the Flutter framework, capable of continuously collecting georeferenced data, including during background operation, and transmitting them to the backend services. Participant anonymity is ensured through the assignment of incremental, non-identifiable user IDs. The server-side infrastructure is implemented using Python-based microservices deployed via Docker containers, with data persistence managed through a NoSQL MongoDB database. MongoDB’s document-oriented BSON data model enables flexible schema evolution, horizontal scalability, and high-throughput read/write operations, thereby facilitating efficient real-time data management and future system expansion. Overall, this architecture ensures modularity, scalability, and robustness of the Care-MOVE platform (
Figure 1). The blocks in
Figure 1 define the implementation pipeline. The first block addresses the design and implementation of the client side. Subsequently, the backend was developed using a Docker-based microservices architecture. The final block describes the MongoDB database implemented on the backend side.
The system employs RabbitMQ as a message broker to enable asynchronous communication between microservices. When the client collects geolocation and activity data, messages are queued and distributed to microservices responsible for processing, filtering, storage, and analysis. Processed data are stored in MongoDB and made available for visualization on the dashboard, ensuring the reliability and scalability of the infrastructure.
The Care-MOVE application monitors multiple user behavior parameters in real-time:
Module 1: GPS Location,
Module 2: Movement State with flutter_background_geolocation (Still, Walking, Running, On_Bicycle, In_Vehicle),
Module 3: Caloric Expenditure,
Module 4: Air Quality.
Modules 1 and 2 utilize the
flutter_background_geolocation plugin, which allows continuous tracking even when the application is closed, automatically detecting the user’s movement state. Module 3 calculates the calories burned using the formula:
MET values are assigned based on activity type (
still = 0, walking = 2, running = 7, on_bicycle = 4, in_vehicle = 0), while user weight is entered during registration, and activity duration is internally calculated. This method provides immediate, personalized estimates of daily caloric expenditure, useful both for individual monitoring and for aggregated analyses. Module 4 integrates environmental data via the public
ARPA Puglia API, including real-time measurements of
PM10,
PM2.5,
CO,
NO2,
C6H6, and
O3. The user’s GPS coordinates allow the system to automatically select the nearest monitoring station, correlating air quality with daily mobility. This integration enables contextualized analyses of environmental and physical conditions, enriching the interpretation of behavioral data. The data were obtained from the official
ARPA Puglia portal. By integrating this information into the Care-MOVE application, it is possible to cross-reference the user’s mobility data with environmental data, enhancing analyses and enabling more comprehensive and informed monitoring of health status. The overall system architecture (
Figure 1) was designed to ensure maximum privacy and security, in full compliance with the General Data Protection Regulation (GDPR). Data were anonymized using an incremental user ID and transmitted to the server in real-time. During access, essential information such as the user’s weight and health references is requested, and informed consent is presented with a link to the privacy policy.
Examining the application in more detail, Care-MOVE primarily features two main screens. Upon first launch, the initial screen displayed is the login screen. On the first screen, the user is prompted to enter their weight (for calorie calculation) and the reference medical ward associated with the patient. The options available are “Internal Medicine—Outpatient” and “Internal Medicine—Inpatient”. Finally, the screen requests informed us of consent. Once login is completed, the application prompts the user to grant a series of permissions required for the proper functioning of Android’s security and privacy policies.
These include:
Permission to receive notifications;
Consent for continuous background geolocation access;
Permission to read physical activity data (e.g., movement);
Permission to run the application persistently (even when not actively used or when the screen is off).
These steps are essential to ensure continuous, accurate, and non-intrusive monitoring of the user’s motor activity habits. After granting all the required consents, the user is directed to a clean and minimalist home page, where they can view their unique ID, real-time physical activity, and air quality data retrieved from the ARPA Puglia station closest to their current location (
Figure 2).
Motor activity data are recorded with a dynamic frequency as soon as the user begins moving. Step data, on the other hand, are automatically synchronized with the server every 15 min, ensuring continuous data collection even during minimal use or when the application is running in the background. At this stage, the configuration is fully completed, and the user can close the application entirely. No further actions are required on their part. From this point onward, the application remains fully operational. If the user turns off the device, the application will resume functioning as soon as movement is detected upon restart. Before initiating the patient recruitment phase and large-scale data collection, the development team conducted an intensive preliminary testing phase to verify system stability, reliability, and compatibility across multiple devices. Specifically, the application was installed and tested on numerous Android smartphone models from various market segments and brands (Motorola, Xiaomi, Samsung, etc.) to evaluate potential behavioral differences and the impact of different Android versions, including manufacturer customizations and background processing limitations. These tests particularly focused on the application’s ability to continuously collect data, even in cases of forced closure or “kill” while maintaining uninterrupted communication with the server. Any anomalies detected (e.g., interruptions in data collection or loss of synchronization) were analyzed, corrected, and verified in subsequent test iterations, progressively improving client-side system performance. In parallel, backend load testing was performed to simulate the behavior of many patients sending data to the server simultaneously. These tests aimed to evaluate the scalability and robustness of the microservices architecture, ensuring that the system could correctly handle the continuous incoming data flow without packet loss, significant slowdowns, or errors during storage, processing, and visualization. To date, an initial cycle of patient recruitment (168) has been completed, enabling the start of the initial data collection phase within the Care-MOVE application. All data generated by registered patients are securely stored on the server, which serves as a central repository for continuous and structured monitoring of physical activity and related environmental contexts.
The currently available Care-MOVE dataset represents an initial experimental version, already sufficiently rich and structured to support exploratory analyses. Each recorded data entry contains several key attributes for subsequent evaluation, including:
UID: unique and anonymized numerical identifier for the patient;
Timestamp: precise date and time of measurement;
Latitude and Longitude: GPS coordinates of the user’s location;
Activity: detected motor state (e.g., walking, running, vehicle, stationary);
Calories (application): estimated calories calculated directly by the application;
Referrer: the medical department or facility associated with the patient;
Weight: value entered during registration for calorie calculation;
ID and name of the nearest ARPA station;
Pollutant values: measurements from that station (PM10, PM2.5, NO2, etc.);
Air Quality Index and class: according to the official classification;
Calories (server-side): recalculated calorie estimates on the backend using more complex models.
This version of the dataset provides a solid foundation for studying correlations between physical activity, environmental exposure, and personal profiles, as well as for training future predictive models.
Table 1 reports the features previously described, extracted from the application for each patient. In this case, the data shown specifically refer to the patient with UID 6.
The development of algorithms and analytical processes focused on creating artificial intelligence models for interpreting sensor data and identifying activity patterns has been carried out. Techniques such as data pre-processing, exploratory data analysis, and machine learning, including regression and decision trees, were applied to build predictive and personalized models. Anomaly detection and trend analysis further support clinical decision-making, with all procedures extensively documented in technical reports.
The components developed during this phase include the human activity recognition module, which was implemented using the activity recognition functionality provided by the flutter_background_geolocation plugin. During user activity monitoring, algorithms were employed to automatically calculate calories burned.
Additionally, the air quality module was developed, enabling data extraction from the public ARPA Puglia API. This API provides a variety of pollution-related measurements, including PM10, PM2.5, CO, NO2, C6H6, and O3, collected from monitoring stations distributed across the Apulia region.
This phase also included testing related to data analysis, where the collected dataset will be applied. Specifically, anomaly detection algorithms will be used to identify irregularities in patient movements as well as fluctuations in air quality.
Modules 1 and 2 currently use the flutter_background_geolocation library, one of the most advanced solutions for automatic mobility data collection via smartphone devices. This library integrates artificial intelligence mechanisms designed to run directly on the user’s device without requiring a constant cloud connection, with strong emphasis on energy optimization. Specifically, the plugin’s algorithm analyzes in real-time the signals from various onboard sensors (accelerometer, gyroscope, GPS module) to automatically detect and classify the user’s movement state. Recognized classes include “still”, “walking”, “running”, “on_bicycle”, and “in_vehicle”, with classification performed by internal models based on temporal patterns in the sensor data. The flutter_background_geolocation library leverages the native Activity Recognition services provided by iOS and Android. On Android, it uses Google Play Services’ Activity Recognition API via ActivityRecognitionClient. That API periodically collects bursts of low-power sensor data (primarily accelerometer and composite sensors like “step detector”, “tilt detector”, etc.) and returns a list of standard activity types IN_VEHICLE, WALKING, RUNNING, ON_BICYCLE, ON_FOOT, STILL, UNKNOWN, each with a confidence level (0–100). Flutter_background_geolocation wraps and orchestrates these system services, allowing for the configuration of parameters such as activityRecognitionInterval (sampling interval), minimumActivityRecognitionConfidence (confidence threshold), stopDetectionDelay, desiredAccuracy, and more. The primary processing for classification is handled by Apple’s Core Motion libraries and the proprietary models provided by Google Play Services.
Module 4 extends the application by integrating real-time environmental data sources, such as the ARPA Puglia public API, to obtain air-quality information from the region’s monitoring stations. While it does not employ AI techniques to extract raw data, the module uses advanced geospatial matching algorithms to accurately associate the user’s GPS coordinates with the nearest ARPA station. To this end, it uses Geopy’s geopy. distance. geodesic function, which computes the geodesic distance on the WGS-84 ellipsoid. Internally, this function relies on Charles Karney’s GeographicLib, which implements a numerically stable algorithm derived from Vincenty’s 1975 formulas and improved by Karney in 2013. The algorithm used is Karney’s robust, iterative version of the Vincenty formulas for solving the inverse geodetic problem (i.e., given two points, compute the distance on the ellipsoid surface). The WGS-84 ellipsoid model, defined by its equatorial and polar semi-axes and flattening, provides sub-millimeter accuracy in many scenarios, far superior to a simple spherical model.
In the experimental context, we have so far collected data from 168 patients and published the preliminary dataset as a CSV file. The currently available dataset represents an initial experimental version, already sufficiently rich and structured to support initial exploratory analyses. Each record contains key attributes for further evaluation.
5. Experimentation
For the experimental phase of the study, a total sample of 168 patients was selected from the Department of Internal Medicine at Policlinico di Bari, distributed across two operational units:
All patients were over 65 years of age and presented with various chronic clinical conditions, primarily cardiovascular, metabolic, and immunological disorders. The gender distribution showed a slight male predominance (52% male, 48% female), like the overall cohort (51.8% male, 48.2% female). Among the inpatients, a higher prevalence of comorbidities and a greater degree of dependence in personal and social activities were observed, whereas outpatients primarily exhibited stabilized chronic conditions and higher functional autonomy.
Of the 168 patients initially recruited, 53 actively participated in the digital monitoring phase by using the Care-MOVE application continuously for a period of five days. Accordingly, the present study focused on this subsample of 53 subjects. This selection was not arbitrary but was motivated by the specific aim of the work, which was conceived and presented as an exploratory feasibility (proof-of-concept) study. The objective was to assess whether passive signals derived from smartphone sensors contain discriminative information related to cognitive status, rather than to provide a definitive estimate of clinical performance. Consistent with this perspective, the results are not intended to be generalizable to the broader clinical population but are instead presented as indicative of the feasibility of the proposed approach within a limited and well-defined cohort. With respect to the limited demographic diversity and the absence of an explicit analysis of comorbidities, these factors may influence motor behavior and act as potential confounders. However, the primary objective of this study was not to model the effects of clinical or demographic variables but to conduct a preliminary evaluation of sensor-derived signals in a controlled setting.
The 5-day observation window represents a compromise between (i) the need to collect sufficient data to estimate routines and intra-day variability, (ii) maintaining adequate participant adherence, and (iii) ensuring the clinical and operational sustainability of the data collection protocol. The obtained results suggest that, despite the limited observation duration, activity patterns are sufficiently stable to enable cognitive state classification, supporting the utility of the proposed approach in an exploratory setting. Previous pilot studies and real-world applications of digital phenotyping indicate that short but continuous observation windows can be sufficient to capture habitual mobility and activity behaviors while simultaneously reducing participant burden and the risk of dropout, an aspect that is particularly relevant in elderly or vulnerable populations. The limited temporal window may be influenced by transient factors, such as acute health conditions, mood state, or environmental variables (e.g., weather conditions), which may introduce variability not directly related to cognitive status. For this reason, the results should be interpreted as indicative of a short-term association, rather than as descriptive of long-term cognitive decline dynamics.
For each participant, all numerical columns available in the CSV file were used across all ML and deep learning models to obtain a representation of the data that was as comprehensive and informative as possible. These typically include:
Geographic coordinates (latitude, longitude), which implicitly describe the evolution of the user’s position over time and, consequently, the associated spatial trajectory;
Physical activity variables (activity), originally categorical and subsequently transformed into numerical values using label encoding. The considered activities include still, in vehicle, in bicycle, running, and walking, allowing different levels and modes of movement to be distinguished;
Energy expenditure indicators, such as calories and calories_server, which provide an estimate of the user’s caloric consumption, computed locally on the smartphone and server-side, respectively. Caloric values computed on the server are generally more accurate;
Static individual characteristics, such as weight, which represent time-invariant personal attributes that are nonetheless relevant for the interpretation of dynamic behavioral data.
The integration of these variables enables the combination of spatial, behavioral, and energetic aspects with individual characteristics, resulting in a rich and structured data representation for subsequent analysis and modeling stages.
The environmental data are incorporated as contextual (spatiotemporal) covariates aligned with mobility data. However, to substantiate the claim of “integration”, it is important to also demonstrate their informational contribution. Although air quality data were collected and associated with the individual users’ trajectories, their integration into the current classification framework remains limited and indirect. In particular, the environmental information available in the original files is largely represented as textual descriptions (e.g., pollutant values and categorical classes), which are not automatically converted into numerical variables and therefore do not explicitly contribute to the training of the predictive models.
Aggregated data analysis revealed that:
Time spent in_vehicle accounted for approximately 54% of monitored time;
Time spent walking averaged 22%, varying according to clinical status and functional autonomy;
On_bicycle usage was minimal (≈2%), while running was almost absent (<0.2%);
Still (inactivity or stationary periods) accounted for approximately 22%.
For each subject, these variables are organized into a temporal sequence, which is either truncated or padded to a predefined maximum length. In deep learning models (CNN, Transformer, LSTM, GRU), the full sequence is provided directly as input, preserving the temporal structure. In contrast, for classical machine learning models, the sequence is flattened into a single feature vector, thereby discarding explicit temporal information while enabling the use of traditional classifiers.
Regarding feature normalization, no single global strategy was applied. Distance- or margin-based models (e.g., SVM and KNN) incorporate feature standardization through StandardScaler, whereas other models (Logistic Regression, Random Forest, Extra Trees, Gradient Boosting) directly operate on raw features. In deep learning models, no explicit pre-normalization of numerical features is applied; however, some architectures include Layer Normalization layers, which help stabilize training but do not replace a full statistical normalization of the input variables. Overall, the proposed approach relies on a raw multivariate temporal representation of user behavior, based on GPS trajectories, encoded activity labels, and energy expenditure indicators, without explicit extraction of higher-order motion features or a full quantitative integration of environmental data. This design choice intentionally reduces preprocessing complexity and preserves the original structure of the data.
Regarding energy expenditure, the average daily caloric consumption ranged between 1500 and 2000 kcal, with substantial individual variation due to age, body weight, and activity levels. Regular use of digital tools (smartphones, monitoring apps, and telemedicine platforms) was associated with:
Improved perceived health status;
Increased personal autonomy;
A broader social network, particularly among patients with medium-to-high socioeconomic status.
Consistent with findings from the broader survey, only a minority (~18%) of patients used digital technologies for continuous health monitoring, although perceived utility remained high (76%). Additionally, 59% of patients reported using the Internet to obtain health information or remote medical consultations, highlighting the growing potential of home-based digital healthcare. These findings highlight both the current engagement of older adults with digital health tools and the potential for more structured monitoring solutions. Building on this context, the experimental phase utilized the Care-MOVE dataset, which comprises a structured collection of files, each corresponding to a single participant identified by a unique ID. Each file contains a time series of data collected during behavioral and environmental monitoring sessions.
The experiment was designed to systematically evaluate the ability of different models, both traditional and deep learning, to correctly classify subjects based on their sensory and behavioral data. This was conducted to verify whether there is a correlation between equivalent cognitive functioning scores and features from the Movement Dataset.
The cognitive assessment indicates that most patients exhibited preserved global cognitive functioning, as reflected by Mini-Mental State Examination (MMSE) scores within the normal range, along with adequate processing speed measured through the Trail Making Test (TMT-A and TMT-B). In contrast, executive functions appeared mildly compromised, with slightly below-normative performance observed on the Frontal Assessment Battery (FAB) and on the TMT B–A difference index. From a clinical perspective, the most prevalent comorbidities involve the cardiovascular system (41%), followed by immune-related (15%) and metabolic conditions (14%). Furthermore, approximately 10% of the sample presented moderate to severe limitations in personal and social autonomy.
The adopted binary classification was constructed based on the Equivalent Scores (ES) of neuropsychological tests, which represent an ordinal measure allowing patient performance to be compared against the normative population. In the analyzed sample, no subjects with ES = 0, corresponding to significantly impaired performance, were present. Therefore, the aggregation of ES = 1–2 and ES = 3–4 does not represent an arbitrary simplification, but rather a choice consistent with both the empirical distribution of the data and the clinical interpretation of the Equivalent Scores. Specifically, ES = 1–2 reflect performance at the lower limits of normality or in the low-average range, whereas ES = 3–4 indicate cognitive functioning in the medium-to-upper range of the normative distribution, thus corresponding to overall preserved cognitive functioning. The binary classification is therefore aligned with the exploratory objective of the study and is not intended to replace a detailed clinical assessment, which may constitute the focus of future work. In this sense, dichotomization was employed as an exploratory step to test the sensitivity of the system to macroscopically distinct cognitive states, representing a realistic scenario for screening or preliminary triage applications. From a computational perspective, the distribution of subjects across the four clinical levels was highly imbalanced, making the training and reliable evaluation of multi-class models challenging and prone to overfitting or highly unstable performance estimates. With the expansion of the dataset, future versions of the model could be formulated as four-class classification tasks or ordinal models, which would be more consistent with the clinical progression of cognitive decline and would enable a more fine-grained and clinically meaningful assessment of the different disease stages [
38].
Each participant was assigned a score (Global_PE) through a psychological assessment that reflects their level of cognitive impairment, divided into four classes:
1: No impairment;
2: Mild impairment;
3: Moderate impairment;
4: Severe impairment.
However, to simplify the experiment and strengthen the analysis, these four classes were grouped into two categories:
The first category combines scores 1 and 2, representing no or mild impairment;
The second category combines scores 3 and 4, representing moderate or severe impairment.
This transformation converts the problem into a binary classification task, which is more suitable for testing our Machine Learning models and assessing whether the data collected from smartphones contain informative signals to distinguish between relatively stable cognitive conditions and more severely compromised states.
The approach adopted was Leave-One-User-Out (LOUO) validation, a technique particularly suited to biomedical or behavioral fields, where the aim is to measure the model’s ability to generalize to patients never seen during training. For each iteration of the experimental process, data from only one user are excluded and used as the test set, while data from all other patients constitute the training set. The cycle is repeated as many times as there are available patients, so each user is used once as an independent test. Before training, each user’s CSV file (53 old patients) is uploaded and preprocessed: activities are numerically encoded using LabelEncoder, non-numeric or missing columns are zero-impacted, and temporal data sequences are truncated or padded to a fixed length of 200 timesteps. On average, each participant contributed approximately 30,318 data rows. Under this setting, there was no overlap between temporal windows or samples from the test participant and those used for training, as data were loaded, indexed, and processed separately at the patient level. This approach is generally regarded as one of the most conservative validation strategies in biomedical applications involving longitudinal signals as it provides a realistic estimate of generalization performance to unseen subjects.
To further minimize the risk of data leakage, we clarify that:
Training and test sets are split by UID prior to any training step;
The activity LabelEncoder is fitted on the complete set of activity labels present in the files, including the “unknown” class; this operation does not rely on clinical information, nor does it incorporate any target-related information;
For models requiring normalization (e.g., SVM), standardization is implemented within a pipeline and performed exclusively on the training set of each fold.
The values are then converted into arrays for use with neural networks. The experimentation includes extensive benchmarking that compares different types of models:
Shallow learning models, such as Logistic Regression, Random Forest, Gradient Boosting, Support Vector Machine (SVM RBF), and K-Nearest Neighbors (KNN), which operate on “flattened” representations of the data (sequences transformed into one-dimensional vectors).
Deep learning models, such as 1D CNN, Transformer, LSTM, and GRU, which instead work directly on temporal sequences, capturing local or long-term relationships in the data.
Consequently, fixed and consistent hyperparameter settings were adopted for all neural architectures across all LOUO folds, selected based on well-established configurations in the literature and considerations of computational stability. All deep learning models take as input multivariate temporal sequences with a maximum length of 200 samples; shorter sequences were standardized via post-padding. Training was performed using the Adam optimizer, with binary cross-entropy as the loss function and a batch size of 8. The maximum number of epochs was set to 40, with 15% of the training data reserved as an internal validation set. To mitigate overfitting, an early stopping mechanism with a patience of 5 epochs was applied, restoring the weights corresponding to the minimum validation loss.
The 1D CNN architecture consisted of two convolutional layers with 64 and 128 filters, respectively, a kernel size of 3, and ReLU activation, each followed by a max-pooling layer with a pooling factor of 2. The convolutional block was followed by a fully connected layer with 64 units, ReLU activation, and a dropout rate of 0.3. The output layer consisted of a single unit with sigmoid activation for binary classification.
The Transformer architecture employed a single self-attention block with 4 attention heads and a key dimensionality equal to the number of input features. This was followed by two dense layers with 128 and 64 units, respectively, both using ReLU activation and dropout regularization (0.3 and 0.2). Temporal aggregation of representations is achieved through global average pooling.
The LSTM and GRU architectures share a similar structure. Both employ two stacked recurrent layers, with 128 units in the first layer (returning the full sequence) and 64 units in the second layer. Dropout is applied at each level (0.3 for the first layer and 0.2 for the second). The final representation is further transformed through a fully connected layer with 64 units, ReLU activation, and a dropout rate of 0.2, prior to the sigmoid-activated output layer.
In all deep learning architectures, class weighting is dynamically computed on the training data of each LOUO fold to compensate for class imbalance. Moreover, the temporal order of the sequences was preserved by disabling data shuffling during training.
Regarding probability calibration, the predictions produced by the deep learning models were not subjected to explicit post-calibration procedures; therefore, the resulting probabilities should primarily be interpreted as discriminative scores. We acknowledge that a formal calibration assessment represents a relevant extension of the present work, particularly for clinical applications. Deep learning models were included primarily due to the high temporal granularity of the data: each participant contributes, on average, 30,317.79 longitudinal observations. However, it is well-known that such models typically require larger sample sizes to learn robust temporal representations without overfitting. For this reason, the present work was conceived and presented as an exploratory feasibility (proof-of-concept) study, rather than a definitive clinical validation. During each iteration, each model was trained on the training set and then evaluated on the excluded user. Predictions and class membership probabilities were saved for each model. At the end of the LOUO validation, all predictions were aggregated, and average performance metrics were calculated: accuracy, precision, recall, and F1-score. These measures provide a comprehensive and robust assessment of each model’s performance, allowing us to identify which approaches are most effective and generalizable in the experimental context.
In the following section, we present a detailed evaluation of the machine learning and deep learning models applied to the Care-MOVE dataset. Model performance was assessed using multiple metrics, including Accuracy, Balanced Accuracy, Precision, Recall, F1-score, and Matthews Correlation Coefficient (MCC), providing a robust and unbiased view of predictive capabilities. To quantify the statistical stability of the results, all metrics will be reported alongside 95% confidence intervals computed via bootstrap resampling, highlighting both the reliability and variability of each model.
6. Results and Observations
The results obtained in
Table 2 highlight a clear distinction between the predictive power of traditional models and that of deep learning models while also revealing the high complexity and structure of the Care-MOVE dataset. The SVM with RBF kernel emerged as the clearly best-performing model, achieving an accuracy of 98.1% and nearly identical precision, recall, and F1 scores, demonstrating an optimal balance between sensitivity and specificity. This result suggests that despite dealing with temporal data, the class separability in the corresponding feature space is highly nonlinear but well-captured by the kernel function. In contrast,
Random Forest showed a performance below 86%, indicating that, although effective at handling complex relationships between static features, it failed to fully model the dynamic component of behavioral data. Deep learning models, particularly
LSTM and Transformer, showed lower performances, ranging from 79% to 81%, but maintained a good balance in metrics, indicating that they can capture some temporal dependencies, although likely suffering from the limited size of the dataset and the strong inter-subject variability typical of LOUO validation.
1D CNNs and
GRUs showed comparable results (75%), suggesting that the sequential structure of the signal does not provide sufficiently regular patterns to take full advantage of convolutional or recurrent architectures (
Table 2).
The analysis confirms that cross-user generalization is a non-trivial task and that the spatial and contextual representation of features plays a more decisive role than temporal dynamics. The superiority of the SVM indicates that correlations between behavior, physiological parameters, and environmental quality manifest themselves in rather stable and distinctive feature configurations, rather than in complex temporal sequences. This suggests that in the domain of cognitive and behavioral monitoring, hybrid models capable of combining the discriminative power of kernel methods with the temporal pattern modeling capability of deep networks may represent the most promising direction for future research.
The high correlation observed between cognitive functioning equivalent scores and movement dataset features in subjects over 65 can be interpreted as the result of the close interconnection between the motor, physiological, and cognitive domains that characterize aging. Behavioral and physiological variables detected by sensors, such as accelerometers, GPS, or wearable devices, provide an indirect but highly sensitive measure of an individual’s global functional status, including motor, motivational, and cognitive components. In older patients, cognitive decline is not manifested exclusively through memory or attention deficits but is often accompanied by subtle and progressive changes in mobility patterns and daily habits. Individuals with impaired cognitive profiles tend, for example, to exhibit less spatial variability in their movements, more repetitive and predictable routes, a reduction in the time spent in dynamic activities, and, in general, a lower overall level of physical activity than their cognitively intact peers [
39,
40].
These differences are directly reflected in the features extracted from the movement data. Parameters such as the quantity and intensity of activities, the frequency of posture changes, and the estimated daily energy expenditure show systematically different trends between the two groups. Temporal analysis of the data sequences also highlighted aspects related to the organization of daily routines: subjects with cognitive impairment exhibited more rigid or, conversely, disorganized patterns, both indicative of reduced planning capacity. The use of the LOUO validation procedure strengthened the robustness and generalizability of the results, demonstrating that differences in movement behaviors do not depend on idiosyncratic individual characteristics but represent consistent patterns within the population studied.
The strong relationship between mobility parameters, energy expenditure, and quality of living environment reflects the systemic nature of human functioning, where body, behavior, and mind interact dynamically and reciprocally. This explains why, in elderly individuals, movement characteristics constitute a particularly sensitive indicator of cognitive status, potentially also useful for the early screening of cognitive decline and for longitudinal monitoring of functional health.
This section reports the performance of the evaluated Machine Learning and Deep Learning models on the proposed classification task. Results are presented using multiple complementary metrics to provide a robust and unbiased assessment, particularly in the presence of class imbalance. Performance variability was quantified by means of 95% confidence intervals computed via bootstrap resampling, allowing for an estimation of the statistical stability of each metric.
Table 3 summarizes the best-performing models according to Balanced Accuracy, together with Accuracy and Matthews Correlation Coefficient (MCC). For each metric, the point estimate and the corresponding 95% confidence interval are reported. Balanced Accuracy was selected as the primary ranking criterion, as it equally weights sensitivity across classes and is therefore more informative than plain Accuracy in imbalanced scenarios. MCC was included as a complementary measure of global prediction quality, as it accounts for all entries of the confusion matrix and provides a correlation-based assessment.
The Support Vector Machine with RBF kernel achieved the highest performance across all reported metrics, substantially outperforming both classical ensemble methods and deep learning architectures. In particular, the near-perfect Balanced Accuracy (0.988) and MCC (0.942) indicate an excellent separation between the two classes and a highly consistent prediction behavior across subjects. The extremely narrow confidence intervals further suggest strong robustness and low variance across bootstrap resamples. This result suggests that the underlying feature space is highly discriminative and well-suited to margin-based classifiers. The RBF kernel appears particularly effective in capturing the nonlinear decision boundaries required by the task, even in the presence of inter-subject variability. Tree-based ensemble models (Random Forest and Extra Trees) showed solid and stable performance, achieving balanced accuracies close to 0.69 and MCC values above 0.50. These models demonstrated strong sensitivity toward the majority class while partially struggling with minority-class recall, as reflected in their confusion matrices. Among the Deep Learning approaches, recurrent architectures (LSTM and GRU) consistently outperformed convolutional networks. This trend suggests that temporal dependencies or sequential patterns within the data play a more relevant role than purely local representations. The Transformer-based model achieved the best overall performance among neural approaches, reaching a Balanced Accuracy of 0.703. This indicates that attention-based mechanisms may provide additional benefits in modeling complex inter-feature relationships, although their performance remained inferior to that of the SVM in this experimental setting.
Models exhibiting high Accuracy but lower Balanced Accuracy (e.g., KNN and Logistic Regression), shown in
Table 3, were observed to favor the majority class, leading to poor discrimination of the minority class. This effect highlights the importance of using class-balanced metrics and confirms that Accuracy alone would provide a misleading assessment of model quality. The MCC metric further corroborates these findings, as models with unbalanced predictions tend to yield MCC values close to zero or even negative, despite acceptable Accuracy levels. An inspection of confusion matrices revealed that misclassifications were not uniformly distributed across subjects. Certain individuals were consistently misclassified across multiple models, suggesting the presence of subject-specific patterns that are harder to generalize. Conversely, the SVM demonstrated a remarkable ability to correctly classify nearly all subjects, with only a single false negative observed. This behavior indicates that model performance is not solely driven by class frequency but also by the intrinsic separability of subject-level representations. High performance was primarily observed for the SVM with RBF kernel, whereas deep learning models (CNN, Transformer, LSTM, GRU) exhibited lower performance. This outcome is consistent with both the relatively simple nature of the binary task (0/1) and the pronounced class imbalance in the dataset, which may lead to high accuracy values even in the presence of a limited number of misclassifications.
7. Conclusions and Future Work
This study demonstrates the feasibility and potential of the Care-MOVE smartphone-based application for continuous, unobtrusive monitoring of mobility, behavioral patterns, and environmental exposure in older adults. By integrating passive sensor data derived from everyday smartphone use with standardized neuropsychological assessments, the proposed framework enables the collection of detailed, contextualized, and ecologically valid information on daily functioning in real-world settings.
The experimental evaluation, conducted on a subsample of 53 participants aged over 65 and monitored continuously for five days, provides clear evidence of a strong association between movement-derived features and cognitive status. Features related to activity distribution (e.g., time spent walking, stationary, or in vehicle), spatial mobility, and estimated energy expenditure emerged as informative indicators of global functional and cognitive condition. Despite the relatively short observation window, the stability of activity patterns proved sufficient to support reliable classification, confirming the sensitivity of mobility-related signals to cognitive differences in elderly populations.
The comparative analysis of Machine Learning and Deep Learning models under a conservative Leave-One-User-Out (LOUO) validation framework highlighted marked differences in generalization performance. Traditional machine learning approaches, and in particular the Support Vector Machine with RBF kernel, substantially outperformed both ensemble-based methods and temporal deep learning architectures. The SVM achieved an accuracy of 98.1%, a balanced accuracy of 0.988, and an F1-score of 0.981, accompanied by a Matthews Correlation Coefficient (MCC) of 0.942, indicating an excellent and well-balanced discrimination between cognitive classes. The narrow confidence intervals associated with these metrics further suggest strong statistical stability and robustness across bootstrap resampling.
Tree-based ensemble models, such as Random Forest and Extra Trees, showed solid but lower performance, with balanced accuracy values around 0.69 and MCC values slightly above 0.50, reflecting good sensitivity to the majority class but reduced capability in minority-class discrimination. Deep Learning architectures, including LSTM, GRU, Transformer, and 1D CNN models, captured aspects of the temporal structure of the data but exhibited lower balanced accuracy and MCC values, ranging approximately between 0.26 and 0.42. These results indicate limited robustness and reduced cross-subject generalization, likely due to the small number of subjects and the high inter-subject variability intrinsic to the Leave-One-User-Out setting.
The performance profile across metrics confirms that the discriminative information required to distinguish cognitive conditions is primarily embedded in stable, nonlinear feature configurations rather than in complex temporal dependencies. The use of Balanced Accuracy and MCC proved essential for a reliable assessment of model quality, avoiding misleading interpretations based solely on overall accuracy in the presence of class imbalance.
For these reasons, the present work is explicitly designed and presented as an exploratory proof-of-concept study rather than a definitive clinical validation. Nevertheless, the integrated Care-MOVE approach not only enables the collection of rich and contextualized data but also opens new perspectives for proactive digital healthcare focused on risk prevention, improvement of quality of life, and promotion of autonomy in elderly patients. The results support the potential role of smartphone-derived mobility data as a complementary digital biomarker to traditional cognitive assessments, contributing objective and continuous information to support personalized and preventive care strategies.
Consistent with its exploratory proof-of-concept nature, several directions for future research emerged directly from the limitations and design choices of the present study. A primary extension concerns the temporal dimension of monitoring. Expanding the observation window from a few days to longer longitudinal periods will be essential to disentangle short-term variability from stable behavioral patterns and to better characterize cognitive trajectories over time. Longer monitoring periods would also mitigate the influence of transient factors such as acute health conditions, mood fluctuations, or environmental contingencies.
A second key direction involves the inclusion of the entire recruited cohort of 168 participants. Increasing the number of subjects will improve the statistical power, enhance model robustness, and enable more advanced analytical formulations. In particular, future studies may move beyond binary classification toward multi-class or ordinal modeling approaches that more closely reflect the clinical progression of cognitive decline, in line with the structure of Equivalent Scores used in neuropsychological assessment.
From a data integration perspective, future work will focus on strengthening the quantitative contribution of environmental information. Although air quality data were collected and associated with mobility trajectories, their use in the current classification framework remains limited. Converting pollutant measurements and indices into fully numerical features and systematically integrating them into predictive models represents a natural and necessary extension of the present work.
Methodologically, future research will explore refined modeling strategies that balance interpretability and temporal sensitivity. The results obtained suggest that hybrid approaches, combining the discriminative power of kernel-based classifiers with lightweight temporal representations, may offer a promising compromise between performance and generalizability. In parallel, improvements in data quality management such as handling missing data, temporal interpolation, and adaptive smoothing will be addressed to enhance robustness in longer-term deployments.
External validation in independent clinical contexts will be required to assess the generalizability and translational potential of the Care-MOVE framework. In this perspective, the system represents a scalable and modular foundation for proactive digital healthcare solutions aimed at continuous monitoring, early risk detection, and the promotion of autonomy and quality of life in aging populations, fully aligned with its role as an exploratory proof-of-concept rather than a definitive clinical validation.