Next Article in Journal
Correction: Chatziioannou et al. Bridging Perceptions: A Comparative Evaluation of Public Space Design Qualities by Experts and Users. Urban Sci. 2025, 9, 412
Previous Article in Journal
Open-Data Decision Support for Critical Medicines Availability in Urban Supply Chains Under Disruptions: Evidence from Kyiv and Lviv
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Multi-Day Activity Pattern Inference Using Constrained Gaussian Mixture Model (GMM) Classification

by
Nikhita Kannam
1,
Mahdieh Allahviranloo
1,* and
Laure Alice Raymonde Vatin
2
1
Department of Civil Engineering, The City College of New York—CUNY, New York, NY 10031, USA
2
ENTPE-Graduate School of Civil, Environmental and Urban Engineering, 69518 Lyon, France
*
Author to whom correspondence should be addressed.
Urban Sci. 2026, 10(6), 331; https://doi.org/10.3390/urbansci10060331
Submission received: 28 April 2026 / Revised: 4 June 2026 / Accepted: 8 June 2026 / Published: 17 June 2026

Abstract

Multi-day travel diaries are often associated with high rates of partial completion, limiting their value for activity-based demand modeling. This paper develops a probabilistic framework that encodes daily activity sequences, clusters them with a Gaussian Mixture Model (GMM) to obtain soft (probabilistic) memberships, and predicts missing days through a constrained Lagrangian regression that guarantees valid probability distributions. Applied to the New York City Citywide Mobility Survey for 2019 and 2022, the soft-clustering approach achieves an RMSE as low as 0.17—substantially outperforming hard-clustering baselines (16–36% accuracy)—and reconstructs population-level time-use profiles with approximately 5–6% mean absolute error. Results show that post-pandemic activity patterns are more home-anchored and less varied, with pronounced socioeconomic divergence in recovery trajectories.

1. Introduction

The COVID-19 pandemic caused major disruptions to urban activity and travel patterns, several of the behavioral shifts that emerged during the pandemic—most notably the widespread move to remote and hybrid work—appear to be lasting rather than being temporary [1,2,3]. These shifts have implications for travel demand and urban mobility that go beyond the pandemic period itself. Comparing daily activity patterns before and after the pandemic is therefore important: it tells us whether cities went back to how things were or whether a new behavioral baseline has emerged that transportation planning needs to accommodate. By looking at citywide mobility survey data from 2019 and 2022, we analyze which aspects of daily behavior shifted, who was affected, and how persistent these changes are. Activity-based travel demand models are based on travel being a derived demand and people travel to participate in activities at different places and times [4,5,6]. A key difficulty in these models is that daily activity patterns vary widely across the population and also change from day to day for the same person [7,8,9].
It is evident that multi-day travel can better capture and predict activity patterns, and that looking at a single day can result in misleading insights [7,8]. Bhat and Koppelman [10] reviewed time-use research in transportation and found that how people spend their discretionary time depends on their social and demographic background and also differs by day of the week. These studies made a clear case for collecting data over multiple days. ALBATROSS model [11] used decision-tree rules learned from observed behavior to simulate daily schedules under spatial, temporal, and social constraints. However, these models used deterministic logic and were not designed to compare observed activity sequences across large populations.
By borrowing the sequence alignment method from bioinformatics transportation modelers, researchers found a way to compare multiple dimensions of activity patterns at once, by slicing the day into several time slots and using letters representing activities performed by individuals [12]. The Levenshtein (edit) distance measures how different two strings (activity chains in case of travel behavior analysis) are by counting the minimum insertions, deletions, or substitutions needed to turn one string into the other [13,14,15]. With these distances in hand, clustering methods such as affinity propagation, k-means, hierarchical clustering, and density-based approaches have been used to find common types of daily patterns in travel diary data [13,16,17,18]. Allahviranloo and Recker [12] applied pattern recognition to personal travel behavior and showed that this data-driven classification can uncover behavioral groups that trip-based summaries miss.
With more mobility data now available, researchers have added richer features and probabilistic models to these frameworks. Dimensionality reduction methods such as PCA, t-SNE, and UMAP have been used to project high-dimensional travel data into low-dimensional spaces for visualization and clustering [19,20]. Gaussian Mixture Models (GMMs) are particularly useful in these reduced spaces because they give membership probabilities rather than hard labels, which better reflects the mixed nature of daily activity schedules [21,22]. Similarly, Sun and Axhausen [23] used probabilistic tensor factorization on transit smart-card data to extract mobility patterns along time-of-day, passenger-type, and spatial dimensions, showing that probabilistic methods can capture mobility structure without forcing people into rigid categories. To analyze the difference in activity patterns across different regions, Allahviranloo and Aissaoui [24] compared time-use across several metropolitan areas in Northern America and found that differences in land use, transit, and demographics lead to distinct population-level activity patterns. This highlights the need for context-specific modeling and motivates our focus on one city observed at two points in time, where the change in context comes not from geography but from the pandemic. Due to significance of pandemic on mobility, several scientists have looked at post-pandemic mobility through the lens of activity participation and daily schedule structure. Su et al. [25] showed that analyzing both motifs and sequences reveals clear differences between telecommuters and regular commuters—about 20% of telecommuters stay home all day on a workday, compared to only 8% of regular commuters, and those who do travel tend to make more trips and follow more complex schedules. Rafiq et al. [26] studied the early pandemic period using an aggregate structural regression model and found that working from home reduced workplace visits, non-work activities tied to work travel, and total miles traveled. Other studies confirm that telework and hybrid work continue to reshape commute choices and raise equity concerns [1,3]. On the methods side, activity-based modeling has moved toward more data-driven approaches, including neural network scheduling frameworks [27], inverse reinforcement learning for sequential activity decisions [28], and hybrid machine-learning extensions of traditional demand models [29]. Bhuiyan and Habib [30] showed that bioinspired sequence alignment with Levenshtein distance can extract representative weekly activity patterns from single-day diaries, while Haghighi and Miller [31] reviewed week-long activity-based modeling and identified the lack of scalable probabilistic frameworks for multi-day pattern recognition as a key gap. These developments motivate a framework that represents activity sequences probabilistically while testing whether post-pandemic changes persist.
Despite this progress, several gaps still remain to be closed and we try to address some of those in this paper. First, most clustering studies use single-day data and hard cluster assignments, which cannot represent days that mix multiple activity types. Multi-week diary studies [32,33] show that while work and school are stable from day to day, discretionary activities vary a lot, making single-label assignments even less suitable. Second, predicting activity patterns for missing survey days—a common problem in multi-day diaries—has received little attention, especially under constraints that guarantee valid probability distributions. Third, the COVID-19 pandemic disrupted urban mobility on a massive scale [34,35,36], with drops in transit ridership, shifts to cars and walking, and the rise of remote work [37,38,39]. In New York City, subway ridership fell by over 90% during the initial lockdown and recovered unevenly across boroughs and demographic groups [40]. Most of this literature focused on aggregate ridership or mode share. Few studies have examined how the full structure of daily activity patterns—the sequence, timing, and mix of activities—changed at the person level, or whether those changes persisted beyond the peak phases of the pandemic [19,36]. We address this last gap by comparing the 2019 and 2022 waves of the New York City Citywide Mobility Survey [41,42]. The comparison follows a before-and-after design: 2019 is the last full pre-pandemic year, with no COVID-19 effects, and serves as the behavioral baseline. The 2022 wave was collected more than two years after the initial lockdowns, by which point emergency-phase disruptions (stay-at-home orders, school closures, service shutdowns) had ended. Any behavioral differences that remain at that point can reasonably be attributed to lasting structural shifts rather than temporary crisis responses. This timing lets us test whether changes in daily activity patterns persisted beyond the acute phase of the pandemic or whether the city went back to its 2019 routines.
This paper develops a probabilistic framework for classifying, imputing, and reconstructing multi-day activity patterns from incomplete travel diary data. We apply it to produce a demographically disaggregated comparison of daily activity structures in New York City before (2019) and after (2022) the COVID-19 pandemic, using the Citywide Mobility Survey data. New York City is a useful setting because it is one of the most transit-dependent cities in the world, with diverse populations and neighborhoods. The availability of multi-day survey data from both periods allows a direct comparison of activity patterns.
The paper is organized around three research questions: (a) How can we effectively impute missing data in multi-day activity diaries while maintaining the probabilistic and multi-pattern nature of daily human behavior? (b) Did the structure of daily activity patterns in NYC (timing, sequence and mix of activities) change between 2019 and 2022, and do these changes indicate a persistent shift toward more home-centered routines? (c) Do the observed changes differ across demographic groups (income, education, gender, age) in ways relevant to transportation planning? Aligned with these questions, the contributions of this paper are threefold:
  • Filling in incomplete survey data. In both surveys, fewer than one-third of participants provided data for all seven days. Instead of dropping these respondents or treating each day separately, we build a sequential prediction model that reconstructs activity profiles for every participant using their observed days and demographic attributes. This produces a complete multi-day dataset that keeps day-to-day dependencies within each person.
  • Soft (probabilistic) clustering with constrained prediction. We fit a K-component GMM to the embedded activity-chain space and keep the full probability vector for every observed day, rather than assigning one hard label. To make sure predicted vectors for missing days are valid probability distributions (non-negative, summing to one), we use a constrained Lagrangian regression model with penalty terms. This soft membership captures the mixed nature of daily activity patterns better than hard clustering.
  • Reconstructing population-level time-use from predicted memberships. We turn each person’s predicted probability vector into a daily activity profile by taking a weighted average of cluster-representative activity chains (encoded as one-hot matrices over 38 half-hour time slots and 11 activity types). Aggregating across people gives population-level time-use distributions with a Mean Absolute Error of about 5–6% compared to observed data.
The remainder of this paper is organized as follows. Section 2 describes the Citywide Mobility Survey data and participant demographics for both years. Section 3 details the methodology, covering activity chain construction, GMM soft clustering, and the constrained Lagrangian prediction model. Section 4 presents results on cluster structures and time-use patterns. We discuss the application and policy implications in Section 5, followed by conclusions in Section 6.

2. Data

This study uses the New York City Citywide Mobility Survey, an annual survey that records travel behavior, daily routines, and recreational activities over multiple days for a representative sample of residents [43]. Each respondent reports trip-level records—origin, destination, mode, purpose, departure and arrival times—along with household and demographic information (age, gender, education, income, etc.) [44]. The 2019 data were collected22 May to 30 June 2019; the 2022 data 28 September to 17 November 2022. The CMS uses address-based sampling, drawing a random sample of residential addresses across 10 survey zones covering all five boroughs [41,42]. In 2019 the survey had 3346 respondents; in 2022, 2966 were recruited through address-based sampling (86%) and re-invitations of 2019 participants (14%). The CMS reports provide unweighted and weighted sample profiles benchmarked against the American Community Survey [41,42].
Respondents could participate via a seven-day smartphone travel diary, a one-day online survey, or a one-day phone interview. Because the imputation framework (Section 3) requires multiple days per person, we keep only smartphone diary participants. For gender, responses coded ‘Other’ or ‘Prefer not to answer’ were retained as separate categories. For education, invalid or missing responses were dropped, while explicit ‘Prefer not to answer’ responses were kept. Records missing other demographic variables were excluded (the number of such cases was small). The final analytic sample has 12,513 person-day records in 2019 and 13,060 in 2022.
Throughout this paper, day 1 through day 7 refer to the order in which a participant reported survey days (participation order), not calendar weekdays. Table 1 summarizes participant demographics. The bottom panel shows how many days each respondent completed. In 2019, 36.3% of respondents completed only one day and 12.5% completed all seven; in 2022, 30.7% completed one day and 15.1% completed all seven. This high rate of partial completion motivates the imputation framework discussed in Section 3.

3. Methodology

Our method has three stages: (1) encode each survey day as an activity chain, (2) fit a Gaussian Mixture Model to find recurring daily patterns, and (3) predict soft-cluster membership vectors for missing survey days using a Lagrangian-constrained optimization. Activity-chain encoding preserves the order and timing of daily activities, so each person’s day is treated as a complete schedule rather than a set of disconnected trips. Levenshtein distance and UMAP embedding let us compare and visualize these sequences while keeping differences in timing and activity ordering intact. This encoding has already been used to compare activity patterns across different metropolitan and to generate time use patterns regions [24] showing that it works with different survey instruments and urban contexts. On top of this established encoding [16], the present paper introduces a GMM soft-clustering layer that gives membership probabilities rather than hard labels—better suited to the mixed nature of daily schedules. The constrained prediction model then allows incomplete multi-day diaries to be kept by imputing missing days while ensuring that the predicted vectors remain valid probability distributions. The encoding layer is proven across regions; the probabilistic clustering and imputation layers introduced here do not depend on any city-specific data structure. Recent studies support this modularity: Ji et al. [45] applied UMAP-based clustering to driver mobility across six US metro areas; Bhuiyan and Habib [30] used Levenshtein-based alignment with hierarchical clustering on multi-day sequences; Purbohadi et al. [46] combined UMAP with density-based clustering on large transit fare-collection data; and Haghighi and Miller [31] reviewed multi-day activity modeling and flagged scalability and transferability as open challenges. The main scalability bottleneck is the pairwise Levenshtein distance matrix, which is O ( N 2 ) . For our dataset (∼13,000 chains) this is fine, but scaling to millions of records requires adaptation—for instance, fitting the pipeline on a representative subsample and then projecting new observations into the learned UMAP space via its transform method, or using approximate nearest-neighbor algorithms to bring the computation down to O ( N log N ) . The prediction model itself is a single linear layer and scales trivially.

3.1. Activity Chain Construction

We convert raw trip records into standardized activity chains following [16]. Each day’s activities are encoded with 11 categories: H (in-home), W (work), S (school), A (escort), M (shopping), L (meal), R (social/recreation), E (errands), C (mode change), N (overnight), and O (other). The day is divided into 30 min intervals from 5:00 AM to 11:30 PM, giving 38 time slots. If more than one activity falls into a slot, the one with the longest duration is kept. Each person who was surveyed over multiple days gets a separate chain for each day. For example, someone surveyed for three days has three chains. We use the term activity chain (or simply chain to mean this single-day record. Records that do not produce exactly 38 valid time slots after this process are excluded; any person-day with incomplete within-day coverage or missing trip-time fields is removed. Figure 1 shows an example.

3.2. Dimensionality Reduction and Distance Computation

Each activity chain is treated as a character string. We compute the Levenshtein (edit) distance for every pair of chains [47]. This distance counts the minimum number of insertions, deletions, or substitutions needed to turn one string into another and is widely used for comparing activity sequences [13,14,15]. The result is an N × N symmetric distance matrix, where N is the total number of chains.
This distance matrix is too high-dimensional for direct clustering. Methods like MDS and t-SNE have been used in the literature, but they can lose local structure at scale or be sensitive to parameter choices. We use UMAP (Uniform Manifold Approximation and Projection) [20], which preserves both local and global structure well [48]. UMAP works in two steps. First, it builds a weighted neighbor graph. For each point x i , it finds nearest neighbors from the distance matrix and sets the probability that x i treats x j as a neighbor to:
p i j = exp d ( x i , x j ) ρ i σ i , i , j
where d ( x i , x j ) is the pairwise distance, ρ i is the distance to the nearest neighbor of x i , and σ i controls the neighborhood size. Since p i j is not symmetric, edge weights are symmetrized as w i j = p i j + p j i p i j p j i . Second, UMAP finds a two-dimensional layout by minimizing the cross-entropy between the high- and low-dimensional graph weights, producing a Euclidean embedding suitable for GMM clustering. We apply UMAP directly to the precomputed Levenshtein distance matrix (metric=‘precomputed’), with n_neighbors  = 50 , min_dist  = 0.1 , and two output dimensions. A fixed random seed of 42 is used for reproducibility.

3.3. Gaussian Mixture Model (GMM)

We cluster the UMAP-embedded points with a Gaussian Mixture Model (GMM) with K components. The GMM models the data as a weighted sum of multivariate normal distributions:
p ( z ) = k = 1 K π k N z μ k , Σ k ,
where π k is the mixing weight of cluster k, μ k is its mean, and Σ k is its covariance matrix. For each chain i, the model returns the posterior probability of
γ i k = Pr ( cluster k z i ) , i , k
which describes how strongly chain i belongs to cluster k. These posteriors are the basis of the soft-clustering representation described next. For each cluster, we also extract a representative chain—the observed chain whose embedding is closest to the cluster mean.
We estimate the GMM parameters ( π k , μ k , Σ k ) using the Expectation–Maximization (EM) algorithm in scikit-learn. The mixing weight π k is the prior probability that a random observation belongs to cluster k, with k = 1 K π k = 1 . Larger clusters get a larger π k . The EM algorithm alternates between two steps: the E-step computes γ i k given current parameters, and the M-step updates π k , μ k , and Σ k to better fit the data. This repeats until the parameters converge. We fit the GMM with full covariance matrices and K = 20 components; the choice of K is discussed in Section 4. The model is fit with five random initializations and up to 300 EM iterations, with the highest-likelihood solution retained. The fitted model produces both hard cluster labels and soft probability vectors for each activity chain.

3.4. Soft Clustering

Hard clustering assigns each day to one cluster. But in practice, many days mix different types of activities. For example, a person may work in the morning, run errands in the afternoon, and stay home in the evening. Hard clustering would label this day as just “work” or just “home,” losing the mixed nature of the schedule. Soft clustering instead represents each day as a probability distribution over all clusters.
From the fitted GMM, each observed chain gets a K-dimensional probability vector γ i = ( γ i 1 , , γ i K ) as in (3). Each element gives the probability that the chain belongs to a given cluster, and the elements sum to one. So, each participant’s seven-day survey becomes a sequence of seven probability vectors instead of seven labels. Figure 2 shows an example.

3.5. Constrained Lagrangian Prediction Model

The predicted probability vector for each person-day must satisfy two constraints: it must sum to one, and all elements must be non-negative. We use a constrained regression model that outputs a K-dimensional vector and is trained with penalty terms that push the output toward a valid probability distribution. The model is:
y ^ target = μ + j L γ j y j + β x
where y ^ target is the predicted probability vector for the target day, μ is the intercept, L is the set of lag days (e.g., days 1, 2, 3), γ j are learned weights for the observed probability vectors y j of previous days, x contains demographic variables (age, gender, education, income, household size), and β maps demographics to cluster probabilities. The loss function has three parts: (i) Mean Squared Error (MSE) between predicted and observed vectors, (ii) a penalty if the K outputs do not sum to one, and (iii) a penalty for any negative prediction:
Loss = MSE ( y target , y ^ target ) + λ k = 1 K y ^ target ( k ) 1 2 + λ neg k = 1 K max 0 , y ^ target ( k ) 2
The model is implemented in PyTorch 2.11.10 as a single linear layer and trained by an optimizer with a learning rate of 0.01 , L 2 weight regularization of 0.01 , and 100 outer optimization steps. The multipliers λ and λ neg are set to large values so that the model treats these constraints as near-hard requirements. Input features are standardized using the training set and then transformed using the same scaler for the test set. We split respondents 80/20 into training and test sets at the person level rather than the row level, so the same respondent never appears in both sets.

3.6. Time-Use Reconstruction

To connect the model output to transportation planning, we reconstruct population-level time-use patterns from the predicted cluster memberships. For each of the K clusters, we take the medoid chain and convert it into a one-hot matrix of size 38 × 11 (30 min time slots by 11 activity types). For person i on day d, the predicted activity profile is a weighted combination of these cluster representatives:
A ^ i , d = k = 1 K p i , d ( k ) · P k , i , d
where p i , d ( k ) is person i’s predicted probability of belonging to cluster k on day d (from the Lagrangian model) and P k is the one-hot representative pattern for cluster k. The resulting 38 × 11 matrix A ^ i , d gives the predicted activity share at each time slot. Averaging across all individuals produces population-level time-use distributions that can be compared to observed data.

4. Results

This section evaluates the framework as an integrated pipeline. The final output—reconstructed population-level time-use schedules—serves as an end-to-end validation because errors at any upstream stage (encoding, embedding, clustering, or prediction) propagate through to the reconstruction. We report three complementary metrics: (1) test-set RMSE for the prediction stage; (2) comparison with hard-clustering baselines, which justifies the soft-clustering design; and (3) time-use reconstruction MAE, which captures the cumulative error of all stages by comparing predicted and observed population-level activity distributions.

4.1. Statistical Analysis of Activity Patterns

To illustrate how the pandemic affects different segments of the population, the activity profiles for each year are analyzed across five activity categories: Home, Work/Education, Maintenance/Shopping, Leisure, and Mixed/Complex. Distributions are then computed across gender, age, education, and income levels, as shown in Figure 3.
The biggest shift is among high-income respondents; the share of Home-category respondents earning $100k+ rises noticeably between 2019 and 2022, while the under-$50k share drops, indicating that home-anchored days became more common among higher-income respondents. Lower-income groups (under $50k) remain in Work/Education and Maintenance/Shopping, reflecting continued in-person employment and errands.
Education shows a similar pattern where respondents with a Bachelor’s degree or higher move into the Home category by 2022 more than other groups. For gender, both men and women show lower Work/Education shares, but women’s Home share reaches 56.1% versus 42.4% for men. Among age groups, adults 65 and older have the highest Home share (nearly 40%), while 18–44-year-olds have the highest share of Work/Education activity.
These patterns point to possible mechanisms behind the shifts we observe. Higher-income and college-educated respondents are likely overrepresented in professional and financial jobs where remote and hybrid arrangements became feasible [1]; the growth of these groups in the Home category between 2019 and 2022 may reflect this fact. Lower-income respondents, by contrast, are probably more concentrated in service, retail, and care jobs that require physical presence—which would explain their continued appearance in the Work/Education and Maintenance/Shopping categories. A similar argument can be made for education; people with bachelor’s or graduate degrees are more likely to hold jobs that can be reorganized around flexible schedules, while those with less education may face more rigid workplace and travel requirements. We cannot confirm these mechanisms with the survey data alone, but the patterns are consistent with what one would expect given the occupational makeup of these groups.

4.2. Activity Chain Analysis

To choose the number of clusters, we first used the Elbow Method and the Davies–Bouldin Index. These suggest 7 clusters for 2019 and 4 for 2022. However, 4 or 7 groups are too few to represent the range of daily behavior across 11 activities and 38 time slots. The medoids at these counts are too generic—for example, one is mostly “work” and another mostly “home”—and the RMSE for probability predictions is high. After an exhaustive testing process, we find that 20 clusters give enough variety to capture diverse daily routines without overfitting. This choice produces lower RMSE, and the 20 medoids for each year are distinct, confirming that choice of 20 clusters is appropriate for our analysis.
Table 2 and Table 3 list the medoid activity chains derived from the 20-cluster GMM solutions for the 2019 and 2022 datasets, as well as a short description of the activity chain, respectively.
In 2019 and 2022, the 20 clusters fall into four broad behavioral families: home-anchored days with little or no out-of-home activity, traditional commuter days with a sustained work block, discretionary days dominated by shopping, recreation, meals, or escort activity, and structured non-work days such as school. The distribution across these families differs by year. In 2019, traditional commuter and discretionary days together account for 60% of person-days, with home-anchored days coming in second. In 2022, there is a shift; home-anchored days alone account for over 40% of person-days, and traditional commuter clusters shrink in both size and number. The GMM clusters represent meaningful daily activity pattern types and the post-pandemic shift reflects a change in the overall mix of daily routines rather than a change in individual cluster sizes.
We apply the GMM pipeline to both years. UMAP plots of the 20-cluster solutions are shown in Figure 4 and Figure 5. In 2019, data points spread broadly across the 20 clusters, showing more variety in daily routines. In 2022, points cluster more tightly around a few dominant groups, suggesting that daily activities became less varied and more home-centered after the pandemic.

4.3. Borough-Level Demographic Patterns

To examine whether these demographic patterns differ across boroughs, we link each person-day to its home borough using the CMS survey zones. Figure 6 shows the percentage-point change between 2019 and 2022 in the share of person-days held by higher-education (Bachelor’s+) and higher-income ($100k+) respondents, by borough and activity category.
Table 4 shows that the borough distribution is stable between 2019 and 2022, with Queens the largest share and Staten Island the smallest. The city-wide shift toward home-anchored patterns is uneven across boroughs. Brooklyn and Queens show the clearest movement: in both, the under-$50k share within the Home-Anchored category declined while the $100k–$200k and $200k+ shares rose (the $100k–$200k Home-Anchored share in Queens went from 21.3% to 30.7%). Manhattan shows smaller income movements; the Bronx the least change. Staten Island has declines in middle-income tiers and a slight rise in $200k+.
Education tells a consistent story. In Brooklyn, Queens, and Manhattan, the Bachelor’s+ share grew within Home-Anchored while lower-education shares declined. Staten Island is similar (about a 9-point increase in the Bachelor’s share). The Bronx is the only borough where lower-education tiers grew within Home-Anchored. Overall, the post-pandemic shift toward home-anchored days is concentrated among higher-income and more-educated residents in boroughs with larger professional workforces (Brooklyn, Queens, Manhattan), while the Bronx’s slower change is consistent with a workforce where more jobs require physical presence.

4.4. Soft Clustering Prediction with Lagrangian Constraints

A hard-clustering baseline (Section 4.5) shows that predicting a single label gives much lower accuracy, which is why we use the probabilistic approach. We apply the Lagrangian model (4) and (5) to predict each target day’s probability vector from preceding days.
Table 5 and Table 6 show test-set RMSE for different input day combinations in 2019 and 2022. Each row is a set of input days; each column is the target day (day 3 through day 7). We test all consecutive sequences: every 2-day pair, every 3-day block, and so on up to the full 6-day history. Any input sequence can predict any later day. Dashes mark cases where the input overlaps the target. For example, predicting day 7 in 2019 (Table 5) with Days 5 and 6 gives RMSE 0.2050, while using all six days gives 0.1978. Longer histories generally improve accuracy.
We also tested non-sequential input combinations—for instance using the day 2 and day 4 pattern to predict the pattern for day 6—but the difference is minimal. For 2022, the best non-sequential combination for day 7 gives RMSE 0.170, versus 0.172 for the sequential one. Since this gap is negligible, sequential days are enough for accurate predictions. They are also more practical, since planners usually work with continuous data-collection windows.

4.5. Benchmarking the Model Output Against Other ML Models

We compare our model against machine learning methods of Random Forest and Neural Network classifiers trained with hard cluster labels ( γ i k { 0 , 1 } ). These models receive the same inputs—previous days and demographics—and predict a single cluster label for the missing day. Table 7 shows that accuracy with hard clustering on 20 clusters is very low (16–36%). This is because daily routines often mix multiple activity types, so forcing a single label means the model gets penalized even when its second-best guess is reasonable. The soft-clustering approach (Table 5 and Table 6) achieves RMSE values around 0.20, for year 2019, and as low as 0.17, for year 2022, confirming the superiority of our proposed GMM-based soft clustering approach.

4.6. Time-Use Reconstruction Results

To check how well the soft-clustering predictions work in practice, we reconstruct time-use schedules using (6). Each person’s predicted probability vector for day 7 is combined with the cluster medoid patterns to get a predicted activity matrix (time slots by activities). We average these across the population. The same procedure on the observed day 7 chains gives the observed matrix. Table 8 and Table 9 compare predicted and observed shares at the hourly level. Each cell shows Predicted% (Observed%), with blue shading for the absolute error, and darker blue refers to a larger gap.
The 2019 reconstruction (Table 8) has mostly small errors. The largest gaps are around midday in the Home column (12–16% predicted vs. 26–32% observed between 11:00 and 14:00), with moderate errors in Escort and Errands. The 2022 reconstruction (Table 9) has larger gaps of 15–25 percentage points in the Home column, and over-predicts Mode Change and Overnight in the afternoon and evening.
The tables also show how the pandemic changed time-use. In 2019, the Work share peaks near 30% between 09:00 and 14:00 and the Home share drops to 26–32% at midday. In 2022, Work is lower (peaking near 27%) and Home stays higher all day (37–46% between 11:00 and 16:00), reflecting more remote and hybrid work after the pandemic. The model captures these shifts in both years.
Table 10 pulls together the headline numbers from the three evaluation stages. The soft-cluster RMSE values come from the full six-day input row of Table 5 and Table 6; the hard-clustering accuracy is the best day 7 result for Random Forest or Neural Network from Table 7; and the time-use MAE is the mean absolute difference between predicted and observed cells in Table 8 and Table 9, averaged across 19 h and 11 activity types. The soft-clustering model achieves test-set RMSE of 0.198 (2019) and 0.173 (2022) on day 7 prediction using the full six-day history, versus hard-clustering baselines that top out at 36% accuracy. The predicted memberships reconstruct population time-use with MAE of 2.44% (2019) and 5.44% (2022). The 2022 error is higher, mainly because of the transition-heavy clusters, but the framework still captures the overall shape of daily time-use in both years.

5. Applications and Policy Implications

The framework can scale to large urban analytics tasks. It compresses each person’s daily activities into a fixed-length probability vector. Aggregating these vectors across a population gives a picture of how the city spends its time across different activity patterns.
The clustering representation also supports tracking changes in urban behavior over time. Our 2019-to-2022 comparison shows not just changes in travel volumes but shifts in how people organize their days: different temporal profiles, different spatial patterns, and different recovery paths for different demographic groups. These results can help policymakers target interventions to specific groups and areas. The time-use reconstruction has direct applications for transportation planning. By combining predicted cluster memberships with 30 min activity distributions, planners can estimate demand by area and demographic group. For example, applying this to MTA ridership data could give zone-level demand forecasts by time of day and population segment. Instead of assuming one peak period for everyone, the model shows when and where different groups generate demand, which can inform pricing and service frequency decisions.
The borough-level breakdowns in Section 4 (Table 4 and Figure 6) show how the same model produces distinct demographic profiles for each borough, which can inform borough-specific planning rather than treating the city as one unit. Brooklyn and Queens, where higher-income and college-educated residents shifted most toward home-based days, may benefit from midday service additions and flexible fare structures in neighborhoods where hybrid-work travel is now more common. Manhattan still has the highest concentration of degree holders across activity categories, so off-peak service investments remain relevant there. The Bronx shows the smallest income shift and higher shares of lower-education and lower-income residents in both Home-Anchored and Work/Education categories—useful context for ensuring reliable peak and off-peak service on routes that serve in-person jobs.
More broadly, the results suggest that the old model of planning around a single morning and evening peak no longer fits the city as well as it once did. The growth of home-anchored days among professional workers means that midday and off-peak travel is no longer just discretionary—it increasingly includes work-related trips by people on hybrid schedules. At the same time, traditional peak demand has not disappeared; it has just become more concentrated among lower-income workers who have fewer alternatives. This creates a service-design tension: agencies need to spread capacity more evenly across the day without cutting peak service for riders who still depend on it. The framework developed here can help quantify that tension by showing, for each borough and demographic group, how many person-days fall into each activity family and when those days generate travel demand. That kind of granularity is hard to get from aggregate ridership counts alone.
There are also equity implications. If fare structures or service cuts are designed around the assumption that midday riders are discretionary, they will disproportionately affect hybrid workers who now travel at those times out of necessity. Conversely, if peak-hour investments are reduced because aggregate ridership is down, the burden falls on in-person workers—disproportionately lower-income and less-educated—who cannot shift their schedules. The demographic disaggregation in this framework makes these trade-offs visible before decisions are made, rather than after.

6. Conclusions

This paper set out to do two things: build a method that can handle incomplete multi-day travel diaries without throwing away the richness of daily activity patterns, and use that method to ask whether the pandemic left a lasting mark on how New Yorkers organize their days.
On the empirical side, the answer is clear. The pandemic did not just reduce travel—it reorganized daily life, and the reorganization stuck. By 2022, home-anchored days had become the dominant pattern in the data, overtaking the commuter and discretionary days that dominated in 2019. This was not a uniform shift. Higher-income and college-educated residents moved toward home-based routines in large numbers, consistent with jobs that allow remote or hybrid work. Lower-income residents, concentrated in jobs that require showing up, continued commuting at roughly the same rates. Women shifted toward home-based schedules more than men. Older adults spent even more time at home than before. The borough-level analysis shows that these shifts are not uniformly distributed across New York City: Brooklyn and Queens show the clearest movement of higher-income and college-educated residents toward home-anchored days, while the Bronx shows the smallest compositional change, consistent with a workforce that has remained more reliant on in-person employment.
What makes these findings useful for planning is that they go beyond aggregate numbers. Knowing that ridership is down 20% is one thing; knowing that the drop is concentrated among professional workers who now travel at different times, while service workers still need the same peak-hour trains, is something planners can act on. The framework gives that level of detail.
On the methods side, three design choices matter. First, soft clustering keeps the mixed nature of daily behavior intact. A day that is 60% commute-like and 30% errand-like gets represented that way, not forced into one bin. This is what makes the time-use reconstruction possible—hard labels cannot support it. Second, the constrained Lagrangian loss function guarantees that every predicted vector is a valid probability distribution without needing post-hoc corrections. This sounds like a technical detail, but it matters at scale: when you are generating predictions for tens of thousands of person-days, you cannot afford to hand-fix outputs that violate basic constraints. Third, the sequential imputation turns sparse diary data into complete multi-day profiles. In both survey waves, fewer than a third of respondents completed all seven days. Without imputation, any analysis of weekly patterns would be limited to that small subset. With it, we can use the full sample.
The 20-dimensional probability vector that comes out of this pipeline is a compact but informative summary of someone’s day. It is small enough to store at scale, detailed enough to reconstruct realistic hourly schedules, and abstract enough to protect individual privacy. Once calibrated on survey data, the same model could be applied to passive data sources—smart cards, phone traces, ride-hailing records—to monitor behavioral shifts city-wide without requiring everyone to fill out a travel diary. Applied across years, it becomes a way to detect when and how populations change the way they spend their time.
Several limitations should be noted. The study covers one city; applying the framework to cities with different transit systems, labor markets, and demographics would test whether the findings generalize or reflect New York’s specific conditions. The data lack fine-grained spatial information, so we cannot tie patterns to specific corridors, stations, or neighborhoods—only to boroughs. Each survey wave is a cross-section, not a panel. We are comparing two populations at two points in time, not tracking the same individuals. This means we can describe what changed in aggregate, but we cannot say with certainty that specific individuals changed their behavior. Finally, the survey was collected in specific months (May–June for 2019, September–November for 2022), which introduces seasonal differences that we cannot fully separate from pandemic effects.
Future work could go in several directions. The most immediate is combining the survey-calibrated GMM with passive mobility data to test whether the model scales from thousands of respondents to millions of residents. If the cluster structure learned from surveys holds up when applied to smart-card or phone-trace data, it would open up near-real-time monitoring of activity patterns at the city level. A second direction is adding spatial resolution—mode choice, origin-destination pairs, and land-use characteristics—so that the framework can support corridor-level planning, not just city-wide or borough-level analysis. Third, applying the approach to multiple cities would allow standardized comparisons of how different populations organize their time and how they responded to the same disruption. Cities with different transit dependency, industry mix, and housing patterns likely experienced different reshufflings of daily routines, and a common methodological framework would make those comparisons meaningful.

Author Contributions

Conceptualization, M.A.; methodology, M.A. and N.K.; software, N.K. and L.A.R.V.; validation, L.A.R.V.; formal analysis, N.K.; investigation, N.K., M.A. and L.A.R.V.; data curation, N.K. and L.A.R.V.; writing—original draft, M.A. and N.K.; writing—review & editing, M.A. and N.K.; visualization, L.A.R.V.; supervision, M.A.; project administration, M.A.; funding acquisition, M.A. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the US Department of Transportation (Award ID: 69A3552344815 and 69A3552348320). The authors sincerely appreciate this support.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are openly available in NYC open data at https://data.cityofnewyork.us/Transportation/Citywide-Mobility-Survey-Household-2022/5mb3-padx/about_data (accessed on 1 April 2024).

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Anik, M.A.H.; Habib, M.A. COVID-19 and Teleworking: Lessons, Current Issues and Future Directions for Transport and Land-Use Planning. Transp. Res. Rec. J. Transp. Res. Board 2024, 2678, 554–566. [Google Scholar] [CrossRef]
  2. Tahlyan, D.; Mahmassani, H.; Stathopoulos, A.; Said, M.; Shaheen, S.; Walker, J.; Johnson, B. In-person, hybrid or remote? Employers’ perspectives on the future of work post-pandemic. Transp. Res. Part A Policy Pract. 2024, 190, 104273. [Google Scholar] [CrossRef]
  3. Ashour, L.; Shen, Q. Unveiling post-pandemic commute choices amidst the rise of telework. Transp. Res. Part D Transp. Environ. 2025, 149, 105068. [Google Scholar] [CrossRef]
  4. McNally, M.G. The Activity-Based Approach. In Handbook of Transport Modelling, 2nd ed.; Hensher, D.A., Button, K.J., Eds.; Emerald Group Publishing: Leeds, UK, 2007; pp. 55–73. [Google Scholar]
  5. Rasouli, S.; Timmermans, H. Activity-Based Models of Travel Demand: Promises, Progress and Prospects. Int. J. Urban Sci. 2014, 18, 31–60. [Google Scholar] [CrossRef]
  6. Bowman, J.L.; Ben-Akiva, M.E. Activity-Based Disaggregate Travel Demand Model System with Activity Schedules. Transp. Res. Part A Policy Pract. 2001, 35, 1–28. [Google Scholar] [CrossRef]
  7. Pas, E.I. Weekly Travel-Activity Behavior. Transportation 1988, 15, 89–109. [Google Scholar] [CrossRef]
  8. Hanson, S.; Huff, J.O. Systematic Variability in Repetitious Travel. Transportation 1988, 15, 111–135. [Google Scholar] [CrossRef]
  9. Susilo, Y.O.; Axhausen, K.W. Repetitions in Individual Daily Activity–Travel–Location Patterns: A Study Using the Herfindahl–Hirschman Index. Transportation 2014, 41, 995–1011. [Google Scholar] [CrossRef]
  10. Bhat, C.R.; Koppelman, F.S. A Retrospective and Prospective Survey of Time-Use Research. Transportation 1999, 26, 119–139. [Google Scholar] [CrossRef]
  11. Arentze, T.A.; Timmermans, H.J. A Learning-Based Transportation Oriented Simulation System. Transp. Res. Part B Methodol. 2004, 38, 613–633. [Google Scholar] [CrossRef]
  12. Allahviranloo, M.; Recker, W. Mining activity pattern trajectories and allocating activities in the network. Transportation 2015, 42, 561–579. [Google Scholar] [CrossRef]
  13. Allahviranloo, M.; Recker, W. Daily Activity Pattern Recognition by Using Support Vector Machines with Multiple Classes. Transp. Res. Part B Methodol. 2013, 58, 16–43. [Google Scholar] [CrossRef]
  14. Joh, C.H.; Arentze, T.; Timmermans, H. Multidimensional Sequence Alignment Methods for Activity–Travel Pattern Analysis: A Comparison of Dynamic Programming and Genetic Algorithms. Geogr. Anal. 2001, 33, 247–270. [Google Scholar] [CrossRef]
  15. Wilson, C. Activity Patterns in Space and Time: Calculating Representative Hagerstrand Trajectories. Transportation 2008, 35, 485–499. [Google Scholar] [CrossRef]
  16. Allahviranloo, M. Pattern Recognition and Personal Travel Behavior. Ph.D. Thesis, University of California, Irvine, CA, USA, 2014. [Google Scholar]
  17. Jiang, S.; Ferreira, J.; Gonzalez, M.C. Clustering Daily Patterns of Human Activities in the City. Data Min. Knowl. Discov. 2012, 25, 478–510. [Google Scholar] [CrossRef]
  18. Goulet-Langlois, G.; Koutsopoulos, H.N.; Zhao, J. Measuring Regularity of Individual Travel Patterns. IEEE Trans. Intell. Transp. Syst. 2018, 19, 1583–1592. [Google Scholar] [CrossRef]
  19. Yao, W.; Yu, J.; Yang, Y.; Chen, N.; Jin, S.; Hu, Y.; Bai, C. Understanding Travel Behavior Adjustment Under COVID-19. Commun. Transp. Res. 2022, 2, 100068. [Google Scholar] [CrossRef]
  20. McInnes, L.; Healy, J.; Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv 2018, arXiv:1802.03426. [Google Scholar] [CrossRef]
  21. Fraley, C.; Raftery, A.E. Model-Based Clustering, Discriminant Analysis, and Density Estimation. J. Am. Stat. Assoc. 2002, 97, 611–631. [Google Scholar] [CrossRef]
  22. Kim, J.; Corcoran, J.; Papamanolis, M. Route Choice Stickiness of Public Transport Passengers: Measuring Habitual Bus Ridership Behaviour Using Smart Card Data. Transp. Res. Part C Emerg. Technol. 2017, 83, 146–163. [Google Scholar] [CrossRef]
  23. Sun, L.; Axhausen, K.W. Understanding Urban Mobility Patterns with a Probabilistic Tensor Factorization Framework. Transp. Res. Part B Methodol. 2016, 91, 511–524. [Google Scholar] [CrossRef]
  24. Allahviranloo, M.; Aissaoui, L. A Comparison of Time-Use Behavior in Metropolitan Areas Using Pattern Recognition Techniques. Transp. Res. Part A Policy Pract. 2019, 129, 271–287. [Google Scholar] [CrossRef]
  25. Su, R.; McBride, E.C.; Goulias, K.G. Unveiling Daily Activity Pattern Differences between Telecommuters and Commuters Using Human Mobility Motifs and Sequence Analysis. Transp. Res. Part A Policy Pract. 2021, 147, 106–132. [Google Scholar] [CrossRef]
  26. Rafiq, R.; McNally, M.G.; Sarwar Uddin, Y.; Ahmed, T. Impact of Working from Home on Activity-Travel Behavior during the COVID-19 Pandemic: An Aggregate Structural Analysis. Transp. Res. Part A Policy Pract. 2022, 159, 35–54. [Google Scholar] [CrossRef] [PubMed]
  27. Fredriksson, J.; Karlström, A. A Fully Neural Network-Based Travel Demand and Scheduling Model, Covering Activities, Destinations, and Modes of Transportation. Data Sci. Transp. 2025, 7, 28. [Google Scholar] [CrossRef]
  28. Song, Y.; Li, D.; Ma, Z.; Liu, D.; Zhang, T. A state-based inverse reinforcement learning approach to model activity-travel choices behavior with reward function recovery. Transp. Res. Part C Emerg. Technol. 2024, 158, 104454. [Google Scholar] [CrossRef]
  29. Miller, E.J. The current state of activity-based travel demand modelling and some possible next steps. Transp. Rev. 2023, 43, 565–570. [Google Scholar] [CrossRef]
  30. Bhuiyan, M.R.H.; Habib, M.A. Modeling weekly representative activity patterns of population groups: A bioinspired multiple sequence alignment approach. Transportation 2025. [Google Scholar] [CrossRef]
  31. Haghighi, M.; Miller, E.J. Week-long activity-based modelling: A review of the existing models and datasets and a comprehensive conceptual framework. Transp. Rev. 2025, 45, 119–148. [Google Scholar] [CrossRef]
  32. Axhausen, K.W.; Zimmermann, A.; Schönfelder, S.; Rindsfüser, G.; Haupt, T. Observing the Rhythms of Daily Life: A Six-Week Travel Diary. Transportation 2002, 29, 95–124. [Google Scholar] [CrossRef]
  33. Schlich, R.; Axhausen, K.W. Habitual Travel Behaviour: Evidence from a Six-Week Travel Diary. Transportation 2003, 30, 13–36. [Google Scholar] [CrossRef]
  34. Bucsky, P. Modal Share Changes Due to COVID-19: The Case of Budapest. Transp. Res. Interdiscip. Perspect. 2020, 8, 100141. [Google Scholar] [CrossRef] [PubMed]
  35. De Vos, J. The Effect of COVID-19 and Subsequent Social Distancing on Travel Behavior. Transp. Res. Interdiscip. Perspect. 2020, 5, 100121. [Google Scholar] [CrossRef] [PubMed]
  36. Zhang, J.; Hayashi, Y.; Frank, L.D. COVID-19 and Transport: Findings from a World-Wide Expert Survey. Transp. Policy 2021, 103, 68–85. [Google Scholar] [CrossRef] [PubMed]
  37. Brough, R.; Freedman, M.; Phillips, D.C. Understanding Socioeconomic Disparities in Travel Behavior During the COVID-19 Pandemic. J. Reg. Sci. 2021, 61, 753–774. [Google Scholar] [CrossRef] [PubMed]
  38. Hu, S.; Chen, P. Who Left Riding Transit? Examining Socioeconomic Disparities in the Impact of COVID-19 on Ridership. Transp. Res. Part D Transp. Environ. 2021, 90, 102654. [Google Scholar] [CrossRef]
  39. Salon, D.; Conway, M.W.; Capasso da Silva, D.; Gehrke, S.R.; Derrible, S.; Mohammadian, A.; Khoeini, S.; Parker, N.; Mirtich, L.; Shamshiripour, A.; et al. The Potential Stickiness of Pandemic-Induced Behavior Changes in the United States. Proc. Natl. Acad. Sci. USA 2021, 118, e2106499118. [Google Scholar] [CrossRef] [PubMed]
  40. Liu, L.; Miller, H.J.; Scheff, J. The Impacts of COVID-19 Pandemic on Public Transit Demand in the United States. PLoS ONE 2020, 15, e0242476. [Google Scholar] [CrossRef] [PubMed]
  41. New York City Department of Transportation. 2019 Citywide Mobility Survey Results; Technical Report; NYC Department of Transportation: New York, NY, USA, 2019.
  42. New York City Department of Transportation. 2022 Citywide Mobility Survey Results; Technical Report; NYC Department of Transportation: New York, NY, USA, 2022.
  43. New York City Department of Transportation. Citywide Mobility Survey. 2022. Available online: https://www.nyc.gov/html/dot/html/about/citywide-mobility-survey.shtml (accessed on 1 April 2023).
  44. New York City Department of Transportation. 2022 Citywide Mobility Survey Questionnaire; Technical Report; NYC Department of Transportation: New York, NY, USA, 2022.
  45. Ji, Y.; Gao, S.; Huynh, T.; Scheele, C.; Triveri, J.; Kruse, J.; Bennett, C.; Wen, Y. Rethinking the regularity in mobility patterns of personal vehicle drivers: A multi-city comparison using a feature engineering approach. Trans. GIS 2023, 27, 1126–1148. [Google Scholar] [CrossRef]
  46. Purbohadi, D.; Azizah, L.M.; Kurniasari, L.; Wulandari, N.D.; Pratiwi, N.; Hastuti, P. Clustering Commuter Behavior based on Automated Fare Collection (AFC). Eng. Technol. Appl. Sci. Res. 2025, 15, 8899. [Google Scholar] [CrossRef]
  47. Levenshtein, V.I. Binary Codes Capable of Correcting Deletions, Insertions, and Reversals. Sov. Phys. Dokl. 1966, 10, 707–710. [Google Scholar]
  48. Diaz-Papkovich, A.; Anderson-Trocmé, L.; Gravel, S. UMAP Reveals Cryptic Population Structure and Phenotype Heterogeneity in Large Genomic Cohorts. PLoS Genet. 2019, 15, e1008432. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Example of activity chain construction.
Figure 1. Example of activity chain construction.
Urbansci 10 00331 g001
Figure 2. Probability vector example showing soft cluster memberships for one activity chain.
Figure 2. Probability vector example showing soft cluster memberships for one activity chain.
Urbansci 10 00331 g002
Figure 3. Cluster distributions by demographic group for 2019 and 2022.
Figure 3. Cluster distributions by demographic group for 2019 and 2022.
Urbansci 10 00331 g003
Figure 4. UMAP embedding with 20 GMM clusters—2019. Data points are more widely dispersed across the 20 clusters, reflecting a broader range of daily routines.
Figure 4. UMAP embedding with 20 GMM clusters—2019. Data points are more widely dispersed across the 20 clusters, reflecting a broader range of daily routines.
Urbansci 10 00331 g004
Figure 5. UMAP embedding with 20 GMM clusters—2022. Points are grouped more tightly around a few dominant clusters, consistent with the interpretation that post-pandemic daily activities became less varied and more concentrated around home-based patterns.
Figure 5. UMAP embedding with 20 GMM clusters—2022. Points are grouped more tightly around a few dominant clusters, consistent with the interpretation that post-pandemic daily activities became less varied and more concentrated around home-based patterns.
Urbansci 10 00331 g005
Figure 6. Percentage-point change (2022 minus 2019) in the share of person-days held by higher-education (Bachelor’s degree or above, (left) panel) and higher-income ($100k+, (right) panel) respondents, by borough and activity-pattern category. Bars above zero indicate that the group’s share grew in that category between 2019 and 2022; bars below zero indicate it shrank.
Figure 6. Percentage-point change (2022 minus 2019) in the share of person-days held by higher-education (Bachelor’s degree or above, (left) panel) and higher-income ($100k+, (right) panel) respondents, by borough and activity-pattern category. Bars above zero indicate that the group’s share grew in that category between 2019 and 2022; bars below zero indicate it shrank.
Urbansci 10 00331 g006
Table 1. Participant demographics and survey completion summary (2019 vs. 2022). Values are n (%).
Table 1. Participant demographics and survey completion summary (2019 vs. 2022). Values are n (%).
VariableCategory20192022
GenderMen1806 (21.8%)1547 (22.5%)
Women1440 (17.4%)1276 (18.5%)
Other genders19 (0.2%)42 (0.6%)
Prefer not to answer81 (1.0%)101 (1.5%)
Age18–241863 (22.5%)1418 (20.5%)
25–341502 (18.1%)1358 (19.7%)
35–441225 (14.8%)1222 (17.7%)
45–541186 (14.3%)876 (12.7%)
55+2105 (25.4%)1701 (24.6%)
EducationLess than HS/Prefer not to answer5063 (61.1%)4000 (58.1%)
Some college713 (8.6%)545 (7.9%)
Bachelor’s degree1003 (12.1%)1034 (15.0%)
Graduate/post-grad874 (10.5%)940 (13.7%)
IncomeUnder $75,0001552 (46.5%)1223 (41.2%)
$75,000–$99,999425 (12.7%)358 (12.1%)
$100,000–$199,999718 (21.5%)699 (23.6%)
$200,000+265 (7.9%)361 (12.3%)
Prefer not to answer386 (11.5%)325 (11.0%)
Household size1 person964 (28.8%)953 (32.1%)
2 people1040 (31.1%)989 (33.3%)
3 people596 (17.8%)474 (16.0%)
4+ people746 (22.3%)550 (18.5%)
Survey days
completed
1 day3007 (36.3%)2114 (30.7%)
2–3 days2145 (25.9%)1780 (25.8%)
4–6 days2093 (25.3%)1958 (28.4%)
All 7 days1036 (12.5%)1039 (15.1%)
Table 2. Cluster centroid summary (20 clusters)—2019.
Table 2. Cluster centroid summary (20 clusters)—2019.
ClusterSize (Days)Centroid ChainPattern Type
0506 (4.86%)AAAAAAAAAAAAAAAMAAAAAAAAAAAAAAAAALLLAAAll-day escort with brief meals
1984 (9.46%)HHHHHHHHHHHHHHHSSSAAAAAAAMMHHHHHHHHHHHHome + morning school escort and shopping
2988 (9.50%)HHHHHHEWWWWWWWWWWWWWWWWWWWWWWWLLLLHHHHTraditional commuter, full work block
3950 (9.13%)HHHHHHHHHHHHHMHHHHHHHHHHHHHHHHAHHHHHHHHome-anchored with brief errand
4731 (7.03%)LLLLLLLLLLLLLLLLLAAAALLLLLLLLLLLLLLLLLAll-day meal/leisure with escort break
5550 (5.29%)WWWWEWWWWWWWMWWWWWWWWWWWWWWWWWWWWWWWWWLong-shift worker with mid-day shopping
6491 (4.72%)MMMMMMMMMMMMMMMMELLLRRRRRRRRREHHHHHHHHShopping morning + recreation afternoon
7350 (3.37%)LLLLLLLLEWWWWWWWWWWWHHHHHHHHHHRRRRRRRRMixed: Meals, midday work, evening recreation
8372 (3.58%)RRRRRRRRRRRRRREERRRRRRMMMMMERRRRRRRRRRAll-day recreation with shopping break
9272 (2.62%)HHHHHHHHHLLLLLLLLLHHHHHHHWWWWWWWWWWWWWHome morning + evening work shift
10336 (3.23%)EEEEEEEEEEEEEEEEEEEAAHHHEWWWWWWWWWWWWWLong errand run + late work shift
111112 (10.69%)HHHHHHEWWWWWWWWWMMWWWWWWWWLLHHHHHHHHHHTraditional commuter with midday shopping
12242 (2.33%)SSSSSSSSSSSSSSSSSSAHHHHHHHSSSHRSSSSSSSStudent/school-dominant day
131515 (14.56%)HHHHHHHHHHHHHLEWWWWLLLLLHHHHHHHHHHHHHHHome-anchored with short work block
14254 (2.44%)HHHHHWWWWWWMMMMMMMMMMMMMMMMMMHHHHHHHHHWork morning + extended shopping
15458 (4.40%)HHHHHHHHHLLLLLHHHHHHHSSSSSSSSHHHHHHHHHHome + afternoon school block
16763 (7.33%)HHHHHHHHHHHMMMHHHLLHEHHHHHHHHHHHHHHHHHHome-anchored with discretionary breaks
17388 (3.73%)AAAAAAAWWWWWWAAAAAAAAAAAAAAHAAAAAAAAAAEscort-dominant with brief work
18386 (3.71%)HHHHHHHHHHHHMRRRRRRRRRRRRRRHHHHHHHHHHHHome morning + sustained recreation block
19865 (8.31%)HWWWWWWWWWWWWWWWWWWWWWHHHHHHHHHHHHHHHHTraditional commuter, full work block
Table 3. Cluster centroid summary (20 clusters)—2022.
Table 3. Cluster centroid summary (20 clusters)—2022.
ClusterSize (Days)Centroid ChainPattern Type
02217 (18.69%)HHHHHHHHHHHHHHHHHHHHHMMHHHHHHHHHHHHHHHHome-anchored, near-constant H
1304 (2.56%)CCCCCCCCCCCWWWWWWWWWWWWWWCCCCCCCCCCCCCTransition-heavy worker (long C blocks)
2902 (7.61%)HHHHHHCWWWWWWWWWWWAHWWWWWWWWHHHHHHHHHHTraditional commuter with mode-change
31634 (13.78%)HHHHHHHHHHHHHHHHHHRRRHHHHHHHHHHHHHHHHHHome-anchored with brief recreation
4800 (6.75%)CCCCCCCCCCCCCCCOMMMMMMMMMMMCCCCCCCCCCCTransition-heavy shopping day
5220 (1.85%)RRRRRRRRRRRRRRREAAAAAARRRRRRRRRRRRRRRRAll-day recreation with escort break
6248 (2.09%)HHHHHHHHHHHHHHHHHHHHHHHHCWWWWWWWWWWWWWHome morning + evening work shift
7570 (4.81%)NNNNNNNNNNNRRRLLLLONNNNNNNNNNNNNNNNNNNOvernight-dominant day
81093 (9.22%)HHHHHHHRHHHHHHHHHHHHHHHHHCCCCHHHHHHHHHHome-anchored with brief recreation
9125 (1.05%)AAAAAAAAALLLLLLLLLLLAAAAAAAAAAAAAAAAAAEscort + meal pattern
101060 (8.94%)HHHHHHHHHHHHCCEECHHHHHHHHHHHHHHHHHHHHHHome + short errand/mode-change
11198 (1.67%)WWWWWWWWWWWWWWWWWMMMMMWWWWWWWWWWWWWWWWLong-shift worker with midday shopping
12158 (1.33%)MMMMMMMMMMARAAAAAAAAAAAMMMMMMMMMMMMMMMAll-day shopping with escort break
13575 (4.85%)HHHHCCCCCCCCCCSSSSSSSSLLLLLMRRRRRHHHHHStudent with mode-change, meals, recreation
14899 (7.58%)HHHHHHHHHHHWWWWWWWWWLLLLLLLLLLLHHHHHHHTraditional commuter with leisure afternoon
15159 (1.34%)LLLLLLAMMMMMMMMMMMMMMMMLLLLLLLLLLLLLLLMeal/leisure with shopping block
16343 (2.89%)HHHHHHHHHHHHHHHHHHRNENNNNNNNNNNNHHHHHHHome + extended overnight period
17574 (4.84%)HHHHHHHHHHHHHHHHHHHHHRRMHHHHHHHHHHHHHHHome-anchored with brief recreation
18898 (7.57%)HHHHHHWWWWWWWWWWWWWWWHHHHHHHHHHHHHHHHHTraditional commuter, full work day
19162 (1.37%)HHHHHNNNNNNNNNNNNNHHWWWWWWWWWWWWWWWWWWOvernight period + late work shift
Table 4. Borough distribution of the analytic sample.
Table 4. Borough distribution of the analytic sample.
Borough20192022
Queens3620 (28.9%)3702 (28.3%)
Bronx2681 (21.4%)2579 (19.7%)
Manhattan2480 (19.8%)2785 (21.5%)
Brooklyn2409 (19.3%)2969 (22.7%)
Staten Island1323 (10.6%)1025 (7.8%)
Total12,513 (100%)13,060 (100%)
Table 5. Test RMSE heatmap for sequential day combinations (2019, 20 clusters).
Table 5. Test RMSE heatmap for sequential day combinations (2019, 20 clusters).
CombinationDay 3Day 4Day 5Day 6Day 7
2-Day Sequences
(1, 2)0.20400.20680.20390.21010.2086
(2, 3)0.20580.20450.21010.2087
(3, 4)0.20260.21100.2131
(4, 5)0.20830.2087
(5, 6)0.2050
3-Day Sequences
(1, 2, 3)0.20490.20290.20900.2063
(2, 3, 4)0.20230.21110.2103
(3, 4, 5)0.21010.2085
(4, 5, 6)0.2047
4-Day Sequences
(1, 2, 3, 4)0.20120.21000.2083
(2, 3, 4, 5)0.20940.2056
(3, 4, 5, 6)0.2029
5-Day Sequences
(1, 2, 3, 4, 5)0.20840.2037
(2, 3, 4, 5, 6)0.2008
6-Day Sequence
(1, 2, 3, 4, 5, 6)0.1978
Table 6. Test RMSE heatmap for sequential day combinations (2022, 20 clusters).
Table 6. Test RMSE heatmap for sequential day combinations (2022, 20 clusters).
CombinationDay 3Day 4Day 5Day 6Day 7
2-Day Sequences
(1, 2)0.18780.19670.19610.19570.1886
(2, 3)0.19340.19460.19560.1904
(3, 4)0.19120.19860.1954
(4, 5)0.19450.1953
(5, 6)0.1876
3-Day Sequences
(1, 2, 3)0.19100.19550.19520.1858
(2, 3, 4)0.19180.19620.1892
(3, 4, 5)0.19280.1928
(4, 5, 6)0.1875
4-Day Sequences
(1, 2, 3, 4)0.19220.19620.1843
(2, 3, 4, 5)0.19040.1859
(3, 4, 5, 6)0.1842
5-Day Sequences
(1, 2, 3, 4, 5)0.19100.1796
(2, 3, 4, 5, 6)0.1770
6-Day Sequence
(1, 2, 3, 4, 5, 6)0.1726
Table 7. Comparative Performance of Random Forest and Neural Network (Hard Clustering) on 20 clusters.
Table 7. Comparative Performance of Random Forest and Neural Network (Hard Clustering) on 20 clusters.
Random ForestNeural Network
2019 2022 2019 2022
Target Day Sequence Accuracy Sequence Accuracy Sequence Accuracy Sequence Accuracy
Day 31, 20.1651, 20.3071, 20.1461, 20.238
Day 41, 20.1731, 2, 30.2922, 30.1732, 30.237
Day 53, 40.1852, 3, 40.2531, 2, 30.1631, 2, 3, 40.243
Day 61, 2, 3, 4, 50.1811, 20.2223, 40.1393, 4, 50.241
Day 73, 4, 5, 60.1731, 20.3635, 60.1762, 3, 4, 5, 60.267
Note: Accuracies reflect test-set performance.
Table 8. Predicted vs. Observed Time-Use Proportions (2019).
Table 8. Predicted vs. Observed Time-Use Proportions (2019).
HourHomeWorkSchoolEscortShoppingMealRecreationErrandsMode ChgOvernightOther
05:0061.0 (65.7)7.4 (5.6)2.4 (2.8)8.4 (8.0)4.6 (4.1)9.6 (8.4)3.7 (3.9)3.0 (1.5)0.0 (0.0)0.0 (0.0)0.0 (0.0)
06:0057.5 (62.8)10.9 (7.5)2.4 (2.7)8.4 (7.8)4.6 (4.1)9.6 (8.3)3.7 (4.1)3.0 (2.7)0.0 (0.0)0.0 (0.0)0.0 (0.0)
07:0056.2 (56.2)10.3 (11.4)2.4 (3.6)8.4 (8.4)4.6 (4.3)9.6 (8.5)3.7 (4.5)4.9 (3.1)0.0 (0.0)0.0 (0.0)0.0 (0.0)
08:0039.4 (44.9)23.3 (18.4)2.4 (5.1)6.4 (9.6)4.6 (5.0)9.6 (9.0)3.7 (3.6)10.7 (4.6)0.0 (0.0)0.0 (0.0)0.0 (0.0)
09:0035.4 (38.7)34.7 (26.8)2.4 (4.4)4.4 (10.0)4.6 (4.8)10.3 (8.5)3.7 (3.2)4.7 (3.6)0.0 (0.0)0.0 (0.0)0.0 (0.0)
10:0029.1 (35.5)35.0 (29.9)2.4 (5.3)4.4 (10.6)8.2 (4.2)14.3 (8.8)3.7 (3.2)3.0 (2.5)0.0 (0.0)0.0 (0.0)0.0 (0.0)
11:0016.0 (31.8)29.7 (31.3)2.4 (5.1)6.4 (12.6)17.6 (4.8)20.0 (9.6)4.9 (3.2)3.0 (1.5)0.0 (0.0)0.0 (0.0)0.0 (0.0)
12:0018.1 (28.9)35.3 (31.1)6.2 (4.4)6.2 (11.9)9.5 (7.1)9.8 (11.0)2.6 (3.3)12.3 (2.2)0.0 (0.0)0.0 (0.0)0.0 (0.0)
13:0012.0 (27.6)33.5 (29.4)10.1 (3.5)11.6 (13.6)10.2 (7.0)11.2 (12.2)6.2 (4.5)5.3 (2.3)0.0 (0.0)0.0 (0.0)0.0 (0.0)
14:0016.7 (26.3)35.3 (30.3)0.0 (4.4)25.1 (14.3)2.7 (7.2)12.5 (10.5)6.2 (4.2)1.5 (2.8)0.0 (0.0)0.0 (0.0)0.0 (0.0)
15:0020.5 (28.8)26.3 (28.4)2.3 (3.7)20.8 (14.7)2.7 (7.1)14.5 (10.3)10.8 (4.7)2.2 (2.3)0.0 (0.0)0.0 (0.0)0.0 (0.0)
16:0029.0 (32.5)19.3 (23.9)4.5 (3.9)16.1 (14.0)6.4 (7.3)17.6 (11.0)7.1 (4.7)0.0 (2.8)0.0 (0.0)0.0 (0.0)0.0 (0.0)
17:0035.6 (37.6)22.5 (18.0)4.5 (3.2)12.3 (14.0)10.2 (8.0)6.2 (11.5)7.1 (4.8)1.5 (2.8)0.0 (0.0)0.0 (0.0)0.0 (0.0)
18:0038.6 (44.0)18.3 (12.9)6.9 (3.9)6.4 (12.9)8.4 (6.3)13.8 (12.4)5.8 (4.6)1.8 (3.0)0.0 (0.0)0.0 (0.0)0.0 (0.0)
19:0054.0 (47.0)18.3 (9.3)3.4 (4.4)8.4 (10.4)1.4 (7.9)6.2 (13.6)5.9 (4.4)2.3 (2.9)0.0 (0.0)0.0 (0.0)0.0 (0.0)
20:0055.1 (52.4)10.4 (7.0)1.2 (3.6)11.1 (9.9)0.0 (7.8)14.1 (13.1)8.2 (4.2)0.0 (2.1)0.0 (0.0)0.0 (0.0)0.0 (0.0)
21:0057.7 (58.7)10.4 (6.2)2.4 (3.9)6.2 (8.4)0.0 (6.4)16.4 (10.7)7.0 (3.9)0.0 (1.8)0.0 (0.0)0.0 (0.0)0.0 (0.0)
22:0065.6 (61.3)10.4 (5.8)2.4 (3.5)4.0 (7.7)0.0 (5.8)10.7 (9.9)7.0 (3.8)0.0 (2.1)0.0 (0.0)0.0 (0.0)0.0 (0.0)
23:0065.6 (62.6)10.4 (5.5)2.4 (3.2)8.4 (7.7)0.0 (5.4)6.2 (9.6)7.0 (3.8)0.0 (2.3)0.0 (0.0)0.0 (0.0)0.0 (0.0)
Table 9. Predicted vs. Observed Time-Use Proportions (2022).
Table 9. Predicted vs. Observed Time-Use Proportions (2022).
HourHomeWorkSchoolEscortShoppingMealRecreationErrandsMode ChgOvernightOther
05:0046.5 (77.3)2.6 (3.1)0.0 (0.1)2.1 (1.2)2.0 (1.1)6.7 (1.8)5.1 (2.3)0.0 (0.3)17.5 (6.3)17.5 (6.0)0.0 (0.3)
06:0046.5 (74.6)2.6 (4.8)0.0 (0.1)2.1 (1.3)2.0 (1.0)6.7 (1.9)5.1 (2.9)0.0 (0.2)17.5 (6.9)17.5 (5.7)0.0 (0.4)
07:0046.0 (66.7)2.6 (9.1)0.0 (0.4)2.1 (2.7)2.0 (1.0)6.7 (2.1)5.1 (3.2)0.0 (0.3)17.9 (8.3)17.5 (5.5)0.0 (0.6)
08:0027.7 (53.4)14.3 (15.6)0.0 (0.7)5.5 (3.8)5.4 (1.6)0.0 (2.5)6.5 (4.0)0.0 (0.7)22.5 (12.1)18.1 (4.6)0.0 (0.7)
09:0029.1 (45.7)18.5 (23.2)0.0 (1.1)1.1 (2.0)8.8 (2.2)1.1 (3.2)5.1 (4.8)0.0 (2.0)18.3 (10.4)18.1 (4.6)0.0 (0.6)
10:0027.5 (41.9)27.1 (27.0)0.0 (1.3)1.0 (1.4)6.7 (3.2)2.1 (2.9)14.8 (6.2)0.0 (1.9)11.3 (9.0)9.3 (4.3)0.0 (0.6)
11:0022.9 (38.6)35.7 (27.1)0.0 (1.5)2.0 (1.2)6.7 (3.4)2.1 (4.6)22.6 (6.6)0.0 (2.4)7.4 (9.9)0.6 (3.9)0.0 (0.6)
12:0022.9 (37.3)35.7 (26.3)0.4 (1.3)2.0 (1.4)6.7 (3.5)19.6 (4.7)2.5 (7.5)5.6 (3.3)2.1 (9.7)0.6 (4.1)1.7 (0.6)
13:0016.5 (37.8)34.4 (25.3)0.9 (1.3)7.1 (1.4)17.3 (3.7)19.6 (5.6)2.1 (7.0)0.0 (3.3)1.5 (9.4)0.6 (4.2)0.0 (0.7)
14:0020.4 (37.6)24.7 (24.8)0.9 (1.3)11.3 (2.4)18.6 (3.9)2.1 (5.9)1.9 (6.9)1.2 (3.1)0.0 (9.3)10.2 (3.9)8.7 (0.8)
15:0026.5 (38.7)26.5 (23.6)0.9 (1.1)9.2 (2.7)12.8 (3.4)3.1 (5.5)1.1 (6.9)0.0 (3.0)0.0 (9.9)19.9 (4.3)0.0 (0.7)
16:0028.9 (40.9)25.6 (21.9)0.4 (1.1)3.1 (2.3)8.9 (4.3)7.0 (4.4)6.1 (7.3)0.0 (2.0)0.0 (10.1)19.9 (4.7)0.0 (0.9)
17:0027.5 (46.2)19.7 (16.4)0.0 (1.2)2.1 (2.9)5.4 (5.3)10.8 (5.0)5.1 (6.3)0.0 (1.6)9.5 (8.9)19.9 (5.0)0.0 (1.0)
18:0026.1 (50.8)13.8 (10.5)0.0 (1.0)2.1 (2.5)3.7 (5.1)10.8 (6.7)5.1 (7.4)0.0 (1.5)18.5 (8.3)19.9 (5.3)0.0 (0.7)
19:0035.9 (58.4)5.3 (6.9)0.0 (0.7)2.1 (2.4)2.5 (3.9)9.9 (5.6)5.5 (7.1)0.0 (1.2)18.8 (7.8)19.9 (5.4)0.0 (0.5)
20:0040.1 (65.2)5.3 (5.3)0.0 (0.7)2.1 (1.7)2.0 (2.2)8.3 (4.2)6.0 (6.7)0.0 (0.8)17.5 (6.7)18.7 (5.9)0.0 (0.3)
21:0042.9 (71.0)5.3 (4.2)0.0 (0.4)2.1 (1.2)2.0 (1.6)6.7 (3.5)6.0 (5.2)0.0 (0.3)17.5 (6.2)17.5 (5.8)0.0 (0.3)
22:00100.0 (74.5)0.0 (3.3)0.0 (0.1)0.0 (1.2)0.0 (1.2)0.0 (2.7)0.0 (3.1)0.0 (0.3)0.0 (7.0)0.0 (6.3)0.0 (0.2)
23:00100.0 (76.4)0.0 (2.9)0.0 (0.1)0.0 (1.4)0.0 (1.1)0.0 (1.7)0.0 (2.6)0.0 (0.3)0.0 (6.6)0.0 (6.6)0.0 (0.2)
Table 10. Summary of framework performance across evaluation dimensions.
Table 10. Summary of framework performance across evaluation dimensions.
Evaluation20192022
Soft-cluster prediction RMSE (Day 7, 6-day input)0.1980.173
Hard-clustering baseline accuracy (best across target days)18.5%36.3%
Time-use reconstruction MAE (population, half-hourly)2.44%5.44%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Kannam, N.; Allahviranloo, M.; Vatin, L.A.R. Multi-Day Activity Pattern Inference Using Constrained Gaussian Mixture Model (GMM) Classification. Urban Sci. 2026, 10, 331. https://doi.org/10.3390/urbansci10060331

AMA Style

Kannam N, Allahviranloo M, Vatin LAR. Multi-Day Activity Pattern Inference Using Constrained Gaussian Mixture Model (GMM) Classification. Urban Science. 2026; 10(6):331. https://doi.org/10.3390/urbansci10060331

Chicago/Turabian Style

Kannam, Nikhita, Mahdieh Allahviranloo, and Laure Alice Raymonde Vatin. 2026. "Multi-Day Activity Pattern Inference Using Constrained Gaussian Mixture Model (GMM) Classification" Urban Science 10, no. 6: 331. https://doi.org/10.3390/urbansci10060331

APA Style

Kannam, N., Allahviranloo, M., & Vatin, L. A. R. (2026). Multi-Day Activity Pattern Inference Using Constrained Gaussian Mixture Model (GMM) Classification. Urban Science, 10(6), 331. https://doi.org/10.3390/urbansci10060331

Article Metrics

Back to TopTop