3.1. Data Collection and Survey Design
The survey was administered in Kerman, Iran, from 5 to 7 November 2023 to residents aged 18 and above. Two ownership-specific stated-preference (SP) questionnaires were developed for car owners and non-car owners. Each instrument covered: (i) socio-economic characteristics (e.g., age, gender, occupation, education, driving licence), (ii) the most recent trip (mode, purpose, start/end times), (iii) four SP choice scenarios based on travel time and cost (tailored by ownership group), and (iv) supplementary items on public-transport use and constraints (e.g., bus-use frequency, acceptable access/waiting times, and transfer tolerance). The design was informed by the literature and calibrated to local travel conditions. To improve the transparency of the stated-preference experiment,
Table 1 summarises the ownership-specific choice sets retained for analysis, the number of tasks, and the main attributes shown to respondents. Two questionnaire versions were prepared because car owners and non-car owners faced different feasible travel alternatives. In each SP task, respondents compared the available alternatives mainly in terms of travel time and monetary cost/fare, and then indicated their preferred alternative. For car owners, the analysed choice set included private car, taxi, bus, the proposed LRT alternative, and app-based taxi/ride-hailing. For non-car owners, the private-car alternative was excluded, while the remaining alternatives followed the same task structure. This design allowed LRT to be evaluated against the alternatives that were behaviourally relevant to each ownership group.
The attribute levels were selected to represent realistic differences in travel time and monetary cost among the available and proposed alternatives in Kerman. The levels were anchored to respondents’ current trip conditions, the expected operational characteristics of the proposed LRT system, and values commonly used in previous LRT stated-preference studies. For each ownership group, the levels were calibrated so that LRT was evaluated against behaviourally relevant alternatives rather than unrealistic or unavailable options. This ownership-specific calibration ensured that the SP tasks reflected plausible trade-offs faced by car owners and non-car owners in the Kerman context. A structured scenario-based SP design was used. The design systematically varied the main policy-relevant attributes, particularly travel time and monetary cost, while keeping the number of choice tasks manageable for respondents. Implausible or clearly dominated alternatives were avoided during scenario construction to ensure that respondents faced meaningful time-cost trade-offs rather than obvious choices. Each respondent evaluated four SP tasks, which provided repeated observations for model estimation while limiting cognitive burden.
To further clarify the SP experiment,
Table 2 and
Table 3 report the four choice tasks retained for analysis for car owners and non-car owners, respectively. Rather than presenting only one illustrative example, these tables report all retained SP tasks used in the final modelling dataset. They show the alternatives, travel-time levels, and monetary cost/fare levels presented to respondents in each task. Monetary values are reported in Tomans, consistent with the original questionnaire cards. For comparability with other monetary values reported elsewhere in the paper, approximate USD equivalents were calculated using a rounded exchange rate of 1 USD = 50,000 Tomans (500,000 IRR), corresponding to the November 2023 survey period.
These tables show that the SP tasks varied both travel time and monetary cost/fare across alternatives. For car owners, the private-car option was included to capture direct competition between LRT and private driving. For non-car owners, the same task logic was used, but the private-car option was excluded because it was not a feasible ownership-based alternative.
A pilot survey involving 60 participants, including 30 transportation engineering graduate students and 30 lay participants, was conducted before the main survey. The pilot was used to assess questionnaire clarity, scenario plausibility, the comprehensibility of attribute levels, and the feasibility of completing four SP tasks without excessive respondent burden. Internal consistency was assessed for the repeated LRT choice indicators across the four SP tasks, because LRT adoption is the focal behavioural outcome of the study. Internal consistency was assessed for the SP block using the four repeated LRT choice indicators. For the pooled pilot sample across the car-owner and non-car owner questionnaire versions, Cronbach’s alpha was 0.926, indicating the very good internal consistency of the SP block. Because these four LRT indicators are binary repeated-choice indicators rather than Likert-type scale items, Cronbach’s alpha is used here only as an exploratory internal-consistency check for the SP block. Pilot responses were also screened for missing values, inconsistent responses, and implausible scenario evaluations. Face validity was assessed through expert review and participant feedback, and minor wording and formatting revisions were made before the main survey.
The main survey was conducted through interviewer-administered questionnaires at 12 strategically selected urban locations using a systematic intercept sampling procedure. The survey locations were selected to cover major urban activity and travel-generation points in Kerman, including areas with different public-transport access conditions and trip purposes. At each location, trained interviewers approached every tenth eligible passer-by aged 18 years and above during the survey period. If the selected person declined to participate or was not eligible, the interviewer proceeded to the next tenth eligible passer-by. Respondents were first screened for car-ownership status and then assigned to the corresponding questionnaire version. This procedure was designed to obtain a diverse sample of urban travellers at selected activity locations rather than a fully random household sample of all Kerman residents.
Of 820 completed interviews, 736 respondents remained after respondent-level data cleaning, including 388 car owners and 348 non-car owners. Thus, 84 questionnaires were excluded at the respondent level because they contained incomplete socio-demographic or trip information required for segmentation, incomplete SP responses, or insufficient usable information for subsequent choice modelling. Each respondent encountered four SP choice scenarios, leading to a potential total of 2944 SP choice observations. After choice-observation-level screening, 1916 valid SP observations were retained for model estimation, including 955 from car owners and 961 from non-car owners. To clarify the data-screening process,
Table 4 summarises the transition from completed interviews to the final respondent sample and the retained SP choice observations used for model estimation.
At the choice-observation level, invalid SP tasks were removed only when the task was incomplete, when the selected alternative was missing, or when more than one alternative was selected in the same choice task, making the dependent variable ambiguous. The latter case is what is meant here by a scenario-specific inconsistency. No additional behavioural exclusion rule, such as speed-based rejection, straight-lining, or dominated-alternative violation, was applied at the choice-observation level. Of the 1028 dropped SP observations, 875 were excluded because of incomplete SP tasks or missing selected alternatives, while 153 were excluded because multiple alternatives were selected in the same task. Because invalid tasks were removed at the task level rather than deleting all observations from the corresponding respondent whenever possible, respondents could contribute between one and four valid SP tasks. The final estimation dataset is therefore an unbalanced repeated-choice dataset. In the car-owner sample, respondents contributed one, two, three, and four valid tasks in 106, 124, 31, and 127 cases, respectively. In the non-car owner sample, the corresponding numbers were 57, 97, 66, and 128 respondents.
The adequacy of the sample size for the ownership-specific choice-modelling task is supported by several methodological and statistical considerations. First, the minimum sample-size requirement was checked using Orme’s rule of thumb for stated-preference investigations:
where
N denotes the minimum required sample size,
c denotes the number of alternatives,
t denotes the number of choice tasks per respondent, and
a denotes the maximum number of attribute levels. In the present study,
c = 5 for car owners and
c = 4 for non-car owners,
t = 4, and
a = 4, yielding minimum required sample sizes of 157 for car owners and 125 for non-car owners. These values are well below the retained respondent samples of 388 car owners and 348 non-car owners. Second, the econometric complexity of the employed Nested Logit (NL) and Mixed Logit (MXL) models necessitates a richer dataset than simple logit models, given the need to estimate random-parameter distributions and correlation structures. The provision of approximately 1000 observations per ownership segment (car owners versus non-owners) supplies sufficient degrees of freedom to achieve conventional levels of statistical significance for the majority of covariates. Third, as a survey-size check, the final sample of 736 respondents exceeds the commonly cited minimum sample size of approximately 384 respondents for large populations under a 95% confidence level and 5% margin of error, assuming maximum variability (
p = 0.5) [
38]. Using the same conservative assumption and applying the finite-population margin-of-error expression from Cochran [
38] to an approximate city population of 750,000 gives an indicative margin of error of about 3.6% for the respondent sample. This calculation should be interpreted as a sampling-precision check rather than as a claim of full demographic representativeness, since the survey was conducted at selected urban locations. The indicative margin of error was computed as Equation (2):
where
Z = 1.96 for a 95% confidence level,
p = 0.5,
n = 736, and
N ≈ 750,000.
Fourth, the segmentation strategy preserves analytical robustness after stratification; each subgroup retains roughly 370 respondents (≈950–960 observations), which is adequate to support rigorous comparative inference without incurring excessive stochastic variability. Finally, the estimation results, including the reported pseudo-R2 values and likelihood-ratio statistics, suggest that the retained data contain sufficient information to estimate the ownership-specific choice models and capture relevant passenger behaviour patterns.
3.2. Model Specification and Variable Definition
The analytical framework of this study is grounded in Random Utility Theory (RUT) [
4], which posits that individuals choose the alternative that maximises their perceived utility. The utility
Unj that individual n derives from choosing travel mode
j is composed of a deterministic component
Vnj and a stochastic error term
εnj (Equation (3)).
The deterministic component Vnj is specified as a linear function of observed attributes Xnj (e.g., travel time, cost, and socio-demographic interactions) with corresponding parameters β to be estimated. To robustly analyse mode-choice behaviour towards the hypothetical Light Rail Transit (LRT) system and to explicitly test for structural differences between car owners and non-car owners, we estimate and compare three discrete choice model specifications separately for each group. The MNL, MXL, and NL models were estimated using PythonBiogeme version 3.2.14 (EPFL, Lausanne, Switzerland).
Socio-economic and trip-related variables were included in the utility functions only when they had a clear behavioural interpretation and improved model interpretability. Their inclusion was guided by the literature, local travel conditions, and theoretical expectations regarding access constraints, service familiarity, transfer tolerance, expenditure-based affordability, and mobility limitations. To reduce overfitting concerns, variables with weak behavioural justification or unstable signs were not retained in the final specifications.
In all estimated models, the dependent variable is the chosen alternative in each SP choice task. For car owners, the choice variable takes one of five alternatives: private car, taxi, bus, LRT, and ride-hailing. For non-car owners, it takes one of four alternatives: taxi, bus, LRT, and ride-hailing. The independent variables consist of alternative-specific service attributes, mainly travel time and monetary cost/fare, and selected socio-economic or trip-related interaction variables. Economic-status variables used in the models are based on the monthly household expenditure categories reported in
Table 5. The expenditure dummies are interpreted relative to the omitted reference category and should not be read as separate income variables. The exact variables included in each utility function are reported in the parameter-estimation tables. Blank cells in these tables indicate that the corresponding variable was not retained in that specification. Thus, the three model structures differ mainly in their behavioural assumptions—IID substitution in MNL, random taste heterogeneity in MXL, and within-nest correlation in NL—rather than in the definition of the dependent variable. Accordingly, coefficient-level comparisons are made primarily between MNL and MXL specifications, which share the closest utility structure. NL results are interpreted mainly as supplementary evidence on structured substitution, rather than as coefficient-by-coefficient replications of the MNL/MXL specifications.
3.2.1. Multinomial Logit Model
The MNL model serves as the baseline, assuming the error terms
εnj are independently and identically distributed (IID) following a Gumbel distribution. This leads to the well-known closed-form probability expression shown in Equation (4):
While computationally efficient, the MNL model imposes the Independence of Irrelevant Alternatives (IIA) property, which may be restrictive if certain modes are perceived as closer substitutes [
6].
3.2.2. Mixed Logit Model
To account for unobserved preference heterogeneity across individuals, we employ a Mixed Logit model (also known as Random Parameters Logit). This flexible model allows parameters to vary randomly across the population [
5]. The utility function in the MXL specification (Equation (5)) modifies the basic RUT framework:
where
βn is a vector of individual-specific parameters with density
f(
βθ). The unconditional choice probability is the integral of the conditional logit probability over this distribution, as given by Equation (6):
We estimate this probability using simulated maximum likelihood with 2000 Halton draws, following Train [
6]; stability was checked by comparing alternative draw settings during model development. Because each respondent could contribute more than one valid SP task and the final dataset was unbalanced, respondent-level panel MXL specifications were also estimated as a diagnostic check using the available number of valid tasks per respondent. These panel estimations explicitly accounted for repeated-choice correlation within respondents and confirmed that the substantive behavioural conclusions were not sensitive to treating the retained SP tasks as an unbalanced repeated-choice dataset. In both ownership segments, the signs and substantive interpretation of the key policy-relevant effects remained unchanged, including the negative effect of LRT travel time, the role of private-car operating cost among car owners, and the main public-transport substitution mechanisms among non-car owners. For this reason, the panel MXL results are used as a diagnostic robustness check rather than reported as additional main specifications.
Candidate random parameters were retained only when the estimated heterogeneity was statistically meaningful, behaviourally interpretable, stable, and did not imply implausible sign reversal for time or cost coefficients. Random coefficients with negligible, unstable, or behaviourally implausible spreads were treated as fixed effects in the final MXL specifications. For the car-owner sample, the retained random parameters mainly capture heterogeneity in selected competing-mode travel-time sensitivities. For the non-car owner sample, the retained random parameters capture heterogeneity in bus schedule-related sensitivity and LRT travel-time sensitivity. A constrained triangular distribution was used for retained random parameters because it allows limited preference heterogeneity to be represented while preserving behaviourally interpretable coefficient signs.
3.2.3. Nested Logit Model
To relax the IIA assumption and to model potential correlation among similar alternatives, we test a two-level Nested Logit structure [
6,
39]. Given the distinct choice sets for car owners and non-car owners (
Figure 1), we specify separate nesting structures for each group, reflecting their different available alternatives and potential substitution patterns:
For car owners (
Figure 1a), the tested nesting structure groups private-like motorised alternatives, namely private vehicle, taxi, and ride-hailing, separately from fixed-route public-transport alternatives, namely bus and LRT. This structure reflects the behavioural distinction between flexible, door-to-door or semi-door-to-door motorised modes and scheduled collective transport modes. This grouping does not imply that private car, taxi, and ride-hailing are behaviourally identical. Rather, in the Kerman context, where existing public transport faces limitations in service quality, reliability, passenger facilities, and information provision, these alternatives may be perceived as sharing unobserved attributes related to flexibility, reduced dependence on fixed schedules, and greater perceived individual control relative to conventional public transport. An alternative three-nest car-owner structure separating private car from hired/on-demand motorised services, namely {private car}, {taxi, ride-hailing}, and {bus, LRT}, was also tested. However, the singleton private-car nest produced unstable inclusive-value estimates and, in an alternative run, a singular variance–covariance matrix. Therefore, this alternative structure was not retained.
For non-car owners (
Figure 1b), the tested nesting structure groups taxi and ride-hailing as private-like motorised alternatives, while bus and LRT are grouped as fixed-route public-transport alternatives. This structure reflects the expectation that non-car owners compare LRT mainly with existing public transport on the one hand and flexible paid motorised services on the other.
In the NL model, the probability of individual n choosing alternative
j in nest
m is given by the product of a conditional and a marginal probability (Equation (7)):
where the conditional probability of choosing
given nest
is expressed in Equation (8):
and the marginal probability of choosing nest
is given by Equation (9):
Nest-level covariates enter the marginal nest utility component of the NL model and affect the marginal probability of selecting nest m. They are not part of the inclusive-value/log-sum term itself, which is computed from the conditional utilities of alternatives within each nest. The inclusive-value parameter separately captures the degree of within-nest correlation. The inclusive-value (IV) parameter,
, measures the degree of independence within the nest
. A value between 0 and 1 indicates that the nested structure is consistent with utility maximisation and that alternatives within the same nest are perceived as closer substitutes [
39].