Paper deep dive
Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes
Farbod Abbasi, Zachary Patterson, Bilal Farooq
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 8/22/2026, 3:11:26 AM
Summary
This paper proposes a regularized two-stage generative framework for synthesizing synthetic populations for activity-based travel demand models. Stage 1 uses a Wasserstein GAN with gradient penalty (WGAN-GP) augmented with three regularization terms (IGP, LDR, CLAP) to generate tabular socio-demographic attributes, addressing feasibility, diversity, and novelty. Stage 2 employs Transformer and LSTM-Attention models to generate sequential travel attributes (departure time, trip purpose, travel mode) conditioned on the tabular profiles. The framework is evaluated on the 2018 Montreal Origin-Destination survey, demonstrating that regularized models outperform vanilla WGAN-GP in recovering sampling zeros and generating realistic population proportions.
Entities (11)
Relation Signals (11)
Framework → evaluatedon → Montreal OD Survey
confidence 98% · The framework is evaluated using the 2018 Montreal Origin-Destination survey...
Regularized Models → outperforms → Vanilla WGAN-GP
confidence 97% · Results show that regularized models outperform the vanilla WGAN-GP across feasibility, diversity, and novelty.
Regularization → addresses → Structural Zeros
confidence 96% · guide the generator toward... fewer infeasible samples [structural zeros].
Regularization → addresses → Sampling Zeros
confidence 96% · regularization refers to additional loss terms that guide the generator toward broader valid coverage... improving sampling-zero recovery
Transformer → generates → Sequential Attributes
confidence 96% · In Stage 2, Transformer and LSTM-Attention models generate sequential travel attributes...
LSTM-Attention → generates → Sequential Attributes
confidence 96% · In Stage 2, Transformer and LSTM-Attention models generate sequential travel attributes...
Transformer → conditionson → Tabular Attributes
confidence 95% · generate sequential travel attributes... conditioned on the synthesized tabular profiles.
LSTM-Attention → conditionson →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging. Small samples miss valid attribute combinations, known as sampling zeros, and generative models may also produce infeasible structural zeros. Moreover, realistic synthetic populations must capture both static socio-demographic attributes and sequential travel behaviour, such as trip chains. This paper proposes a regularized two-stage generative framework to address these challenges, where regularization refers to additional loss terms that guide the generator toward broader valid coverage and fewer infeasible samples. In Stage 1, a Wasserstein GAN with gradient penalty is augmented with three regularization terms, IGP, LDR, and CLAP, to improve feasibility, diversity, and novelty in tabular population synthesis. In Stage 2, Transformer and LSTM-Attention models generate sequential travel attributes, including departure time, trip purpose, and travel mode, conditioned on the synthesized tabular profiles. We also introduce novelty and count-aware metrics to evaluate whether valid unseen combinations are recovered and generated in realistic proportions. Results show that regularized models outperform the vanilla WGAN-GP across feasibility, diversity, and novelty. Regularization increases feasibility by 2.1 to 3.7 percentage points and novelty by 6.6 to 10.0 percentage points, improving sampling-zero recovery without sacrificing feasibility. The F1 score improves by 6.3 to 8.6 percentage points. For sequential attributes, LSTM-Attention best matches the trip-length distribution, while Transformer achieves higher overall sequential F1, 90.6\% versus 89.1\%. Cross-stage validation confirms strong consistency between generated mobility status and generated trip chains.
Tags
Links
- Source: https://arxiv.org/abs/2608.15867v1
- Canonical: https://arxiv.org/abs/2608.15867v1
Trouble viewing inline? Open PDF directly →
Full Text
85,681 characters extracted from source content.
Expand or collapse full text
Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel AttributesJournal: Nuclear Physics B Farbod Abbasi Email: farbod.abbasi@mail.concordia.ca Corresponding author: Corresponding author Affiliation: Concordia University, Montreal, Quebec, Canada Zachary Patterson Email: Zachary.Patterson@concordia.ca Affiliation: Concordia University, Montreal, Quebec, Canada Bilal Farooq Email: bilal.farooq@torontomu.ca Affiliation: Toronto Metropolitan University, Toronto, Ontario, Canada Abstract Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging. Small samples miss valid attribute combinations, known as sampling zeros, and generative models may also produce infeasible structural zeros. Moreover, realistic synthetic populations must capture both static socio-demographic attributes and sequential travel behaviour, such as trip chains. This paper proposes a regularized two-stage generative framework to address these challenges, where regularization refers to additional loss terms that guide the generator toward broader valid coverage and fewer infeasible samples. In Stage 1, a Wasserstein GAN with gradient penalty is augmented with three regularization terms, IGP, LDR, and CLAP, to improve feasibility, diversity, and novelty in tabular population synthesis. In Stage 2, Transformer and LSTM-Attention models generate sequential travel attributes, including departure time, trip purpose, and travel mode, conditioned on the synthesized tabular profiles. We also introduce novelty and count-aware metrics to evaluate whether valid unseen combinations are recovered and generated in realistic proportions. Results show that regularized models outperform the vanilla WGAN-GP across feasibility, diversity, and novelty. Regularization increases feasibility by 2.1 to 3.7 percentage points and novelty by 6.6 to 10.0 percentage points, improving sampling-zero recovery without sacrificing feasibility. The F1 score improves by 6.3 to 8.6 percentage points. For sequential attributes, LSTM-Attention best matches the trip-length distribution, while Transformer achieves higher overall sequential F1, 90.6% versus 89.1%. Cross-stage validation confirms strong consistency between generated mobility status and generated trip chains. Keywords: Population synthesis , Tabular attributes , Sequential attributes , Regularization term , Transformer , Generative adversarial networks 1 Introduction Transportation planning increasingly relies on activity-based models (ABMs) to simulate individual-level travel behavior by representing daily activity and trip schedules at a disaggregate level (11; 14). ABMs capture the complex interdependencies between socio-demographic characteristics and mobility decisions, enabling more accurate forecasts of travel demand in support of transportation planning and decision-making (10). A fundamental requirement of ABMs is a synthetic population that is statistically representative of the true population in the modelled region (3). The quality of this synthetic population directly conditions the realism and reliability of any downstream simulation (8). However, generating a reliable synthetic population is a fundamentally difficult problem. Due to privacy concerns and the high cost of data collection 6, researchers must rely on limited survey samples that typically cover only one to five percent of the actual population (18). These samples, such as public use microdata samples (PUMS) or regional household travel surveys (HTS), provide rich disaggregate information on socio-demographic and behavioural attributes. Crucially, capturing the joint distribution of individual attributes from such data, rather than merely reproducing marginal distributions, is essential for behavioural realism (12; 26). Even when disaggregate data are available, three interrelated challenges limit the quality of synthesized populations. The first concerns the achievement of feasibility, diversity, and novelty in tabular attribute synthesis (12; 8; 13). Feasibility requires that every generated individual maps to a logically possible combination of attributes. Diversity requires that the generator reproduce the full range of valid attribute combinations present in the true population. Novelty requires that the generator recover valid combinations that exist in the true population but are absent from the training sample due to its limited coverage. These three objectives are in natural tension. A model that aggressively pursues diversity and novelty risks generating infeasible combinations, while a model that is overly conservative may collapse onto a subset of common training patterns and fail to recover unseen combinations. Two key concepts capture this challenge. Sampling zeros are valid attribute combinations that exist in the true population but are absent from the training sample due to limited coverage; a model that fails to recover them produces a population that lacks both diversity and novelty. Structural zeros are logically impossible combinations, such as a child holding a driver’s license or a young person in full retirement; a model that generates them produces an infeasible and unrealistic population. Addressing feasibility, diversity, and novelty therefore requires explicit regularization terms added to the loss function to recover sampling zeros while suppressing structural zeros. The second challenge concerns the type and structure of attributes that a synthetic population must represent (7). Real-world individuals are characterized not only by static tabular attributes, such as age, gender, employment status, and vehicle ownership, but also by sequential behavioural attributes, including trip purpose sequences, departure time patterns, and travel mode chains. These sequential attributes are fundamentally different in structure from tabular ones. They carry temporal dependencies, exhibit variable length, and are closely linked to individual socio-demographic characteristics. A synthetic population that captures only tabular attributes is incomplete and cannot serve as a credible input to activity-based modeling. The third challenge concerns how the quality of synthetic populations is evaluated. Existing evaluation frameworks primarily assess feasibility and diversity (23). While feasibility is a necessary and valid criterion, relying on diversity alone provides an incomplete picture of generative model performance. Diversity does not reveal where a model succeeds or fails. It does not distinguish between reproducing combinations already seen during training and recovering sampling zeros. A model that achieves high diversity by reproducing common training combinations without recovering sampling zeros may appear to perform well while failing to generalize. We therefore introduce novelty, defined as the share of valid combinations absent from the training sample but present in real population, providing a measure of generalization capacity that diversity alone cannot offer. A further limitation of conventional metrics is that they treat all attribute combinations as equally important regardless of frequency, allowing a model that misrepresents population proportions to still score well. We address this by introducing count-aware variants of both diversity and novelty, which assess whether combinations are generated in proportions consistent with the true population. Together, feasibility, diversity, novelty, and their count-aware counterparts contain a more complete basis for evaluating synthetic population quality. No existing framework addresses these limitations simultaneously. Studies that propose regularization strategies focus exclusively on tabular attribute generation and do not model sequential behaviour (23). Conversely, studies that generate sequential attributes alongside tabular ones do not incorporate explicit mechanisms to address sampling zeros or structural zeros (7; 6). Moreover, existing evaluation frameworks remain limited, assessing feasibility and conventional diversity while providing no means to separately measure novelty or the distributional accuracy of generated combinations. To the best of our knowledge, no existing framework addresses these three challenges jointly. This paper addresses all three challenges within a unified two-stage generative framework. In the first stage, a Wasserstein Generative Adversarial Network (WGAN) (5) with gradient penalty (17) is used to synthesize tabular attributes. To improve feasibility, diversity, and novelty together, we introduce and compare three soft regularization mechanisms incorporated into the generator loss. Unlike hard constraints, these terms do not remove all violations by design; instead, they guide the generator toward a better balance between recovering valid unseen combinations and reducing structurally invalid profiles. In the second stage, a Transformer-based model (38) and an LSTM with attention mechanism are trained to generate sequential behavioural attributes, namely trip purpose, departure time, and travel mode, conditioned on the tabular attributes. These two architectures are compared to assess which better captures the transition structures and temporal dependencies of real behavioural sequences. The framework is evaluated using the 2018 Montreal Origin-Destination survey, which covers four percent of the regional population and provides both rich socio-demographic and detailed travel diary information. The main contributions of this study are threefold. First, we propose and evaluate three soft regularization terms for GAN-based tabular population synthesis to improve the balance between feasibility, diversity, and novelty. Second, we extend synthetic population generation beyond static tabular attributes by generating conditional sequential travel attributes using Transformer and LSTM-attention models. Third, we propose a comprehensive evaluation framework that assesses the quality of the generated population across both tabular and sequential dimensions. Together, these contributions provide a more complete basis for generating and evaluating synthetic populations for activity-based travel demand models. The remainder of this paper is organized as follows. Section 2 reviews the relevant literature on population synthesis methods, deep generative models, diversity and feasibility in synthetic data, and sequential attribute generation. Section 3 describes the Montreal Origin-Destination survey and the case study setup. Section 4 presents the proposed methodology, including the regularized GAN framework and the Transformer and LSTM-based sequence generation models. Section 5 reports and discusses the empirical results across all evaluation dimensions. Section 6 concludes the paper with key insights and directions for future research. 2 Literature Review We first discuss traditional and deep learning approaches for tabular attribute synthesis, then examine recent efforts to incorporate sequential behavioural features. Finally, we review strategies for improving diversity and feasibility and identify remaining research gaps. 2.1 From Traditional Methods to Deep Generative Models Early population synthesis methods relied on marginal-fitting techniques, most notably iterative proportional fitting (IPF), which adjust sample weights to match known aggregate totals from census or administrative data(40; 30). While computationally efficient and easy to implement, these methods are fundamentally limited to reproducing marginal distributions and cannot capture the joint dependencies between individual attributes that are essential for behavioural realism. To address this limitation, probabilistic simulation methods were introduced, including Markov Chain Monte Carlo (MCMC) and Hidden Markov Model (HMM) based approaches, which approximate joint distributions by drawing samples from conditional probability structures (14; 33; 25). These methods offer greater flexibility and can generate new individuals beyond those observed in the sample, but they rely on strong structural assumptions and their scalability deteriorates as the number of attributes increases. Bayesian network approaches further advanced joint distribution modelling by explicitly learning dependency structures among attributes35; 31, while hierarchical and multilevel models extended this to capture household and individual level relationships simultaneously34; 25. More recently, integrated pipelines have combined synthesis with spatial assignment and dynamic updating to improve practical applicability 15; 19; 24. Despite these advances, traditional methods share a common limitation: none explicitly addresses sampling zeros and structural zeros, and most are constrained to reproducing combinations already present in the training sample. 2.2 Diversity and Feasibility in Generative Population Synthesis Deep generative models, including VAEs, GANs, diffusion models, and LLMs, have emerged as a promising alternative to traditional methods, owing to their ability to learn complex joint distributions and generate novel individuals beyond those observed in the training sample. However, the extent to which these models address sampling zeros and structural zeros varies considerably. To characterize this variation, we organize existing studies into four groups. Studies in G1 ignore both challenges entirely, treating synthesis as a pure distribution-matching problem (39; 29). Studies in G2 recognize the existence of these challenges but propose no formal metrics or mechanisms to address them (1; 9). Studies in G3 formally measure sampling and structural zeros, establishing useful evaluation frameworks, but they do not introduce any mechanism to control this during training (16; 21; 22; 20; 36; 32). Only studies in G4 explicitly control and improve both challenges through mechanisms incorporated into the generative process: 23 proposed the first regularization framework combining diversity and feasibility objectives, while 27 leveraged LLM-based temperature scaling for controllable generation. However, even within G4, existing evaluation frameworks remain incomplete in two ways. First, they measure coverage of observed attribute combinations but do not distinguish between combinations seen in the training sample and valid unseen combinations, making it difficult to assess generalization. Second, they often treat all combinations equally and do not evaluate whether generated data reflect realistic population proportions. This study addresses these two limitations by introducing novelty and count-aware metrics. Novelty measures the recovery of valid combinations that are absent from the training sample but present in the full population. The count-aware versions of diversity and novelty further evaluate whether these recovered combinations are generated in proportions that are consistent with the real population. Crucially, all studies across G1 to G4, regardless of their level of engagement with sampling zeros and structural zeros, focus exclusively on tabular attribute generation and do not model sequential behavioral attributes, a limitation that we address in the following section. 2.3 Sequential Attribute Generation in Synthetic Populations There has been limited research to extend synthetic population generation beyond tabular attributes to include sequential behavioral features such as trip purpose, departure time, and travel mode chains. 7 introduced the Composite Travel GAN, the first joint generative framework combining a GAN for tabular attributes with a SeqGAN reinforcement learning module for sequential generation, establishing the foundational architecture for this line of work. 6 proposed a modular pipeline integrating CTGAN for tabular synthesis with an RNN-based model for trip sequence generation, connected through a Hungarian matching algorithm. 2 focused specifically on spatio-temporal sequence generation using a multi-headed GAN with GRU, jointly modeling location and time without a tabular component. More recently, 28 introduced a multimodal conditional framework combining CTGAN, VQ-VAE, and Transformer architectures with contrastive learning to link individual attributes to activity sequences and locations. Despite these advances, existing studies do not explicitly address sampling zeros or structural zeros in the sequential generation process. This study addresses this gap by proposing a unified two-stage framework for tabular and sequential generation. The framework incorporates explicit mechanisms to improve novelty, diversity and feasibility in the tabular stage, while also introducing novelty and count-aware evaluation metrics to assess how well the generated sequential attributes remain realistic and consistent with the tabular population. Recent studies have made important progress in synthetic population generation, but several gaps remain. First, most tabular synthesis methods either ignore sampling zeros and structural zeros or only evaluate them after generation, without incorporating mechanisms to control them during training. Second, studies that do address these issues focus mainly on static tabular attributes and do not generate sequential travel behaviour. Third, existing sequential generation approaches model trip chains but do not explicitly address sampling zero recovery and structural zero reduction. Finally, current evaluation frameworks remain limited because they do not fully assess novelty and count-aware performance. These gaps motivate the proposed framework, and Table 1 summarizes how the present study differs from the existing literature. Table 1: Summary of related work on synthetic population generation. ✓ = addressed, × = not addressed, ∼ = partially addressed. Samp. = Sampling Zeros; Struct. = Structural Zeros; Seq. = Sequential generation; Eval. = Evaluation completeness (novelty and count-aware metrics). Study Model Seq. Samp. Zeros Struct. Zeros Grp. Eval. Traditional Methods 40; 30 IPF / QISI × × × Trad. × 14; 33; 25 MCMC / HMM × × × Trad. × 35; 31 Bayesian Network × × × Trad. × 34; 25 Hierarchical / Multilevel × × × Trad. × 19; 15 Integrated Pipeline × × × Trad. × G1: Ignore Sampling & Structural Zeros 29 CTGAN / VAE × × × G1 × 39 GAN + DAG × × × G1 × G2: Recognize (No Metrics, No Control) 9 VAE × Recognized Recognized G2 × 1 VAE + CVAE × Recognized × G2 × G3: Measure (No Training-Time Control) 16 WGAN / VAE × Measured Measured G3 × 21 Copula + GAN × Measured Measured G3 × 22 Diffusion × Measured ∼ G3 × 36 Diffusion × Measured Measured G3 × 32 CT-GAN × Measured ∼ G3 × 20 CVAE / CGAN × Measured × G3 × G4: Control & Improve 23 Reg. GAN / VAE × Controlled Controlled G4 × 27 LLM + BN × Controlled Controlled G4 × Sequential Generation Studies 7 GAN + SeqGAN ✓ × × Seq. × 6 CTGAN + RNN ✓ × × Seq. × 2 GAN + GRU ✓ × × Seq. × 28 CTGAN + VQ-VAE + Transformer ✓ × × Seq. × Present Study Present Study Reg. GAN + Transformer / LSTM ✓ Controlled Controlled Ours ✓ 3 Data and Case study Since access to data on an entire population is not feasible, to perform population synthesis we use the 2018 Montreal Origin Destination (OD) survey, which draws a sample of approximately 4% of the total population. The OD dataset encompasses detailed travel information for 162,588 individuals. Table 2 provides descriptive statistics of the OD dataset. It contains 7 individual attributes with a total of 35 categorical classes. Table 2: Tabular socio-demographic attributes and category proportions in the 2018 Montreal OD dataset. Attribute (Dimensions) Category Proportion (%) 1. Number of vehicles (m_auto) 0 9.49 1 36.41 2 39.95 3+ 14.15 2. Household size (m_pers) 1 13.03 2 34.92 3 18.14 4 24.03 5 9.88 3. Gender (p_sexe) Male 48.77 Female 51.23 4. Age group (p_grage) 0–4 years 3.63 5–9 years 5.07 10–14 years 5.48 15–19 years 5.12 20–24 years 4.60 25–34 years 8.71 35–44 years 12.66 45–54 years 14.39 55–64 years 18.13 65–74 years 14.03 75 years and older 8.19 5. Employment status (p_statut) Full-time worker 40.71 Part-time worker 4.67 Student 19.55 Retired 26.43 Other 2.74 children under 4 years 3.63 At home 2.27 6. Driver’s license (p_permis) Yes 72.52 No 12.30 Not applicable 15.19 7. Mobility (p_mobil) Yes 78.71 No 17.66 children under 4 years 3.63 In addition to these static attributes, the dataset contains three sequential trip-level features: departure time, trip purpose, and trip mode. These sequential variables represent each individual’s daily travel behaviour as an ordered trip chain. The sequence length represents the total number of trips made by an individual. Figure 8 illustrates the distribution of trip sequence lengths in the OD dataset. As shown, the majority of individuals make between zero and four trips per day, while longer trip chains occur less frequently. Figure 1: Trip Sequence Length Distribution From this point onward, the OD dataset is treated as a representative sample of the full population. A subset of this dataset is designated as training data, and our objective is to generate synthetic samples that can replicate the complete set of individuals present in the full OD dataset. We can reasonably assume that if we are able to accurately reconstruct the full OD dataset using only a subset of it as training data, then it would also be feasible to generate the entire population through the OD dataset itself. The concept of sampling zeros refers to valid feature combinations that exist in the full population but are missing from the training dataset due to limited sample size. These are not infeasible combinations, and they simply do not appear in the sample, which poses a challenge for generative models attempting to replicate the full population accurately. Figure 2 illustrates this issue. The blue bars represent unique combination coverage, which is the percentage of unique feature combinations in the full dataset captured at each sample size. As the sample size increases, a larger proportion of the unique combinations present in the full dataset is recovered. Correspondingly, The green bars represent population mass coverage, defined as the proportion of total individuals in the OD dataset who fall into these captured combinations. For example, at a 1% sample size, 14.7% of the unique feature combinations are observed. However, these combinations account for 78.2% of the individuals in the dataset. This highlights that the training data lacks diversity, and therefore, a synthetic model trained on it must be capable of generating plausible combinations not seen in the training samples. Figure 2: Relationship between sample size and sampling zeros We focus our analysis on the 1% sample size. To claim that a synthetic population is truly reliable, it must increase the number of unique feature combinations beyond what is observed in the training data and recover the 11.8% of individual records from the full OD dataset that are not included in the training set, by generating samples that match with missing combinations. 4 Methodology 4.1 Framework Overview This paper proposes a two-stage generative framework for synthesizing individual-level travel behavior profiles consistent with the distributional properties of an observed OD survey. Let each individual in the population be represented by a pair (X,S)(X,S), where X∈ℝdX ^d denotes a vector of tabular attributes and S=(s1,s2,…,sT)S=(s_1,s_2,…,s_T) denotes a sequence of length T. The objective of synthetic population generation is to learn the joint distribution P(X,S)P(X,S) from observed data and generate new synthetic samples (X~,S~)( X, S) that preserve statistical realism, novelty, and behavioural feasibility. Directly modelling the joint distribution of tabular and sequential attributes is challenging due to their different structures. To address this, we decompose the joint distribution as: P(X,S)=P(X)P(S∣X).P(X,S)=P(X)P(S X). (1) Based on Figure 3, two generative models are trained to learn the distribution of tabular attributes. In the first stage, a Wasserstein GAN (5) with Gradient Penalty (17) (WGAN-GP), enhanced with three regularization terms to improve the coverage of valid attribute combinations. In the second stage, a Transformer model and an LSTM model augmented with Attention are trained to generate sequences attributes conditioned on tabular attributes. Figure 3: Two stage framework for synthetic population generation 4.2 Stage 1: Tabular Attribute Synthesis 4.2.1 WGAN-GP Architecture In the first stage of the proposed framework, a WGAN-GP is used to generate tabular socio-demographic attributes. WGAN-GP is adopted because it provides a more stable training process than the standard GAN (4) and reduces the risk of mode collapse by optimizing the Wasserstein distance between the real and generated data distributions (37). A GAN consists of two competing neural networks: a generator and a critic. The generator G receives a random noise vector z∼pz(z)z p_z(z) from the latent space and maps it into a synthetic individual profile x~=G(z) x=G(z). In this study, each generated profile represents a combination of categorical tabular attributes such as household size, vehicle ownership, age group, employment status, driver’s license status, gender, and mobility status. The critic D, instead of classifying samples as real or fake, assigns a scalar score to each sample and learns to distinguish the distribution of real individuals from the distribution of generated individuals. The WGAN objective encourages the generator to produce synthetic samples that receive high critic scores, while the critic is trained to assign higher scores to real samples than to generated ones. To improve training stability and enforce the Lipschitz constraint required by WGAN, a gradient penalty term is added to the critic loss. The critic loss is defined as: ℒD=z∼pz(z)[D(G(z))]−x∼pr(x)[D(x)]+λGPx^[(‖∇x^D(x^)‖2−1)2],L_D=E_z p_z(z)[D(G(z))]-E_x p_r(x)[D(x)]+ _GPE_ x [ (\| _ xD( x)\|_2-1 )^2 ], (2) where x denotes a real sample, G(z)G(z) denotes a generated sample, x x is an interpolated sample between real and generated data, and λGP _GP controls the strength of the gradient penalty. The standard generator loss in WGAN-GP is defined as: ℒG=−z∼pz(z)[D(G(z))].L_G=-E_z p_z(z)[D(G(z))]. (3) This adversarial loss encourages the generator to create synthetic samples that are increasingly similar to the real population distribution. However, matching the overall distribution alone does not guarantee that the generated population is sufficiently diverse, feasible, or novel. In particular, a generator may still fail to recover valid combinations that are absent from the training sample, or it may generate unrealistic attribute combinations. To address this limitation, we extend the generator objective by incorporating an additional regularization term: ℒGreg=−z∼pz(z)[D(G(z))]+λregℒreg,L_G^reg=-E_z p_z(z)[D(G(z))]+ _regL_reg, (4) where ℒregL_reg denotes one of the proposed regularization terms and λreg _reg controls its contribution to the generator loss. These regularization terms are designed to guide the generator toward better exploration of the latent space, improving the recovery of rare but valid combinations while preserving feasibility. The following section describes the proposed regularization terms used to enhance diversity, novelty, and feasibility in tabular population synthesis. 4.2.2 Regularization Terms A key challenge in tabular population synthesis is that a small training sample cannot represent all valid combinations that exist in the full population. In this study, the model is trained using only a 1% sample of the OD dataset. As a result, many valid but low-frequency combinations may be absent from the training data. A standard WGAN-GP may therefore learn to reproduce only the most frequent combinations and fail to generate rare but realistic individuals. This limits the diversity and novelty of the generated population. To address this issue, three regularization terms are incorporated into the WGAN-GP generator loss introduced in the previous section. These terms are designed to encourage the generator to make better use of the latent space and produce a wider range of valid attribute combinations. Each regularization term is tested separately in order to evaluate its effect on diversity, novelty, and feasibility. Figure 4 provides a schematic overview of the regularized WGAN-GP used in Stage 1 for tabular attribute synthesis. Figure 4: Overview of the proposed regularized WGAN-GP framework for tabular attribute synthesis Regularization provides a flexible way to guide the generator during training, but it should not be interpreted as a hard feasibility constraint. The proposed terms cannot guarantee that all structurally invalid profiles will be removed or that all missing valid combinations will be recovered. Instead, they act as soft penalties in the generator loss and encourage more useful exploration of the latent space. This is important because the valid population space is only partially observed in the training sample. Therefore, strict constraints may improve feasibility but can also prevent the model from recovering rare but valid combinations. Compared with existing distance-based regularization approaches (23), the proposed terms place more emphasis on controlled exploration and novelty. Prior work mainly uses regularization to reduce infeasible generation by discouraging samples that move far from valid regions of the data space. In contrast, our approach encourages the generator to explore broader latent-space mappings, to recover valid combinations absent from the limited training sample. Therefore, The goal is to improve novelty and diversity without reducing feasibility. Inverse Gradient Penalty (IGP) The IGP term encourages the generator to be sensitive to changes in the latent space. The motivation is that if two latent vectors are different, their generated outputs should also be sufficiently different. Otherwise, the generator may ignore parts of the latent space and map many different latent inputs to the same or very similar outputs. For two latent vectors z1z_1 and z2z_2, IGP compares the distance between their generated outputs with the distance between the latent vectors themselves. The loss is defined as: ℒIGP=−z1,z2[min(‖G(z1)−G(z2)‖2‖z1−z2‖2,τ)],L_IGP=-E_z_1,z_2 [ ( \|G(z_1)-G(z_2) \|_2 \|z_1-z_2 \|_2,τ ) ], (5) where τ is a threshold that limits the maximum value of the ratio. A larger ratio means that changes in the latent space lead to meaningful changes in the generated output. Because the loss has a negative sign, minimizing it encourages the generator to increase this ratio up to the threshold. In this way, IGP helps the generator use the latent space more effectively and produce a more diverse and novel samples. Latent Diversity Regularization (LDR) The LDR term directly encourages nearby points in the latent space to generate different outputs. Without this regularization, small changes in the latent vector may produce almost identical individuals, which indicates local mode collapse. LDR reduces this problem by rewarding the generator when a small movement in the latent space leads to a noticeable change in the generated output. For a latent vector z and a small perturbation δ, the LDR loss is defined as: ℒLDR=−z,δ[‖G(z+δ)−G(z)‖2],L_LDR=-E_z,δ [ \|G(z+δ)-G(z) \|_2 ], (6) The negative sign means that minimizing this loss encourages the distance between G(z+δ)G(z+δ) and G(z)G(z) to become larger. In other words, the generator is encouraged to produce more distinct samples for nearby latent inputs. This improves local diversity and helps the model explore more valid combinations. Cross-Latent Agreement Penalty (CLAP) The CLAP term promotes global diversity by discouraging the generator from producing similar outputs for different latent vectors. While LDR focuses on nearby points in the latent space, CLAP compares outputs generated from two independently sampled latent vectors. If two different latent vectors produce very similar synthetic individuals, the model receives a larger penalty. For two latent vectors z1z_1 and z2z_2, the CLAP loss is defined as: ℒCLAP=z1,z2[exp(−‖G(z1)−G(z2)‖2)].L_CLAP=E_z_1,z_2 [ (- \|G(z_1)-G(z_2) \|_2 ) ]. (7) This penalty is large when the generated outputs are close to each other and becomes smaller as the outputs become more different. Therefore, CLAP encourages the generator to spread generated samples across the output space and reduces duplication among generated individuals. Overall, the four regularization terms target diversity and novelty from different perspectives. These regularized WGAN-GP variants are compared with the vanilla WGAN-GP to assess whether they improve the recovery of valid missing combinations while maintaining feasibility. 4.3 Stage 2: Sequential Attribute Generation Stage 2 generates sequential behavioural attributes conditioned on the tabular profile produced in Stage 1. Each individual is associated with three aligned categorical sequences: departure time group, trip purpose, and main travel mode. Let X denote the tabular attribute vector, and let StimeS^time, SpurposeS^purpose, and SmodeS^mode denote the three behavioral sequences. The conditional joint distribution of these aligned sequences is modelled autoregressively as: P(Stime,Spurpose,Smode∣X)=∏t=1TP(sttime,stpurpose,stmode∣s<ttime,s<tpurpose,s<tmode,X).P(S^time,S^purpose,S^mode X)= _t=1^TP\! (s_t^time,s_t^purpose,s_t^mode s_<t^time,s_<t^purpose,s_<t^mode,X ). (8) This formulation reflects that, at each trip index t, the model predicts three aligned behavioural attributes rather than a single token. The history available to the model consists of all previously generated departure time, purpose, and mode tokens, together with the tabular attributes. 4.3.1 Transformer Model The Transformer model conditions on the tabular attributes by mapping X into a context representation, which is inserted as a prefix token at the beginning of the sequence. The embedded trip tokens are then appended to this context token and passed through a causally masked Transformer encoder. The causal mask ensures that each trip position can attend only to the tabular context and previous trips, preserving the autoregressive structure of the generation process. At each trip index t, the Transformer produces a shared hidden representation hth_t. From this shared representation, three separate output heads generate probability distributions over departure time, trip purpose, and travel mode: P(sttime∣s<ttime,s<tpurpose,s<tmode,X)=Softmax(Wtimeht),P\! (s_t^time s_<t^time,s_<t^purpose,s_<t^mode,X )=Softmax(W_timeh_t), (9) P(stpurpose∣s<ttime,s<tpurpose,s<tmode,X)=Softmax(Wpurposeht),P\! (s_t^purpose s_<t^time,s_<t^purpose,s_<t^mode,X )=Softmax(W_purposeh_t), (10) P(stmode∣s<ttime,s<tpurpose,s<tmode,X)=Softmax(Wmodeht),P\! (s_t^mode s_<t^time,s_<t^purpose,s_<t^mode,X )=Softmax(W_modeh_t), (11) where WtimeW_time, WpurposeW_purpose, and WmodeW_mode are head-specific projection matrices corresponding to their respective vocabularies. Let =time,purpose,modeK=\time,purpose,mode\. The negative log-likelihood sequence objective is written as: ℒseq=−∑t=1T∑k∈logP(stk∣s<ttime,s<tpurpose,s<tmode,X).L_seq=- _t=1^T _k P\! (s_t^k s_<t^time,s_<t^purpose,s_<t^mode,X ). (12) This shared-representation and multi-head output structure allows the model to capture dependencies among departure time, trip purpose, and travel mode while preserving the categorical structure of each behavioural attribute. During inference, the Transformer generates the sequence autoregressively, sampling the three trip attributes at each step from their predicted distributions. 4.3.2 LSTM with Attention Model Although the Transformer model provides a flexible architecture for capturing long-range dependencies through self-attention, it is also more complex and computationally demanding. Therefore, we also implement an LSTM with attention as a second sequence-generation model. This model provides a recurrent alternative that is simpler in structure while still allowing the use of attention over previous trips. Comparing the Transformer and LSTM-attention models allows us to evaluate whether the additional complexity of the Transformer leads to better sequential attribute generation performance. The LSTM-attention model conditions on the tabular attributes by initializing the recurrent hidden and cell states. The embedded trip sequence is then processed by the LSTM: O=LSTM([e1,…,eT],h0,c0),O=LSTM ([e_1,…,e_T];h_0,c_0 ), (13) where O=(o1,…,oT)O=(o_1,…,o_T) denotes the sequence of LSTM outputs, and h0h_0 and c0c_0 are initialized from the tabular attributes. A causal attention layer is applied to the LSTM outputs so that each trip position can selectively use information from previous trips while preventing access to future trips: A=softmax(OOTdmodel+M)O,A=softmax ( O^T d_model+M )O, (14) where M is a causal mask that blocks attention to future positions. The attended representation is then refined using residual connections, layer normalization, and a feedforward layer. The final hidden representation at each trip index is passed through the same three-head output structure used in the Transformer model to predict departure time, trip purpose, and travel mode. During inference, the model generates the sequence one trip at a time while maintaining its recurrent state and applying attention over the generated history. Overall, the Transformer relies fully on self-attention to model dependencies across the trip chain, while the LSTM-attention model combines recurrent memory with attention-based refinement. Comparing these two models allows us to evaluate which architecture better captures sequential travel behavior conditioned on tabular attributes. 4.4 Evaluation Framework To evaluate the quality of the synthetic population, we use a framework that considers both tabular and sequential attributes. For tabular attributes, we evaluate feasibility, diversity, novelty, and the count-aware variants of diversity and novelty. For sequential attributes, we assess whether the generated trip chains reproduce the temporal patterns observed in the real OD data. Let CrealC_real denote the set of unique tabular attribute combinations observed in the full OD dataset, CsampleC_sample the set of combinations observed in the training sample, and CfakeC_fake the set of combinations observed in the synthetic population. Similarly, let nreal(c)n_real(c), nsample(c)n_sample(c), and nfake(c)n_fake(c) denote the number of individuals with combination c in the full OD dataset, training sample, and synthetic population, respectively. The set of valid combinations that are absent from the training sample is defined as: Cunseen=Creal−Csample.C_unseen=C_real-C_sample. (15) This set represents sampling zero combinations. Combinations that exist in the full population but are not observed in the limited training sample. 4.4.1 Feasibility Feasibility measures whether the generated individuals correspond to valid attribute combinations. In this study, a generated combination is considered feasible if it appears at least once in the full OD dataset. Feasibility is defined as: Feasibility=∑c∈Crealnfake(c)Nfake.Feasibility= _c∈ C_realn_fake(c)N_fake. (16) A value of 1 indicates that all generated individuals belong to combinations observed in the full OD dataset. Lower values indicate that the model has generated invalid or structurally unrealistic tabular profiles. 4.4.2 Diversity and Count-Aware Diversity Diversity evaluates how much of the full population’s combinatorial support is recovered by the synthetic population. It is measured as the fraction of real combinations that also appear in the generated data: Diversity=|Creal∩Cfake||Creal|.Diversity= |C_real∩ C_fake||C_real|. (17) This metric captures whether the model can generate a wide range of valid attribute combinations. However, it only considers whether a combination appears or not, and does not account for how frequently each combination is generated. To address this limitation, we also use count-aware diversity, which compares the generated and real counts for each valid combination: Diversityw=∑c∈Crealmin(nreal(c),nfake(c))∑c∈Crealnreal(c).Diversity_w= _c∈ C_real (n_real(c),n_fake(c)) _c∈ C_realn_real(c). (18) This metric rewards the model not only for recovering valid combinations, but also for generating them in proportions that are consistent with the full OD dataset. Therefore, the count-aware diversity score provides a more informative and stricter evaluation than the standard diversity score, because it accounts not only for whether combinations are recovered, but also for whether they are generated with realistic frequencies. 4.4.3 Novelty and Count-Aware Novelty Novelty measures the ability of the model to recover valid combinations that were not present in the training sample. These combinations are important because they correspond to sampling zeros. Novelty is defined as: Novelty=|Cunseen∩Cfake||Cunseen|.Novelty= |C_unseen∩ C_fake||C_unseen|. (19) A higher novelty score indicates that the model is better able to generalize beyond the observed training combinations and recover valid but unseen profiles. As with diversity, the standard novelty metric only measures whether unseen combinations are recovered, regardless of their frequency. Therefore, we also compute count-aware novelty: Noveltyw=∑c∈Cunseenmin(nreal(c),nfake(c))∑c∈Cunseennreal(c).Novelty_w= _c∈ C_unseen (n_real(c),n_fake(c)) _c∈ C_unseenn_real(c). (20) This metric evaluates whether the model recovers unseen valid combinations in proportions that are consistent with their frequency in the full OD dataset. Together, novelty and count-aware novelty provide a direct measure of the model’s capacity to recover sampling-zero combinations. 4.4.4 Sequential Evaluation Metrics In addition to tabular quality, the generated population must reproduce realistic trip chain behaviour. We therefore evaluate the sequential attributes from two complementary perspectives. First, we compare the trip length distribution of the generated data against the real OD dataset. For each individual sequence, the trip length is defined as the number of trips in the sequence. This evaluation measures whether the generated population reproduces the overall number of trips per individual. Let LR(k)L_R(k) and LG(k)L_G(k) denote the number of real and generated records with trip length k, respectively. The trip length distribution is then compared visually using bar charts across the real, and generated datasets. Furthermore, we assess whether the generated population reproduces complete and interpretable activity-sequence patterns. For this evaluation, we identify the most frequent trip-purpose sequences in the real OD dataset and compare their relative shares with the corresponding shares in the generated population. To further evaluate the sequential quality of the generated data, we analyze the diversity, novelty, and feasibility of the generated sequence attributes. For each sequential attribute, we extract unigrams, bigrams, and trigrams from the real data, the generated data, and the random sample. Unigrams represent individual sequence elements, bigrams capture pairwise transitions, and trigrams represent longer sequential patterns. Diversity is evaluated by comparing the number of unique n-gram patterns produced by the generated data with those observed in the real data and the sample. This indicates whether the generated data can reproduce a broad range of sequential patterns rather than only repeating the limited sample. Novelty is assessed by identifying generated n-grams that do not appear in the sample dataset. Some of these novel patterns may correspond to valid patterns that exist in the full real dataset but were missed by the small sample. Therefore, we also examine how many missing real patterns are recovered by the generated data. Feasibility is evaluated by identifying generated patterns that are not supported by the real data. These generated-only patterns may indicate unrealistic or invalid sequential combinations. Figure 5, illustrates the overall structure of the evaluation framework. Full OD Data1% Training SampleGenerated PopulationEvaluation FrameworkTabular QualitySequential QualityCross-StageConsistencyFeasibilityDiversity / DiversitywNovelty / NoveltywF1 scores Trip lengthN-gramsDiversity, Novelty, Feasibility Mobility statusvs.Trip chains Synthetic Population Quality Figure 5: Overview of the evaluation framework for assessing tabular quality, sequential quality, and cross-stage consistency. 5 Results and discussion The empirical evaluation is conducted using the Montreal OD survey, described in Section 3, which serves as the true population. A 1% random sample of the OD dataset is designated as training data. This sample captures 628 unique tabular attribute combinations, representing only 14.7% of the 4,282 unique combinations present in the full dataset. The remaining 3,654 combinations constitute sampling zeros A model that merely memorises the training distribution cannot recover these missing profiles, making their recovery the central challenge addressed by the proposed regularization framework. Stage 1 is evaluated across four generative model variants: a vanilla WGAN-GP and three regularized extensions that incorporate CLAP, IGP, and LDR. All generated populations are evaluated against the full OD dataset using the feasibility, diversity, novelty, and count-aware metrics defined in Section 4. Stage 2 compares the Transformer and LSTM-Attention models for sequential behavioural attribute generation. Sequential quality is assessed through three complementary metrics. First, the trip-length distribution of the generated population is compared against the full OD dataset to verify that the models reproduce realistic daily travel volumes. Second, to assess whether the models reproduce complete daily activity patterns, we compare the distribution of the most frequent real trip purpose sequences with their corresponding shares in the generated population. Third, the internal structure of the generated trip chains is evaluated by extracting unigrams, bigrams, and trigrams from the departure time, trip purpose, and travel mode sequences, and measuring their feasibility, diversity, and novelty against the real data. Finally, to assess the coherence between the two stages of the framework, we conduct a validation based on the mobility status attribute (p_mobil). Because mobility status is generated as part of the tabular profile in Stage 1 and directly conditions whether an individual should produce any trips in Stage 2, it serves as a natural bridge between the two stages. Specifically, we examine whether individuals generated as non-mobile in Stage 1 are consistently assigned empty or null trip sequences in Stage 2, and whether mobile individuals receive plausible trip chains. This consistency check provides a direct measure of how well the joint framework preserves the relationship between socio-demographic attributes and sequential travel behaviour. 5.1 Stage 1: Tabular Attribute Synthesis 5.1.1 Distributional Similarity We begin by assessing how well the best-performing model reproduces the marginal attribute distributions of the full OD dataset. Figure 6 presents side-by-side bar charts comparing the proportions of real (blue) and CLAP-generated (red) individuals across all seven tabular attributes. Across all attributes, the generated data closely follows the real distribution. There is no clear pattern of over-representing or under-representing any category. The model also matches simple binary attributes, such as gender and mobility status, very well. It also captures the uneven distributions of age and employment. These results show that CLAP can learn the main marginal patterns of the population even when it is trained on only a 1% sample. Figure 6: Comparison of marginal probability distributions. To quantify distributional similarity across model variants and increasing levels of attribute interaction, Figure 7 reports the mean Standardized Root Mean Square Error (SRMSE) for the vanilla WGAN-GP and the three regularized variants (CLAP, IGP, and LDR) at the univariate, bivariate, and trivariate levels. Overall, the regularized models consistently outperform the vanilla baseline. Among the regularized variants, CLAP achieves the lowest SRMSE across all interaction levels, indicating a stronger ability to reproduce both marginal distributions and higher-order attribute associations. Figure 7: Mean SRMSE across univariate, bivariate, and trivariate distributions for all model variants. However, it is important to note that distributional similarity provides only an initial evaluation of synthetic population quality. Although the SRMSE results show that the regularized models, especially CLAP, are better able to match key distributional patterns, further evaluation is required to assess whether the generated data are feasible, novel, diverse, and useful for downstream applications. 5.1.2 Feasibility Table 3 presents the performance of the Vanilla WGAN-GP and the three regularized variants. A first important observation is that all models generate a larger number of unique attribute combinations than those observed in the real dataset. The real population contains 4,282 unique combinations, while the generated datasets contain between 5,412 and 6,408 unique combinations. Even in the most conservative case, the generator produces more combinations than the full real data. This increase in the number of generated combinations is not necessarily a negative result. However, it is important to check whether this extra diversity leads to unrealistic samples. Feasibility is the first metric used to examine this issue. This metric measures the proportion of generated samples that correspond to valid combinations in the real population. As shown in Table 3, all models achieve high feasibility. The Vanilla model obtains a feasibility score of 0.898, while the regularized models achieve higher values: 0.935 for CLAP, 0.919 for IGP, and 0.934 for LDR. This means that approximately 90% to 94% of the generated individuals represent combinations that genuinely exist in the real data. The regularized models all outperform the Vanilla model in terms of feasibility. This result shows that adding regularization does not push the generator toward unrealistic attribute profiles. Instead, the regularized models are able to improve the validity of the generated population. The difference between the best and worst feasibility scores is relatively small, around three to four percentage points, which suggests that all models learn the basic structure of the population reasonably well. The relationship between feasibility and exploration is particularly important. In general, when a model is encouraged to generate more diverse samples, there is a risk that it may also generate more infeasible combinations. However, the results show that the regularized models achieve both higher diversity and higher feasibility than the Vanilla model. This indicates that regularization improves the quality of exploration rather than simply increasing the number of generated combinations. Overall, feasibility is not a major problem for any of the models when training is effective. Most generated samples correspond to valid real-population combinations. 5.1.3 Diversity As shown in Table 3, all regularized models achieve higher diversity than the Vanilla model. Vanilla obtains a diversity score of 0.638, while CLAP, IGP, and LDR improve this value to 0.678, 0.741, and 0.714, respectively. Among all models, IGP achieves the highest diversity score, meaning that it recovers 74.1% of all real combinations. This result shows that IGP encourages the generator to explore a wider part of the attribute space compared with the unregularized baseline. However, the standard diversity score only checks whether a combination appears or not. It does not consider whether the generated frequency of that combination is close to its true frequency in the real population. For this reason, the weighted diversity metric is more informative. It evaluates not only whether the model recovers real combinations, but also whether it generates them in more realistic proportions. When weighted diversity is considered, the difference between the regularized models becomes much smaller. IGP and CLAP both achieve a weighted diversity score of 0.684. Vanilla remains the lowest, with a score of 0.666. Although IGP achieves the highest standard diversity score, its advantage largely disappears once population frequencies are taken into account. This suggests that IGP is very effective at recovering a large number of unique combinations, but is less successful at generating them in proportions that match the real population. In contrast, CLAP is able to maintain its performance more consistently when moving from standard diversity to weighted diversity. While its standard diversity score is lower than that of IGP, its weighted diversity remains competitive, indicating that the combinations it recovers are generated in more realistic proportions. 5.1.4 Novelty Novelty focuses on an even more challenging task. It measures how many valid combinations that were absent from the 1% training sample are recovered by the generated data. In this experiment, there are 3,654 unseen combinations. These combinations exist in the full real population but are not present in the small training sample. Therefore, a high novelty score indicates that the model is not only memorizing the training sample, but is also able to recover valid profiles beyond what it directly observed during training. The novelty results show that all regularized models outperform the Vanilla model. Vanilla achieves a novelty score of 0.600, while CLAP, IGP, and LDR improve this value to 0.637, 0.706, and 0.676, respectively. IGP again achieves the highest score, recovering more than 70% of the unseen valid combinations. This confirms that IGP has the strongest ability to explore beyond the limited 1% sample and recover combinations that were missing from the training data. The weighted novelty results provide a stricter evaluation. Vanilla has the lowest weighted novelty score, with a value of 0.479. The regularized models perform better, with CLAP reaching 0.548, LDR reaching 0.545, and IGP achieving the highest value of 0.579. These results show that the regularized models are not only better at finding unseen valid combinations, but also better at generating them in more realistic frequencies. However, the drop from standard novelty to weighted novelty also shows that recovering unseen combinations is easier than reproducing their correct population proportions. These results highlight the importance of using both standard and weighted metrics. Standard diversity and novelty are useful for measuring coverage, but they can make a model appear stronger if it generates many combinations without matching their real frequencies. Weighted diversity and weighted novelty provide a stricter and more realistic evaluation because they consider both coverage and proportional accuracy. Therefore, the weighted metrics are especially important for judging the quality of the synthetic population. Table 3: Combination-level validity, diversity, and novelty results for the generated datasets. Method nrealn_real nsamplen_sample nunseenn_unseen nfaken_fake Feasibility Diversity Div. (w) Novelty Nov. (w) Vanilla 4282 628 3654 6351 0.898 0.638 0.666 0.600 0.479 CLAP 5412 0.935 0.678 0.684 0.637 0.548 IGP 6408 0.919 0.741 0.684 0.706 0.579 LDR 5460 0.934 0.714 0.690 0.676 0.545 5.1.5 Comparison of Regularization Terms The final F1 scores provide a combined evaluation of feasibility and coverage. Since the weighted metrics are stricter and more informative than their standard versions, the weighted F1 scores, especially F1DwF1_D_w and F1NwF1_N_w, are the primary basis for comparison in this section. As shown in Table 4, all regularized models achieve higher F1 scores than the Vanilla model in both diversity and novelty evaluations. An important question is which F1 score should be used as the main criterion for comparing the models. One possible option is F1DF1_D, which combines feasibility with standard diversity. A stricter comparison can be made using the weighted F1 scores. Among these, F1NwF1_N_w is the most important metric for this study. This is because novelty is more closely related to the sampling-zero problem than diversity. Diversity evaluates coverage over all real combinations, including those that may already be present in the 1% training sample. In contrast, novelty focuses specifically on valid real combinations that were absent from the training sample. Therefore, novelty directly measures whether the model can recover sampling-zero combinations rather than simply reproduce what it has already seen. For this reason, the most reliable validation should make the evaluation stricter in two ways. First, it should use the weighted version of the metric. Second, it should focus on novelty rather than diversity. Based on F1NwF1_N_w, the ranking of the models is IGP, CLAP, LDR, and Vanilla. IGP achieves the highest score of 0.711, showing that it provides the best trade-off between generating valid combinations and recovering unseen combinations in realistic proportions. For completeness, all F1 variants are reported in Table 4, but the weighted novelty F1 score is treated as the main reference for selecting the best model. Table 4: F1-based evaluation combining feasibility with diversity and novelty metrics. Method F1DF1_D F1_D_w F1_N F1_N_w Vanilla 0.746 0.765 0.719 0.625 CLAP 0.786 0.790 0.758 0.691 IGP 0.820 0.784 0.799 0.711 LDR 0.809 0.793 0.784 0.688 5.2 Stage 2 — Sequence Attribute Synthesis 5.2.1 Trip Length and Activity Pattern The trip-length distribution shows how many trips each individual makes per day. If the generated population behaves similarly to the real population, the generated trip length distribution should be close to the real one. Figure 8 compares the real OD distribution with the outputs of the LSTM-Attention and Transformer sequence models for each of the four tabular models: CLAP, IGP, LDR, and Vanilla. The real distribution is shown in blue, the LSTM-Attention output in orange, and the Transformer output in green. Two metrics are used to measure distributional similarity: Jensen-Shannon Divergence (JSD) and Wasserstein-1 distance (W1). For both metrics, lower values indicate a closer match to the real distribution. Figure 8: Trip-length distribution comparison between the real OD data and the generated outputs. The first comparison is between the two sequence models. As shown in Figure 8, LSTM-Attention gives lower JSD values than the Transformer for the regularized tabular models. This means that LSTM-Attention matches the general shape of the real trip-length distribution more closely. However, the W1 values are very close between the two sequence models, and in some cases the Transformer obtains a slightly lower W1. This suggests that although LSTM-Attention better captures the overall probability pattern, the two models are quite similar in terms of how far the generated trip counts are from the real ones. The second comparison is the effect of the Stage 1 tabular model. The choice of tabular model clearly affects how well the sequence model reproduces the trip-length distribution. Among the regularized models, CLAP and IGP produce the strongest results. In particular, CLAP combined with LSTM-Attention achieves the best W1 score, showing the closest match to the overall shape of the real distribution. Compared with the Vanilla baseline, CLAP provides a clear improvement, with an approximate W1 difference of 0.058. This indicates that a better tabular generation stage can also improve the quality of the sequence-generation stage. To further assess whether the generated data reproduce complete and interpretable activity chain patterns, we compare the distribution of the ten most frequent real activity sequences with their corresponding shares in the CLAP-LSTM generated population, as shown in Figure 9. Figure 9: Distribution of the ten most frequent real activity sequences and their corresponding shares in the CLAP-LSTM generated population. These dominant sequences account for a substantial share of the observed activity patterns and therefore provide a useful benchmark for evaluating whether the model captures the main structure of daily travel behaviour. The real dataset contains 162,588 records and 6,389 unique activity sequences, while the CLAP-LSTM generated population contains the same number of records and 6,058 unique sequences. This indicates that the generated population preserves a comparable level of sequence diversity. As shown in Figure 9, the generated shares closely follow the real distribution across the major activity-sequence patterns. 5.2.2 N-gram Analysis of Trip Chains This section evaluates the sequential quality of the generated trip chains using n-gram analysis. Three levels of sequential structure are considered. As the n-gram level increases, the evaluation becomes stricter because the model must reproduce more complex sequential dependencies. For each feature and n-gram level, the generated sequences are compared with the real OD data using diversity, novelty, and feasibility. Figure 10 compares the n-gram performance of LSTM-Attention and Transformer across the three sequence features: departure time, trip purpose, and travel mode. Overall, both models perform well at the unigram and bigram levels, but their behaviour becomes more different at the trigram level, where the task is more difficult. In terms of diversity and novelty, the Transformer also performs well for departure time and trip purpose. For these two features, it generally matches or improves over LSTM-Attention, especially at the trigram level. This indicates that the Transformer is effective not only at generating valid sequences, but also at recovering a wider range of real and unseen sequential patterns. However, the pattern is different for travel mode. In this feature, LSTM-Attention performs better than the Transformer in diversity and novelty, showing that it recovers more travel-mode combinations, although with lower feasibility. The decrease in performance at the trigram level is expected. Trigrams represent longer and more specific sequential patterns, so they are harder to reproduce than unigrams or bigrams. This is especially clear when comparing the three features. Departure time has the best trigram performance in diversity and novelty because it has a smaller number of possible combinations, with only 79 trigram patterns. In contrast, travel mode has 875 trigram combinations, and trip purpose has 1,322 combinations. As the number of possible combinations increases, it becomes more difficult for the model to recover all valid patterns, especially unseen ones. Therefore, the lower diversity and novelty scores for trip purpose and travel mode are mainly due to the larger and more complex combinatorial space. One important observation is that the Transformer generally maintains higher feasibility than LSTM-Attention, especially at the trigram level. This difference is most visible for travel mode. Based on Table 5, the feasibility of LSTM-Attention drops to around 0.84 for travel-mode trigrams, while the Transformer keeps feasibility close to 0.96. A similar pattern can be observed for trip purpose, where the LSTM-Attention feasibility is around 0.92, while the Transformer reaches approximately 0.97. This suggests that the Transformer is better at keeping longer generated sequences valid, even when the sequential patterns become more complex. The detailed numerical results for all features, n-gram levels, and model combinations are provided in Table 5. Figure 10: N-gram evaluation of LSTM-Attention and Transformer. The shaded area shows the range of results across the Stage 1 tabular model variants. Table 5: Combined N-gram evaluation results for all features using a 1% sample. Results are shown separately for each feature and model family. Feature Method N-gram LSTM Transformer Diversity Novelty Feasibility Diversity Novelty Feasibility d_grhre_vec vanilla unigram 1.000 – 1.000 1.000 – 1.000 bigram 1.000 1.000 0.969 1.000 1.000 0.968 trigram 0.911 0.781 0.907 0.949 0.875 0.906 IGP unigram 1.000 – 1.000 1.000 – 1.000 bigram 1.000 1.000 0.969 1.000 1.000 0.965 trigram 0.911 0.781 0.906 0.949 0.875 0.899 LDR unigram 1.000 – 1.000 1.000 – 1.000 bigram 1.000 1.000 0.969 1.000 1.000 0.965 trigram 0.924 0.813 0.909 0.962 0.906 0.902 CLAP unigram 1.000 – 1.000 1.000 – 1.000 bigram 1.000 1.000 0.969 1.000 1.000 0.967 trigram 0.975 0.938 0.904 0.949 0.875 0.903 d_motif_vec vanilla unigram 1.000 – 1.000 1.000 – 1.000 bigram 0.957 0.870 0.977 0.957 0.870 0.988 trigram 0.439 0.301 0.921 0.530 0.416 0.966 IGP unigram 1.000 – 1.000 1.000 – 1.000 bigram 0.963 0.889 0.977 0.988 0.963 0.987 trigram 0.458 0.323 0.918 0.533 0.417 0.964 LDR unigram 1.000 – 1.000 1.000 – 1.000 bigram 0.975 0.926 0.978 0.969 0.907 0.987 trigram 0.445 0.309 0.924 0.523 0.404 0.966 CLAP unigram 1.000 – 1.000 1.000 – 1.000 bigram 0.957 0.870 0.977 0.975 0.926 0.989 trigram 0.449 0.314 0.921 0.548 0.438 0.966 d_mode_main vanilla unigram 1.000 – 0.996 1.000 – 0.999 bigram 0.996 0.993 0.972 0.933 0.890 0.989 trigram 0.696 0.648 0.857 0.648 0.589 0.962 IGP unigram 1.000 – 0.996 1.000 – 0.998 bigram 0.992 0.986 0.970 0.937 0.897 0.986 trigram 0.710 0.663 0.841 0.672 0.619 0.955 LDR unigram 1.000 – 0.996 1.000 – 0.999 bigram 0.983 0.973 0.972 0.920 0.870 0.988 trigram 0.686 0.636 0.862 0.635 0.576 0.962 CLAP unigram 1.000 – 0.996 1.000 – 0.999 bigram 0.983 0.973 0.971 0.954 0.925 0.989 trigram 0.728 0.684 0.849 0.688 0.637 0.960 To identify the best overall model combination, the n-gram results are further summarized using an F1 score. For unigrams, the F1 score is computed using feasibility and diversity, because unigrams do not have a novelty value. For bigrams and trigrams, the F1 score is computed using feasibility and novelty, because these are more directly related to the recovery of unseen sequential patterns. Figure 11 reports the mean F1 score for all eight model combinations. The final score is averaged across the three sequence features, departure time, trip purpose, and travel mode, and across the three n-gram levels. This provides a single overall measure of sequential quality. The ranking shows that Transformer-based combinations achieve the highest overall scores. The best-performing model is CLAP + Transformer, with a mean F1 score of 0.906. This confirms that the Transformer provides a consistent improvement in overall sequential quality. The figure also shows that the Stage 1 tabular model still affects the final sequence quality. CLAP gives the best result for both sequence architectures. Compared with the Vanilla baseline, the regularized tabular models generally improve the final F1 score. Figure 11: Overall sequential quality ranking of the eight model combinations based on the mean F1 score. 5.3 Cross-Stage Consistency: Mobility Validation To assess whether the two stages of the framework are coherent with each other, the mobility status attribute, pmobilp_mobil, is used as a bridge between the tabular and sequential outputs. In the real OD dataset, individuals with pmobil=2p_mobil=2 are non-mobile and should not have any trips. Therefore, all three sequential attributes, departure time, trip purpose, and travel mode, should be empty sequences represented as [0][0]. In contrast, individuals with pmobil=1p_mobil=1 are mobile and should have at least one trip. Based on this rule, two types of violations are counted: non-mobile individuals who are incorrectly assigned trips, and mobile individuals who are incorrectly assigned no trips. The purpose of this analysis is not to compare or rank the models, but to check whether the generated tabular and sequential components remain semantically consistent after the two-stage generation process. As shown in Table 6, all model combinations achieve very high consistency, with accuracy values above 98% for LSTM-Attention and above 99% for Transformer. This indicates that the generated mobility status is largely consistent with the generated trip sequences. These results confirm that the two-stage framework preserves meaningful semantic coherence between the tabular and sequential outputs. In other words, the sequence-generation stage does not behave independently of the tabular attributes; instead, it generally produces trip chains that are logically compatible with the mobility status assigned in the first stage. Table 6: Consistency check between p_mobil and the three trip-related features. Method LSTM Transformer Violations all 3 = [0] Violations all 3 ≠ [0] Total violations Accuracy (%) Violations all 3 = [0] Violations all 3 ≠ [0] Total violations Accuracy (%) IGP 2278 16 2294 98.59 771 2 773 99.52 LDR 2804 9 2813 98.27 647 5 652 99.60 vanilla 1343 22 1365 99.13 466 8 474 99.71 CLAP 1089 17 1106 99.32 316 5 321 99.80 Real data baseline: 0 violations and 100% consistency accuracy. Total number of samples: 162,588. 6 Conclusions This paper proposed a two-stage framework for generating synthetic populations with both tabular socio-demographic attributes and sequential travel-behaviour attributes. The framework addresses three key limitations in existing population synthesis methods: achieving feasibility, diversity, and novelty in tabular attribute synthesis, integrating the modelling of sequential behavioural attributes such as trip purpose, departure time, and travel mode, and improving the evaluation of synthetic population quality beyond conventional metrics. Together, these challenges are addressed through a framework that recovers valid but unseen combinations, avoids structurally invalid profiles, generates behaviorally realistic trip chains, and evaluates synthetic populations using feasibility, diversity, novelty, and count-aware measures. In the first stage, a WGAN-GP model was extended with regularization terms designed to improve the recovery of sampling-zero combinations while reducing the generation of structurally invalid profiles. The results show that regularization improves the quality of tabular synthesis compared with the vanilla WGAN-GP baseline. All regularized models achieved higher feasibility, diversity, novelty, and count-aware scores than the unregularized model. Among the proposed regularization methods, IGP produced the strongest overall performance when evaluated using the count-aware novelty F1 score, indicating that it provides the best balance between generating feasible individuals and recovering valid combinations absent from the 1% training sample. In the second stage, Transformer and LSTM-Attention models were used to generate sequential behavioural attributes conditioned on the synthesized tabular profiles. The results show that both models are able to reproduce key characteristics of daily trip chains, including trip-length distributions, activity patterns, and n-gram patterns in departure time, trip purpose, and travel mode sequences. While LSTM-Attention performs competitively in matching some trip-length distributions, the Transformer generally provides stronger sequential validity, especially for higher-order trip-chain patterns. The cross-stage mobility validation further confirms the coherence of the proposed framework. Generated mobility status from the tabular stage was highly consistent with the generated trip sequences in the sequential stage, with consistency rates above 98% for LSTM-Attention and above 99% for Transformer-based models. This result indicates that the sequential generator does not operate independently of the tabular attributes, but instead produces behavioural outputs that remain logically compatible with individual socio-demographic profiles. While the proposed regularization terms improve the balance between novelty, diversity, and feasibility, they remain soft training mechanisms rather than hard constraints. Therefore, they cannot fully prevent structural violations or guarantee the recovery of all sampling zero combinations. Their performance also depends on hyperparameter choices and the quality of the limited training sample. Several directions remain for future work. First, the framework should be tested on additional regions and travel survey datasets to assess its transferability and robustness across different urban contexts. Second, future research should investigate household-level constraints and interactions, since many travel decisions are shaped by relationships among household members. Third, the sequence-generation stage could be extended to include richer spatial information, such as activity locations or origin-destination patterns. Fourth, the framework could be extended from a static synthesis approach to a dynamic population synthesis model by incorporating demographic changes over time, such as ageing, household formation, employment transitions, vehicle ownership changes, and shifts in mobility status. Finally, future work should evaluate the generated synthetic populations within full activity-based modelling pipelines to measure how improvements in population synthesis affect downstream travel-demand forecasts. Acknowledgments This study is funded by the Canada First Research Excellence Fund under the Bridging Divides program. Declaration of generative AI During the preparation of this work, the author(s) used ChatGPT to assist with language editing, clarity improvement, manuscript formatting, and the preparation of tables and figures. After using this tool, the author(s) reviewed and edited the content as needed and take full responsibility for the content of the published article. Data availability The data that support the findings of this study are from the 2018 Montreal Origin-Destination survey. Restrictions apply to the availability of these data, which were used under license for the current study and are not publicly available from the authors. References Aemmer and MacKenzie (2022) Z. Aemmer and D. MacKenzie Generative population synthesis for joint household and individual characteristics. Computers, Environment and Urban Systems 96, p. 101852. Cited by: §2.2, Table 1. Agrawal et al. (2025) P. Agrawal, V. Jayaraman, J. Oon, F. Ling, and M. A. B. Ramli Generation of mobility patterns for private vehicles using multi-headed sequence generative adversarial networks. Transportation Research Procedia 82, p. 908–922. Cited by: §2.3, Table 1. Agriesti et al. (2024) S. Agriesti, C. Roncoli, and B. Nahmias-Biran Assignment of a synthetic population for activity-based modeling employing publicly available data (vol 11, 148, 2022). Isprs International Journal Of Geo-Information 13 (8). Cited by: §1. Al Maawali and Amna (2024) R. Al Maawali and A. Amna Optimization algorithms in generative ai for enhanced gan stability and performance. Applied Computing Journal, p. 359–371. Cited by: §4.2.1. Arjovsky et al. (2017) M. Arjovsky, S. Chintala, and L. Bottou Wasserstein generative adversarial networks. In International conference on machine learning, p. 214–223. Cited by: §1, §4.1. Arkangil et al. (2022) E. Arkangil, M. Yildirimoglu, J. Kim, and C. Prato A deep learning framework to generate realistic population and mobility data. arXiv preprint arXiv:2211.07369. Cited by: §1, §1, §2.3, Table 1. Badu-Marfo et al. (2022) G. Badu-Marfo, B. Farooq, and Z. Patterson Composite travel generative adversarial networks for tabular and sequential population synthesis. IEEE Transactions on Intelligent Transportation Systems 23 (10), p. 17976–17985. Cited by: §1, §1, §2.3, Table 1. Bigi et al. (2024) F. Bigi, T. H. Rashidi, and F. Viti Synthetic population: a reliable framework for analysis for agent-based modeling in mobility. Transportation Research Record 2678 (11), p. 1–15. Cited by: §1, §1. Borysov et al. (2019) S. S. Borysov, J. Rich, and F. C. Pereira How to generate micro-agents? a deep generative modeling approach to population synthesis. Transportation Research Part C: Emerging Technologies 106, p. 73–97. Cited by: §2.2, Table 1. Both et al. (2021) A. Both, D. Singh, A. Jafari, B. Giles-Corti, and L. Gunn An activity-based model of transport demand for greater melbourne. arXiv preprint arXiv:2111.10061. Cited by: §1. Castiglione et al. (2015) J. Castiglione, M. Bradley, and J. Gliebe Activity-based travel demand models: a primer. Cited by: §1. Darsel et al. (2026) V. Darsel, E. Come, and L. Oukhellou Robust and reproducible evaluation framework for population synthesis models—application to probabilistic and deep generative models. Available at SSRN 5295092. Cited by: §1, §1. Eigenschink et al. (2023) P. Eigenschink, T. Reutterer, S. Vamosi, R. Vamosi, C. Sun, and K. Kalcher Deep generative models for synthetic data: a survey. IEEE access 11, p. 47304–47320. Cited by: §1. Farooq et al. (2013) B. Farooq, M. Bierlaire, R. Hurtubia, and G. Flötteröd Simulation based population synthesis. Transportation Research Part B: Methodological 58, p. 243–263. Cited by: §1, §2.1, Table 1. Fournier et al. (2021) N. Fournier, E. Christofa, A. P. Akkinepally, and C. L. Azevedo Integrated population synthesis and workplace assignment using an efficient optimization-based person-household matching method. Transportation 48 (2), p. 1061–1087. Cited by: §2.1, Table 1. Garrido et al. (2020) S. Garrido, S. S. Borysov, F. C. Pereira, and J. Rich Prediction of rare feature combinations in population synthesis: application of deep generative modelling. Transportation Research Part C: Emerging Technologies 120, p. 102787. Cited by: §2.2, Table 1. Gulrajani et al. (2017) I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville Improved training of wasserstein gans. Advances in neural information processing systems 30. Cited by: §1, §4.1. Habib et al. (2020) K. N. Habib, W. El-Assi, and T. Lin How large is too large? a review of the issues related to sample size requirements of regional household travel surveys with a case study on the greater toronto and hamilton area (gtha). arXiv preprint arXiv:2005.00563. Cited by: §1. Hörl and Balac (2021) S. Hörl and M. Balac Synthetic population and travel demand for paris and Île-de-france based on open and publicly available data. Transportation Research Part C: Emerging Technologies 130, p. 103291. Cited by: §2.1, Table 1. Johnsen et al. (2022) M. Johnsen, O. Brandt, S. Garrido, and F. Pereira Population synthesis for urban resident modeling using deep generative models. Neural Computing and Applications 34 (6), p. 4677–4692. Cited by: §2.2, Table 1. Jutras-Dubé et al. (2024) P. Jutras-Dubé, M. B. Al-Khasawneh, Z. Yang, J. Bas, F. Bastin, and C. Cirillo Copula-based transferable models for synthetic population generation. Transportation Research Part C: Emerging Technologies 169, p. 104830. Cited by: §2.2, Table 1. Kang et al. (2023) J. Kang, Y. Kim, M. M. Imran, G. Jung, and Y. B. Kim Generating population synthesis using a diffusion model. In 2023 Winter Simulation Conference (WSC), p. 2944–2955. Cited by: §2.2, Table 1. Kim and Bansal (2023) E. Kim and P. Bansal A deep generative model for feasible and diverse population synthesis. Transportation Research Part C: Emerging Technologies 148, p. 104053. Cited by: §1, §1, §2.2, Table 1, §4.2.2. Kukic et al. (2023) M. Kukic, S. Benchelabi, and M. Bierlaire Hybrid simulator for capturing dynamics of synthetic populations. In 2023 IEEE 26th International Conference on Intelligent Transportation Systems (ITSC), p. 2646–2651. Cited by: §2.1. Kukic et al. (2024) M. Kukic, X. Li, and M. Bierlaire One-step gibbs sampling for the generation of synthetic households. Transportation Research Part C: Emerging Technologies 166, p. 104770. Cited by: §2.1, Table 1, Table 1. La et al. (2025) D. M. La, H. L. Vu, L. Kamruzzaman, and E. Miller Population synthesis: a problem-based review. Transport Reviews 45 (3), p. 366–389. Cited by: §1. Lim et al. (2025) S. Y. Lim, H. Yun, P. Bansal, D. Kim, and E. Kim A large language model for feasible and diverse population synthesis. arXiv preprint arXiv:2505.04196. Cited by: §2.2, Table 1. Lu et al. (2026) Y. Lu, G. Liu, X. Li, and Z. Jin Generate individual spatiotemporal activity sequences from population synthesis via deep learning approaches. Engineering Applications of Artificial Intelligence 167, p. 113759. Cited by: §2.3, Table 1. Mensah et al. (2025) D. O. Mensah, G. Badu-Marfo, and B. Farooq Robustness analysis of deep learning models for population synthesis. Transportation Research Procedia 82, p. 3790–3806. Cited by: §2.2, Table 1. Prédhumeau and Manley (2023) M. Prédhumeau and E. Manley A synthetic population for agent-based modelling in canada. Scientific Data 10 (1), p. 148. Cited by: §2.1, Table 1. Rahman and Fatmi (2023) M. N. Rahman and M. R. Fatmi Population synthesis accommodating heterogeneity: a bayesian network and generalized raking technique. Transportation research record 2677 (6), p. 41–57. Cited by: §2.1, Table 1. Rastogi et al. (2025) T. Rastogi, D. Jonsson, and A. Karlström Population synthesis using incomplete microsample. Transportation Research Procedia 86, p. 80–87. Cited by: §2.2, Table 1. Saadi et al. (2016) I. Saadi, A. Mustafa, J. Teller, B. Farooq, and M. Cools Hidden markov model-based population synthesis. Transportation Research Part B: Methodological 90, p. 1–21. Cited by: §2.1, Table 1. Sun et al. (2018) L. Sun, A. Erath, and M. Cai A hierarchical mixture modeling framework for population synthesis. Transportation Research Part B: Methodological 114, p. 199–212. Cited by: §2.1, Table 1. Sun and Erath (2015) L. Sun and A. Erath A bayesian network approach for population synthesis. Transportation Research Part C: Emerging Technologies 61, p. 49–62. Cited by: §2.1, Table 1. Tang et al. (2025) M. Tang, P. Lu, and Q. Feng Generating feasible and diverse synthetic populations using diffusion models. arXiv preprint arXiv:2508.09164. Cited by: §2.2, Table 1. Thanh-Tung and Tran (2020) H. Thanh-Tung and T. Tran Catastrophic forgetting and mode collapse in gans. In 2020 international joint conference on neural networks (ijcnn), p. 1–10. Cited by: §4.2.1. Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. Advances in neural information processing systems 30. Cited by: §1. Yang et al. (2025) H. Yang, H. Wu, L. Yuan, X. Ren, J. Y. Chow, J. Gao, and K. Ozbay Deep and diverse population synthesis for multi-person households using generative models. arXiv preprint arXiv:2508.09964. Cited by: §2.2, Table 1. Zhu and Ferreira Jr (2014) Y. Zhu and J. Ferreira Jr Synthetic population generation at disaggregated spatial scales for land use and transportation microsimulation. Transportation Research Record 2429 (1), p. 168–177. Cited by: §2.1, Table 1.