Paper deep dive
An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction
Narges Ahmadi, Yubo Jiao, Jônatas Augusto Manzolli, Jiangbo Yu, Luis Miranda-Moreno
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/21/2026, 4:41:20 AM
Summary
This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction to analyze weather-sensitive travel mode choices. Using a chatbot-administered, image-augmented stated-preference survey with 92 student commuters, the research compares multinomial logit models, machine learning benchmarks (logistic regression, random forest), and nine locally deployed large language models (LLMs). Results indicate that random forest achieved 69.6% accuracy, while the best text-only zero-shot LLM reached 69.9%. Vision-based LLM configurations utilizing weather images achieved the highest accuracy at 71.5%. The study demonstrates that habitual travel information, specific prompting strategies, and visual context significantly improve LLM-based travel-choice prediction within an auditable multi-agent framework.
Entities (10)
Relation Signals (7)
Narges Ahmadi → affiliatedwith → McGill University
confidence 99% · Narges Ahmadi... Department of Civil Engineering, McGill University
Random Forest → achievedaccuracy → 69.6%
confidence 95% · Random forest achieved 69.6% five-class accuracy
Vision-Based Configuration → achievedaccuracy → 71.5%
confidence 95% · the best vision-based configuration reached the highest observed five-class accuracy of 71.5%
McGill University → locatedin → Montreal
confidence 95% · McGill University, Montreal, Quebec, Canada
Multinomial Logit Model → analyzes → Weather-related associations
confidence 92% · Weather-related associations were analyzed using a multinomial logit model
Chatbot → administers → Stated-Preference Survey
confidence 90% · A chatbot-administered, image-augmented stated-preference survey
Large Language Models → outperformedby → Vision-Based Configuration
confidence 85% · the best vision-based configuration reached the highest observed five-class accuracy... indicating that visual context may provide additional predictive information
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather scenarios, yielding 454 respondent-scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and extended through persona, few-shot, and vision-based configurations. Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting. Habitual travel information produced the most consistent gains, Expert framing generally outperformed Role-Play, and persona information was most useful when habitual travel information was unavailable. Few-shot prompting improved prediction for several models, with gains stabilizing after a small number of examples. Using the same weather images shown to respondents, the best vision-based configuration reached 71.5% five-class accuracy, indicating that visual context may provide additional predictive information for selected models. Overall, the study shows how conversational surveys, structured data processing, conventional behavioral modeling, machine learning, and multimodal LLM prediction can be coordinated within an auditable multi-agent workflow.
Tags
Links
- Source: https://arxiv.org/abs/2608.20320v1
- Canonical: https://arxiv.org/abs/2608.20320v1
Trouble viewing inline? Open PDF directly →
Full Text
80,190 characters extracted from source content.
Expand or collapse full text
An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction Narges Ahmadia,*, Yubo Jiaoa, Jônatas Augusto Manzollia, Jiangbo Yua, Luis Miranda-Morenoa aDepartment of Civil Engineering, McGill University, Montreal, Quebec, Canada Abstract Travel-behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes and applies a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from 92 student commuters across five predefined weather scenarios, yielding 454 respondent–scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and further extended through persona, few-shot, and vision-based configurations. Cycling was particularly sensitive to adverse weather, while public transit use increased substantially under the Snowy scenario. Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting. Habitual travel information produced the most consistent improvement in LLM prediction, Expert framing generally outperformed Role-Play, and persona information was most beneficial when habitual travel information was unavailable. Few-shot prompting further improved five-class prediction for several models, with gains generally stabilizing after a small number of examples. Vision-capable models were also evaluated using the same weather images presented to respondents, and the best vision-based configuration reached the highest observed five-class accuracy of 71.5%, indicating that visual context can provide additional predictive information for selected models. Taken together, the findings show that habitual travel information, prompting strategy, demonstrations, and visual context can meaningfully shape LLM-based travel-choice prediction. More broadly, the study demonstrates how conversational survey administration, structured data processing, conventional behavioral modeling, machine learning, and multimodal LLM prediction can be coordinated within a single auditable multi-agent workflow, while broader application requires validation with larger and more representative traveler samples. Keywords: Large language models, travel behavior, travel survey, conversational agents, mode choice, multi-agent system framework *Corresponding author. Email: narges.ahmadi@mail.mcgill.ca 1 Introduction Understanding and predicting travel behavior is a fundamental component of transportation planning, encompassing interrelated decisions regarding when, where, and how individuals travel. Among these decisions, mode choice is particularly consequential because it determines how demand is distributed across transportation systems and influences infrastructure needs, environmental impacts, and the effectiveness of policies aimed at encouraging sustainable mobility (31; 4). Mode choice is shaped not only by individual and alternative-specific characteristics but also by the context in which travel occurs. Weather is an important example, as precipitation, extreme temperatures, and adverse road conditions can alter the attractiveness of walking and cycling and shift travelers toward transit or private vehicles, with effects that vary across individuals and seasons (21; 7; 37). Understanding such context-dependent responses therefore requires approaches capable of capturing heterogeneous travel behavior under conditions that may not be adequately represented in observed travel data. Collecting behavioral data that capture these contextual variations remains challenging. Traditional travel surveys, whether administered through interviews, paper questionnaires, or static web forms, can be costly and time-consuming to implement and scale (24; 3). Stated-preference (SP) surveys provide an important alternative by allowing researchers to examine choices under controlled or hypothetical conditions, including weather scenarios that may be difficult to observe systematically in revealed-preference data. However, conventional SP instruments typically describe such situations using fixed textual descriptions, requiring respondents to construct the intended context themselves and limiting the incorporation of richer contextual information (11). This is particularly relevant for weather-dependent travel choices, for which visual conditions such as precipitation, visibility, snow accumulation, or road conditions may form an important part of the decision environment. These limitations motivate data-collection approaches that can combine scalable survey deployment with richer and more consistent representations of travel contexts. Recent advances in artificial intelligence (AI) provide new possibilities for addressing these limitations. Conversational chatbots can administer surveys through guided interactions, incorporate respondents’ previous answers, and embed multimedia content directly within the survey experience, potentially enabling more flexible and engaging forms of behavioral data collection (44; 35; 23). At the same time, generative AI tools can support the construction of predefined visual scenarios, while large language models (LLMs) enable natural-language interaction and the processing of heterogeneous information. Together, these capabilities create opportunities to move beyond static, text-based questionnaires toward conversational and context-rich behavioral surveys. Despite growing applications of conversational AI in survey research, however, chatbot-based data collection remains largely unexplored in travel-behavior and mode-choice research, particularly for stated-preference experiments that integrate image-augmented contextual information. Parallel advances have occurred in travel-choice modeling. Discrete choice models (DCMs), grounded in random-utility theory, remain central to mode-choice analysis because of their interpretability and ability to estimate behaviorally meaningful relationships (31; 4; 39). Machine-learning (ML) methods have subsequently expanded the predictive toolkit by capturing nonlinear relationships and complex interactions that may be difficult to specify parametrically (18; 1). More recently, LLMs have introduced a different paradigm in which, rather than estimating a fixed mathematical mapping between explanatory variables and choices, they can interpret traveler characteristics and contextual information expressed through natural-language prompts and generate individual-level predictions with limited or no task-specific training (45; 10). This flexibility also makes it possible to vary the information provided to the model; for example, demographic characteristics, travel history, examples of previous decisions, or contextual descriptions, through alternative prompting configurations. Yet two important questions remain insufficiently examined in travel-behavior prediction. The first concerns how prediction performance changes across systematically designed LLM configurations and prompting strategies. The second is whether vision-capable LLMs can directly incorporate visual representations of travel conditions, such as the same weather images presented to human respondents, rather than relying exclusively on textual descriptions. Beyond prediction, LLMs have increasingly been studied as synthetic respondents, serving as computational representations of individuals that generate decisions conditioned on specified personas and choice contexts. This line of research initially developed in economics, political science, and survey research, where LLM-generated responses have been compared with human decisions (19; 2). Recent transportation studies similarly suggest that conditioning LLMs on traveler characteristics, behavioral histories, or latent attitudes can improve their correspondence with observed travel choices (29; 36). A related development is the emergence of multi-agent systems in which specialized or persona-based LLM agents perform complementary functions or represent heterogeneous decision-makers at scale (29; 28; 13). Nevertheless, LLM-generated behavior does not necessarily reproduce human decision-making reliably, particularly under limited-context or zero-shot settings (26; 6; 20). Despite these developments across other behavioral domains, the integration of specialized agents for data collection, data processing, and LLM-based behavioral prediction within a common multi-agent workflow remains largely unexplored for travel behavior and individual mode-choice modeling. This study addresses these gaps through a multi-agent workflow that connects AI-assisted behavioral data collection with travel-choice modeling and LLM-based prediction. The paper makes four main contributions. First, it develops a conversational, image-augmented SP survey in which an AI-based chatbot collects traveler characteristics and repeated mode choices across predefined weather scenarios, providing a structured dataset specifically designed for downstream behavioral prediction. Second, it provides empirical evidence on weather-related variation in commuter mode choice by examining how the same individuals change their stated choices across systematically varied weather scenarios. Third, it systematically evaluates LLMs as mode-choice predictors under multiple prediction configurations, including zero-shot, few-shot, persona- and travel-history-enhanced prompting, as well as vision-based prediction using the same weather images presented to respondents, and benchmarks their performance against conventional DCM and ML approaches. Fourth, it demonstrates a multi-agent workflow that links conversational survey administration, structured data processing, and parallel discrete-choice, machine-learning, and LLM-based modeling within an integrated data-collection-to-prediction process, providing a reproducible basis for exploring AI-assisted travel-behavior research. The remainder of the paper is organized as follows. Section 2 reviews related work on travel-behavior data collection and conversational agents, mode-choice modeling, LLM-based choice prediction, and multi-agent applications. Section 3 presents the proposed multi-agent methodology. Section 4 describes the case-study implementation, and Section 5 reports the descriptive and behavioral-modeling results and the LLM prediction experiments. The subsequent sections discuss the implications and limitations of the findings and conclude with directions for future research. 2 Background and Related Work 2.1 Travel Behavior and Mode Choice Mode choice is one of the oldest and most heavily modeled problems in transportation research, and the multinomial logit (MNL) model has been its dominant framework for decades (31; 4). Grounded in random utility theory, MNL and its extensions such as nested and mixed logit remain the standard tool for estimating how travel time, cost, and individual characteristics translate into a discrete travel choice, in large part because their coefficients admit a direct behavioral interpretation (39). Machine learning (ML) classifiers, including random forests, gradient boosting, and support vector machines, have become an increasingly common complement to discrete choice models. Comparative studies have repeatedly found that ML classifiers can match or exceed MNL on predictive accuracy, particularly when the feature set is large or the choice set is imbalanced, though the resulting models are harder to interpret in behavioral terms (18; 1). Both model families face a common challenge: traveler preferences are not fixed but vary with weather, season, the built environment, and individual heterogeneity. Weather in particular has a well-documented association with active-mode choice, with cold, wet, and low-visibility conditions suppressing walking and cycling and shifting travelers toward transit or driving (7; 37; 21). Capturing this sensitivity often requires stated-preference (SP) methods that vary conditions directly, because revealed-preference data may not span the full range of weather a respondent might encounter (3; 24). Both MNL and ML models, however, require labeled data, and their predictive performance may deteriorate when application conditions differ from those represented during model development. This limitation motivates the growing interest, reviewed in Section 2.3, in LLM-based choice prediction. 2.2 Data Collection Methods in Travel Behavior Research Travel behavior data has traditionally been collected through paper diaries, telephone interviews, and static web-based surveys. Each instrument trades off cost, reach, and depth; for example, paper surveys are labor-intensive to process, telephone interviews are expensive to scale, and static web forms, while cheap to distribute, present every respondent with the same fixed sequence of questions regardless of their answers (8). Conversational and digital instruments relax some of these constraints. Chatbot-administered surveys can branch dynamically on prior responses, and early evaluations report that respondents find conversational interfaces more engaging than static forms without a loss in data quality (11). Recent transportation-specific work has begun to apply modular AI agents to survey administration and interviewing, arguing that conversational agents can improve engagement, transparency, and cost efficiency relative to conventional instruments (44), and adaptive, reinforcement-learning-driven conversational survey designs have been proposed to further personalize the interview flow (38). Stated-preference tasks have also begun to incorporate richer visual aids, from virtual reality environments to video and photorealistic imagery, on the premise that a visual scenario elicits a more ecologically valid response than a text description alone (16). Pilot studies using video-based conversational chatbots to elicit perceived cycling safety (35) and photorealistic embodied conversational agents in survey research (23) point in the same direction, but remain limited in scale. Despite this progress, most chatbot instruments in transportation are still used for information delivery, such as trip planning or transit information, rather than for behavioral data collection, and few stated-preference studies embed rich media such as generated images directly into a chatbot-administered choice task. This is the gap the data collection agent in Section 4.1 is designed to address. 2.3 LLMs for Behavioral Simulation Large language models have recently been explored as tools for behavioral simulation across several fields. In economics, they have been proposed as simulated agents (19); in political science, they have been used to reproduce survey samples (2) and support population-scale agent-based simulations (34); and in market research and consumer behavior, they have been applied to segmentation and behavioral simulation (25; 9; 15). Related studies have examined trust behavior (22) and large-scale replication of psychology and management experiments (14), illustrating the broader potential of LLMs as synthetic participants. At the same time, this literature also raises important concerns. Replication studies show that LLM-generated survey responses may diverge from human data in ways that are difficult to detect (6), while methodological critiques caution against treating simulated responses as equivalent to human-subject evidence without independent validation (20). Research on persona-based and character-consistent prompting similarly suggests that adopting a simulated identity may either improve or distort behavioral responses (12; 42). Within transportation, a small but growing body of work investigates whether LLMs can predict individual travel choices without being trained on the target dataset. Applications include mode choice prediction (32), alternative-set evaluation (33), and behaviorally informed ridesourcing choice modeling (36). Other studies examine whether persona-based representations improve agreement with observed travel choices (27), whether LLMs can recover travelers’ valuation of time (43), and whether they can reproduce stated airline passenger preferences (40). An early working paper further suggests that LLMs can capture a meaningful share of the predictive signal in observed mode choice data without dataset-specific training (26). Despite these advances, existing studies generally examine individual prompting or modeling choices in isolation. A systematic comparison of context richness, model scale, persona information, few-shot learning, and visual inputs within a single travel-behavior prediction task remains limited. In particular, the use of the same visual travel conditions presented to human respondents as direct inputs to vision-capable LLMs remains unexplored, constituting an additional gap addressed by this study. 2.4 Multi-Agent Systems and LLM-Agent Frameworks Outside transportation, LLM-based multi-agent systems have been proposed for tasks ranging from simulating marketing and consumer behavior (13) to sandboxed economic modeling of consumer preference alignment (41) and constructing virtual organizational decision-makers for management research (17). These systems typically decompose a task among specialized agents that exchange intermediate results, drawing on the broader software-engineering literature on tool use and orchestration in autonomous LLM agents (45). Within transportation, this framing remains largely conceptual. Recent work has proposed frameworks for LLM-agent-based modeling of transportation systems (29) and dual-agent approaches for aligning LLM agents with human learning and adjustment behavior (28). Few studies, however, have implemented and evaluated an end-to-end workflow from survey design through data processing to behavioral prediction on a common empirical dataset. The present paper addresses this gap by instantiating the three stages as specialized agents and applying them to the same commuter mode-choice dataset. 3 Methodology 3.1 Framework Overview We propose a multi-agent framework that integrates data collection, data processing, and behavioral modeling for commuter mode-choice research. The framework consists of three specialized agents: a Data Collection Agent, a Data Processing Agent, and a Data Modeling Agent. The agents perform distinct research functions and exchange information through standardized data objects and explicit input–output interfaces. Data transfer, version control, and execution records are organized through a shared workspace and predefined workflow rules. These supporting functions are system infrastructure rather than additional agents. Figure 1 illustrates the forward workflow and the associated refinement paths. Figure 1: Multi-agent framework for commuter travel-behavior research. Solid arrows denote forward data and artifact flows, whereas dashed arrows denote refinements implemented in a subsequent workflow iteration. In this study, an agent is defined as a goal-directed and researcher-supervised software component that can interpret user instructions, maintain task state, select from an authorized set of tools, and produce structured and auditable outputs. An agent is therefore not restricted to an LLM. Depending on the task, it may combine an LLM with deterministic rules, database operations, statistical software, machine-learning algorithms, or general-purpose programming tools. The proposed framework is consequently a multi-agent research workflow rather than an LLM-only pipeline. The term agent in this framework refers to a software agent that supports the research process and should not be confused with a traveler agent in conventional agent-based transport simulation. Human respondents remain the source of stated survey choices, while researchers retain control over the survey constructs, data-processing requirements, model specifications, and interpretation of results. Throughout this paper, active data collection refers to stateful conversational survey administration with predefined interaction logic and structured response capture; it does not refer to active learning or adaptive sampling. Weather-sensitive demand prediction refers to individual mode-choice prediction across the five predefined weather scenarios. The framework follows the information flow: →col(⋅,col)(raw,ℬcol)→pro(⋅,ℰ,pro)(pro,pro)→mod(⋅,mod)(ℳ^,ℛmod)Q A_col (·;u_col ) (D^raw,B_col ) A_pro (·,E;u_pro ) (D^pro,Z_pro ) A_mod (·;u_mod ) ( M,R_mod ) (1) where Q denotes the researcher-defined survey specification; colu_col, prou_pro, and modu_mod denote the user instructions provided to the three agents; rawD^raw is the collected raw dataset; ℬcolB_col is the survey deployment package; ℰE represents optional external data; proD^pro is the processed dataset; proZ_pro contains supplementary processing outputs; ℳ M is the fitted model or set of models; and ℛmodR_mod contains model estimates, predictions, diagnostics, and evaluation results. Equation (1) represents one forward execution pass. The dashed arrows in Figure 1 indicate diagnostic feedback that may initiate a revised survey, processing, or modeling specification. Such revisions do not automatically modify an active study; they are reviewed by the researchers and implemented as a new workflow version. The framework treats raw responses as immutable and reserves held-out test outcomes for evaluation rather than revision of the configuration assessed on that test set. This separation assigns a distinct responsibility to each agent. The Data Collection Agent converts a survey specification into a deployable conversational interface and collects responses. The Data Processing Agent converts raw records into analysis-ready data. The Data Modeling Agent translates the research objective into an executable modeling workflow and returns fitted models and associated results. The decomposition reduces the need for a single general-purpose agent to perform all tasks and makes errors easier to identify at their source. Table 1: Roles, interfaces, and supported functions of the three agents. Agent Main inputs Authorized functions Main outputs Data Collection Agent Survey specification and collection instructions Chatbot construction, survey logic implementation, procedural question answering, response validation, clarification, deployment, and response aggregation Deployable survey interface, distribution link or QR code, raw responses, interaction logs, and collection metadata Data Processing Agent Raw survey data, optional external data, codebook, and processing instructions De-identification, validation, recoding, missing-data handling, text coding, data integration, feature construction, and quality assessment Processed dataset, quality flags, derived variables, supplementary data, and processing records Data Modeling Agent Processed data, modeling instructions, candidate model library, and evaluation requirements Model specification, code generation, statistical estimation, machine-learning training, LLM-based choice prediction, validation, and result organization Fitted model, parameter estimates, predictions, uncertainty measures, diagnostics, evaluation metrics, and structured results Note: The table summarizes functions supported by the general framework. The subset implemented in the present case study is described in Section 4. 3.2 General Agent Representation and Workflow Orchestration The following formulation describes functions supported by the general framework. Section 4 identifies the subset instantiated in the present case study; not every supported function was exercised or empirically evaluated in the current application. Let k∈col,pro,modk∈\col,pro,mod\ index the three agents. Each agent is represented as a mapping: k:(k,k,k,k)↦(k,k′,ℓk)A_k: (x_k,u_k,s_k; _k ) (y_k,s _k, _k ) (2) where kx_k denotes task-specific input artifacts, ku_k denotes the user instruction, ks_k is the current task state, and k _k contains the agent configuration, such as its prompt templates, model settings, tool permissions, and termination rules. The outputs include the task result ky_k, the updated state k′s _k, and an execution log ℓk _k. At execution step t, the agent selects a tool or operation according to: τk,t=πk(k,t,k,t,k,k),τk,t∈k, _k,t= _k (s_k,t,x_k,t,u_k; _k ), _k,t _k, (3) where πk _k is the agent policy and kT_k is its authorized tool library. The selected tool produces an intermediate output and updates the task state: (k,t,k,t+1)=gτk,t(k,t,k,t) (o_k,t,s_k,t+1 )=g_ _k,t (x_k,t,s_k,t ) (4) This formulation allows an agent to use different tools for different subtasks. For example, the Data Processing Agent may use deterministic Python code for range checks, an LLM for coding open-ended responses, and a database operation for joining external data. Tool selection is constrained by the agent’s permissions and by the output schema required by the next stage. The workflow orchestrator initiates each agent, verifies whether its required inputs are available, and passes only validated artifacts to the next stage. Direct, unrestricted natural-language communication between agents is avoided. Instead, the agents communicate through typed data objects, including datasets, codebooks, configuration files, model specifications, and execution records. This design improves interoperability and limits the propagation of unsupported agent outputs. 3.3 Data Collection Agent The Data Collection Agent converts a researcher-defined questionnaire into a deployable conversational survey, manages survey distribution, collects respondent inputs, and aggregates the resulting records. Its design builds on modular conversational survey systems in which the user interface, prompts, knowledge resources, session variables, and conversational logic are explicitly configured. 3.3.1 Survey specification The survey remains a researcher-defined measurement instrument. Let =qℓ=1LQ= \q_ \_ =1^L denote a survey containing L question objects. Each question is represented as: qℓ=(cℓ,wℓ,rℓ,ℓ,bℓ,vℓ)q_ = (c_ ,w_ ,r_ ,O_ ,b_ ,v_ ) (5) where cℓc_ is the behavioral construct, wℓw_ is the question wording, rℓr_ is the response type, ℓO_ is the set of allowable options when applicable, bℓb_ defines branching and display conditions, and vℓv_ defines validation rules. Response types may include single-choice, multiple-choice, numerical input, Likert-scale, text, voice, or image-assisted questions. The deployable chatbot is generated as: =Γcol(,col,col)C= _col (Q,u_col, _col ) (6) where Γcol _col is the chatbot-construction procedure and col _col contains the available interface components, prompt templates, knowledge resources, image assets, and deployment tools. The output C contains the survey interface, question sequence, branching logic, data-storage schema, and interaction rules. The agent may use an LLM to generate interface code, reformulate researcher instructions into chatbot logic, or produce an initial implementation for review. However, constructs, response scales, experimental attributes, and mandatory questions cannot be changed without explicit researcher approval. This restriction preserves measurement consistency across respondents. 3.3.2 Conversational interaction and response validation For respondent i and question ℓ , let hiℓ(r)h_i ^(r) denote the interaction history after clarification round r. The agent evaluates whether the accumulated response satisfies the question requirement: ξiℓ(r)=Vcol(qℓ,hiℓ(r),col),ξiℓ(r)∈0,1 _i ^(r)=V_col (q_ ,h_i ^(r); _col ), _i ^(r)∈\0,1\ (7) where ξiℓ(r)=1 _i ^(r)=1 indicates that the response is sufficiently complete and valid. If ξiℓ(r)=0 _i ^(r)=0 and the maximum number of clarification rounds has not been reached, the agent generates a question-specific clarification: qiℓc,(r+1)=Fcol(qℓ,hiℓ(r),col)q_i ^c,(r+1)=F_col (q_ ,h_i ^(r),u_col ) (8) The process continues until the response is accepted or the predefined clarification limit RmaxR_ is reached. The final structured answer is: xiℓcol=Ecol(hiℓ(Riℓ),qℓ),Riℓ≤Rmaxx_i ^col=E_col (h_i ^(R_i ),q_ ), R_i ≤ R_ (9) where EcolE_col maps the interaction history to the survey schema. For fixed-choice questions, deterministic option matching is used whenever possible. LLM-based mapping is used only when an answer is provided in unrestricted text or speech and cannot be mapped using deterministic rules. The agent may also provide procedural assistance, such as explaining a survey term, repeating a question, changing the interaction language, or allowing a respondent to revise an earlier answer. Such assistance must not recommend a travel mode, signal a preferred response, or otherwise alter the intended behavioral task. 3.3.3 Deployment and raw-data output After the survey has been approved, the agent generates a deployment bundle: ℬcol=(URL,QR,version,recruitmentmaterial),B_col= (URL,QR,version,recruitment\ material ), (10) which may contain a public or tokenized survey link, a QR code, survey version information, and standardized recruitment materials. For respondent i, the collection record is represented as: diraw=(icol,i,i,i),d_i^raw= (x_i^col,h_i,m_i,p_i ), (11) where icolx_i^col contains the extracted answers, ih_i contains the raw interaction history, im_i contains response and session metadata, and ip_i contains provenance information such as the survey version, prompt version, interface version, and timestamps. The complete raw dataset is: raw=⋃i=1Ndiraw.D^raw= _i=1^Nd_i^raw. (12) The raw dataset is treated as immutable. Later corrections, exclusions, or recoding decisions are stored as separate processing operations rather than overwriting the original responses. 3.4 Data Processing Agent The Data Processing Agent converts the raw records into an analysis-ready dataset based on the user instruction prou_pro. Its functions may include de-identification, schema validation, data-type conversion, missing-data handling, text coding, consistency checking, feature construction, and integration with external datasets. It may also produce supplementary outputs such as a codebook, data-quality report, descriptive summaries, or a list of unresolved records. Let ℰE denote optional external data, such as weather records, spatial attributes, transit service information, or zonal characteristics. The initial processing object is: (0)=J(raw,ℰ,join)D^(0)=J (D^raw,E;K_join ) (13) where J(⋅)J(·) is a controlled integration operator and joinK_join specifies the permitted identifiers, temporal alignment rules, spatial matching rules, and source metadata. External data are not merged unless their source and matching procedure can be recorded. The processing agent constructs an ordered sequence of operations: =(T1,T2,…,TH),Th∈pro,P= (T_1,T_2,…,T_H ), T_h _pro, (14) where proT_pro is the authorized processing-tool library. At step h, the agent selects an operation according to: zh=πpro(pro,h,ν((h−1)))z_h= _pro (u_pro,s_h,ν (D^(h-1) ) ) (15) where ν(⋅)ν(·) returns the current dataset profile, including variable types, missingness, range violations, duplicate records, and unresolved text fields. The dataset is then updated as: (h)=Tzh((h−1);h),h=1,…,HD^(h)=T_z_h (D^(h-1); θ_h ), h=1,…,H (16) where h θ_h contains the operation-specific parameters. The final processed dataset is: pro=(H)D^pro=D^(H) (17) Deterministic operations are preferred when the processing rule can be stated explicitly. Examples include numerical range checks, category recoding, duplicate removal, unit conversion, date processing, and table joins. LLMs are used for language-dependent tasks for which fixed rules are insufficient, such as mapping an open-ended explanation to a predefined coding scheme or extracting a stated constraint from a conversation. For an unstructured response riℓr_i and an admissible code set ℓC_ , an LLM-assisted coding operation returns: (x~iℓ,γiℓ)=Hψ(riℓ,ℓ,pro) ( x_i , _i )=H_ψ (r_i ,C_ ,u_pro ) (18) where x~iℓ x_i is the proposed code and γiℓ _i is a task-specific confidence or validation score. The stored value is determined by: xiℓpro=gℓ(riℓ),if deterministic parsing succeeds,x~iℓ,if x~iℓ∈ℓ and γiℓ≥δℓ,NA,otherwisex_i ^pro= casesg_ (r_i ),&if deterministic parsing succeeds,\\[4.0pt] x_i ,&if x_i _ and _i ≥ _ ,\\[4.0pt] NA,&otherwise cases (19) where gℓg_ is the deterministic coding rule and δℓ _ is the acceptance threshold. Records that do not satisfy the rule are flagged for human review instead of being assigned an unverified value. The processed dataset may be represented at the respondent–scenario level. For respondent i and commuting scenario s, define: dispro=(i,is,is,yis,is,is)d_is^pro= (p_i,z_is,a_is,y_is,q_is,v_is ) (20) where ip_i contains respondent characteristics, isz_is contains scenario and alternative attributes, isa_is contains alternative-availability indicators, yisy_is is the observed mode choice, isq_is contains data-quality flags, and isv_is contains data provenance. The resulting dataset is: pro=dispro:i=1,…,N,s=1,…,SiD^pro= \d_is^pro:i=1,…,N,\;s=1,…,S_i \ (21) In addition to proD^pro, the agent returns: pro=(book,ℱquality,summary,unresolved)Z_pro= (C_book,F_quality,S_summary,U_unresolved ) (22) where bookC_book is the final codebook, ℱqualityF_quality contains quality flags, summaryS_summary contains optional descriptive outputs, and unresolvedU_unresolved identifies records requiring manual review. 3.5 Data Modeling Agent The Data Modeling Agent constructs, estimates, and evaluates behavioral models based on the processed data and the user instruction modu_mod. The instruction may specify the prediction target, explanatory variables, candidate model families, validation scheme, evaluation metrics, and required outputs. The agent may use an LLM to translate the instruction into a formal model specification or executable code, but numerical estimation is delegated to the selected statistical or machine-learning tool. A modeling request is represented as: mod=(,,,,)G_mod= (Y,X,M,V,C ) (23) where Y defines the target outcome, X defines the candidate predictors, M is the authorized model library, V defines the validation design and performance metrics, and C contains modeling constraints. The candidate library may include random-utility models, statistical classifiers, machine-learning models, neural networks, and LLM-based role-play predictors. Let trD_tr, vaD_va, and teD_te denote the training, validation, and test partitions. The agent selects a model family m and configuration λ by solving: (m⋆,⋆)∈argminm∈,∈Λmℛ^va(m,,tr,va)+ηΩ(m,) (m , λ )∈ _m , λ∈ _m \ R_va (m, λ;D_tr,D_va )+η (m, λ ) \ (24) where ℛ^va R_va is the validation loss, Ω(⋅) (·) is an optional complexity or constraint penalty, and η controls the penalty weight. The selected model is then fitted as: ℳ^=Train(m⋆,⋆,tr∪va) M=Train (m , λ ;D_tr _va ) (25) For commuter mode choice, the general prediction function can be written as: y^is=f^(i,is,is) y_is=f_ θ (p_i,z_is,a_is ) (26) where ip_i contains individual characteristics, isz_is contains scenario-specific attributes, isa_is defines the available choice set, and θ denotes the fitted parameters or learned model state. When a random-utility model is selected, the utility of mode j may be represented as: Uisj=Visj+εisj,Visj=⊤isjU_isj=V_isj+ _isj, V_isj= β x_isj (27) where isjx_isj contains individual-, scenario-, and alternative-specific variables. Under a multinomial logit specification, the choice probability is: Pisj=aisjexp(Visj)∑k∈aiskexp(Visk)P_isj= a_isj (V_isj ) _k a_isk (V_isk ) (28) where J is the set of travel modes and aisj∈0,1a_isj∈\0,1\ indicates whether mode j is available (39). More flexible specifications, such as mixed logit or latent-class models, may be selected when the user requests explicit treatment of taste heterogeneity (30). The framework also permits LLM-based choice prediction through alternative prompt framings. Let ρ(⋅)ρ(·) denote a prompt-construction function and ψ denote the LLM configuration. For repeated run r, the predicted choice is: y^is(r)=Fψ[ρ(i,is,is,mod);ϵr],r=1,…,R y_is^(r)=F_ψ [ρ (p_i,z_is,a_is,u_mod ); _r ], r=1,…,R (29) where ϵr _r represents stochastic variation across repeated runs. The empirical LLM-based probability of choosing mode j is: P^isjLLM=1R∑r=1R(y^is(r)=j) P_isj^LLM= 1R _r=1^RI ( y_is^(r)=j ) (30) The LLM predictor may operate in zero-shot or few-shot mode. In zero-shot prediction, the prompt contains the respondent profile, commuting scenario, available modes, and output requirements but no labeled examples. In few-shot prediction, labeled examples are added from the training data. Examples from validation or test respondents are not permitted. When each respondent contributes multiple scenarios, data partitioning is conducted at the respondent level rather than the record level. Therefore, ℐtr∩ℐva=ℐtr∩ℐte=ℐva∩ℐte=∅I_tr _va=I_tr _te=I_va _te= (31) where ℐtrI_tr, ℐvaI_va, and ℐteI_te are the respondent sets assigned to the three partitions. This rule prevents information from the same respondent from appearing in both model development and evaluation. The modeling output is represented as: ℛmod=(^,^,ℐeffect,model,diagnostic,evaluation)R_mod= ( θ, Y,I_effect,U_model,G_diagnostic,V_evaluation ) (32) where θ contains the estimated parameters or fitted model state, Y contains predictions, ℐeffectI_effect contains effect estimates or variable importance measures, modelU_model contains uncertainty information, diagnosticG_diagnostic contains diagnostic outputs, and evaluationV_evaluation contains validation metrics. The agent must distinguish predictive associations from causal effects unless the study design supports causal identification. Each workflow version follows a forward execution path, with each stage producing a structured output that the next agent consumes through a defined interface: a survey specification is provided to the Data Collection Agent, which returns raw data; the raw data are provided to the Data Processing Agent, which returns a processed dataset; and the processed dataset is provided to the Data Modeling Agent, which returns fitted models, predictions, diagnostics, and evaluation metrics. As shown by the dashed arrows in Figure 1, diagnostics from one stage may motivate a researcher-approved revision of an earlier survey, processing, or modeling specification in a subsequent workflow version. The present case study implemented one formal forward pass; the dashed paths represent supported refinement mechanisms for subsequent versions. Section 4 describes how the three agents were instantiated for the McGill commuter mode-choice case study, and Section 5 reports the resulting analyses and predictions. 4 Case Study The framework in Section 3 was instantiated on a commuter mode choice case study at McGill University in Montreal, Canada, a setting where students commute by walking, cycling, transit, and driving across a full range of seasonal weather, from summer heat to winter snow. This section describes how each of the three agents was implemented for that case study; detailed results appear in Section 5. 4.1 Data Collection Agent: Chatbot Survey and Deployment The Data Collection Agent was implemented as a purpose-built conversational chatbot, TravelBehaviorSurveyBot, developed in Voiceflow and deployed as a web application accessible through a public link and QR code (Figure 2). Respondents completed the survey on their own devices through a guided dialogue that collected demographic and mobility-resource information, usual summer and winter travel patterns, weekly mode-use frequencies, and ratings of factors influencing mode choice, followed by five SP weather scenarios. Each scenario presented one photorealistic image generated using defined prompts to represent Sunny, Hot–humid, Rainy, Foggy/cold, and Snowy commuting conditions. All respondents were shown the same fixed image for each scenario. For each scenario, respondents selected among five modeled alternatives: walking, cycling, public transit, bike and public transit combined, and driving. The chatbot followed predefined survey logic and did not alter the experimental attributes or choice set during deployment. Figure 2 illustrates the chatbot interface and the weather images presented during the survey. This design required neither app installation nor an interviewer and ensured consistent presentation of the scenarios across respondents. The Phase 1 sample included 92 McGill student commuters. Figure 2: Chatbot survey interface and sample weather-scenario question with embedded generated images. 4.2 Data Processing Agent The Data Processing Agent converted the chatbot’s raw, JSON-encoded response export into an analysis-ready dataset. Processing steps included parsing nested list- and slider-valued fields, validating and recoding categorical responses, constructing derived features, and flagging incomplete or inconsistent records for exclusion. Of the 460 potential observations from 92 respondents and five scenarios, six were excluded because the recorded mode choice was missing or unparseable. The final processed dataset therefore contained 454 valid respondent–scenario observations. Key derived features include weekly mode-use frequency by season, the seven factor-importance scores, binary vehicle-, license-, and bicycle-ownership flags, and years of cycling experience, which are reported in Section 5.1. 4.3 Data Modeling Agent: Modeling and LLM Setup The Data Modeling Agent implemented two families of models on the processed dataset. The traditional family comprises an MNL estimated in Biogeme (5) and ML classifiers, principally logistic regression and random forest. The LLM family comprises nine locally run models ranging from 2 to 35 billion parameters. The traditional and LLM approaches were executed as parallel model families rather than as sequential modeling stages. For each respondent–scenario pair, an LLM receives a natural-language representation of the respondent profile together with the corresponding weather condition and is asked to predict the selected travel mode. The prompting experiments follow a structured progression in which two prompt framings, Expert (EXP) and Role-Play (RP), are crossed with two levels of contextual information, Base Context (BC) and Richer Context (RC); the richer setting incorporates additional habitual travel information from respondents’ profiles. These four base conditions are then extended through persona-augmented prompting based on respondent personas extracted using a three-class latent class analysis, few-shot in-context learning with labeled examples at several values of k, and direct image input for vision-capable models using the same weather images shown to survey respondents. Predictions are evaluated for both the five-alternative mode-choice task and a binary active-versus-non-active classification, with all LLM and traditional models assessed on the same 454 respondent–scenario observations. Accuracy is reported and interpreted relative to random-guessing baselines of 20.0% for the five-class task and 50.0% for the binary task. The full traditional-model estimates and LLM experiments are reported in Section 5. 5 Results 5.1 Sample Profile and Traditional Model Results 5.1.1 Sample Profile and Weather-Related Mode Choice The Phase 1 sample comprises 92 McGill student commuters, predominantly aged 18–24, with a nearly balanced gender distribution and heterogeneous mobility resources (Figure 3). Although 83% of respondents hold a driver’s license, 36% live in households without a motor vehicle and approximately half own a personal bicycle. Self-reported habitual commute patterns also vary considerably by season: cycling declines from 13.0% of primary summer commutes to 2.2% in winter, while public transit increases from 20.7% to 43.5%, indicating substantial seasonal adaptation in travel behavior. A three-class latent class analysis further identifies a small group of year-round cyclists and two larger groups characterized by active–transit and car–transit travel patterns; these behavioral personas are subsequently used in the persona-based LLM experiments. Figure 3: Sample profile of survey respondents. The stated-preference choices vary systematically across the five weather scenarios (Figure 4). Walking declines moderately across conditions, whereas cycling and bike+transit show the largest reductions and approach minimal shares under the Snowy scenario. In contrast, public transit gains share, increasing from 17.6% under the Sunny scenario to 45.1% under the Snowy scenario, while driving changes comparatively little. Overall, the descriptive results indicate that active and multimodal alternatives are more sensitive to the predefined scenarios than the motorized alternatives, with public transit serving as the primary substitute as conditions become less favorable. Figure 4: Stated mode share across weather scenarios (left) and change in mode share relative to the Sunny scenario, in percentage points (right). 5.1.2 Discrete Choice and Machine-Learning Models The MNL model provides an interpretable benchmark for the weather-related mode-choice patterns observed descriptively. The expanded specification includes weather-scenario indicators, mobility resources, cycling experience, age, household income, and self-rated weather importance, and significantly improves model fit relative to the baseline specification (LR =53.92=53.92, df=8df=8, p<0.0001p<0.0001). With Walking and Sunny weather as the reference categories, the coefficients indicate lower relative utility for cycling and higher relative utility for public transit and driving under adverse scenarios. The Cycling coefficients are negative under Rainy (β=−0.856β=-0.856, p=0.044p=0.044), Foggy/cold (β=−1.166β=-1.166, p=0.012p=0.012), and Snowy conditions (β=−1.820β=-1.820, p=0.009p=0.009), with the largest reduction under the Snowy scenario. Bike+Transit also has a negative coefficient under Snowy conditions (β=−1.695β=-1.695, p=0.050p=0.050). In contrast, Public Transit has positive coefficients under Hot–humid (β=0.518β=0.518, p=0.030p=0.030), Rainy (β=1.070β=1.070, p<0.001p<0.001), Foggy/cold (β=1.024β=1.024, p<0.001p<0.001), and Snowy conditions (β=1.324β=1.324, p<0.001p<0.001). Driving similarly has positive coefficients under Rainy (β=0.949β=0.949, p=0.007p=0.007), Foggy/cold (β=0.723β=0.723, p=0.031p=0.031), and Snowy conditions (β=1.042β=1.042, p=0.003p=0.003). Bicycle ownership and cycling experience are positively associated with the cycling-related alternatives, whereas household motor-vehicle availability and holding a driver’s license are positively associated with driving. Overall, the MNL estimates are consistent with the descriptive results in Figure 4. Table 2: Selected Multinomial Logit Coefficients (Reference Mode: Walking; Reference Weather: Sunny) Mode Covariate Coef. Rob. t p Cycling Rainy (vs. Sunny) −-0.856 −-2.01 0.044* Cycling Foggy/cold (vs. Sunny) −-1.166 −-2.52 0.012* Cycling Snowy (vs. Sunny) −-1.820 −-2.63 0.009** Cycling Owns bicycle 1.219 2.06 0.039* Cycling Cycling experience (yrs.) 0.341 2.67 0.008** Public Transit Hot–humid (vs. Sunny) 0.518 2.17 0.030* Public Transit Rainy (vs. Sunny) 1.070 3.68 <0.001*** Public Transit Foggy/cold (vs. Sunny) 1.024 4.08 <0.001*** Public Transit Snowy (vs. Sunny) 1.324 4.59 <0.001*** Bike+Transit Snowy (vs. Sunny) −-1.695 −-1.96 0.050* Bike+Transit Owns bicycle 1.572 2.64 0.008** Bike+Transit Cycling experience (yrs.) 0.281 2.49 0.013* Bike+Transit Self-rated weather importance 0.471 2.54 0.011* Driving Rainy (vs. Sunny) 0.949 2.68 0.007** Driving Foggy/cold (vs. Sunny) 0.723 2.16 0.031* Driving Snowy (vs. Sunny) 1.042 3.01 0.003** Driving Household motor vehicles 0.767 2.67 0.007** Driving Driver’s license 1.769 2.13 0.033* * p≤0.05p≤ 0.05; ** p<0.01p<0.01; *** p<0.001p<0.001. For predictive benchmarking, the MNL, logistic regression, and random forest were evaluated using the same five-fold cross-validation procedure grouped by respondent, ensuring that observations from the same individual did not appear in both training and test sets. Random forest produced the highest accuracy, reaching 69.6% for the five-class mode-choice task and 88.8% for the binary active-versus-non-active task. Logistic regression reached 60.2% and 85.1%, respectively, while the MNL achieved 44.7% five-class and 81.2% binary accuracy. The five-class results show a predictive advantage for the machine-learning models, particularly random forest, whereas the smaller differences in binary accuracy should be interpreted in light of the strong majority-class imbalance. These findings highlight complementary roles: the MNL provides behaviorally interpretable evidence on weather scenarios and traveler characteristics, whereas the ML models provide stronger data-trained benchmarks for the LLM experiments that follow. 5.2 LLM Zero-Shot The zero-shot experiments evaluate two dimensions of prompt design, context richness and prompt framing. Base Context (BC) includes respondent demographics, mobility resources, factor-importance ratings, and the weather scenario, whereas Richer Context (RC) additionally incorporates additional habitual travel information. Expert (EXP) framing asks the LLM to predict the respondent’s choice as a transportation expert, while Role-Play (RP) framing asks the model to simulate the respondent directly. The four resulting conditions, EXP-RC, EXP-BC, RP-RC, and RP-BC, were evaluated across nine locally deployed LLMs ranging from 2 to 35 billion parameters (Table 3). Figure 5 reveals two consistent patterns. First, incorporating additional habitual travel information substantially improves zero-shot prediction. For Gemma 3:4B, five-class accuracy increases from 41.3% under EXP-BC to 64.2% under EXP-RC, indicating that habitual travel information provides a strong predictive signal beyond demographics and mobility resources. Second, Expert framing generally outperforms Role-Play framing at the same context level, with the difference more pronounced among smaller models. Larger models in the evaluated set generally produce higher five-class accuracy than the smallest models, although the pattern is not monotonic and model size is confounded with model family and architecture. The highest zero-shot five-class point estimate is achieved by Gemma 4:12B at approximately 70%, whereas binary accuracy varies less across models. Together, these results indicate that the information supplied to the model and the framing of the prediction task can be as important as the selected model. Table 3: LLM Zero-Shot Five-Class and Binary Accuracy EXP-RC EXP-BC RP-RC RP-BC Model 5-cl. Binary 5-cl. Binary 5-cl. Binary 5-cl. Binary Gemma 4:e2B 44.0 ± 4.5 68.7 ± 4.1 31.6 ± 4.1 66.8 ± 4.3 41.2 ± 5.0 67.9 ± 4.2 32.3 ± 4.3 65.8 ± 4.1 Llama 3.2:3B 41.2 ± 4.3 62.5 ± 4.5 12.7 ± 3.6 56.7 ± 4.4 26.7 ± 3.7 60.7 ± 4.3 18.6 ± 3.6 60.3 ± 4.5 Gemma 3:4B 64.2 ± 4.4 79.1 ± 3.3 41.3 ± 4.0 64.1 ± 4.4 54.5 ± 4.2 75.9 ± 3.9 36.7 ± 4.2 55.6 ± 4.4 Gemma 4:e4B 42.8 ± 4.4 67.9 ± 4.4 40.4 ± 4.5 66.4 ± 4.2 40.1 ± 4.5 68.3 ± 4.0 35.7 ± 4.5 66.0 ± 4.2 Gemma 4:12B 69.9 ± 3.9 73.5 ± 4.2 64.6 ± 4.3 73.0 ± 4.2 69.8 ± 4.4 71.5 ± 4.1 63.9 ± 5.0 70.6 ± 4.1 Gemma 4:26B 61.2 ± 4.6 71.1 ± 4.0 59.6 ± 4.7 69.6 ± 4.4 59.3 ± 4.5 69.7 ± 4.0 57.4 ± 4.3 69.7 ± 4.5 Gemma 3:27B 60.7 ± 4.6 71.0 ± 4.1 54.1 ± 4.1 69.5 ± 4.2 49.1 ± 4.5 71.5 ± 4.2 47.1 ± 4.6 68.3 ± 4.2 Gemma 4:31B 63.1 ± 4.4 69.5 ± 4.0 62.2 ± 4.7 68.6 ± 4.0 59.2 ± 4.6 70.4 ± 4.0 56.7 ± 4.9 70.3 ± 4.3 Qwen 3.6:35B 64.3 ± 4.3 71.9 ± 4.0 61.5 ± 4.4 68.4 ± 4.1 56.4 ± 4.6 71.0 ± 4.3 52.1 ± 4.7 68.5 ± 4.2 Figure 5: Zero-shot accuracy across all four base conditions (Expert/Role-Play × Base/Richer Context) for nine models, shown separately for five-class and binary accuracy. Prediction difficulty also varies systematically across weather conditions and travel modes. As shown in Figure 6, Snowy and Rainy scenarios are generally easier to predict because they produce stronger and more consistent shifts toward public transit and away from cycling. In contrast, Sunny and Warm conditions yield more heterogeneous choices and lower prediction accuracy. At the mode level, Public Transit is predicted relatively well across models, whereas Bike+Transit is consistently the most difficult alternative. These differences suggest that LLM accuracy depends not only on model and prompt configuration but also on the behavioral regularity of the scenario and alternative being predicted. Figure 6: Accuracy patterns by weather condition and mode across all evaluated models, with a bar-chart breakdown of weather- and mode-level accuracy under the EXP-RC experiment. 5.3 Advanced LLM Configurations: Persona, Few-Shot, and Vision The advanced experiments examine whether additional behavioral and contextual information improves prediction beyond the base zero-shot configurations. Persona augmentation is most beneficial when explicit travel history is unavailable. Under Base Context, adding a behavioral persona substantially improves some smaller models, particularly Llama 3.2:3B, whereas its effect is limited when Richer Context already contains respondents’ habitual seasonal travel modes. This pattern suggests that persona information mainly summarizes behavioral information when direct travel-history information is absent, rather than adding substantial information once habitual travel information is already available. Few-shot prompting provides further improvements in five-class prediction, although the magnitude of the gain varies across models and diminishes as additional examples are introduced. Figure 7 shows that a small number of labeled examples improves performance for most models, with proportionally larger gains among smaller models. Accuracy generally stabilizes after approximately ten examples, while additional demonstrations provide limited or inconsistent improvements. The alternative example-selection strategies considered produce broadly comparable performance, suggesting that the presence of informative demonstrations is more important than a complex retrieval strategy. These improvements are more evident for the five-class task than for the binary active-versus-non-active classification, indicating that few-shot prompting is particularly effective for predicting specific travel modes. Vision-capable models were further evaluated using the same generated weather images presented to survey respondents. Visual inputs provide additional contextual information for several models, although the effect is not uniform. For some larger models, vision-based prompting approaches or exceeds the corresponding text-based Richer Context performance, indicating that the specific scenario images provide predictive signal that can complement or partially substitute for text-based contextual descriptions. Qwen 3.6:35B, for example, improves when the textual weather representation is replaced by the scenario image, while Gemma 4:26B benefits further when vision input is combined with few-shot examples. Other models show smaller gains or declines, indicating that multimodal capability alone does not guarantee improved mode-choice prediction. Overall, the advanced experiments indicate that additional information is most valuable when it contributes behavioral or contextual signals not already represented in the prompt. Habitual travel information remains the most consistent source of predictive information, persona descriptions are particularly useful when that history is unavailable, and few-shot examples provide useful task guidance with rapidly diminishing returns. Vision inputs provide a distinct source of contextual information and can improve prediction for selected models, supporting further evaluation of multimodal LLMs in travel-behavior applications. Figure 7: Combined LLM results. Top row: (left) zero-shot versus peak few-shot five-class accuracy per model, and (right) few-shot scaling curves across the six models. Bottom row: (left) experiment spread across all conditions, and (right) vision augmentation results comparing text-only baselines (EXP-RC and EXP-BC) against vision input (EXP-Img) and vision with three-shot prompting (EXP-ImgFS3). 6 Discussion The results demonstrate how conversational data collection, conventional behavioral models, and LLM-based prediction can be connected within a common multi-agent workflow for travel-behavior research. Beyond the performance of any individual model, the study provides evidence on three broader questions: whether an AI-assisted survey can capture meaningful context-dependent travel behavior, which forms of information are most useful for LLM-based mode-choice prediction, and how these models compare with established DCMs and ML approaches. Multi-agent workflow and behavioral evidence. The workflow links the Data Collection, Data Processing, and Data Modeling Agents through structured outputs, creating a transparent and modular process from survey administration to behavioral prediction. This organization does not improve accuracy by itself; rather, it enables consistent data transfer and systematic evaluation of where performance gains arise. The chatbot survey produced coherent differences across the weather scenarios: cycling and bike+transit declined as conditions worsened, whereas public transit gained substantial share. The MNL estimates were consistent with these descriptive patterns, showing lower relative utility for cycling-related modes and higher relative utility for transit and driving under adverse scenarios. This agreement supports the internal coherence of the survey responses, although it does not by itself establish external validity or actual travel behavior under observed weather. LLM configuration and comparison with conventional models. LLM performance was strongly influenced by the information included in the prompt. RC, which adds habitual travel information to BC, produced the most consistent improvement, showing that habitual travel behavior contains predictive information not fully captured by demographics and mobility resources. EXP framing generally outperformed RP, particularly for smaller models, while gains from increasing model size diminished among larger models. Persona augmentation and few-shot prompting were most useful when they supplied missing information. Persona descriptions improved prediction mainly under BC, where habitual travel information was unavailable, but added little under RC. Few-shot examples improved five-class accuracy for several models, particularly smaller ones, although the gains generally stabilized after a small number of demonstrations. The conventional and LLM-based models therefore play complementary roles. The MNL provides behavioral interpretation, while logistic regression and random forest offer data-trained predictive benchmarks. Random forest achieved the highest trained-model accuracy, while the best zero-shot LLM produced a comparable five-class point estimate without task-specific fitting to the survey records. This result suggests potential value for LLM-based prediction when labeled travel-behavior data are limited, although the small sample and multiple evaluated configurations require cautious interpretation. The five-class task provides a clearer assessment of whether models distinguish among individual modes. Vision-based prediction. The vision experiments directly connect the survey and prediction stages of the multi-agent workflow by providing the models with the same generated weather images shown to respondents. The best vision-based configuration achieved a five-class accuracy point estimate of 71.5%, compared with 69.9% for the highest text-only zero-shot LLM result and 69.6% for random forest. These small differences should be interpreted as descriptive point-estimate differences rather than evidence of statistical superiority. Vision combined with few-shot prompting also produced gains for selected models. Because each weather condition was represented by one fixed image, the results indicate that these specific images provided additional predictive signal for selected models; they do not establish generalization to unseen weather images or visual environments. Limitations and future research. The findings should be interpreted in light of several limitations. The sample consists of 92 students from one university and is not representative of the wider population. The analysis is based on stated choices rather than observed travel under actual weather conditions, and uncontrolled visual features within the generated images may have influenced both respondents and models. Each weather condition was represented by one fixed image, so weather condition and image identity were not separately identified. The relatively small and imbalanced dataset also limits detailed analysis of minority modes and traveler subgroups, while the results may vary across other LLM families and inference settings. The framework itself was not experimentally compared with a conventional research workflow; its contribution therefore concerns integration, traceability, and modularity rather than a measured reduction in survey cost or processing effort. In addition, the configuration comparisons were exploratory, and the highest reported accuracy may partly reflect the evaluation of multiple models and prompt settings on the same sample. Future research could evaluate the workflow with larger and more diverse samples, additional trip purposes and locations, and revealed-preference validation. Controlled image experiments could isolate the effects of weather, road condition, and surrounding environment. Further work could also examine prediction calibration, subgroup performance, and transferability across populations. The multi-agent workflow could support adaptive surveys and carefully validated synthetic respondents, but such outputs should be validated against human behavior before being used in transportation planning or policy analysis. 7 Conclusions This study developed and evaluated a multi-agent workflow that integrates conversational survey design, AI-generated visual scenarios, structured data processing, conventional choice and machine-learning models, and LLM-based prediction within a unified travel-behavior research framework. The workflow was demonstrated using stated mode choices from 92 McGill student commuters across five image-augmented weather scenarios. The empirical results support three main conclusions. First, the conversational survey produced behaviorally coherent responses across weather conditions, with adverse weather reducing cycling-related choices and increasing reliance on public transit. Second, the comparison of MNL, random forest, and LLM-based predictors highlighted their complementary roles. MNL provided interpretable behavioral relationships, random forest offered a strong trained predictive benchmark, and LLMs enabled flexible incorporation of habitual travel information, persona descriptions, examples, and visual context without requiring a separate model specification for each information type. Third, LLM performance depended strongly on prompt design and input configuration. Habitual travel information and expert-oriented framing generally improved prediction, while selected vision-capable models benefited from the inclusion of the same weather images shown to respondents. The best zero-shot LLM achieved a point estimate comparable to the trained machine-learning benchmark without task-specific fitting to the survey observations. The main contribution of this study is therefore not the replacement of established travel-choice models, but the demonstration of an integrated and extensible workflow in which data collection, processing, behavioral analysis, predictive modeling, and multimodal reasoning can be coordinated through specialized agents. This framework provides a practical basis for incorporating heterogeneous behavioral information into travel-choice analysis while retaining conventional models for interpretation, benchmarking, and validation. More broadly, the findings indicate that LLMs may expand how traveler characteristics and contextual conditions are represented in transportation research, provided that their outputs are systematically evaluated against observed human behavior. 8 Acknowledgments The authors used OpenAI ChatGPT and Claude (Anthropic) to assist with manuscript language editing and restructuring. All research-design decisions, data analyses, numerical results, references, interpretations, and final wording were reviewed by the authors. The LLM-based choice predictions evaluated in this study form part of the reported research methodology. 9 Author Contributions All authors contributed to the study conception, analysis, and writing. All authors reviewed the results and approved the final version of the manuscript. 10 Declaration of Conflicting Interests The authors declare that there is no conflict of interest regarding the publication of this paper. References Ali and Fissha (2026) M. Ali and Y. Fissha Propose adjustable support vector machine approach for classifying imbalanced work travel mode choice data. Transportation Research Interdisciplinary Perspectives 35, p. 101786. Cited by: §1, §2.1. Argyle et al. (2023) L. P. Argyle, E. C. Busby, N. Fulda, et al. Out of one, many: using language models to simulate human samples. Political Analysis 31 (3), p. 337–351. Cited by: §1, §2.3. Bansal and Kockelman (2018) P. Bansal and K. M. Kockelman Are we ready to embrace connected and self-driving vehicles? a case study of Texans. Transportation 45 (2), p. 641–675. Cited by: §1, §2.1. Ben-Akiva and Lerman (1985) M. Ben-Akiva and S. R. Lerman Discrete choice analysis: theory and application to travel demand. MIT Press, Cambridge, MA. Cited by: §1, §1, §2.1. Bierlaire (2003) M. Bierlaire BIOGEME: a free package for the estimation of discrete choice models. In Proceedings of the 3rd Swiss Transportation Research Conference, Ascona, Switzerland. Cited by: §4.3. Bisbee et al. (2024) J. Bisbee, J. D. Clinton, C. Dorff, B. Kenkel, and J. M. Larson Synthetic replacements for human survey data? the perils of large language models. Political Analysis 32 (4), p. 401–416. Cited by: §1, §2.3. Böcker et al. (2013) L. Böcker, M. Dijst, and J. Prillwitz Impact of everyday weather on individual daily travel behaviours in perspective: a literature review. Transport Reviews 33 (1), p. 71–91. Cited by: §1, §2.1. Bonnel and Munizaga (2018) P. Bonnel and M. A. Munizaga Transport survey methods—in the era of big data facing new and old challenges. Transportation Research Procedia 32, p. 1–15. Note: Transport Survey Methods in the Era of Big Data: Facing the Challenges External Links: ISSN 2352-1465, Document, Link Cited by: §2.2. Brand et al. (2023) J. Brand, A. Israeli, and D. Ngwe Using AI for market research. Technical report Technical Report 23-062, Harvard Business School. Cited by: §2.3. Brown et al. (2020) T. Brown, B. Mann, N. Ryder, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems 33, p. 1877–1901. Cited by: §1. Celino and Calegari (2020) I. Celino and G. R. Calegari Submitting surveys via a conversational interface: an evaluation of user acceptance and approach effectiveness. International Journal of Human-Computer Studies 139, p. 102410. Cited by: §1, §2.2. Chen et al. (2024) N. Chen, Y. Wang, Y. Deng, and J. Li The oscars of AI theater: a survey on role-playing with language models. arXiv preprint arXiv:2407.11484. Cited by: §2.3. Chu et al. (2025) M. Chu, L. Terhorst, K. Reed, et al. LLM-based multi-agent system for simulating and analyzing marketing and consumer behavior. In 2025 IEEE International Conference on E-Business Engineering (ICEBE), p. 72–79. Cited by: §1, §2.4. Cui et al. (2025) Z. Cui, N. Li, and H. Zhou A large-scale replication of scenario-based experiments in psychology and management using large language models. Nature Computational Science 5 (8), p. 627–634. Cited by: §2.3. Estevez et al. (2025) M. Estevez, M. T. Ballestar, and J. Sainz Market research and knowledge using generative AI: the power of large language models. Journal of Innovation & Knowledge 10 (5), p. 100796. Cited by: §2.3. Farooq and Cherchi (2024) B. Farooq and E. Cherchi Workshop synthesis: virtual reality, visualization and interactivity in travel survey. Transportation Research Procedia 76, p. 686–691. Cited by: §2.2. Garzon-Vico et al. (2026) A. Garzon-Vico, K. S. Komalapati, A. Shahid, and J. Rosier Using large language models to construct virtual top managers: a method for organizational research. arXiv preprint arXiv:2601.18512. Cited by: §2.4. Hensher and Ton (2000) D. A. Hensher and T. T. Ton A comparison of the predictive potential of artificial neural networks and nested logit models for commuter mode choice. Transportation Research Part E 36 (3), p. 155–172. Cited by: §1, §2.1. Horton (2023) J. J. Horton Large language models as simulated economic agents: what can we learn from homo silicus?. Technical report Technical Report 31122, National Bureau of Economic Research. Cited by: §1, §2.3. Hullman et al. (2026) J. Hullman, D. Broska, H. Sun, and A. Shaw This human study did not involve human subjects: validating LLM simulations as behavioral evidence. arXiv preprint arXiv:2602.15785. Cited by: §1, §2.3. Hyland et al. (2018) M. Hyland, C. Frei, A. Frei, and H. S. Mahmassani Riders on the storm: exploring weather and seasonality effects on commute mode choice in Chicago. Travel Behaviour and Society 13, p. 44–60. Cited by: §1, §2.1. Jia et al. (2024) F. Jia, Z. Ye, S. Lai, et al. Can large language model agents simulate human trust behavior?. Advances in Neural Information Processing Systems 37, p. 15674–15729. Cited by: §2.3. Krajcovic et al. (2026) M. Krajcovic, P. Demcak, and E. Kuric Talking surveys: how photorealistic embodied conversational agents shape response quality, engagement, and satisfaction. Behavior Research Methods 58 (8), p. 212. Cited by: §1, §2.2. Krueger et al. (2016) R. Krueger, T. H. Rashidi, and J. M. Rose Preferences for shared autonomous vehicles. Transportation Research Part C 69, p. 343–355. Cited by: §1, §2.1. Li et al. (2025) Y. Li, Y. Liu, and M. Yu Consumer segmentation with large language models. Journal of Retailing and Consumer Services 82, p. 104078. Cited by: §2.3. Liu et al. (2024) T. Liu, M. Li, and Y. Yin Can large language models capture human travel behavior? evidence and insights on mode choice. Technical report SSRN Working Paper 4937575. Cited by: §1, §2.3. Liu et al. (2025a) T. Liu, M. Li, and Y. Yin Aligning LLM with human travel choices: a persona-based embedding learning approach. arXiv preprint arXiv:2505.19003. Cited by: §2.3. Liu et al. (2026) T. Liu, J. Yang, Y. Yin, M. Li, L. Wang, and Z. Zhu Aligning LLM agents with human learning and adjustment behavior: a dual agent approach. Transportation Research Part C 191, p. 105818. Cited by: §1, §2.4. Liu et al. (2025b) T. Liu, J. Yang, and Y. Yin Toward LLM-agent-based modeling of transportation systems: a conceptual framework. Artificial Intelligence for Transportation 1, p. 100001. Cited by: §1, §2.4. McFadden and Train (2000) D. McFadden and K. Train Mixed MNL models for discrete response. Journal of Applied Econometrics 15 (5), p. 447–470. Cited by: §3.5. McFadden (1974) D. McFadden Conditional logit analysis of qualitative choice behavior. In Frontiers in Econometrics, P. Zarembka (Ed.), p. 105–142. Cited by: §1, §1, §2.1. Mo et al. (2023) B. Mo, H. Xu, R. Ma, et al. Large language models for travel behavior prediction. arXiv preprint arXiv:2312.00819. Cited by: §2.3. Nishida et al. (2025) R. Nishida, T. Ishigaki, and M. Onishi Large language models predict transportation mode choice behavior for a variety of alternative sets. Transportation Research Record. Cited by: §2.3. Park et al. (2024) J. S. Park, C. Q. Zou, A. Shaw, et al. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109. Cited by: §2.3. Ren et al. (2026) F. Ren, Z. Zhang, T. Mendel, and T. Yabe Assessing the feasibility of a video-based conversational chatbot survey for measuring perceived cycling safety: a pilot study in New York City. arXiv preprint arXiv:2604.07375. Cited by: §1, §2.2. Sameen et al. (2025) M. Sameen, X. Zhang, and X. Zhao Synthesizing attitudes, predicting actions (SAPA): behavioral theory-guided LLMs for ridesourcing mode choice modeling. arXiv preprint arXiv:2509.18181. Cited by: §1, §2.3. Saneinejad et al. (2012) S. Saneinejad, M. J. Roorda, and C. Kennedy Modelling the impact of weather conditions on active transportation travel behaviour. Transportation Research Part D 17 (2), p. 129–137. Cited by: §1, §2.1. Tang and Shang (2025) J. Tang and Y. Shang AURA: a reinforcement learning framework for AI-driven adaptive conversational surveys. arXiv preprint arXiv:2510.27126. Cited by: §2.2. Train (2009) K. E. Train Discrete choice methods with simulation. 2nd edition, Cambridge University Press. Cited by: §1, §2.1, §3.5. Voltes-Dorta and Suau-Sanchez (2025) A. Voltes-Dorta and P. Suau-Sanchez Can large language models mimic airline passenger preferences?. Artificial Intelligence for Transportation 3, p. 100034. Cited by: §2.3. Wu et al. (2026) Y. Wu, Y. Liu, and X. Deng MALLES: a multi-agent LLMs-based economic sandbox with consumer preference alignment. arXiv preprint arXiv:2603.17694. Cited by: §2.4. Xu et al. (2024) R. Xu, X. Wang, J. Chen, et al. Character is destiny: can role-playing language agents make persona-driven decisions. Note: Preprint Cited by: §2.3. Yan et al. (2026) Y. Yan, T. Liu, and Y. Yin Valuing time in silicon: can large language models replicate human value of travel time. Travel Behaviour and Society 44, p. 101245. Cited by: §2.3. Yu et al. (2025) J. Yu, J. Zhao, L. Miranda-Moreno, and M. Korp Modular AI agents for transportation surveys and interviews: advancing engagement, transparency, and cost efficiency. Communications in Transportation Research 5 (1), p. 100172. Cited by: §1, §2.2. Zhao et al. (2023) W. X. Zhao, K. Zhou, J. Li, et al. A survey of large language models. arXiv preprint arXiv:2303.18223. Cited by: §1, §2.4.