Paper deep dive
User Review Writing via Interview with Dialogue Systems
Yoshiki Tanaka, Michimasa Inaba
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/13/2026, 12:30:29 AM
Summary
The paper proposes a novel dialogue-based system using GPT-4 to assist users in writing detailed e-commerce reviews. The system conducts an interview to elicit product experiences, generates a review text based on the dialogue history, and predicts a corresponding rating. Evaluations show that the system produces more helpful reviews than human-written ones and requires less editing than baseline systems, though it faces challenges regarding response latency.
Entities (5)
Relation Signals (4)
Review Text Generator â generates â User Review
confidence 95% ¡ the review text generator generates review text based on the dialogue history
Rating Predictor â predicts â Rating
confidence 95% ¡ the rating predictor predicts a rating consistent with the generated review text
Interview Dialogue System â elicitsinformationfrom â User
confidence 90% ¡ the dialogue system acts as an interviewer, eliciting user opinions on products
GPT-4 â powers â Interview Dialogue System
confidence 90% ¡ We use the gpt-4-0613 model to implement our system.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:User reviews on e-commerce and review sites are crucial for making purchase decisions, although creating detailed reviews is time-consuming and labor-intensive. In this study, we propose a novel use of dialogue systems to facilitate user review creation by generating reviews from information gathered during interview dialogues with users. To validate our approach, we implemented our system using GPT-4 and conducted comparative experiments from the perspectives of system users and review readers. The results indicate that participants who used our system rated their interactions positively. Additionally, reviews generated by our system required less editing to achieve user satisfaction compared to those by the baseline. We also evaluated the reviews from the reader' perspective and found that our system-generated reviews are more helpful than those written by humans. Despite challenges with the fluency of the generated reviews, our method offers a promising new approach to review writing.
Tags
Links
- Source: https://arxiv.org/abs/2603.07070v1
- Canonical: https://arxiv.org/abs/2603.07070v1
Trouble viewing inline? Open PDF directly â
Full Text
47,635 characters extracted from source content.
Expand or collapse full text
User Review Writing via Interview with Dialogue Systems Yoshiki TanakaandMichimasa Inaba The University of Electro-Communications y-tanaka, m-inaba@uec.ac.jp Abstract User reviews on e-commerce and review sites are crucial for making purchase decisions, although creating detailed reviews is time- consuming and labor-intensive. In this study, we propose a novel use of dialogue systems to facilitate user review creation by generat- ing reviews from information gathered during interview dialogues with users. To validate our approach, we implemented our system us- ing GPT-4 and conducted comparative exper- iments from the perspectives of system users and review readers. The results indicate that participants who used our system rated their interactions positively. Additionally, reviews generated by our system required less editing to achieve user satisfaction compared to those by the baseline. We also evaluated the reviews from the readersâ perspective and found that our system-generated reviews are more helpful than those written by humans. Despite chal- lenges with the fluency of the generated re- views, our method offers a promising new ap- proach to review writing. 1 Introduction User reviews posted on e-commerce and review sites benefit both buyers and sellers. For users con- sidering a purchase, these reviews provide prod- uct details and insights, thereby supporting the decision-making process of purchase. Addition- ally, they act as feedback for sellers, enabling prod- uct quality improvement. Reviews that contain specific product details or are based on personal user experiences are especially valuable. How- ever, writing these insightful reviews can be time- consuming and labor-intensive, as the task re- quires careful language and proofreading for po- tential buyers. To reduce usersâ burden of writing reviews, re- searchers (e.g., ( Dong et al.,2017;Li and Tuzhilin, 2019)) have focused on automated review genera- (...) Could you elaborate on what specific features of the Sony PS5 Playstation 5 Console Disc Version you found most impressive? I find that its loading speed, the graphics, and the ablility to play PS4 games. I also have to say I am very impressed buy theduel sense controller and how it feels in the hand. It sounds like the speed and graphics really stood out to you. Could you tell me more about your experience with the loading speed? How did it enhance your gaming experience? Loading screens in games are only a few seconds instead of a minute or more like the older gen consoles and the graphics really shine with colors and detail. (...) Now, let's talk about theDualSense controller. Could you describe how it feels in your hand and how it contributes to your gaming experience? It is very comfortable in that hand and does not feel cramped at all . It has haptic feedback that reallyhelp with immersion in the games, and ... ... Dialogue System The Sony PS5 Playstation 5 Console Disc Version is an impressive piece of gaming technology that has exceeded my expectations in every way. The loading speed is a game-changer, reducing waiting times to mere seconds, which is a significant improvement over older generation consoles . (...) The DualSense controller is another standout feature. It's comfortable to hold and doesn't feel cramped , even during extended gaming sessions. (...) Interview Dialogue Generated Review and Rating User ... Figure 1: Example of review creation supported by the proposed method. In the dialogue and review, the con- tents of the same-colored text correspond to each other. tion, thereby making the review creation process more efficient. Despite these efforts, the ability of existing re- view generation systems to personalize for users and products remains limited. Previous studies have explored utilizing a wide range of informa- tion about products and users for generating re- views. However, available subjective information is limited. Specifically, without direct input from the user, incorporating the userâs actual experi- ences with the product into the generated review is challenging. This constraint significantly limits the systemâs ability to personalize for the user. To overcome this problem, we focused on supporting the creation of reviews by directly eliciting infor- arXiv:2603.07070v1 [cs.HC] 7 Mar 2026 mation about products from users. In this study, we propose the novel utilization of dialogue systems for creating user reviews. Fig- ure1shows an example of the review creation pro- cess supported by the dialogue system according to our proposed method. First, the dialogue system acts as an interviewer, eliciting user opinions on products through interview dialogues. Second, the review text generator generates review text based on the dialogue history. Finally, the rating predic- tor predicts a rating consistent with the generated review text. Our method allows users to easily create reviews by simply interacting with the sys- tem, thus reducing the effort involved in review creation. To evaluate our method, we implemented a sys- tem incorporating our approach using GPT-4. Sub- sequently, we conducted experiments using our system, collecting data on dialogues between the system and users, the generated reviews, predicted ratings, and participantsâ feedback on our system. We discuss the effectiveness of our method after analyzing the collected data. In summary, our main contributions are as follows: 1.As a novel application of dialogue systems, we propose a method for supporting user re- view creation. Furthermore, we developed a system incorporating our approach using GPT-4. 2.We conducted a comprehensive survey from the perspectives of system users and review readers, showing that our method can provide high-quality and helpful reviews for both par- ties. 2 Related Work 2.1 Interview Dialogue System and Dataset The interview dialogues are aimed at eliciting in- formation from the interviewees. Prior research suggests that surveys conducted on chatbot plat- forms yield higher-quality responses than web sur- vey platforms (Kim et al.,2019). This finding in- dicates that employing dialogue systems to collect user opinions and impressions is a promising ap- proach. Researchers have collected interview dialogue data on various topics, including radio ( Majumder et al.,2020), news (Zhu et al.,2021), sports (Sun et al.,2022), and cooking (Okahisa et al.,2022). The objectives of these collections vary from an- alyzing dialogue patterns (Majumder et al.,2020; Okahisa et al.,2022) to dialogue summarization (Zhu et al.,2021). Here, we utilize the interview dialogue system to support the creation of helpful reviews. 2.2 Review Generation User reviews reflect userâs opinions and requests regarding a product. These insights benefit buyers and sellers. Additionally, user reviews have a wide range of applications. Previous research has ap- plied reviews to natural language processing tasks such as recommendations ( Qiu et al.,2021), opin- ion summarization (BraĹžinskas et al.,2020), and task-oriented dialogue (Zhao et al.,2023). User reviews that include detailed information about the product and user experiences are use- ful. However, writing these reviews is a labor- intensive task for humans. To increase the effi- ciency of this process, researchers have proposed automated review generation models, enhancing their review generation capabilities by utilizing information such as ratings ( Dong et al.,2017; Sharma et al.,2018;Li et al.,2019;Kim et al., 2020), images (Truong and Lauw,2019;Vu et al., 2020), past reviews written by the user (Li and Tuzhilin ,2019), and aspect-oriented features (Li and Tuzhilin ,2019). Unlike these studies, we fo- cus on the collaborative writing of user reviews with the support of the dialogue system. Some researchers have focused on supporting users in creating reviews, similar to our approach ( Ni and McAuley,2018;Bhat et al.,2023). For ex- ample, Ni and McAuley proposed utilizing short phrases related to products that are provided by customers, such as review summaries and prod- uct titles, as auxiliary data for generating reviews ( Ni and McAuley,2018). In their system, the user provides information in a unidirectional man- ner. In contrast, we utilize an interview-specific dialogue system to collect information from the user through interactive interaction. The dialogue system can ask follow-up questions to obtain ad- ditional details regarding a product although this information may be ambiguous. This capability supports the creation of detailed reviews. 2.3 Dialogue Summarization In our method, we proposed to convert con- versational data (i.e., interview dialogue history) into non-conversational data (i.e., review texts). Therefore, our work is closely related to dia- logue summarization research. To build an ef- fective model for dialogue summarization, re- searchers have proposed diverse approaches to learning methods (Zou et al.,2021;Li et al.,2023; Zhong et al.,2022;Zhang et al.,2021). Addition- ally, researchers have built dialogue summariza- tion datasets that can be used for training mod- els; these datasets cover daily life conversations (Gliwa et al.,2019;Chen et al.,2021), meetings (Carletta et al.,2006;Zhong et al.,2021), TV se- ries (Chen et al.,2022), and media dialogue (Zhu et al. ,2021). While these studies aim to condense dialogue histories into brief texts, our approach takes a different direction. We focus on extract- ing useful product information for readers from interview dialogues and organizing it into a non- conversational data format, rather than compress- ing it into shorter text. 3 Methodology To create useful reviews, reviewers must provide detailed product information. Interview dialogue systems are employed to effectively elicit this in- formation. To enhance readability, we propose or- ganizing the dialogue history into a non-dialogue format. Our method comprises three processes: in- terview dialogue, review text generation, and rat- ing prediction. In this paper, the systems that per- form these processes are referred to as the âinter- view dialogue system,â the âreview text generator,â and the ârating predictor,â respectively. Our sys- tem utilizes these components in sequence to gen- erate reviews as the output. An overview of our system is shown in Figure 2. We use the gpt-4- 0613 model to implement our system. 3.1 Interview Dialogue System To assist potential buyers in making purchase de- cisions, guiding users to create helpful reviews is crucial. In our approach, therefore, our system should be designed to effectively collect informa- tion from the user. To achieve this, we propose uti- lizing an interview dialogue system. For the inter- view dialogue system, it is desirable to elicit both the pros and cons of a product in a balanced and detailed manner. Specifically, the system should be capable of asking follow-up questions about the content mentioned by the user or changing the topic to inquire about different aspects of the prod- uct. Rating Predictor Interview Dialogue System Review Text Generator S : Could you tell me more ... U: Loading screens in games are only a few seconds instead of ... ... 5 Input: Output: The Sony PS5 Playstation 5 Console Disc Version is an impressive piece of gaming technology that ... Dialogue History Review Text Rating Figure 2: Overview of our system. First, the inter- view dialogue system interviews the user to elicit their impressions and requests about the product they used. Next, the review text generator uses the dialogue his- tory as input to generate a review text. Finally, the rat- ing predictor predicts a rating consistent with the senti- ment of the generated review text. We designed a prompt that incorporated instruc- tions for the system to perform these behaviors. Moreover, aiming to both collect sufficient infor- mation for creating reviews and ensure users donât become bored, we added constraints regarding the number of turns to the prompt. In our experiments, we adopted instructions to ask at least 8 questions and conclude the interview within 15 turns. Addi- tionally, to ensure the interview does not continue indefinitely, we externally implemented a setting in the interview dialogue system to end the dia- logue after 15 turns. The prompt template for the interview dialogue is shown in Appendix A.1. 3.2 Review Text Generator Although the dialogue history between interview dialogue systems and users offers useful and de- tailed product information, it often contains redun- dancies. Consequently, it is not appropriate to post it directly as a user review. Therefore, we pro- pose transforming the dialogue history into a for- mat suitable for reviews. Our review text genera- tor aims to capture the essence of the interview di- alogue history while generating review texts from the perspective of the user. To generate reviews that align with the userâs feedback, the system must faithfully reflect the content of the dialogue history in the review text. Our prompts include in- structions to concisely summarize important infor- mation mentioned during the interview and gener- ate the main body of the user review. The prompt template for the review generation is shown in Ap- pendix A.2. 3.3 Rating Predictor In e-commerce and review platforms, customer ratings are aggregated into a single score, provid- ing other users with an initial impression of the product. For an aggregated score to be reliable, reviewers must assign ratings that accurately re- flect the content of their review text. While the ratings impact the reputation widely, considering the potential for human error in assigning ratings, automating the task might be an effective solution. Our rating predictor automatically outputs a rating consistent with the sentiment of the input review text, ranging from 1 to 5 as an integer. Ratings con- sistent with the content of the review texts could reduce exaggerated scoring caused by user subjec- tivity. As a result, this can improve the reliability of the ratings. We utilized GPT-4 to implement a rating pre- dictor. To enhance predictive performance, we designed prompts that apply chain-of-thought prompting ( Wei et al.,2022;Wang et al.,2023; Kojima et al.,2022), that feeds large language models not only examples of question-and-answer pairs but also examples of the thought processes leading to those answers. In this study, we col- lected five sets of product titles, review texts, and ratings from Amazon.com to create output exem- plars, each corresponding to ratings from 1 to 5. Subsequently, for each set, we wrote descriptions of the reasoning paths leading to the prediction of the rating from the product title and review text. We used these as few-shot exemplars within the context. Similarly, for target reviews, GPT-4 is en- couraged to generate a reasoning path and an an- swer. 4 Experiments We aim to facilitate the review-writing process for reviewers and provide helpful reviews to read- ers. To investigate the practicality of our method, we conducted evaluations from the perspectives of system users (Section 4.1) and review readers (Section 4.2). 4.1 Participant Evaluation To evaluate our method, we collected feedback through interviews, generated reviews, and ques- tionnaires. Data collection was conducted through Amazon Mechanical Turk (MTurk) 1 . 1 https://w.mturk.com Figure 3: Participant responses to âIn the past year, how often have you posted reviews?â 4.1.1 Experimental Setup We tuned the temperature for each system. For the interview dialogue system, the temperature was set to 0.2. The review text generator and rating predictor generate outputs that are faithful to the input. Therefore, we set the temperature to 0 for these systems to suppress the diversity of the gen- erated text. 4.1.2 Baseline System To demonstrate the effectiveness of using inter- view dialogue systems that adapt questions based on the context, we constructed a baseline system. The baseline system replaces the interview dia- logue system with one that asks manually created questions in a fixed order. To construct the base- line system, we manually created nine questions on topics such as the reason for purchasing the product and the evaluation of the product in com- parison with other products. All questions asked by the baseline system are listed in Appendix B. We collected data using this system in the same manner as with our proposed system. 4.1.3 Evaluation Procedure Initially, participants conducted an interview di- alogue with our interview dialogue system. Af- ter the interview, they were presented with the generated reviews and ratings. Participants then completed a post-interview survey comprising multiple-choice and open-ended questions. For each setting, we recruited 100 participants located in AU, CA, NZ, GB, or the US and had a 95% ap- proval rate with at least 500 previously approved HITs. 4.1.4 Post-Interview Survey After the interview, participants responded to a post-interview survey. Several questions in this Table 1: Likert Items in Post-interview Survey DimensionLabels in Figure4Statements Interview EnjoyableHow fun was your interaction with the chatbot interviewer? SkillfulThe interviewer skillfully elicited your impressions or opinions. In-depthThe chatbot interviewer attempted to elicit your impressions or opinions in depth. Review FaithfulThe system-generated review faithfully reflects what you said during your interviews. ConciseThe system-generated review offers a concise summary of the points you mentioned during the interview. System QualityPlease rate the overall quality of the system. Burdened(I)I felt burdened to have an interview chat about the product. Burdened(R)Writing a review with the support of the system is more burdensome than writing a review yourself. Figure 4: Participant responses to questions on a Likert scale from 1 (Strongly disagree) to 5 (Strongly agree) in a post-interview survey. For each question, the upper bar shows the results from our system and the lower bar shows the baseline results. survey were answered using a 5-point Likert scale. These questions are related to the interview dia- logue, the generated reviews, and the overall sys- tem (See Table 1). We also asked participants how frequently they post reviews to compare with their usual review- writing experiences. As shown in Figure 3, 95% of the participants posted at least one user review in the past year. Additionally, participants were asked to rate the product they selected by respond- ing to the question: "If you were to rate this prod- uct again, what rating would you give it?" and pro- vided a rating from 1 to 5. 4.1.5 Participant Feedback Analysis Figure4illustrates the distribution of responses to eight questions 2 . Regarding the dimensions of 2 For Burdened(R), we excluded responses from partici- pants who selected the âHave never written any reviews be- fore in my lifeâ option to the question in Figure3. the interview and review, most participants evalu- ated two components positively: our interview dia- logue system and our review text generator. Partic- ipants showed a similar positive trend across two settings for the four items: In-depth, Faithful, Con- cise, and Quality. Notably, for Quality, 90% or more of the participants rated the overall quality of the systems positively. Our system provided users with more enjoy- able interviews and higher satisfaction regarding the generated reviews compared with the base- line system. As shown in Figure4, when us- ing our system based on GPT-4, more participants agreed that interacting with the system was fun. Moreover, the difference in the methods used to elicit informationour interview dialogue system and the baselineimpacts usersâ enjoyment, with statistically significant differences (MannWhitney U test,p<0.05). Participants also responded to the multiple-choice question, âIf you had to edit and post a system-generated review to your satisfaction, how much of it would you need to rewrite?â. Figure 5shows that different types of systems resulted in varied response distributions. In particular, 38% and 27% of participants using the baseline system and our system, respectively, responded that they needed to rewrite more than 50% of the review. These results indicate that our system can provide reviews with higher satisfac- tion than the baseline. Our system imposed a greater burden on par- ticipants. Figure 4shows that a higher percent- age of participants agreed thatwriting a review with the support of our system is more burden- some than writing alone, compared to the baseline. We argue that the response time of the system is one of the reasons for this difference. Our GPT- 4-based system, which generates responses based on usersâ utterances, takes a longer time to gener- Figure 5: Participant responses to âIf you had to edit and post a system-generated review to your satisfaction, how much of it would you need to rewrite?â ate responses than a baseline that asks predefined questions. Notably, several participants suggested that the response speed of our system should be improved. In response to the free-form question âWhat is one enhancement that can be made to im- prove this system?â, we received answers such as âmore fast repliesâ and âneed quick reply.â In our experiments, unlike the ChatGPT interface 3 , we did not employ real-time response generation us- ing streaming functionality. Adding this feature would be an effective modification to enhance our systemâs response speed, which is expected to sig- nificantly improve user experience. 4.1.6 Case Study Our interview dialogue system can generate follow-up questions that explore the content of usersâ ambiguous responses in depth. Table 2 shows an example of the data collected, com- prising the dialogue history regarding an electric shaver and the corresponding review text gener- ated. During the interview, our system initially asked about the participantâs overall satisfaction with the product, to which the participant replied, â... well satisfied but with few minor issues.â Based on this response, our system posed follow- up questions to clarify the aspects that the partic- ipant was satisfied with and the issues they men- tioned. As demonstrated in this example, our sys- tem can elicit deeper information about products from users compared with the baseline system. Additionally, our interview dialogue system can change topics during an interview to collect infor- mation on various aspects of a product. For the in- terview dialogues shown in Table 2, the first three turns focus on the participantâs satisfaction with the product. Subsequently, the system changes the 3 https://chatgpt.com/ topic by saying âNow, letâs go back to the issueâ and thus shifting focus to the issues that the user mentioned in the first turn. In the next turn, our system shifts the topic again to highlight the prod- uctâs impressive features. By switching topics in this manner, our system can acquire information on various aspects of the product. Our system can also generate review texts and ratings that faithfully reflect the content mentioned in the interview. As shown in the lower part of Ta- ble 2, the generated review text closely aligns with user comments such as âi would say well satisfied ...â and âsome times small hair from the beard gets stucks.â Moreover, our system accounts for issues mentioned by the user and predicts a rating of 4, one less than the maximum. 4.2 Third-Party Evaluation Helpful user reviews offer effective decision- making support for potential buyers and sellers. To assess the quality of the reviews generated by our system from the readerâs perspective, we con- ducted multidimensional evaluations. 4.2.1 Experimental Procedure This evaluation was conducted in a pairwise set- ting. In this experiment, annotators were pre- sented with two reviews: one generated by our system and another for comparison. We com- pared our system-generated reviews with Amazon reviews to evaluate the quality difference from human-written reviews. Additionally, to assess the differences from the perspective of the readers, we compared our system-generated reviews with those generated by a baseline system, which were collected in Section 4.1. Annotators chose either the superior option or a tie for each of the eight dimensions, as shown in Table3. We recruited the annotators through MTurk. 4.2.2 Selection Criteria for Comparative Reviews We collected human-written reviews from the Amazon Review Dataset (Ni et al.,2019). To fa- cilitate comparison for the annotators, we initially removed any human-written reviews where the cat- egory or rating did not match those generated by our system. After this filtering, to avoid selecting poor-quality reviews, we retained reviews in the top 5% of human-written reviews based on helpful- ness votes, using these as comparison candidates. Table 2: Example of a dialogue between our system and the user (top, S: system, U: user) and generated review text (bottom). In this example, our rating predictor predicted a rating of 4. S (...) To start off, could you tell me about your overall satisfaction with this product? Ui would say well satisfiedbut with few minor issues. S I see, youâre generally satisfied but there are a few minor issues. Could you elaborate on what you particularly like about the Braun Series 9 9370c Electric Shaver? Usome times small hair from the beard gets stucksin particular spot on the blade area in the machine. S I understand that small hairs getting stuck in the blade area is one of the issues youâve encountered. But before we delve into that, could you first share what you find satisfying or impressive about the product? What are its positive points in your opinion? U the postivie sides is smooth shaving experience and battery capacity which is good S Itâs great to hear that you appreciate the smooth shaving experience and the battery capacity. Now, letâs go back to the issue you mentioned earlier about small hairs getting stuck. Could you tell me more about this problem? How often does it happen and how does it affect your shaving experience? U i wouls say not often, ifaced this issue only once. S I see, (...) Thatâs good to know. Now, aside from the smooth shaving experience and the battery capacity, are there any other features of the Braun Series 9 9370c Electric Shaver that you find beneficial or impressive? ... I am generally satisfiedwith the Braun Series 9 9370c Rechargeable Wet & Dry Menâs Electric Shaver with Clean & Charge Station. The product offers a smooth shaving experience, which I find impressive. The battery capacity is also commendable, lasting for a good amount of time. I appreciate the versatility of the product, as it can be used for both wet and dry shaving. The Clean & Charge Station that comes with the shaver is quite efficient and useful. It effectively cleans and charges the shaver, making it more convenient compared to other products Iâve used before.However, I did encounter a minor issue where small hairs from my beard got stuckin a particular spot on the blade area. This happened only once and did not significantly affect my overall shaving experience. Considering all the features and my experience with the product, I believe it offers good value for its price. I would recommend the Braun Series 9 9370c Electric Shaver to others, mainly for its battery performance and its variety of features, including its wet and dry usage. Finally, among these candidates, we selected the review of the product whose title achieved the highest ROUGE-L score (Lin,2004). We applied the same process when matching re- views generated by the baseline system; however, for some reviews collected using our system, there were no corresponding reviews in the baseline data that matched both category and rating. There- fore, we excluded these unmatched data. Conse- quently, while we conducted comparisons for all 100 reviews written by humans, only 96 baseline- generated reviews met the criteria. 4.2.3 Results and Discussions The overall results are presented in Table 4. The annotators prefer the reviews generated by our sys- tem to those written by humans or generated by the baseline system. Notably, the reviews gener- ated by our system are helpful, provide a balanced view of pros and cons, and offer comprehensive information. These findings indicate that our inter- view dialogue system is capable of eliciting a wide range of information about products from users through topic transitions. The reviews generated by our system lack the fluency of human-written reviews. For instance, our review text generator tends to use the formal product title when referring to the product. Addi- tionally, human-written reviews contain more in- dividual experiences compared with those gener- ated by our system. Despite these limitations, our system has high scalability, offering the potential for improvement. Specically, our systemâs output could be enhanced by rening the prompts to gen- erate texts that are more human-like and elicit de- tailed usage experiences from users. By replacing the baseline system, which uses fixed questions, with our interview dialogue sys- tem, we observe improvements across all met- rics. Notably, our system can generate reviews that are rich in experience-based information, contain more detailed information, and cover a broader range of topics. This demonstrates that our system can elicit more detailed and extensive information from users through follow-up questions and topic transitions. 4.3 Discussion on Predicted Ratings To further explore the characteristics of the re- views and ratings generated by our system, we an- alyze them along two axes: the difference based on the source of the ratings (comparing ratings as- signed by humans to those predicted by our sys- tem) and the difference based on the annotators (comparing the ratings given by system users to those assigned by third parties). To obtain ratings Table 3: Questions in comparative evaluation Labels in Table4Questions HelpfulnessWhich review would be more helpful for making a purchase decision? FluencyWhich review exhibits a more fluent and human-like writing style? ConcisenessWhich review is more concise and to the point? ExperienceWhich review provides more information based on the actual usage experience of the product? BalanceWhich review presents a more balanced view of the productâs pros and cons? DepthWhich review provides more in-depth information about any specific aspect of the product? CoverageWhich review mentions a more comprehensive range of product aspects? OverallWhich review is overall more preferable? Table 4: Results of third-party evaluation. The values represent the percentage of votes each received. ReviewsHelpfulness Fluency Conciseness Experience BalanceDepthCoverageOverall Human38.047.037.057.037.043.040.041.0 Tie6.015.06.09.015.010.05.07.0 Ours56.038.057.034.048.047.055.052.0 Baseline38.5 28.1 45.8 21.9 34.4 35.4 35.4 37.5 Tie12.533.36.216.715.612.510.417.7 Ours49.038.547.961.550.052.154.244.8 Table 5: Average absolute difference in ratings be- tween Amazon customers and Turkers (top-left), be- tween system-predicted ratings and Turkersâ ratings for system-generated reviews (top-right), and be- tween system-predicted ratings and participantsâ rat- ings (bottom-right, see Section4.1.4). Annotator/SourceHuman- written System- generated Turkers0.590.12 Participants in Section4.1-0.57 assigned by third parties, we newly recruited an- notators from MTurk and asked them to assign rat- ings to both the human-written reviews (left col- umn) 4 and those generated by our system (right column). We also collected ratings assigned by participants from the experiments in Section4.1. Note that these participants, unlike the Turkers, had seen the ratings predicted by our system. The results in the top row of Table5demon- strate that the difference between the ratings pre- dicted by our system and those assigned by third parties is remarkably smaller than the difference found in human-written reviews. This finding in- dicates that the sentiment of the reviews generated by our system is easily comprehensible to readers. The ratings predicted by our system, as shown in the right column of Table 5, align more closely with those assigned by third-party annotators than with those of system users. This finding indicates that our system emphasizes objectivity over sub- 4 For the annotations, we used 100 human-written reviews selected in Section4.1. jectivity in its ratings. The aforementioned observations indicate that our system generates review texts that are easy for humans to understand and provide more objective ratings. This finding suggests that our interview di- alogue system and review text generator can gener- ate reviews that accurately capture reviewersâ sen- timents, thereby supporting informed purchasing decisions, while the rating predictor also provides highly objective and reliable ratings. 5 Conclusion In this study, we present a novel method for utiliz- ing dialogue systems to facilitate user review cre- ation. Our approach involves three processes: in- terview dialogue, review text generation, and rat- ing prediction. Although ensuring the fluency of the system-generated reviews remains a challenge, our method provides high-quality and helpful re- views for both reviewers and their readers. Our method possesses high scalability. For in- stance, feeding product descriptions into our inter- view dialogue system could lead to deeper inter- view dialogues about more detailed information. However, our experiments have shown that even without such extensions, our system is capable of providing reviews that are more helpful than human-written ones. Furthermore, adapting our di- alogue systemâs strategies to user preferences dur- ing review writing could improve user experience. Further research can accomplish this objective by conducting a more detailed analysis of user prefer- ences. References Advait Bhat, Saaket Agashe, Parth Oberoi, Niharika Mohile, Ravi Jangir, and Anirudha Joshi. 2023.In- teracting with next-phrase suggestions: How sug- gestion systems aid and influence the cognitive pro- cesses of writing. InProceedings of the 28th Inter- national Conference on Intelligent User Interfaces, IUI â23, page 436â Ě A ̧S452, New York, NY, USA. As- sociation for Computing Machinery. Arthur BraĹžinskas, Mirella Lapata, and Ivan Titov. 2020.Unsupervised opinion summarization as copycat-review generation. InProceedings of the 58th Annual Meeting of the Association for Compu- tational Linguistics, pages 5151â5169, Online. As- sociation for Computational Linguistics. Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner. 2006. The ami meeting corpus: A pre-announcement. InMachine Learning for Multimodal Interaction, pages 28â39, Berlin, Heidelberg. Springer Berlin Heidelberg. Mingda Chen, Zewei Chu, Sam Wiseman, and Kevin Gimpel. 2022.SummScreen: A dataset for ab- stractive screenplay summarization. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8602â8615, Dublin, Ireland. Association for Computational Linguistics. Yulong Chen, Yang Liu, and Yue Zhang. 2021.Dialog- Sum challenge: Summarizing real-life scenario di- alogues. InProceedings of the 14th International Conference on Natural Language Generation, pages 308â313, Aberdeen, Scotland, UK. Association for Computational Linguistics. Li Dong, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou, and Ke Xu. 2017.Learning to generate product reviews from attributes. InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pages 623â632, Valencia, Spain. As- sociation for Computational Linguistics. Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019.SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization . InProceedings of the 2nd Workshop on New Frontiers in Summarization, pages 70â79, Hong Kong, China. Association for Computational Linguistics. Jihyeok Kim, Seungtaek Choi, Reinald Kim Amplayo, and Seung-won Hwang. 2020.Retrieval-augmented controllable review generation . InProceedings of the 28th International Conference on Compu- tational Linguistics, pages 2284â2295, Barcelona, Spain (Online). International Committee on Compu- tational Linguistics. Soomin Kim, Joonhwan Lee, and Gahgene Gweon. 2019.Comparing data from chatbot and web sur- veys: Effects of platform and conversational style on survey response quality. InProceedings of the 2019 CHI Conference on Human Factors in Com- puting Systems, CHI â19, page 1â Ě A ̧S12, New York, NY, USA. Association for Computing Machinery. Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022.Large language models are zero-shot reasoners. InAd- vances in Neural Information Processing Systems, volume 35, pages 22199â22213. Curran Associates, Inc. Junyi Li, Wayne Xin Zhao, Ji-Rong Wen, and Yang Song. 2019.Generating long and informative re- views with aspect-aware coarse-to-fine decoding. In Proceedings of the 57th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 1969â 1979, Florence, Italy. Association for Computational Linguistics. Pan Li and Alexander Tuzhilin. 2019.Towards con- trollable and personalized review generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Lan- guage Processing (EMNLP-IJCNLP), pages 3237â 3245, Hong Kong, China. Association for Computa- tional Linguistics. Yu Li, Baolin Peng, Pengcheng He, Michel Galley, Zhou Yu, and Jianfeng Gao. 2023.DIONYSUS: A pre-trained model for low-resource dialogue sum- marization. InProceedings of the 61st Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 1368â 1386, Toronto, Canada. Association for Computa- tional Linguistics. Chin-Yew Lin. 2004.ROUGE: A package for auto- matic evaluation of summaries. InText Summariza- tion Branches Out, pages 74â81, Barcelona, Spain. Association for Computational Linguistics. Bodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni, and Julian McAuley. 2020. Interview: Large-scale modeling of media dialog with discourse patterns and knowledge grounding . InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8129â8141, Online. Association for Computational Linguistics. Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects . InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Interna- tional Joint Conference on Natural Language Pro- cessing (EMNLP-IJCNLP), pages 188â197, Hong Kong, China. Association for Computational Lin- guistics. Jianmo Ni and Julian McAuley. 2018.Personalized re- view generation by expanding phrases and attend- ing on aspect-aware representations. InProceed- ings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Pa- pers), pages 706â711, Melbourne, Australia. Asso- ciation for Computational Linguistics. Taro Okahisa, Ribeka Tanaka, Takashi Kodama, Yin Jou Huang, and Sadao Kurohashi. 2022. Con- structing a culinary interview dialogue corpus with video conferencing tool. InProceedings of the Thir- teenth Language Resources and Evaluation Confer- ence, pages 3131â3139, Marseille, France. Euro- pean Language Resources Association. Zhaopeng Qiu, Xian Wu, Jingyue Gao, and Wei Fan. 2021.U-bert: Pre-training user representations for improved recommendation. volume 35, pages 4320â 4327. Vasu Sharma, Harsh Sharma, Ankita Bishnu, and Lab- hesh Patel. 2018.Cyclegen: Cyclic consistency based product review generator from attributes. In Proceedings of the 11th International Conference on Natural Language Generation, pages 426â430, Tilburg University, The Netherlands. Association for Computational Linguistics. Hanfei Sun, Ziyuan Cao, and Diyi Yang. 2022. SPORTSINTERVIEW: A large-scale sports inter- view benchmark for entity-centric dialogues. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 5821â5828, Mar- seille, France. European Language Resources Asso- ciation. Quoc-Tuan Truong and Hady Lauw. 2019.Multimodal review generation for recommender systems. In The World Wide Web Conference, W â19, page 1864â Ě A ̧S1874, New York, NY, USA. Association for Computing Machinery. Xuan-Son Vu, Thanh-Son Nguyen, Duc-Trong Le, and Lili Jiang. 2020.Multimodal review generation with privacy and fairness awareness. InProceed- ings of the 28th International Conference on Com- putational Linguistics, pages 414â425, Barcelona, Spain (Online). International Committee on Compu- tational Linguistics. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowd- hery, and Denny Zhou. 2023. Self-consistency im- proves chain of thought reasoning in language mod- els . InThe Eleventh International Conference on Learning Representations. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompt- ing elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824â24837. Curran Associates, Inc. Xinyuan Zhang, Ruiyi Zhang, Manzil Zaheer, and Amr Ahmed. 2021.Unsupervised abstractive dialogue summarization for tete-a-tetes. InProceedings of the AAAI Conference on Artificial Intelligence, vol- ume 35, pages 14489â14497. Chao Zhao, Spandana Gella, Seokhwan Kim, Di Jin, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu, and Dilek Hakkani-Tur. 2023.âwhat do others think?â: Task-oriented conversational modeling with subjective knowledge. InProceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 309â323, Prague, Czechia. Association for Computational Linguistics. Ming Zhong, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2022. Dialoglm: Pre-trained model for long dialogue understanding and summa- rization. InProceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11765â 11773. Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev. 2021. QMSum: A new benchmark for query- based multi-domain meeting summarization. InPro- ceedings of the 2021 Conference of the North Amer- ican Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5905â5921, Online. Association for Computational Linguistics. Chenguang Zhu, Yang Liu, Jie Mei, and Michael Zeng. 2021. MediaSum: A large-scale media interview dataset for dialogue summarization. InProceedings of the 2021 Conference of the North American Chap- ter of the Association for Computational Linguistics: Human Language Technologies, pages 5927â5934, Online. Association for Computational Linguistics. Yicheng Zou, Bolin Zhu, Xingwu Hu, Tao Gui, and Qi Zhang. 2021.Low-resource dialogue summariza- tion with domain-agnostic multi-source pretraining. InProceedings of the 2021 Conference on Empiri- cal Methods in Natural Language Processing, pages 80â91, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. A Prompt Template A.1 Prompt for Interview Dialogue Table6shows a prompt template for interview di- alogues. [PRODUCT_NAME] is a placeholder for the product title, which will be replaced with the product title selected by the participant. [MAX_QUESTION] and [MIN_QUESTION] are placeholders for the maximum and minimum num- ber of dialogue turns. In our experiments, we used 15 and 8, respectively. A.2 Prompt for Review Generation Table7shows a prompt template for review gen- eration. Similar to that for interviewing, [PROD- UCT_NAME] is a placeholder for the product ti- tle. [DIALOGUE] is a placeholder for the dia- logue history, into which the interview dialogue history between our system and the participants is inserted. B Baseline Details Table8shows a prompt template for review gen- eration. Similar to that for interviewing, [PROD- UCT_NAME] is a placeholder for the product ti- tle. [DIALOGUE] is a placeholder for the dia- logue history, into which the interview dialogue history between our system and the participants is inserted. Table 6: Prompt template for interviewing. Your role is âinterviewerâ and my role is âintervieweeâ. About the product I am going to present, please elicit my impressions and opinions from me when I have touched it. Note the following statements. 1. The interviewer elicits the intervieweeâs satisfaction and dissatisfaction (the positive and negative points) with the product in a well-balanced and detailed. 2. In response to the intervieweeâs response, the interviewer asks more in-depth questions about the aspect or elicits feedback about other aspects of the product. 3. Be sure to attach the name of your role at the beginning of your utterance. Since your role is âinterviewerâ, your generation should begin with âInterviewer:â. 4. Donât generate intervieweeâs utterances. 5. Add â[Wait_for_Response]â at the end of your utterance and wait for my response. 6. You must ask at least [MIN_QUESTION] questions. In other words, the dialogue must continue for [MIN_QUESTION] or more turns. 7. Having fulfilled the 6th statement, you can terminate the interview at your discretion. However, the interview must be completed within [MAX_QUESTION] turns. 8. When you terminate the intervew, add â[End_of_Interview]â at the end of your utterance. Now, please elicit my impressions and opinions about the following product from me. [PRODUCT_NAME] Table 7: Prompt template for review generation. [DIALOGUE] The above is a dialogue about â[PRODUCT_NAME]â between the interviewer and the interviewee who has touched on this product. Write a customer review about the product as if written by the interviewee, by briefly summarizing the important information mentioned in the above interview, such as the good and bad points of the product and the intervieweeâs experience with it. Do not output the reviewâs title. The following is a body of the product review of the product written by the interviewee: Table 8: Questions asked by the baseline system Q-1 First, could you tell me about the features and functions of this product? What kind of product is this? Q-2 What made you decide to purchase this product? Q-3 If you have any points that you like or are satisfied with this product, please tell me in detail. Q-4 What are the advantages of this product compared to other products? Q-5 If you have any dissatisfaction with this product or areas for improvement for this product, please tell me in detail. Q-6 What are the disadvantages of this product compared to other products? Q-7 Who would this product be suitable for? Q-8 Is this product worth the price? Also, why do you think so? Q-9 Finally, do you have any requests or impressions about the product?