Paper deep dive
What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews
Jonathan Ivey, Anjalie Field, Ziang Xiao
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/10/2026, 3:03:21 AM
Summary
This paper presents an empirical analysis of qualitative interview response quality by introducing the 'Qualitative Interview Corpus' (343 transcripts, 16,940 responses). The authors evaluate 10 proposed quality metrics, finding that 'research question relevance' is the strongest predictor of response quality, while common NLP metrics like 'clarity' and 'surprisal-based informativeness' are not predictive.
Entities (5)
Relation Signals (3)
Qualitative Interview Corpus â contains â 16,940 participant responses
confidence 100% ¡ a newly constructed dataset of 343 interview transcripts with 16,940 participant responses
Research Question Relevance â predicts â Response Quality
confidence 95% ¡ direct relevance to a key research question is the strongest predictor of response quality
Clarity â doesnotpredict â Response Quality
confidence 90% ¡ two measures commonly used to evaluate NLP interview systems, clarity and surprisal-based informativeness, are not predictive of response quality
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Qualitative interviews provide essential insights into human experiences when they elicit high-quality responses. While qualitative and NLP researchers have proposed various measures of interview quality, these measures lack validation that high-scoring responses actually contribute to the study's goals. In this work, we identify, implement, and evaluate 10 proposed measures of interview response quality to determine which are actually predictive of a response's contribution to the study findings. To conduct our analysis, we introduce the Qualitative Interview Corpus, a newly constructed dataset of 343 interview transcripts with 16,940 participant responses from 14 real research projects. We find that direct relevance to a key research question is the strongest predictor of response quality. We additionally find that two measures commonly used to evaluate NLP interview systems, clarity and surprisal-based informativeness, are not predictive of response quality. Our work provides analytic insights and grounded, scalable metrics to inform the design of qualitative studies and the evaluation of automated interview systems.
Tags
Links
- Source: https://arxiv.org/abs/2604.05163v1
- Canonical: https://arxiv.org/abs/2604.05163v1
Trouble viewing inline? Open PDF directly â
Full Text
66,287 characters extracted from source content.
Expand or collapse full text
What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews Jonathan Ivey Johns Hopkins University jivey6@jhu.edu Anjalie Field Johns Hopkins University anjalief@jhu.edu Ziang Xiao Johns Hopkins University ziang.xiao@jhu.edu Abstract Qualitative interviews provide essential in- sights into human experiences when they elicit high-quality responses. While qualitative and NLP researchers have proposed various mea- sures of interview quality, these measures lack validation that high-scoring responses actually contribute to the studyâs goals. In this work, we identify, implement, and evaluate 10 pro- posed measures of interview response quality to determine which are actually predictive of a responseâs contribution to the study findings. To conduct our analysis, we introduce the Qual- itative Interview Corpus, a newly constructed dataset of 343 interview transcripts with 16,940 participant responses from 14 real research projects. We find that direct relevance to a key research question is the strongest predictor of response quality. We additionally find that two measures commonly used to evaluate NLP interview systems, clarity and surprisal-based informativeness, are not predictive of response quality. Our work provides analytic insights and grounded, scalable metrics to inform the design of qualitative studies and the evaluation of automated interview systems. 1 Introduction Qualitative interviews are a primary method for sur- facing insights into experiences, motivations, and behaviors that quantitative methods cannot capture. However, the value of what insights an interview produces depends directly on the quality of the re- sponses it elicits, and our understanding of what makes a response high-quality rests almost entirely on theoretical intuition. Qualitative researchers have proposed characteristics of high-quality inter- view responses, such as spontaneity and relevance (Kvale and Brinkmann, 2009; Charmaz, 2014; Pat- ton, 2015; Small and Calarco, 2022), but these frameworks disagree substantially on which charac- teristics matter, and none offer empirical evidence that responses with these characteristics actually contribute to a studyâs findings. Such evidence is necessary for determining which measures should guide interview practices. Recent interest in AI has accelerated the need to understand and quantify interview data qual- ity. NLP systems are increasingly being used to conduct or assist human interviews. For exam- ple, Anthropic recently deployed a system to au- tonomously collect qualitative responses to investi- gate how professionals use AI (Handa et al., 2025). Other applications include academic research (Liu et al., 2025), market research (Anugraha et al., 2026), preference elicitation (Choudhury et al., 2025), and gathering public feedback (Jiang et al., 2023). Current interview systems commonly use proxy criteria for judging elicited response quality like specificity, clarity, and relevance (Xiao et al., 2020a,b; Jiang et al., 2023; Hu et al., 2024; Jacob- sen et al., 2025), but these measures similarly lack validation that high-scoring responses contribute to study findings. Without validated evaluation metrics, building and evaluating AI systems for qualitative research remains untenable. In this work, we investigate characteristics of in- terview response quality through the identification and implementation of proposed quality metrics and empirical analysis of a new dataset. First, we identify 10 measures of interview response qual- ity through a review of qualitative literature and research studies on NLP interview systems. We empirically assess these 10 measures over a newly constructed dataset of 343 transcripts from 14 real qualitative research studies. Our analysis of 16,940 participant responses reveals which measures are actually predictive of a responseâs contribution to the study findings, our criterion for overall quality. From our analysis, we find that the measure most predictive of response quality is relevance to a key research question of the study. We also find that responses containing the kind of insights unique to qualitative studies are more likely to be high 1 arXiv:2604.05163v1 [cs.CL] 6 Apr 2026 quality, for example, responses that explain why a belief or experience matters personally to the participant. Finally, we find that two measures commonly used to evaluate interviewer systems, clarity and surprisal-based informativeness, are not significantly predictive of response quality. As the end goal of quality measures is to inform interview strategies, we further use our measures to conduct a case study of how time and interview techniques affect response quality. Our contributions in this work include (1) the Qualitative Interview Corpus, 1 a newly constructed dataset of 343 transcripts from 14 qualitative re- search projects that enables empirical analysis of qualitative interviews, (2) the creation and vali- dation of automated measures of qualitative in- terview characteristics, (3) empirical analysis of which characteristics of participant responses are predictive of overall response quality, and (4) an example use case of how these measures can in- form interview strategies. Our work offers the first empirical analysis of response characteristics in qualitative interviews, offering grounded metrics that can inform both the design of qualitative stud- ies and the evaluation of NLP interview systems. 2 2 Methods To investigate which response characteristics are indicative of their contribution to the study find- ings, we identify 10 proposed quality measures from qualitative research literature and evaluations of interview systems. We then design a quality criterion based on the extent to which a response contributes to the study findings. Finally, we create an automated measure of these response character- istics and our quality criterion to enable large-scale empirical analysis. 2.1 Proposed Characteristics of High-Quality Responses We review qualitative literature and evaluations of NLP interview systems to identify the character- istics of participant responses that are commonly used as quality metrics. In qualitative literature, Kvale and Brinkmann (2009) propose the most robust set of measures, including richness, relevance to the research ques- tion, spontaneity, self-reportedness, and the ratio 1 Dataset to be released athttps://doi.org/10.5064/ F6JWVCH6 2 Fullcodeavailableathttps://github.com/ jonathanivey/interview-quality. of the length of the participant utterance to the length of the interviewer utterance. Patton (2015) further identifies relevance to the research ques- tion and relevance to the exact question asked by the interviewer as key aspects of interview quality. Charmaz (2014) indicates that quality data will be ârich, substantial, and relevant.â Small and Calarco (2022) propose an alternative view of qualitative re- search quality based on five key constructs. Two of these constructs, cognitive empathy and palpability, are characteristics of participant responses. Note that we refer to cognitive empathy as "attributed meaning" to better align with its definition and dis- tinguish it from other characteristics. In NLP interview systems, the most common measures for response quality are based on Gricean maxims (Grice, 1975). This approach identifies quality responses as those with specificity, clar- ity, relevance, and surprisal-based informativeness (Xiao et al., 2020a,b; Jiang et al., 2023; Hu et al., 2024; Jacobsen et al., 2025). These characteristics are often considered alongside measures of user engagement such as response length. Other work has combined these with the above measures from qualitative literature (Cuevas et al., 2025). To ensure that our final set of characteristics is sufficiently distinct, we identify definitions from each of the original sources and merge character- istics with exceptionally similar definitions, such as specificity and palpability. We choose not to include an explicit measure of richness because, based on the existing definitions, we consider rich- ness to be a combination of other characteristics such as specificity, self-reportedness, and attributed meaning. Finally, to reduce multicollinearity be- tween suprisal-based informativeness and response length, we instead use the average word-level sur- prisal rather than the total word-level surprisal. The final set of characteristics, definitions, and their ori- gins is outlined in Table 1. 2.2 Our Criterion for Response Quality To identify high-quality responses, we develop a criterion based on the extent to which a response contributes to the results of a study. Unlike the previously identified characteristics, our criterion is grounded in research outcomes (i.e., the final re- sults) and cannot be measured during the data col- lection phase. However, it can be used to compare and validate the other characteristics that can be measured from responses alone, as demonstrated in §4.2. Our criterion uses the following scoring 2 CharacteristicDefinitionSource Specificity (Palpability) The extent to which a response provides detailed exam- ples rather than abstract generalizations. Kvale and Brinkmann (2009); Charmaz (2014); Xiao et al. (2020b); Small and Calarco (2022) ClarityHow clear and understandable a response is.Kvale and Brinkmann (2009); Xiao et al. (2020b) Immediate RelevanceHow relevant the response is to the specific question asked by the interviewer. Patton (2015); Xiao et al. (2020b) Research Question Relevance How relevant the response is to the overall research question. Kvale and Brinkmann (2009); Charmaz (2014); Patton (2015) SpontaneityThe extent to which the response provides information beyond what is provided in the question. Kvale and Brinkmann (2009) Self-reportednessHow understandable a response is if taken out of con- text. Kvale and Brinkmann (2009) Attributed Meaning (Cognitive Empathy) The extent to which a response demonstrates the per- sonal significance of a belief or experience to the par- ticipant. Small and Calarco (2022); Charmaz (2014) Average SurprisalThe average word-level surprisal of the response.Xiao et al. (2020b) Response Length Ratio Ratio of the length of the participant response to the length of the interviewer question Kvale and Brinkmann (2009) Response LengthLength of the response.Xiao et al. (2020b) Table 1: We identify 10 key characteristics of interview responses from qualitative literature and NLP interview systems, shown here with their definitions and where they were proposed. rubric to estimate the likelihood that a response contributed to the goals of the study: 1. The response is unrelated or contradictory to the results section. 2.The response is tangentially related to the re- sults section with no specific substance. 3.The response aligns with the results section but is general or vague. 4. The response provides an example or senti- ment matching the results sectionâs conclu- sions. 5. The response appears in the results section or is a primary source for it. 2.3 Automatically Identifying Response Characteristics To conduct our analysis across a large dataset, we implement automated measures for the 10 response characteristics and our quality criterion. Three of the response characteristics can be computed di- rectly: we compute response length and response length ratio based on the number of tokens, and we compute the average word-level surprisal using Oh et al.âs (2024) implementation based on token counts from the Pile (Gao et al., 2020). Conceptual Measures The remaining 7 charac- teristics and our quality criterion require concep- tual judgments that we obtain using an LLM judge. For the quality criterion, we prompt the model to rate the participant response on a scale from 1 to 5 according to the rubric in §2.2. For the other measures, we create rubrics from the definitions in the original sources and use them to prompt the model to rate responses on a scale of 1 to 3. In addition to the prompt, we provide the mod- els with (1) the current interview excerpt that the model is rating, (2) the interview excerpt imme- diately preceding the current excerpt to provide conversational context, and (3) 1â2 sentences pro- viding broad context for how the interviews were conducted and the general goals of the project. For our quality criterion, we additionally provide the model with a segment of the results section of the paper. For each excerpte, letSbe the set of all segments in the results section of the correspond- ing study, and letq(e,s)represent the estimated likelihood that excerptecontributed to segment sâ S. We evaluate our quality criterion across all segments and take the maximum value to determine the final score QI(e): QI(e) = max sâS q(e,s) This single score measures the extent to which a participant response contributed to any of the 3 study results. We use the same process for research question relevance. LettingQbe the set of all key research questions andr(e,q)be the estimated relevance of excerpteto a single questionqâ Q, the final relevance score RQ(e) is calculated as: RQ(e) = max qâQ r(e,q) The full prompts used for our measures are pro- vided in Section A. Human Validation To validate whether LLM judges can estimate these conceptual measures, we compare their outputs to human judgments on 100 interview excerpts from 5 representative projects in our dataset. For each excerpt, we have three differ- ent annotators with experience analyzing qualita- tive interviews rate the 7 conceptual characteristics and quality criterion for the participant response, re- sulting in 2,400 total annotations. The 100 excerpts were selected from a random sample that was then balanced to have equal distributions of each rating for each characteristic. We provide annotators with the same information as the LLM with only minor formatting changes, like highlighting participant statements, to reduce cognitive load. An example of the annotation setup is provided in Section B.1. 3 Dataset To our knowledge, there is no openly available dataset for analyzing qualitative interviews across multiple domains. To enable empirical analysis of qualitative interviews, we introduce The Qualita- tive Interview Corpus: a dataset of 343 qualitative interviews and their corresponding papers from 14 research projects across a diverse set of domains (Table 2; see Section B.2 for more details). Data Curation We construct the corpus from deposits to The Qualitative Data Repository, 3 an archive for storing and sharing digital data col- lected through qualitative and multi-method re- search. We select data deposits from projects that conducted English qualitative interviews, provided anonymized transcripts, have openly accessible data, and have a corresponding research paper with the results of their study. We manually review the data and exclude projects that do not elicit open- ended participant responses (e.g., surveys that were conducted orally and then transcribed). Our final 3 https://qdr.syr.edu Research Project# Interviews Mindfulness for Firefighters and EMS Work- ers (Steinberg et al., 2024) 11 Drug Shortage Management (Shuman, 2021)16 GhanaianHealthcareWorkersDuring COVID-19 (Alvarez, 2025) 20 Socializing Policy Feedback (Micatka, 2025)30 Perspectives on Political Representation (Ruedin and Murahwa, 2025) 23 Nutrition Interventions in Rural Ethiopia (Mersha, 2025) 21 Marine Corps Education Project (Fosher, 2020) 32 Intergovernmental Coordination Mechanisms (Milman, 2023) 43 Models of Delivery for Online Spiritual Care (Bezabih and Smith, 2025) 21 Partnership between Kidney Disease Patients and Caregivers (Gazaway et al., 2024) 25 Shared Data for Learning Qualitative Data Analysis (Furlong et al., 2025) 9 Advance Care Planning in Hospice Organiza- tions (Harrison, 2021) 50 Food Retail and Service Workers during COVID-19 (Vignola et al., 2024) 23 High-performance school-age athletes at Aus- tralian schools (OâNeill, 2017) 19 Table 2: The Qualitative Interview Corpus is built from 14 research projects across a diverse set of domains. This table lists each project and the number of interviews that it contributed to the corpus. dataset contains 14 research projects with a total of 343 interviews. PreprocessingTo make the data suitable for com- putational analysis, we first extract 58,688 utter- ances from the PDFs of the interview transcripts. We use speaker tags from the transcripts to assign each utterance to either the participant (31,434 ut- terances) or the interviewer (27,254 utterances). We then manually extract results sections from each research paper. In the case of mixed-methods stud- ies, we limit our results to those that came from the qualitative interviews. We partition each results section into segments that represent the different findings from the paper. Using the interviews, research papers, and sup- plemental documents like data narratives and inter- view plans, we add two pieces of additional data. First, we write a 1â2 sentence summary that briefly explains the overarching goals and context of the 4 Conceptual Measure Human Agreement Human-LLM Agreement Attributed Meaning0.7500.868 Spontaneity0.7160.797 Specificity0.7320.789 Immediate Relevance0.6020.764 Response Quality0.7540.757 Research Question Relevance 0.7140.690 Self-reportedness0.7540.679 Clarity0.5980.606 Table 3: From 2,400 human judgments of our concep- tual measures, we find strong agreement between human ratings (Human) and between the median human ratings and LLM judge ratings (Human-LLM) as measured with Krippendorffâs alpha. project. Then, we identify 3â5 key research ques- tions that the study was trying to answer. Because we use the summary and research questions to ana- lyze the characteristics of responses, as described in §2.3, we ensure that they do not contain infor- mation from the final results of the paper. Instead we align them with the initial goals of the project, as described in the interviews, research paper, and supplemental documents. Excerpts Because qualitative interviews are di- alogues, they often contain overlapping speech. For example, an interviewer may say, âmâ or âyesâ in the middle of a participant response to encourage them to continue speaking. To differen- tiate between a continuing participant response and a new participant response, we combine utterances into sets of excerpts. The first excerpt for each in- terview begins with the beginning of the transcript. Then, new excerpts are determined based on when the interviewer says more than four words. We con- struct 16,940 excerpts, where each excerpt begins with an interviewer utterance (most commonly a question) and contains a full participant response, occasionally interrupted by short interviewer utter- ances. 4 Results 4.1 Can Automated Measures Capture Response Characteristics? Annotator Agreement To validate whether our LLM judges can accurately estimate the conceptual measures, we compare their outputs to 2,400 hu- man judgments over 100 interview excerpts. First, we compare Krippendorffâs alpha between human annotators. Then, we take the median of the human labels for each excerpt and compare it to the LLM judge label using Krippendorffâs alpha. Findings Across the conceptual measures, we find strong agreement between human ratings and equally strong agreement between the median hu- man ratings and the LLM judge ratings (Table 3). These results show that we can automatically mea- sure interview response characteristics and our quality criterion at scale using our LLM judge setup. This finding supports the validity of our findings in §4.2 and enables future applications of our measures, including evaluating interview sys- tems and informing qualitative methodology. 4.2 What Characteristics Are Predictive of Response Quality? To understand what makes a high-quality interview response, we evaluate which characteristics of par- ticipant responses are predictive of the responseâs contribution to the study findings, as measured with our response quality criterion. Mixed-Effects Model Because our data has a nested structure where multiple responses come from a single participant and multiple participants come from a single research project, we cannot assume independence between responses. To ac- count for this, we use a linear mixed-effects model where the outcome is our response quality criterion, the fixed effects are the response characteristics, and the random effects are the participant and the project that the response originates from. The full equation is provided in Section C.1. Our model has a marginalR 2 of0.506, indicating that50.6%of the variation in response quality is explained by the characteristics we identify. We additionally find low multicollinearity and variance inflation factors, which support the reliability and interpretability of our model (details in Section C.2). Findings We find that research question rel- evance, attributed meaning, spontaneity, speci- ficity, immediate relevance, response length, and self-reportedness are significantly predictive of re- sponse quality (Table 4). Of these characteristics, research question relevance has the strongest cor- relation with a standard coefficient more than 3 times larger than any other covariate. This finding suggests that the most important characteristic 5 CharacteristicStd. Coef.P-Value Research Question Relevance0.536<0.001 Attributed Meaning0.137<0.001 Specificity0.056<0.001 Response Length0.048<0.001 Immediate Relevance0.039<0.001 Spontaneity0.037<0.001 Self-reportedness0.0160.026 Response Length Ratio0.0020.059 Clarity0.0010.346 Average Surprisal-0.0090.352 Table 4: Using a linear mixed-effect model, we identify which characteristics of participant responses are most predictive of response quality. The model has marginal R 2 = 0.507 and conditional R 2 = 0.583. of high-quality responses is direct relation to a key research question of the study. The second strongest coefficient is for attributed meaning. Attributed meaning indicates that a re- sponse demonstrates significance or meaning to a participant. These characteristics represent a unique strength of qualitative methods that give re- searchers access to participantsâ lived experiences. Together, these two attributes demonstrate that re- sponses are most valuable when they contribute to the overarching goals of qualitative research: answering research questions through personal in- sights that quantitative methods cannot capture. Five other characteristics have statistically significant coefficients:specificity, response length, immediate relevance, spontaneity, and self- reportedness. Many of these have to do with the form of responses and flow of conversation, and their weaker correlations align with theory from qualitative literature that well-spoken participants may be easier to interview, but they are not guar- anteed to provide more useful answers (Kvale and Brinkmann, 2009). Notably, we do not find statistically significant correlations for clarity, average surprisal, and re- sponse length ratio. This finding contradicts cur- rent practice for NLP interview systems that frequently evaluate with measures of clarity and surprisal-based informativeness. Note that us- ing total surprisal instead of average surprisal and response length results in a standard coefficient of0.041without changing the marginalR 2 or the other standard coefficients. This indicates that the Follow-up Direct Question Indirect Question Interpreting Specifying Structuring Support & Rapport Intro & Context 0 25 50 75 100 Percentage G1G2G3G4G5 Response Quality Criterion 5.04.03.02.01.0 Figure 1: We compare the distributions of quality in responses elicited using different interview techniques. We group techniques into five distinct groups (G1âG5) that correspond to the theoretical function of the tech- niques in the group. Each member of a group has a statistically significant difference in median with the members of all other groups. surprisal measure itself does not provide predictive power beyond being a proxy for response length. 4.3 Case Study: How Do Techniques and Time Affect Quality? The ultimate goal of assessing response quality is to inform decisions about interview strategies and interviewer system design. To highlight the poten- tial for our methods to inform those decisions, we conduct a case study of how interview techniques affect response quality and how response quality changes over time. Interview Techniques Using a similar LLM judge setup as §2.3, we prompt the model to iden- tify relevant techniques that the interviewer used in an excerpt based on Kvale and Brinkmannâs (2009)âs taxonomy of interview techniques (Ta- ble 8). We use the same annotation setup as before to validate these judgments by comparing them to 300 human judgments over 100 excerpts. Because excerpts can contain multiple techniques, we com- pare the average Jaccard similarity. We find that the similarity between pairs of human annotators is0.51, compared to0.5between the LLM judge and the human annotators. We classify excerpts based on the interview tech- niques used in them and then use a Kruskal-Wallis test to find a statistically significant difference in 6 020406080100 Interview Progress (%) 0% 20% 40% 60% 80% 100% Proportion of Responses Response Quality Criterion 5.04.03.02.01.0 Figure 2: We compare the distribution of quality re- sponses over the course of the interviews. Our quality criterion reveals a temporal trend where participants tend to provide the lowest quality responses during the beginnings and ends of interviews. median response quality for responses obtained with the different techniques (p < 0.001). We then use Dunnâs post-hoc test with Bonferroni correc- tion to identify statistically significant differences in medians between pairs and use those to identify groups of techniques with similar response quality (full details of the test are provided in Section D.2). Figure 1 displays the results. Each member of a group has a statistically significant difference in median response quality as compared to members of all other groups. We find that techniques that are used to elicit information core to the research project (Group 1: Follow-up, Direct Questioning, Indirect Questioning) result in interview responses with the highest quality ratings. The group with the second-highest quality responses represents tech- niques that are designed to clarify responses or reach common ground with a participant (Group 2: Specifying, Interpreting). The group with the third-highest quality responses are techniques that guide or direct the attention of participants (Group 3: Structuring). The group with the fourth-highest quality responses aims to build rapport with the participant to elicit higher quality responses later in the interview (Group 4: Support & Rapport Build- ing). The final group represents techniques that explain the project to the participant and collect background information (Group 5: Introduction & Contextualization). Time We also analyze the effects of time on the quality of responses. Because our dataset has in- terviews of varying lengths, we normalize the time as the progress through the total length of the in- terview from 0 to 100% and compare the distribu- tion of response quality over the normalized time (Figure 2). Our results reveal a temporal trend where participants tend to provide the lowest qual- ity responses during the beginnings and ends of interviews. This finding is consistent with com- mon interview timelines where interviewers reserve the beginnings and ends of interviews for logistics, small talk, and winding down. These results demonstrate that our measures cap- ture meaningful differences in response quality that reflect both the different functions of various inter- view techniques and common timelines of inter- views. Future work could use our measures to in- vestigate how interview techniques affect interview quality more deeply, such as if conducting rapport building and contextualization early in an interview improves later responses to direct questions. 5 Discussion Our results provide insights to inform designers of NLP interview systems and qualitative researchers. For interview system designers, we provide em- pirically grounded metrics to evaluate the quality of the data that NLP interview systems collect. Designers should evaluate relevance to a research question, as it is the most predictive of contribu- tion to a studyâs findings. In contrast, researchers should not emphasize clarity and surprisal-based informativeness, as they are not useful for predict- ing contribution to a studyâs findings. We further show that all the metrics we investigate can be measured at scale with LLMs, thus facilitating au- tomated evaluation. These metrics could also be used as a reinforcement learning objective to train an interview system. For qualitative researchers, we validate theoret- ical frameworks of response quality that can be used to guide qualitative studies (e.g., helping re- searchers identify when they need to modify inter- view plans), to train qualitative researchers, and to conduct new studies, like evaluating the effects of different interview techniques or participant selec- tion methods on response quality. Future work can build on our methods by going beyond individual responses and analyzing quality across a full interview context to capture the effects of time-dependent techniques like rapport building. They can also explore connections between our 7 work and data saturation to better quantify not just whether responses contribute to the results, but how they contribute in comparison to one another. 6 Related Work 6.1 Empirical Analysis of Qualitative Interviews Limited prior work has empirically analyzed qual- itative interviews. Surveys on qualitative method- ology have been conducted to understand perspec- tives and common practices (Muthanna and Aldu- ais, 2023; Salet et al., 2025), but these focus on the methodology rather than the data that is collected. One barrier to this type of empirical analysis is the availability of qualitative data. To our knowl- edge, there is no openly available dataset for analyz- ing qualitative interview projects across multiple domains. We address this gap by introducing the Qualitative Interview Corpus, a dataset of 343 qual- itative interviews across a diverse set of domains that will enable further empirical analysis of the data and results from qualitative research projects. 6.2 Data Quality in Qualitative Research Discussions of quality in qualitative literature pri- marily focus on methodological rigor rather than evaluating the collected data (Tracy, 2010; Roul- ston, 2010; Cope, 2014; Korstjens and Moser, 2018) because there is an assumption that a hu- man investigator is directing the research project towards its objectives. However, there are cases that require evaluating the quality of the data itself. For example, experimenting with new interview strategies or evaluating NLP interview systems, which are capable of conducting methodologically rigorous interviews that do not collect any useful data for the projectâs goals. As detailed in §2.1, some work in qualitative re- search has proposed characteristics of high-quality interview responses (Kvale and Brinkmann, 2009; Charmaz, 2014; Patton, 2015; Small and Calarco, 2022), but these frameworks are based on personal experience and lack empirical evidence to validate them, leading to a lack of clarity in evaluating in- terview response quality. Our work addresses this gap by empirically evaluating the extent to which responses with these characteristics actually con- tribute to the results of the study. 6.3 Evaluating Interview Systems The lack of clarity in measuring interview data quality has translated to a lack of clarity in evalu- ating NLP interview systems. Some system objec- tives, like engaging participants (Xiao et al., 2020b; Cuevas et al., 2025) or maintaining coherent con- versation (Guo et al., 2024; Liu and Yu, 2025) have intuitive measures, but there is a lack of empirically grounded measures for data quality. Work in information elicitation has attempted to measure response quality by explicitly modeling belief distributions and information gain (Handa et al., 2024; Choudhury et al., 2025). However, these methods are not appropriate for most qualita- tive interviews, where investigators aim to collect nuanced insights that cannot be clearly mapped onto a probability distribution. Other work has measured quality with domain expert judgments of the insights revealed in the interviews (Anugraha et al., 2026). Though this is a robust method for measuring the final output of an interview system, it is expensive and impractical for many important tasks like intermediate evaluations, defining an objective function, or comparing large numbers of systems. The most popular method for measuring re- sponse quality is using other characteristics of par- ticipant responses, such as specificity, clarity, and relevance, as proxy criteria. These characteristics are chosen with theoretical justifications coming from the Gricean maxims (Xiao et al., 2020a,b; Jiang et al., 2023; Hu et al., 2024; Jacobsen et al., 2025) or qualitative literature (Cuevas et al., 2025). Our work provides the empirical validation miss- ing in prior studies, showing which metrics actually translate to achieving the goals of the study. 7 Conclusion In this work, we introduce the Qualitative Interview Corpus and use it to empirically evaluate which proposed measures of interview response quality are actually predictive of a responseâs contribution to the study findings. We find that the strongest predictor of response quality is relevance to a key research question, and we show that two commonly used metrics, clarity and surprisal-based informa- tiveness, are not predictive of response quality. Our work highlights the importance of empirically val- idating theoretical frameworks in qualitative re- search and enables future research to understand qualitative interviews and evaluate interview sys- 8 tems. Limitations The primary limitation of our work is that our analysis is conducted over 14 qualitative interview projects. While we collect projects that cover a range of topics and study populations (Table 5), we cannot be certain that our results would generalize to new studies. To mitigate this limitation, we de- scribe our framework in detail and release our code to support running our evaluation on other studies. Our work additionally focuses on the perspec- tive of the researcher and interviewer, in that our assessment of interview quality is focused on what aspects of the interview contributed to the final results of the paper. They do not capture the inter- vieweeâs perspective, such as whether or not the interviewee felt comfortable and engaged. Finally, we treat inclusion in paper results as a âground truthâ metric of interview quality, which assumes that researchers correctly analyzed inter- view content. In practice, qualitative researchers may have missed relevant content provided by the interviewee. Ethical Considerations We have coordinated with the Qualitative Data Repository to ensure our use of this data is within the terms of service of their platform and abides by the user agreements that researchers agreed to when uploading data to the platform. We have further established a data release plan with the Qualita- tive Data Repository, through which our processed data will be housed on their platform under the same terms of use as the original unprocessed tran- scripts. As our work constitutes secondary analysis of publicly available de-identified data collected for research purposes, there are no risks that we know of to study participants or researchers included in this data. References Carmen Alvarez. 2025. Data for: Experiences of Ghana- ian Frontline Healthcare Workers During the COVID- 19 Pandemic and Healthcare Leadership Recommen- dations. David Anugraha, Vishakh Padmakumar, and Diyi Yang. 2026. SparkMe: Adaptive Semi-Structured Inter- viewing for Qualitative Insight Discovery. Alemitu Mequanint Bezabih and C. Estelle Smith. 2025. Expanding Models of Delivery for Online Spiritual Care. Kathy Charmaz. 2014. Constructing Grounded Theory. SAGE Publications Ltd, London ; Thousand Oaks, Calif. Deepro Choudhury, Sinead Williamson, Adam Goli Ě nski, Ning Miao, Freddie Bickford Smith, Michael Kirch- hof, Yizhe Zhang, and Tom Rainforth. 2025. BED- LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design. Version Number: 2. Diane G. Cope. 2014. Methods and Meanings: Cred- ibility and Trustworthiness of Qualitative Research. Oncology Nursing Forum, 41(1):89â91. Alejandro Cuevas, Jennifer V. Scurrell, Eva M. Brown, Jason Entenmann, and Madeleine I. G. Daepp. 2025. Collecting Qualitative Data at Scale with Large Lan- guage Models: A Case Study. Proceedings of the ACM on Human-Computer Interaction, 9(2):1â27. Kerry Fosher. 2020. Marine Corps Staff Noncommis- sioned Officer Enlisted Education Project. Darcy E. Furlong, Anna Romero, Kirstin HelstrĂśm, Jes- sica Nina Lester, and Sebastian Karcher. 2025. Data for: Teaching with Shared Data for Learning Qual- itative Data Analysis: A Multi-Sited Case Study of Instructor and Student Experiences. Leo Gao, Stella Biderman, Sid Black, Laurence Gold- ing, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800GB Dataset of Diverse Text for Language Model- ing. arXiv preprint arXiv:2101.00027. Shena Gazaway, Rachel Wells, John Haley, Orlando M. Gutierrez, Tamara Nix-Parker, Isaac Martinez, Clare Lyas, Katina Lang-Lindsey, Richard Knight, and J. Nicholas Odom. 2024. Exploring the Acceptabil- ity of a Community-Enhanced Intervention to Im- prove Decision Support Partnership between Patients with Chronic Kidney Disease and Their Family Care- givers. H. Paul Grice. 1975. Logic and Conversation. In Don- ald Davidson, editor, The logic of grammar, pages 64â75. Dickenson Pub. Co. Shasha Guo, Lizi Liao, Jing Zhang, Cuiping Li, and Hong Chen. 2024. PCQPR: Proactive Conversational Question Planning with Reflection. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 11266â11278, Miami, Florida, USA. Association for Computational Linguistics. Kunal Handa, Yarin Gal, Ellie Pavlick, Noah Good- man, Jacob Andreas, Alex Tamkin, and Belinda Z. Li. 2024. Bayesian Preference Elicitation with Lan- guage Models. Version Number: 1. 9 Kunal Handa, Michael Stern, Saffron Huang, Jerry Hong, Esin Durmus, Miles McCain, Grace Yun, A. J. Alt, Thomas Millar, Alex Tamkin, Jane Leibrock, Stuart Ritchie, and Deep Ganguli. 2025. Introduc- ing Anthropic Interviewer: What 1,250 professionals told us about working with AI. Krista Harrison. 2021. Advance Care Planning in Hos- pice Organizations: A Qualitative Pilot Study. Jiaxiong Hu, Jingya Guo, Ningjing Tang, Xiaojuan Ma, Yuan Yao, Changyuan Yang, and Yingqing Xu. 2024. Designing the Conversational Agent: Asking Follow- up Questions for Information Elicitation. Proceed- ings of the ACM on Human-Computer Interaction, 8(CSCW1):1â30. Rune Møberg Jacobsen, Samuel Rhys Cox, Carla F. Griggio, and Niels Van Berkel. 2025. Chatbots for Data Collection in Surveys: A Comparison of Four Theory-Based Interview Probes. Proceedings of the 2025 CHI Conference on Human Factors in Comput- ing Systems, pages 1â21. Conference Name: CHI 2025: CHI Conference on Human Factors in Com- puting Systems ISBN: 9798400713941. Zhiqiu Jiang, Mashrur Rashik, Kunjal Panchal, Mah- mood Jasim, Ali Sarvghad, Pari Riahi, Erica DeWitt, Fey Thurber, and Narges Mahyar. 2023. Commu- nityBots: Creating and Evaluating A Multi-Agent Chatbot Platform for Public Input Elicitation. Pro- ceedings of the ACM on Human-Computer Interac- tion, 7(CSCW1):1â32. Irene Korstjens and Albine Moser. 2018. Series: Practi- cal guidance to qualitative research. Part 4: Trustwor- thiness and publishing. European Journal of General Practice, 24(1):120â124. Steinar Kvale and Svend Brinkmann. 2009. InterViews: Learning the Craft of Qualitative Research Interview- ing. SAGE Publications, Inc, Los Angeles. Fengming Liu and Shubin Yu. 2025. MimiTalk: Rev- olutionizing Qualitative Research with Dual-Agent AI. Version Number: 1. Zhe Liu, Jiamin Dai, Cristina Conati, and Joanna Mc- Grenere. 2025. Envisioning AI Support during Semi- Structured Interviews Across the Expertise Spectrum. Proceedings of the ACM on Human-Computer Inter- action, 9(2):1â29. Girmay Ayana Mersha. 2025.Data for: Lessons Learned from Operationalizing the Integration of Nutrition-Specific and Nutrition-Sensitive Interven- tions in Rural Ethiopia. Nathan K. Micatka. 2025. Data for: Socializing Policy Feedback: The Persistent Effects of Adolescent Pol- icy Program Use on Political Behaviors and Attitudes in Adulthood. Anita Milman. 2023. Ascertaining Intergovernmental Coordination Mechanisms. Abdulghani Muthanna and Ahmed Alduais. 2023. The Interrelationship of Reflexivity, Sensitivity and In- tegrity in Conducting Interviews. Behavioral Sci- ences, 13(3):218. Byung-Doh Oh, Shisen Yue, and William Schuler. 2024. Frequency Explains the Inverse Correlation of Large Language Modelsâ Size, Training Data Amount, and Surprisalâs Fit to Reading Times. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2644â2663, St. Julianâs, Malta. Association for Computational Linguistics. Maureen OâNeill. 2017. High performance school-age athletes at Australian schools: A study of conflicting demands. Michael Quinn Patton. 2015. Qualitative Research & Evaluation Methods: Integrating Theory and Prac- tice. SAGE Publications, Inc, Los Angeles London New Delhi Singapore Washington DC. Kathryn Roulston. 2010. Considering quality in qualita- tive interviewing. Qualitative Research, 10(2):199â 228. Didier Ruedin and Brian Murahwa. 2025. Perspectives on Political Representation. Xavier Salet, John Gelissen, Guy Moors, and Jelte Wicherts. 2025. Good, bad, different or something else? A scoping review of the convictions, conven- tions and developments around quality in qualitative research. Royal Society Open Science, 12(6):242001. Andrew Shuman. 2021. Data for: Drug Shortage Man- agement: A Qualitative Assessment of a Collabora- tive Approach. Mario Luis Small and Jessica McCrory Calarco. 2022. Qualitative Literacy: A Guide to Evaluating Ethno- graphic and Interview Research. University of Cali- fornia Press, Oakland, California. Beth Steinberg, Yulia Mulugeta, Catherine Quatman- Yates, Maeghan Williams, Anvitha Gogineni, and Maryanna Klatt. 2024. Data for: Barriers and Fa- cilitators to Implementation of Mindfulness in Mo- tion for Firefighters and Emergency Medical Service Providers. Sarah J. Tracy. 2010. Qualitative Quality: Eight âBig- Tentâ Criteria for Excellent Qualitative Research. Qualitative Inquiry, 16(10):837â851. Emilia F. Vignola, Emily Q. Ahonen, and Anjum Hajat. 2024. Data for: What Extraordinary Times Tell Us about Ordinary Ones: A Multiple Case Study of Pre- cariously Employed Food Retail and Service Work- ers in Two U.S. State Contexts during the COVID-19 Pandemic. Ziang Xiao, Michelle X. Zhou, Wenxi Chen, Huahai Yang, and Changyan Chi. 2020a. If I Hear You Cor- rectly: Building and Evaluating Interview Chatbots 10 with Active Listening Skills. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, pages 1â14. Conference Name: CHI â20: CHI Conference on Human Factors in Computing Systems ISBN: 9781450367080. Ziang Xiao, Michelle X. Zhou, Q. Vera Liao, Glo- ria Mark, Changyan Chi, Wenxi Chen, and Huahai Yang. 2020b. Tell Me About Yourself: Using an AI- Powered Chatbot to Conduct Conversational Surveys with Open-ended Questions. ACM Transactions on Computer-Human Interaction, 27(3):1â37. A Model Prompts We provide the exact prompts used in the LLM judges for the response characteristics, quality cri- terion, and interview techniques (Figures 3â11). 11 Prompt: Attributed Meaning You are an expert qualitative researcher analyzing interview data. Rate the level of attributed meaning of the participant statement from in the current interview excerpt below on a scale from 1 to 3. In addition to the current excerpt, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for participant statement in the current excerpt. Scoring Rubric: 1. The participant statement does not provide information indicating how significant an action or experience is to the participant. 2. The participant statement indicates some level of significance of an action or experience to a participant but does not provide information about what that significance is. 3. The participant statement directly shows the significance of an action or experience to the participant and explains what that significance is. CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): excerpt CRITICAL: Output only a single digit (1, 2, or 3). Do not write any additional text. Figure 3: LLM Judge prompt used to evaluate the attributed meaning of participant statements. Prompt: Clarity You are an expert qualitative researcher analyzing interview data. Rate the clarity of the participant statement from in the current interview excerpt below on a scale from 1 to 3. In addition to the current excerpt, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for participant statement in the current excerpt. Scoring Rubric: 1. The participant statement is incoherent, or the meaning is completely unclear. 2. The participant statement is ambiguous or requires guessing to understand. 3. The participant statement is clear to read and understand. CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): excerpt CRITICAL: Output only a single digit (1, 2, or 3). Do not write any additional text. Figure 4: LLM Judge prompt used to evaluate the clarity of participant statements. 12 Prompt: Immediate Relevance You are an expert qualitative researcher analyzing interview data. Rate how relevant the participantâs statement is to the specific question asked by the interviewer in the current interview excerpt below on a scale from 1 to 3. In addition to the current excerpt, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for participant statement in the current excerpt. Scoring Rubric: 1. The participant statement is completely unrelated to the question asked, avoids the question entirely, or addresses a totally different topic. 2. The participant statement is related to the general topic of the question but drifts or answers a different question than the one posed. 3. The participant statement directly answers the specific question posed by the interviewer. CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): excerpt CRITICAL: Output only a single digit (1, 2, or 3). Do not write any additional text. Figure 5: LLM Judge prompt used to evaluate the relevance of participant statements to the interviewerâs question. Prompt: Self-Reportedness You are an expert qualitative researcher analyzing interview data. Rate the self-reportedness of the participant statement from in the current interview excerpt below on a scale from 1 to 3. In addition to the current excerpt, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for participant statement in the current excerpt. Scoring Rubric: 1. The participant statement is not interpretable without additional context. 2. The core idea of the participant statement is understandable but may require additional context for full understanding. 3. The participant statement is fully self-contained, not requiring any additional context to be interpretable. CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): excerpt CRITICAL: Output only a single digit (1, 2, or 3). Do not write any additional text. Figure 6: LLM Judge prompt used to evaluate the self-reportedness of participant statements. 13 Prompt: Specificity You are an expert qualitative researcher analyzing interview data. Rate the specificity of the participant statement from in the current interview excerpt below on a scale from 1 to 3. In addition to the current excerpt, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for participant statement in the current excerpt. Scoring Rubric: 1. The participant statement is generic or abstract, providing only high-level summaries or vague descriptions. 2. The participant statement describes a particular event, action, or opinion without concrete details or examples. 3. The participant statement describes a particular event, action, or opinion with concrete details or examples. CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): excerpt CRITICAL: Output only a single digit (1, 2, or 3). Do not write any additional text. Figure 7: LLM Judge prompt used to evaluate the specificity of participant statements. Prompt: Spontaneity You are an expert qualitative researcher analyzing interview data. Rate the spontaneity of the participant statement from in the current interview excerpt below on a scale from 1 to 3. In addition to the current excerpt, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for participant statement in the current excerpt. Scoring Rubric: 1. The participant statement only confirms or reiterates information provided in the interviewerâs statement. 2. The participant statement adds additional information beyond what was provided in the interviewerâs statement but remains within the topic posed. 3. The participant statement introduces a new topic that may be related but was not introduced in the interviewerâs statement. CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): excerpt CRITICAL: Output only a single digit (1, 2, or 3). Do not write any additional text. Figure 8: LLM Judge prompt used to evaluate the spontaneity of participant statements. 14 Prompt: Research Question Relevance You are an expert qualitative researcher analyzing interview data. Estimate how relevant the participant statement from the current interview excerpt below is to the provided research question on a scale from 1 to 3. In addition to the current excerpt and research question, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for participant statement in the current excerpt. Scoring Rubric: 1. The participant statement is unrelated to the research question or discusses a completely different topic. 2. The participant statement is tangentially related to the topic of the research question. 3. The participant statement directly addresses the research question. RESEARCH QUESTION: research_question CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): excerpt CRITICAL: Output only a single digit (1, 2, or 3). Do not write any additional text. Figure 9: LLM Judge prompt used to evaluate the relevance of participant statements to a key research question. 15 Prompt: Quality Criterion You are an expert qualitative researcher analyzing interview data. Your task is to rate the likelihood that the participant statement from the current interview excerpt below contributed to the provided results section on a scale from 1-5. In addition to the current excerpt, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for participant statement in the current excerpt. Scoring Rubric: 1. The statement is unrelated to the results section or contradicts it. 2. Tangential relation; discusses the topic but offers no specific substance. 3. Aligns with the results section but is general or vague. 4. Provides an example or sentiment that matches the results sectionâs conclusion. 5. Appears in the results section and likely served as a primary source for it. RESULTS SECTION: result CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): excerpt CRITICAL: Output only a single digit (1, 2, 3, 4, or 5). Do not write any additional text. Figure 10: LLM Judge prompt used to evaluate the likelihood that participant statements contributed to the results section. 16 Prompt: Interview Techniques You are an expert qualitative researcher analyzing interview data. Your task is to analyze the interviewer statement in the current interview excerpt below and determine which of the following categories it fits in based on the strategy that the interviewer is using. In addition to the current excerpt, you are also provided with a short context blurb and the interview excerpt that immediately preceded the current excerpt in the transcript. These two sections are only to understand the context of the current excerpt, and your rating should be for interviewer statement in the current excerpt. Possible Categories: 1. Introductory/Contextualization Questions: Open-ended questions that may be unrelated to the overall research questions but are designed to give an understanding of the participant or context. 2. Support and Rapport Building: A statement designed to make a connection with the participant, provide support, or let the participant know that the purpose of the interview is being fulfilled. 3. Follow-up / Elaboration Probe: A statement designed to encourage a participant to continue talking. It may be a simple âuh-huhâ or âm,â or it could also be a direct call such as, âCould you say some more about that?â 4. Specifying / Detail-Oriented Probe: Questions that follow up to ask who, where, what, when, or how to obtain a complete and detailed picture of an activity or experience. 5. Direct Questioning: A question that directly introduces topics or dimensions and asks the respondent about them. 6. Indirect / Projective Questioning: Indirect questions that may ask about the attitudes of others or encourage an indirect statement of the participantâs own motivations, attitudes, or emotions. 7. Structuring: A statement that controls the structure of the interview by transitioning topics, redirecting respondents, or breaking off participant answers that may be irrelevant to the purpose of the interview. 8. Interpreting: A statement that rephrases or interprets answers provided by the participant to get clarification or reach common ground with the participant. CONTEXT BLURB (context only): context_blurb PREVIOUS INTERVIEW EXCERPT (context_only): previous CURRENT INTERVIEW EXCERPT (rate this): current_excerpt CRITICAL: Output only digits (1, 2, 3, 4, 5, 6, 7, 8) separated by commas. Do not write any additional text. Figure 11: LLM Judge prompt used to identify the techniques used in interviewer statements. 17 B Qualitative Interview Corpus Construction B.1 Annotation Setup To validate whether LLM judgments can be used to operationalize our conceptual measures, we de- signed an annotation task, as displayed in Figure 12. We recruited five graduate students with experience analyzing qualitative interviews to rate either 50 or 100 excerpts and compensated them at $20 per hour. B.2 Details of the Qualitative Interview Corpus The Qualitative Interview Corpus is composed of 343 qualitative interviews and their corresponding papers from 14 research projects. In this section, we provide more details about the projects used in the corpus (Table 5) and the composition of the projects, interviews, and excerpts (Table 6). 18 Figure 12: An example of the interface used to collect human annotations. The participant response is redacted to comply with the Qualitative Data Repositoryâs terms of use. 19 Research ProjectSubjectsKeywords# InterviewsAvg. Word Count Mindfulness for Firefighters and EMS Workers (Steinberg et al., 2024) Medicine, Health and Life Sciences firefighters, mindfulness, emergency medical service (EMS) providers, barriers, facilitators, implementation 115,520 Drug Shortage Management (Shuman, 2021) Medicine, Health and Life Sciences pharmacy, inventory control, inventory shortages, cooperation, drug shortages 165,462 Ghanaian Healthcare Workers During COVID-19 (Alvarez, 2025) Medicine, Health and Life Sciences COVID-19, healthcare worker203,891 Socializing Policy Feedback (Micatka, 2025) Social Sciencesadolescence, welfare, politics, attitude, civic, government, youth, policy 306,039 Perspectives on Political Representation (Ruedin and Murahwa, 2025) Social Sciencespolitical representation, politics, voting232,588 Nutrition Interventions in Rural Ethiopia (Mersha, 2025) Medicine, Health and Life Sciences nutrition, nutrition-sensitive, nutrition-specific, community health, agriculture, multisectoral 211,535 Marine Corps Education Project (Fosher, 2020) Medicine, Health and Life Sciences; Social Sciences stress, resilience, training and education, organizational values, biological determination, armed forces, applied social science, combat stress 327,527 Intergovernmental Coordination Mechanisms (Milman, 2023) Earth and Environmental Sciences; Social Sciences coordination, groundwater, sustainability, inter-organizational relationships, water utilities 439,198 Models of Delivery for Online Spiritual Care (Bezabih and Smith, 2025) Computer and Information Science spiritual care, chaplaincy, healthcare, nursing, palliative care, mental health, religion, spirituality 2112,123 Partnership between Kidney Disease Patients and Caregivers (Gazaway et al., 2024) Medicine, Health and Life Sciences decision making, training, program evaluation, chronic illnesses, renal disease, healthcare delivery 252,647 Shared Data for Learning Qualitative Data Analysis (Furlong et al., 2025) Social Sciencesactive learning, teaching methods, college students, college faculty, qualitative research 96,082 Advance Care Planning in Hospice Organizations (Harrison, 2021) Medicine, Health and Life Sciences; Social Sciences hospices, life care planning, palliative treatment, goals of care 506,828 Food Retail and Service Workers during COVID-19 (Vignola et al., 2024) Medicine, Health and Life Sciences; Social Sciences precarious employment, employment quality, fundamental causes, constrained choices, policy, COVID-19 2311,369 High-performance school-age athletes at Australian schools (OâNeill, 2017) Social Sciencesathlete, bullying, high performance, NVivo, parent, school age, schools, student-athlete, teacher 192,021 Table 5: A detailed view of the research projects used in the Qualitative Interview Corpus including their self- identified subjects and keywords from the Qualitative Data Repository, the total number of interviews that they contributed to the corpus, and the average length of their interviews measured in number of words. 20 MetricWord CountInterviewer UtterancesParticipant UtterancesExcerpts Total2,157,93927,25431,43416,940 Average Per Project154,138.501,946.712,245.291,210 Average Per Interview6,147.9779.4691.6449.39 Average Per Excerpt127.391.611.86â Table 6: The word count, number of interviewer utterances, number participant utterances, and number of excerpts in the Qualitative Interview Corpus. 21 C Mixed-Effects Model C.1 Model Equation Because our data has a nested structure where mul- tiple responses come from a single participant and multiple participants come from a single research project, we cannot assume independence between responses. To account for this, we use a linear mixed-effects model given by Equation 1: Y ijk = β 0 + 10 X p=1 β p X pijk + u k + v jk + Îľ ijk (1) Y ijk represents the observed response quality criterion for thei-th response provided by thej- th participant in thek-th research project.β 0 is the overall fixed intercept of the model.X pijk de- notes the value of thep-th fixed-effect predictor for a response. The corresponding fixed-effect coeffi- cient,β p , captures the relationship between thep-th predictor and the response quality. To model the nested variance,u k represents the random intercept for thek-th project, accounting for differences in projects. Similarly,v jk represents the random inter- cept for thej-th participant nested within thek-th project, accounting for differences in participants. Finally,Îľ ijk is the residual error capturing the re- maining unexplained variance for each response. C.2 Multicollinearity and Variance Inflation We design our framework with distinct character- istics of participant responses to minimize multi- collinearity and ensure stable coefficients in our regression. This choice results in low variance inflation factors, which support the stability and interpretability of our mixed-effects modelâs coef- ficients (Table 7). The full correlation among all predictors is provided in Figure 14. Predictor VariableVIF Response Length2.25 Specificity2.21 Spontaneity1.86 Attributed Meaning1.75 Self-reportedness1.69 Response Length Ratio1.64 Research Question Relevance1.60 Immediate Relevance1.39 Clarity1.25 Average Surprisal1.06 Table 7: Variance Inflation Factors (VIF) for our mixed- effects model. TechniqueDescription Introduction & Contextualization Open-ended questions designed to un- derstand the participant or context, often unrelated to core research questions. Support & Rapport Building Statements designed to build a connec- tion, provide support, or validate the par- ticipantâs contribution. Follow-upBrief interjections (e.g., "uh-huh") or di- rect calls to encourage the participant to continue talking. SpecifyingFollow-up questions (who, what, where, when, how) to obtain a detailed picture of an experience. Direction Questioning Questions that directly introduce spe- cific topics or dimensions to the respon- dent. Indirect Questioning Questions about othersâ attitudes to indi- rectly surface the participantâs own mo- tivations or emotions. StructuringStatements used to transition topics, redi- rect respondents, or interrupt irrelevant answers. InterpretingRephrasing or interpreting answers to seek clarification or reach common ground. Table 8: Taxonomy of interview techniques from Kvale and Brinkmann (2009) D Interview Techniques D.1 Technique Taxonomy We use Kvale and Brinkmannâs (2009) taxonomy of interview techniques to conduct our analysis (Table 8). D.2 Dunnâs Post-hoc Test In §4.3, we identify techniques used in interview excerpts and use Dunnâs post-hoc test with Bonfer- roni correction to identify statistically significant differences in medians between pairs of techniques. Figure 13 shows the full set of p-values for Dunnâs post-hoc test. 22 Intro & Context Support & Rapport Structuring Specifying Interpreting Indirect Question Direct Question Follow-Up Intro & Context Support & Rapport Structuring Specifying Interpreting Indirect Question Direct Question Follow-Up 1.0000.0000.0000.0000.0000.0000.0000.000 0.0001.0000.0110.0000.0000.0000.0000.000 0.0000.0111.0000.0000.0000.0000.0000.000 0.0000.0000.0001.0000.5201.0000.0000.000 0.0000.0000.0000.5201.0001.0000.0000.000 0.0000.0000.0001.0001.0001.0001.0000.179 0.0000.0000.0000.0000.0001.0001.0000.453 0.0000.0000.0000.0000.0000.1790.4531.000 0.0 0.2 0.4 0.6 0.8 1.0 P-Value Figure 13: P-values from Dunnâs post-hoc test for difference in median response quality between pairs of interview techniques. 23 Clarity Relevance Specificity Attributed Meaning Self-Reportedness Spontaneity Response Length Response Length Ratio RQ Relevance Information Density Clarity Relevance Specificity Attributed Meaning Self-Reportedness Spontaneity Response Length Response Length Ratio RQ Relevance Information Density 1.000.310.070.040.42-0.05-0.05-0.030.03-0.05 0.311.000.310.190.500.160.110.090.32-0.05 0.070.311.000.480.440.630.570.440.46-0.07 0.040.190.481.000.310.500.540.400.50-0.11 0.420.500.440.311.000.350.230.180.29-0.08 -0.050.160.630.500.351.000.460.360.45-0.09 -0.050.110.570.540.230.461.000.630.410.04 -0.030.090.440.400.180.360.631.000.290.12 0.030.320.460.500.290.450.410.291.00-0.11 -0.05-0.05-0.07-0.11-0.08-0.090.040.12-0.111.00 0.0 0.2 0.4 0.6 0.8 1.0 Figure 14: Correlations observed in the Qualitative Interview Corpus between each pair of characteristics in our framework. 24