Paper deep dive
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
Argyrios Papoudakis, Mirella Lapata, Frank Keller
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/14/2026, 2:34:27 AM
Summary
The paper introduces a modular training framework for character description generation in long-form narratives. It decouples reasoning from generation by using a reasoning model to produce structured QA traces, which are then used to guide a generation model. This approach improves faithfulness and informativeness compared to standard long-context LLMs, which often suffer from performance degradation when built-in reasoning is enabled. The framework is model-agnostic and can be integrated with various long-context strategies like retrieval, hierarchical processing, and incremental updating, with the reasoning model optimized via Group Relative Policy Optimization (GRPO).
Entities (5)
Relation Signals (3)
Argyrios Papoudakis → authored → Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
confidence 100% · Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books Argyrios Papoudakis
QA-Guided Reasoning → evaluatedon → BookWorm
confidence 100% · We evaluate our approach on two character understanding datasets: BookWorm
GRPO → trains → Reasoning Model
confidence 100% · We optimize R φ with GRPO (Shao et al., 2024)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Character description generation is an important capability for narrative-focused applications such as summarization, story analysis, and character-driven simulations. However, generating accurate character descriptions from long-form narratives (e.g., novels) is challenging: models must track evolving attributes (e.g., relationships and events), integrate evidence scattered across the text, and infer implicit details. Despite the success of reasoning-enabled LLMs on many benchmarks, we find that for character description generation their performance improves when built-in reasoning is disabled (i.e., an empty reasoning trace). Motivated by this, we propose a training framework that decouples reasoning from generation. Our approach, which can be applied on top of long-context LLMs or chunk-based methods, consists of a reasoning model that produces a structured QA reasoning trace and a generation model that conditions on this trace to produce the final character description. Experiments on two datasets (BookWorm and CroSS) show that QA-guided reasoning improves faithfulness, informativeness, and grounding over strong long-context baselines.
Tags
Links
- Source: https://arxiv.org/abs/2604.11435v1
- Canonical: https://arxiv.org/abs/2604.11435v1
Trouble viewing inline? Open PDF directly →
Full Text
84,014 characters extracted from source content.
Expand or collapse full text
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books Argyrios PapoudakisMirella LapataFrank Keller Institute of Language, Cognition and Computation School of Informatics, University of Edinburgh 10 Crichton Street, Edinburgh EH8 9AB a.papoudakis@sms.ed.ac.uk, mlap, keller@inf.ed.ac.uk Abstract Character description generation is an impor- tant capability for narrative-focused applica- tions such as summarization, story analysis, and character-driven simulations. However, generating accurate character descriptions from long-form narratives (e.g., novels) is challeng- ing: models must track evolving attributes (e.g., relationships and events), integrate evidence scattered across the text, and infer implicit de- tails. Despite the success of reasoning-enabled LLMs on many benchmarks, we find that for character description generation their perfor- mance improves when built-in reasoning is disabled (i.e., an empty reasoning trace). Mo- tivated by this, we propose a training frame- work that decouples reasoning from generation. Our approach, which can be applied on top of long-context LLMs or chunk-based methods, consists of a reasoning model that produces a structured QA reasoning trace and a generation model that conditions on this trace to produce the final character description. Experiments on two datasets (BookWorm and CroSS) show that QA-guided reasoning improves faithfulness, in- formativeness, and grounding over strong long- context baselines 1 . 1 Introduction Writers craft characters to create engaging stories that invite readers to experience actions, emotions, and goals from the characters’ perspectives. Char- acters form the core of the narrative, with other elements such as plot, conflict, setting, and theme, built around them. A growing body of work in nat- ural language processing aims to model characters for story analysis (Zhu et al., 2023), narrative gen- eration and summarization (Fan et al., 2018), and persona simulation within character-driven interac- tions (Shao et al., 2023). Yet despite their impor- tance, characters remain difficult to model computa- 1 We release our code and data athttps://github.com/ apapoudakis/qa-guided-reasoning Book (full text) Character name QA Reasoning Model R φ QA Trace Q 1 : What is X’s role? A 1 : Protagonist Q 2 : Who is X’s ally? A 2 : Character Y . . . Generation Model G θ <think> QA trace </think> Character Description: X’s adventures begin with her fateful jump down the rabbit hole, and the tale is an extended metaphor for the challenges she will face as she grows into an adult. She possesses unusual composure for a child, and she seems bright but makes many charming mistakes... GRPO training SFT training Figure 1: Overview of QA-based guided reasoning for character description generation. Given a book and char- acter name, the reasoning model generates a structured QA trace capturing salient character information. This trace is injected into the thinking tokens of the genera- tion model, which produces the final character descrip- tion. The reasoning model can be trained with GRPO and the generation model with SFT. tionally, particularly in long-form narratives where their traits, motivations, and relationships evolve over extended spans of text and through complex and often non-linear plots (Chaturvedi et al., 2015; Vishnubhotla et al., 2024). Consequently, a central difficulty in character- centric understanding stems from the sheer length of the input. To address this, earlier work adopted chunk-based approaches using LLMs with short- context abilities to process stories either hierarchi- cally (Wu et al., 2021) or incrementally (Chang et al., 2023). While these methods are computa- tionally efficient, they often struggle to capture all relevant information within individual chunks and to integrate evidence reliably across chunks. Retrieval-augmented methods have also been pro- posed to select salient context (Xu et al., 2023); arXiv:2604.11435v1 [cs.CL] 13 Apr 2026 however, the retrieved passages are frequently dis- joint and may omit critical details, undermining global narrative coherence. More recently, long- context LLMs, capable of processing up to mil- lions of tokens (Team et al., 2024), have been used to analyze stories in a single pass. Nevertheless, even these models face known limitations, includ- ing position bias (Liu et al., 2024), difficulty in exploiting in-context examples (Li et al., 2024) and challenges in integrating information across distant spans (Tian et al., 2025). Beyond context scaling, long narratives present an additional challenge: much information is im- plied rather than explicitly stated. Models must therefore reason over long inputs (>100k tokens), integrating evidence from different text passages. This requires inferring implicit details, combining information which is being revealed piecemeal, and filtering out irrelevant content. Recent work has enabled explicit reasoning via post-training rein- forcement learning (Shao et al., 2024), achieving strong results on tasks such as mathematics (Yang et al., 2024) and short-form question answering (Rein et al., 2024). However, extending these ap- proaches to non-verifiable settings like text gener- ation remains challenging, as reward computation is difficult when many valid outputs exist across multiple evaluation dimensions. We address these limitations by focusing on char- acter description generation in books (Brahman et al., 2021), where models must produce factual descriptions given a character and the full story. We propose a two-stage training framework that decouples reasoning from generation: (1) a reason- ing model produces an explicit reasoning trace, and (2) a generation model conditions on this trace to produce the final description (see Figure 1). Impor- tantly, our reasoning model uses question-answer pairs — which have been previously shown to pro- vide a useful abstraction for isolating salient con- tent in summarization (Huot et al., 2023; Narayan et al., 2023; Liu et al., 2025) — as a scaffold to guide downstream description generation, forcing the model to “reason through” salient evidence be- fore writing. As our reasoning model is input ag- nostic, it can be integrated on top of different long- context methods, including hierarchical merging, incremental updating, and direct processing with long-context LLMs. To train the reasoning model, we use Group Rel- ative Policy Optimization (GRPO) (Shao et al., 2024) directly on gold-standard character descrip- tions and silver-standard reasoning traces derived from them, thereby avoiding the need to simulate rewards based on the final generated outputs. The generation model can either be trained with super- vised fine-tuning (SFT) on these reasoning traces or applied in zero-shot mode. We evaluate our approach on two character understanding datasets: BookWorm (Papoudakis et al., 2024), which contains books of various genres (primarily novels and plays) from Project Gutenberg, and CroSS (Yuan et al., 2024), which comprises recently published novels. Both datasets pair full-text books with character descriptions. Our experiments reveal that enabling the built-in reason- ing capabilities of current LLMs actually degrades performance compared to using an empty reason- ing trace. In contrast, our approach effectively im- proves both the faithfulness and informativeness of character descriptions across both datasets, en- hancing the performance of existing long-context methods. In summary, we make the following con- tributions: •We show that the built-in reasoning mode of current LLMs reduces faithfulness for charac- ter description generation, with empty traces yielding more accurate outputs. •We introduce QA-based guided reasoning, a modular approach that decouples reasoning from generation, producing a structured QA trace and then generating the final character description conditioned on that trace. •We propose a training strategy that optimizes the reasoning model with GRPO using gold- standard descriptions and automatically gen- erated QA traces, avoiding reward design over open-ended generated descriptions. • We show that guided reasoning yields consis- tent improvements in faithfulness and infor- mativeness across two datasets and can be in- tegrated with various long-context strategies. 2 Related Work Character Understanding Work in computa- tional narrative understanding has focused on char- acters to support story summarization (Zhang et al., 2019), generation (Liu et al., 2020), analysis (La- batut and Bost, 2019), and the simulatation of in- teractions among different personas (Shao et al., 2023). Existing studies have examined multiple character dimensions, including roles (Skowron et al., 2016), relationships (Iyyer et al., 2016), per- sonality (Bamman et al., 2013), and emotions (Kim and Klinger, 2019). More recently, attention has shifted to characters in longer-form narratives (e.g., screenplays, books), where increased com- plexity (e.g., longer plots, evolving relationships) raises modeling challenges and motivates evalua- tion of long-context systems. In this setting, several prediction-based tasks have been introduced, such as personality prediction (Sang et al., 2022; Yu et al., 2023) and assigning characters to tropes in scripts (Baruah and Narayanan, 2025). Previous work has also explored richer character represen- tations, e.g., by constructing character sheets (Gu- rung and Lapata, 2024) or learning character em- beddings (Inoue et al., 2022) from books. Another line of research addresses character-based corefer- ence resolution in novels (Martinelli et al., 2025) and screenplays (Baruah and Narayanan, 2023). In parallel, recent work has framed character un- derstanding as a text-generation problem, produc- ing character descriptions and analyses (Brahman et al., 2021; Papoudakis et al., 2024), structured pro- files (Yuan et al., 2024), and persona-conditioned dialogues (Chen et al., 2023). We adopt a similar setting, focusing on free-form character description generation in book-length narratives. Unlike Yuan et al. (2024), we target unstructured descriptions, as existing datasets provide human-written text re- quiring no additional processing. Long-context Models and ReasoningLanguage models can now process extremely long sequences, up to millions of tokens (Team et al., 2024), thanks to advances in sparse attention (Beltagy et al., 2020), long-context post-training (Gao et al., 2025), scalable positional embeddings (Ding et al., 2024), and system (Kwon et al., 2023) and hardware- level (Dao et al., 2022) optimizations. However, greater context length does not reliably translate into better long-document reasoning; recent studies report persistent failure modes, such as position bias (Liu et al., 2024) and difficulty exploiting in- context examples (Li et al., 2024). As the training and inference costs of long-context modeling re- main a major barrier (Fu, 2024), previous work has also explored efficient alternatives that handle long documents with short-context models by in- corporating retrieval, hierarchical processing, or incremental updating (Chang et al., 2023). By oper- ating on chunks or selectively attending to limited context, these methods can be competitive across a range of long-context settings (Xu et al., 2023). In this paper, we show how guided reasoning can be layered on top of both chunk-based approaches (retrieval, hierarchical processing, and incremental updating) and long-context LLMs. Effective reasoning is also crucial for inferring implicit information, filtering irrelevant content, and integrating evidence across long texts. Early work elicited reasoning via prompting (Wei et al., 2022) or fine-tuning (Wei et al., 2021). More re- cently, post-training with reinforcement learning has become a dominant paradigm, improving per- formance across a range of tasks (Ouyang et al., 2022), such as mathematics (Uesato et al., 2022) and coding (Shojaee et al., 2023). However, train- ing reasoning models for non-verifiable tasks (e.g., long-form QA and summarisation) remains un- derexplored: reward design is challenging when many outputs are acceptable and quality depends on multiple criteria. Recent work address this by using BLUE (Chang et al., 2025), perplexity completion (Gurung and Lapata, 2025), or self- certainty (Zhao et al., 2025) to train without exter- nal verifiers. In contrast, we decouple reasoning from genera- tion using a separate reasoning model whose output trace is injected into the generator. This allows us to optimize the reasoner directly with reinforce- ment learning on its own outputs, without defining rewards over the downstream generated text. In- spired by work showing QA pairs serve as effective planning proxies for summarization (Narayan et al., 2023), we use structured QA traces to guide gener- ation. However, our QA pairs function as reason- ing scaffolds that consolidate distributed evidence about characters, rather than as ordering constraints for generation. 3 QA-Guided Reasoning We propose a modular approach that decouples reasoning from surface realization for character description generation. Given a book (or long nar- rative)xand target characterc, a reasoning model R φ produces an intermediate QA-based reasoning traceT, and a generation modelG θ conditions onT to produce a free-form character descriptionˆy. The key idea is to represent intermediate reasoning as typed question-answer items that explicitly surface salient character facts before the final description is written. Figure 1 illustrates this architecture; we provide details below. 3.1 QA Reasoning Letxdenote a story (e.g., a book) andca target character. Whenxexceeds the context length of the reasoning model, we split it intomchunksx = [χ 1 ,..., χ m ].Given a chunkχ i , the reasoning model R φ outputs eitherNone(ifcis not mentioned inχ i ) or a set of QA tuples: T i = R φ (c, χ i ) = (q j , e j , a j , t j ) n i j=1 ,(1) whereq j is a question aboutc,e j is a short support- ing explanation (1–2 sentences),a j is a short an- swer (typically 1–4 words), andt j ∈Tis a question type. Following previous work (Papoudakis et al., 2024), we use the question types role, relationship, personality, event and other, which encourage the model to generate questions with diverse topics without focusing only on specific aspects (see Ap- pendix B). We do not set a maximum number of questions that the model has to generate, but in- stead ask it to include all the relevant information for the character of interest (orNoneotherwise). The final reasoning trace is the concatenation of chunk-level traces, T = CONCAT(T 1 ,..., T m ). We experimented with several question- generationstrategies,includinggenerating questions directly from the full context, generating questions per chunk and then concatenating them, or multi-step pipelines that first extract topics (Noorbakhsh et al., 2025) or create plans (Li and Zhang, 2024) (see Appendix B). Since these variants performed similarly or worse, we adopt chunk-based QA generation for simplicity and to facilitate the training of R φ (Section 3.3). 3.2 Generation Model The generation modelG θ produces a free-form character description conditioned on both the input context and the reasoning trace: ˆy = G θ (c, x, T).(2) Operationally, we injectTinto the model input between special thinking tokens (e.g.,<think>· </think>), and prompt the model to return a single- paragraph description. This design is compatible with both reasoning-enabled LLMs and standard instruction-tuned models that do not explicitly emit reasoning traces. 3.3 Training We finetune theR φ andG θ separately. Since we do not have gold-standard reasoning traces, we derive silver-standard supervision from the gold descrip- tionyby extracting a reference set of QA pairs. The structured format of these reasoning traces al- lows us to parse the generated QA-pairs, verify each of them, and provide a combined reward for the entire sequence. We employ an LLM-as-judge model to compute precision, the percentage of gen- erated QA-pairs that can be verified from the gold- standard description, and recall, the percentage of gold questions (extracted from descriptions) that the sequence of generated QA-pairs can verify. We use the F1 score as a reward signal. Precision and recall are defined as follows: R precision = 1 |T| ∑ (q i ,a i )∈T (VERIFY(q i , a i , P))(3) R recall = 1 |P| ∑ (q i ,a i )∈P (VERIFY(q i , a i , T))(4) We optimizeR φ with GRPO (Shao et al., 2024), which samples multiple traces per prompt and per- forms a relative policy update without a critic, re- ducing computational cost. While we useG θ in a zero-shot setting in sev- eral of our experiments, the generator can also be fine-tuned to better exploit guided-QA traces. Concretely, we perform SFT on tuples(c, x, T, y) whereTis produced by the guided-QA reasoner (either zero-shot or GRPO-trained), and optimize log p θ (y| c, x, T)(optionally applying loss on the injected trace tokens as well). At inference time, the reasoning trace is always supplied byR φ , and G θ generates only the final description. 3.4 Integrating Guided Reasoning with Long-Context Methods BecauseR φ operates on chunks, our QA-guided pipeline is model- and strategy-agnostic: it can be combined with several long-context processing methods by producing tracesT i for the same chunks χ i that these methods already use. Concretely, we investigate the following settings: Retrieval-based Methods (Papoudakis et al., 2024) first select a character-relevant subset of the input (e.g., paragraphs mentioningc), yield- ingx ′ ⊆ x. We simply runR φ over chunks ofx ′ and condition the generator on the resulting traceTand retrieved context. Hierarchical Methods (Chang et al., 2023) split the input into chunks and generate intermediate de- scriptions for each chunk, which are subsequently BookWormCroSS Books324126 Characters9.746.56 Samples5,869824 Input Length97,685129,113 Output Length88.79295.27 Table 1: Dataset statistics: unique books, average char- acters per book, total examples (book–description pairs), and average input/output length in words. merged by a second-stage model. We apply guided reasoning only at the first level: for each chunkχ i , we generate a traceT i and produce an intermedi- ate description conditioned on(χ i , T i ). Merging then operates over the intermediate descriptions without additional reasoning. We employ the rea- soning model in a zero-shot setting and trained with GRPO, while the generation model operates in zero-shot mode. We do not finetune the genera- tion model via SFT, as intermediate descriptions for individual chunks are not available, and the model must both generate and merge descriptions, which makes fine-tuning impractical. Incremental Methods (Chang et al., 2023) pro- cess the input chunk-by-chunk, while maintaining a running global description. At stepi, the genera- tion model updates the current description using the new chunk and its trace, i.e.,ˆy i = G θ (c, χ i , T i , ˆy i−1 ). This allows newly observed evidence about the character to be integrated as it appears in the nar- rative. Similarly to the hierarchical approach, we only train the reasoning model with GRPO, while the generation model operates in zero-shot mode. Long-context LLMs can consume the full inputx (within their context window) for generation. In this case, we still produce traces chunk-wise with R φ and inject their concatenationTinto the long- context generator, i.e., ˆy = G θ (c, x, T). 4 Experimental Setting Datasets We use the BookWorm dataset (Pa- poudakis et al., 2024), which contains books of var- ious genres (mostly novels and plays) from Project Gutenberg paired with characters descriptions from literature websites. 2 For out-of-distribution evalu- ation, we use CroSS (Yuan et al., 2024), a dataset of novels published in 2022–2023 paired with char- acter descriptions. Dataset statistics are in Table 1, examples in Appendix C. 2 BookWorm has two tasks (description and analysis); we use the description partition, which contains more examples. Model Comparisons We use Qwen-3-8b (Yang et al., 2025) as the backbone model for all our ex- periments. We report a No Context baseline, where the model is prompted to generate a character de- scription given the character name and book title (i.e., without any story text). This baseline esti- mates how much parametric knowledge the model already has about the character, serving as a proxy for potential contamination. We also include a Lead baseline, truncating the input story to the maximum length supported by the backbone model. We evaluate our QA-guided approach on top of retrieval-augmented methods that select character- relevant evidence and condition generation on the retrieved context. We consider two retrieval strate- gies. (a) Coref runs a coreference resolver to iden- tify chunks in which the target character is men- tioned, concatenates the selected chunks, and trun- cates if the resulting context exceeds the model’s input budget. We use BookNLP 3 , which prior re- search has shown to achieve strong off-the-shelf performance. (b) BM25 uses the character name as a query and retrieves chunks by BM25 score, con- catenating them until the model’s context limit is reached (32k for our experiments with Qwen-3-8b). For both retrieval approaches, we use 512-token length chunks. We further integrate QA-guided reasoning with two chunk-based long-context strategies, viz., Hier- archical processing and Incremental updating (see Section 3.4). For both methods, we use 16k-token chunks. Additional implementation details and re- sults are in Appendix A and B. All methods are evaluated in three settings: with- out reasoning (empty trace), with the model’s built-in reasoning, and with our proposed guided- QA trace. For guided-QA, we compare zero-shot and GRPO-optimized versions. For retrieval-based methods, we also train an SFT variant, computing the loss only on the target descriptions. Evaluation Metrics We use a set of automatic evaluation metrics to assess the quality of generated character descriptions following previous work (Pa- poudakis et al., 2024). PRISMA (Mahon and Lapata, 2024) is an LLM- as-a-judge metric that evaluates factual accuracy. It extracts facts from the generated description and assesses their correctness against the gold stan- dard, then extracts facts from the gold standard and checks whether the generated output supports 3 https://github.com/booknlp/booknlp them. These precision and recall scores are com- bined into PRISMA F1. We use GPT-4o-mini for fact extraction and Bespoke-MiniCheck-7B (Tang et al., 2024) for fact validation, which achieves SotA performance on LLM-AggreFact. We also evaluate factual precision against the input story using an entailment-based NLI (natural language inference) metric. For each extracted fact, we calculate entailment scores against story chunks, taking the maximum score across all chunks. A fact is considered grounded if its score exceeds 0.5. We use Bespoke-MiniCheck-7B with 1,024-token chunks for this evaluation. We complement NLI metrics with QA-based evaluation (Deutsch et al., 2021) which measures whether the generated de- scription contains the key information needed to answer questions derived from the reference. We use GPT-4o-mini to generate QA-pairs from the reference and DeBERTaV3 (He et al., 2023) fine- tuned on SQuADv2 (Rajpurkar et al., 2018) for question-answering. We further report the entity-mention F1 score as a measure entity-level coverage. Specifically, we compute precision as the proportion of the entities mentioned in the generated description that also ap- pear in the gold-standard description, and recall as the proportion of gold-standard entities recovered by the generated output; these then combine into entity-mention F1. We also report Rouge-L (Lin, 2004), which measures the longest common subse- quence with the reference descriptions. To evaluate the quality of the QA-guided trace, we compare QA-pairs produced by the reasoning model against QA-pairs derived from reference descriptions. We measure whether each generated pair is supported by the gold QA set (precision), and how many gold QA pairs are covered by the generated set (recall). We use GPT-4o-mini in an LLM-as-a-judge setting for this evaluation. Our prompts and further details are in Appendix A. 5 Results Table 2 reports performance on BookWorm across retrieval, hierarchical, incremental, and long- context settings, with and without reasoning. Built-in reasoning hurts faithfulness. Across settings, enabling the model’s built-in reasoning consistently reduces overall quality compared to an empty trace. For example, under Coref-32k, built-in reasoning lowers Rouge-L (18.44→16.54) and NLI (52.58→49.94). A similar pattern holds for Hierarchical-16k, where built-in reasoning de- creases QA (15.79→15.16) and Rouge-L (18.22 →17.33). The effect is also present in the long- context setting for Lead-128k, with NLI dropping from 69.29 to 67.17 and Rouge-L from 18.11 to 16.43. These results suggest that default “thinking” traces introduce verbosity or unsupported details, harming grounding in long narratives (see Table 10 in Appendix B for discussion on reasoning traces). QA-guided reasoning improves grounding and coverage. In contrast, our guided-QA traces im- prove evidence-sensitive metrics, particularly QA F1 and entity coverage. For BM25-32k, guided- QA yields substantial gains in EntMent (34.06→ 36.23) and QA (13.56→14.66), while also im- proving NLI (58.58→59.96), indicating better grounding in the retrieved evidence. For Coref-32k, guided-QA improves QA (13.96→15.01) while maintaining comparable EntMent. For hierarchical processing, guided-QA provides gains in EntMent (36.56→37.67), maintains comparable QA F1 and PRISMA, demonstrating that reasoning helps even when the model already aggregates chunk-level summaries. In the long-context setting, guided-QA with GRPO-trained traces improves over the base- line across all metrics except R-L (PRISMA: 17.02 →19.30, QA: 13.68→16.22). Overall, guided- QA is most beneficial when the context selection step surfaces relevant evidence, but the generator needs help integrating it into a coherent, grounded description. Trace quality and generator training affect dif- ferent metrics.Comparing zero-shot and GRPO- trained traces reveals that better traces translate into better descriptions, though gains are metric- dependent. Under BM25-32k, GRPO traces im- prove PRISMA (17.09→18.38) while preserving high EntMent, whereas SFT on descriptions primar- ily benefits surface-overlap metrics (e.g., Rouge- L: 18.18→18.99) with weaker effects on factual scores. A similar pattern emerges for Coref-32k, where SFT achieves the highest Rouge-L (19.26) but GRPO-trained traces yield the best QA F1 (14.72). This suggests that trace-level optimization and generation-level training target complementary aspects of output quality. Benefits in incremental settings are limited.In- cremental processing presents a challenging setting where guided-QA shows limited benefit. Both built- in and guided reasoning degrade PRISMA com- MethodTraceDesc PRISMAQANLIEntMentRouge-L Qwen-3-8b No Context—ZS11.9710.31—25.9717.55 Lead-32k—ZS15.7212.3561.0130.6017.79 BM25-32k—ZS17.0913.5658.5834.0618.18 BM25-32k—SFT16.17 † 13.2660.00 † 33.11 † 18.99 † + reasoningbuilt-inZS17.0613.5058.4435.1316.51 † + guided-QAZSZS17.8414.66 † 59.96 † 36.23 † 17.43 † + guided-QAGRPOZS18.38 † 15.18 † 59.90 † 37.05 † 17.59 † + guided-QAGRPOSFT18.27 † 15.15 † 60.46 † 36.60 † 17.87 † Coref-32k—ZS19.1113.9652.5836.1218.44 Coref-32k—SFT17.46 † 14.3652.5435.5919.26 † + reasoningbuilt-inZS18.59 † 13.9449.94 † 35.8816.54 † + guided-QAZSZS18.63 † 15.01 † 50.60 † 36.9517.50 † + guided-QAGRPOZS19.4914.86 † 51.23 † 37.66 † 18.01 † + guided-QAGRPOSFT19.3215.10 † 52.0037.58 † 18.42 Hierarchical-16k—ZS19.9915.7970.5736.5618.22 + reasoningbuilt-inZS19.7415.16 † 71.1535.8217.33 † + guided-QAZSZS19.8216.2070.7437.67 † 17.90 † + guided-QAGRPOZS18.94 † 15.3667.92 † 36.4918.19 Incremental-16k—ZS17.6414.4269.6835.5917.58 + reasoningbuilt-inZS16.03 † 13.36 † 66.46 † 33.37 † 15.76 † + guided-QAZSZS16.80 † 14.5668.20 † 34.9216.98 † + guided-QAGRPOZS16.59 † 13.63 † 66.05 † 34.51 † 17.42 Lead-128k (w/ YaRN)—ZS17.0213.6869.2934.9418.11 + reasoningbuilt-inZS16.9313.6567.17 † 34.2016.43 † + guided-QAZSZS19.32 † 15.77 † 71.55 † 37.80 † 17.77 † + guided-QAGRPOZS19.30 † 16.22 † 70.51 † 37.12 † 18.39 † GPT-4.1 mini No context—ZS17.3612.49—29.3517.41 Full context—ZS22.0516.6775.5335.4117.71 Table 2: Results on BookWorm. Trace: — (none), built-in (model’s default), ZS (zero-shot guided-QA), GRPO (GRPO-trained guided-QA). Desc: ZS (zero-shot) or SFT (supervised fine-tuning). † indicates statistically significant difference from the corresponding baseline without reasoning (i.e., first row for each method group); bold indicates best per metric. pared to the baseline (17.64→16.03 and 16.80, respectively). We hypothesize that the sequential updating mechanism struggles to incorporate rea- soning traces effectively, as each update step must reconcile new QA pairs with a partial description. This suggests that our approach is best suited for settings where evidence can be aggregated before generation rather than integrated incrementally. Entity coverage improves even against propri- etary models. For reference, we include GPT- 4.1-mini with full context access, which achieves the highest PRISMA (22.05), QA F1 (16.67) and NLI (75.53). However, guided-QA with the smaller Qwen-3-8b model narrows this gap, achieving the best EntMent score overall (37.80 vs. 35.41). This indicates that structured reasoning traces improve entity coverage even compared to larger proprietary models with longer context windows. Faithful QA traces improve descriptions across QA strategies. Table 4 compares alternative question-generation strategies for the reasoning model and relates trace quality to downstream de- scription performance (the generator is kept zero- shot in all settings). We also report an Oracle upper bound, where QA pairs are extracted from the gold descriptions and used directly as traces, highlight- ing the remaining headroom when perfect interme- diate signals are available. Overall, optimizing the reasoner improves both trace quality and description quality. In particu- lar, GRPO yields the best trace recall and the highest trace F1 (16.96), outperforming both zero- shot guided reasoning and SFT on traces. This improvement translates downstream: conditioning on GRPO traces produces the best QA F1 (14.79) and PRISMA score (15.25), though the PRISMA gains over no reasoning are only marginal (15.20 →15.25). By contrast, increasing the number of ModelMethodTraceDesc PRISMAQANLIEntMentRouge-L Qwen-3-8b No Context—ZS7.692.92—19.8316.40 Coref-32k—ZS23.1313.0350.0135.9017.44 + reasoningbuilt-inZS22.26 † 12.5246.84 † 33.50 † 16.02 † + guided-QAZSZS21.88 † 13.1149.6335.4617.24 † + guided-QAGRPOZS22.63 † 13.08 † 49.7035.38 † 17.33 † Lead-128k (w/ YaRN)—ZS22.2213.7066.3935.4817.21 + reasoningbuilt-inZS18.08 † 10.97 † 66.9831.18 † 15.75 † + guided-QAZSZS23.72 † 14.62 † 68.57 † 36.84 † 17.47 † + guided-QAGRPOZS23.78 † 14.42 † 68.09 † 36.18 † 17.49 † GPT-4.1-mini No context—ZS10.013.98—21.4616.10 Full context—ZS28.4916.0872.6728.4916.80 Table 3: Results on CroSS. Trace: — (none), built-in (model’s default), ZS (zero-shot guided-QA), GRPO (GRPO- trained guided-QA). Desc: ZS (zero-shot) or SFT (supervised fine-tuning). † indicates statistically significant difference from the corresponding baseline without reasoning (i.e., first row for each method group); bold indicates best per metric. QA reasoningDescription Method# QA Prec RecF1PRISMA QA No reasoning—15.2014.42 No chunking6.50 16.10 16.93 16.50 14.0214.03 guided-QA8.60 15.07 17.68 16.2714.4514.45 guided-QA (SFT)6.60 14.84 15.98 15.3814.1114.08 guided-QA (GRPO) 11.07 15.25 19.10 16.9615.2514.79 Oracle7.51— 44.8140.61 Table 4: Comparison of question generation methods for the reasoning model on the BookWorm dataset (vali- dation set) with coreference-based retrieval. The gener- ation model is zero-shot in all experiments. QA pairs alone is not sufficient: guided-QA gener- ates more questions than the no chunking approach but has lower precision and total F1 score. These results suggest that downstream gains are driven by faithful and informative traces rather than by trace length (see question generation experiments in Ap- pendix B). Finally, the oracle traces dramatically outperform all automatic traces, indicating substan- tial room for improvement in trace generation and verification. Guided reasoning improves transfer to CroSS. Table 3 evaluates whether models tuned on Book- Worm transfer to CroSS (Yuan et al., 2024). Using Qwen-3-8b, character-focused context strategies yield substantial gains over no context (e.g., Coref- 32k PRISMA: 7.69→23.13). Consistent with Book- Worm, built-in reasoning degrades grounding met- rics (Coref-32k NLI: 50.01→46.84; Lead-128k EntMent: 35.48→31.18), while guided-QA im- proves or keeps comparable performance across settings. For Lead-128k, guided-QA with GRPO traces achieves the best PRISMA (23.78) and QA F1 (14.42) among Qwen-3-8b configurations, demonstrating that our approach benefits long- context processing on unseen data. GPT-4.1-mini with full context provides a strong upper bound (PRISMA 28.49, NLI 72.67), though Qwen-3-8b with guided reasoning achieves competitive or su- perior entity coverage (EntMent: 36.84 vs. 28.49), confirming that structured traces improve ground- ing even against larger models. 6 Conclusion We address the challenge of generating accurate character descriptions from book-length narratives by proposing a modular framework that decouples reasoning from generation. Our experiments reveal that enabling the built-in reasoning mode of current LLMs often degrades performance on character de- scription generation, with empty reasoning traces yielding more faithful outputs across multiple long- context settings. To address this, we introduce QA- guided reasoning, a two-stage approach consisting of (1) a reasoning model that generates structured question-answer traces capturing salient character information, and (2) a generation model that condi- tions on these traces to produce final descriptions. A key advantage of our approach is its training strategy: the reasoning model is optimized directly with Group Relative Policy Optimization on gold- standard character descriptions and automatically derived QA traces, eliminating the need to define rewards over the final generated text, which is a particularly challenging problem for open-ended generation tasks. The generation model can be trained with supervised fine-tuning by injecting reasoning traces between thinking tokens. Exper- iments on two datasets (BookWorm and CroSS) demonstrate that QA-guided reasoning improves faithfulness and informativeness across multiple evaluation metrics. Our analysis shows that trace quality directly impacts downstream description quality, with GRPO-optimized traces yielding the best performance. The approach also transfers ef- fectively to out-of-distribution data, suggesting that the learned reasoning patterns generalize beyond the training distribution. Future work should explore alternative reasoning structures beyond QA pairs, investigate other rein- forcement learning algorithms and reward formu- lations, and develop character-specific evaluation metrics that better capture narrative understanding. Limitations Our evaluation relies on automatic metrics fol- lowing established practices in prior work. We employ question-answering and fact-based met- rics (PRISMA) to assess descriptions against gold standards, entailment-based metrics (NLI) to mea- sure grounding in the input story, and standard surface-level metrics such as entity-mention F1 and Rouge-L. While these metrics provide useful sig- nals, they have important limitations for evaluating long-context generation. Human evaluation, though more reliable, remains prohibitively expensive and difficult to scale to the thousands of examples re- quired for robust assessment. Automatic metrics, conversely, struggle to capture nuanced aspects of character understanding, such as narrative coher- ence, implicit trait inference, and the integration of evidence across distant spans. Future work should develop character-specific evaluation frameworks that better capture narrative understanding and can be applied at scale. Our training approach uses GRPO to optimize the reasoning model. While GRPO is computa- tionally efficient and performs well in our exper- iments, other reinforcement learning algorithms (e.g., PPO, DPO) or alternative reward formula- tions may yield further improvements. We leave exploration of these alternatives to future work. Finally, our approach requires silver-standard QA traces derived from gold descriptions during training. In settings where high-quality character descriptions are unavailable, alternative supervi- sion strategies (e.g., weak supervision from plot summaries or character wikis) may be necessary. Investigating such strategies would broaden the ap- plicability of our framework. Acknowledgments This work was supported in part by the UKRI Cen- tre for Doctoral Training in Natural Language Pro- cessing, funded by the UKRI (grant EP/S022481/1) and the University of Edinburgh, School of Infor- matics and School of Philosophy, Psychology & Language Sciences. Computing resources were pro- vided by the Edinburgh International Data Facility (EIDF) and the Data-Driven Innovation Programme at the University of Edinburgh. Access to EIDF was facilitated through the University of Edinburgh’s Generative AI Laboratory GAIL Fellow scheme. References David Bamman, Brendan O’Connor, and Noah A. Smith. 2013. Learning latent personas of film characters. In Proceedings of the 51st Annual Meeting of the Associ- ation for Computational Linguistics (Volume 1: Long Papers), pages 352–361, Sofia, Bulgaria. Association for Computational Linguistics. Sabyasachee Baruah and Shrikanth Narayanan. 2023. Character coreference resolution in movie screen- plays. In Findings of the Association for Computa- tional Linguistics: ACL 2023, pages 10300–10313, Toronto, Canada. Association for Computational Lin- guistics. Sabyasachee Baruah and Shrikanth Narayanan. 2025. CHATTER: A character-attribution dataset for nar- rative understanding. In Proceedings of the The 7th Workshop on Narrative Understanding, pages 52–63, Albuquerque, New Mexico. Association for Compu- tational Linguistics. Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150. Faeze Brahman, Meng Huang, Oyvind Tafjord, Chao Zhao, Mrinmaya Sachan, and Snigdha Chaturvedi. 2021. “let your characters tell their story”: A dataset for character-centric narrative understanding. In Findings of the Association for Computational Lin- guistics: EMNLP 2021, pages 1734–1752, Punta Cana, Dominican Republic. Association for Com- putational Linguistics. Yapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh, Chuan Li, Chris Tanner, and Mohit Iyyer. 2025. Bleuberi: Bleu is a surprisingly effective reward for instruction following. arXiv preprint arXiv:2505.11080. Yapei Chang, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2023. Booookscore: A systematic exploration of book-length summarization in the era of llms. In In- ternational Conference on Learning Representations. Snigdha Chaturvedi, Shashank Srivastava, Hal Daume I, and Chris Dyer. 2015.Modeling dynamic relationships between characters in literary novels. arXiv preprint arXiv:1511.09376. Nuo Chen, Yan Wang, Haiyun Jiang, Deng Cai, Yuhan Li, Ziyang Chen, Longyue Wang, and Jia Li. 2023. Large language models meet harry potter: A dataset for aligning dialogue agents with characters. In Find- ings of the Association for Computational Linguistics: EMNLP 2023, pages 8506–8520, Singapore. Associ- ation for Computational Linguistics. Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022.Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in neural information processing systems, 35:16344–16359. Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth. 2021. Towards question-answering as an automatic metric for evaluating the content quality of a sum- mary. Transactions of the Association for Computa- tional Linguistics, 9:774–789. Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. 2024. Longrope: Extending llm context window beyond 2 million tokens. Preprint, arXiv:2402.13753. Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hi- erarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics. Yao Fu. 2024. Challenges in deploying long-context transformers: A theoretical peak performance analy- sis. arXiv preprint arXiv:2405.08944. Tianyu Gao, Alexander Wettig, Howard Yen, and Danqi Chen. 2025. How to train long-context language models (effectively). In Proceedings of the 63rd Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), pages 7376–7399, Vienna, Austria. Association for Compu- tational Linguistics. Alexander Gurung and Mirella Lapata. 2024. CHIRON: Rich character representations in long-form narra- tives. In Findings of the Association for Computa- tional Linguistics: EMNLP 2024, pages 8523–8547, Miami, Florida, USA. Association for Computational Linguistics. Alexander Gurung and Mirella Lapata. 2025. Learn- ing to reason for long-form story generation. arXiv preprint arXiv:2503.22828. Daniel Han, Michael Han, and Unsloth Team. 2023. Unsloth. Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023. DeBERTav3: Improving deBERTa using ELECTRA- style pre-training with gradient-disentangled embed- ding sharing. In The Eleventh International Confer- ence on Learning Representations. Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, and 1 others. 2022. Lora: Low-rank adaptation of large language models. ICLR, 1(2):3. Jian Hu, Xibin Wu, Zilin Zhu, Weixun Wang, De- hao Zhang, Yu Cao, and 1 others. 2024. Openrlhf: An easy-to-use, scalable and high-performance rlhf framework. arXiv preprint arXiv:2405.11143. Fantine Huot, Joshua Maynez, Shashi Narayan, Reinald Kim Amplayo, Kuzman Ganchev, An- nie Priyadarshini Louis, Anders Sandholm, Dipanjan Das, and Mirella Lapata. 2023. Text-blueprint: An interactive platform for plan-based conditional gen- eration. In Proceedings of the 17th Conference of the European Chapter of the Association for Compu- tational Linguistics: System Demonstrations, pages 105–116, Dubrovnik, Croatia. Association for Com- putational Linguistics. Naoya Inoue, Charuta Pethe, Allen Kim, and Steven Skiena. 2022. Learning and evaluating character representations in novels. In Findings of the Asso- ciation for Computational Linguistics: ACL 2022, pages 1008–1019, Dublin, Ireland. Association for Computational Linguistics. Mohit Iyyer, Anupam Guha, Snigdha Chaturvedi, Jor- dan Boyd-Graber, and Hal Daumé I. 2016. Feuding families and former Friends: Unsupervised learning for dynamic fictional relationships. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Hu- man Language Technologies, pages 1534–1544, San Diego, California. Association for Computational Linguistics. Evgeny Kim and Roman Klinger. 2019. Frowning Frodo, wincing Leia, and a seriously great friend- ship: Learning to classify emotional relationships of fictional characters. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 647–653, Minneapolis, Minnesota. Association for Computational Linguistics. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serv- ing with pagedattention. In Proceedings of the 29th symposium on operating systems principles, pages 611–626. Vincent Labatut and Xavier Bost. 2019. Extraction and analysis of fictional character networks: A survey. ACM Comput. Surv., 52(5). Kunze Li and Yu Zhang. 2024. Planning first, ques- tion second: An LLM-guided method for control- lable question generation. In Findings of the Asso- ciation for Computational Linguistics: ACL 2024, pages 4715–4729, Bangkok, Thailand. Association for Computational Linguistics. Tianle Li, Ge Zhang, Quy Duc Do, Xiang Yue, and Wenhu Chen. 2024.Long-context llms strug- gle with long in-context learning. arXiv preprint arXiv:2404.02060. Chin-Yew Lin. 2004. ROUGE: A package for auto- matic evaluation of summaries. In Text Summariza- tion Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics. Danyang Liu, Juntao Li, Meng-Hsuan Yu, Ziming Huang, Gongshen Liu, Dongyan Zhao, and Rui Yan. 2020. A character-centric neural model for auto- mated story generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 1725–1732. Dongqi Liu, Xi Yu, Vera Demberg, and Mirella Lapata. 2025. Explanatory summarization with discourse- driven planning. Transactions of the Association for Computational Linguistics, 13:1146–1170. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paran- jape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language mod- els use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173. Louis Mahon and Mirella Lapata. 2024. A modular ap- proach for multimodal summarization of TV shows. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8272–8291, Bangkok, Thailand. Association for Computational Linguistics. Giuliano Martinelli, Tommaso Bonomo, Pere-Lluís Huguet Cabot, and Roberto Navigli. 2025. BOOK- COREF: Coreference resolution at book scale. In Proceedings of the 63rd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 24526–24544, Vienna, Austria. Association for Computational Linguistics. Shashi Narayan, Joshua Maynez, Reinald Kim Am- playo, Kuzman Ganchev, Annie Louis, Fantine Huot, Anders Sandholm, Dipanjan Das, and Mirella Lap- ata. 2023. Conditional generation with a question- answering blueprint. Transactions of the Association for Computational Linguistics, 11:974–996. Kimia Noorbakhsh, Joseph Chandler, Pantea Karimi, Mohammad Alizadeh, and Hari Balakrishnan. 2025. Savaal: Scalable concept-driven question genera- tion to enhance human learning. arXiv preprint arXiv:2502.12477. Eric W Noreen. 1989. Computer intensive methods for hypothesis testing: An introduction. Wiley, New York, 19:21. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow in- structions with human feedback. Advances in neural information processing systems, 35:27730–27744. Argyrios Papoudakis, Mirella Lapata, and Frank Keller. 2024. BookWorm: A dataset for character descrip- tion and analysis. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 4471–4500, Miami, Florida, USA. Association for Computational Linguistics. Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable ques- tions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Lin- guistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia. Association for Computational Linguistics. David Rein, Betty Li Hou, Asa Cooper Stickland, Jack- son Petty, Richard Yuanzhe Pang, Julien Dirani, Ju- lian Michael, and Samuel R Bowman. 2024. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling. Yisi Sang, Xiangyang Mou, Mo Yu, Dakuo Wang, Jing Li, and Jeffrey Stanton. 2022. MBTI personality pre- diction for fictional characters using movie scripts. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 6715–6724, Abu Dhabi, United Arab Emirates. Association for Com- putational Linguistics. Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. Character-LLM: A trainable agent for role- playing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Process- ing, pages 13153–13187, Singapore. Association for Computational Linguistics. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. Parshin Shojaee, Aneesh Jain, Sindhu Tipirneni, and Chandan K Reddy. 2023. Execution-based code gen- eration using deep reinforcement learning. arXiv preprint arXiv:2301.13816. Marcin Skowron, Martin Trapp, Sabine Payr, and Robert Trappl. 2016. Automatic identification of character types from film dialogs. Applied Artificial Intelli- gence, 30(10):942–973. Yixiao Song, Yekyung Kim, and Mohit Iyyer. 2024. VeriScore: Evaluating the factuality of verifiable claims in long-form text generation. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 9447–9474, Miami, Florida, USA. Association for Computational Linguistics. Liyan Tang, Philippe Laban, and Greg Durrett. 2024. MiniCheck: Efficient fact-checking of LLMs on grounding documents. In Proceedings of the 2024 Conference on Empirical Methods in Natural Lan- guage Processing, pages 8818–8847, Miami, Florida, USA. Association for Computational Linguistics. Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, and 1 others. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, and 1 others. 2025. Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Runchu Tian, Yanghao Li, Yuepeng Fu, Siyang Deng, Qinyu Luo, Cheng Qian, Shuo Wang, Xin Cong, Zhong Zhang, Yesai Wu, Yankai Lin, Huadong Wang, and Xiaojiang Liu. 2025. Distance between relevant information pieces causes bias in long-context LLMs. In Findings of the Association for Computational Lin- guistics: ACL 2025, pages 521–533, Vienna, Austria. Association for Computational Linguistics. Jonathan Uesato, Nate Kushman, Ramana Kumar, Fran- cis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. 2022. Solv- ing math word problems with process-and outcome- based feedback. arXiv preprint arXiv:2211.14275. Krishnapriya Vishnubhotla, Adam Hammond, Graeme Hirst, and Saif Mohammad. 2024. The emotion dy- namics of literary novels. In Findings of the Asso- ciation for Computational Linguistics: ACL 2024, pages 2557–2574, Bangkok, Thailand. Association for Computational Linguistics. Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, An- drew M Dai, and Quoc V Le. 2021. Finetuned lan- guage models are zero-shot learners. arXiv preprint arXiv:2109.01652. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elic- its reasoning in large language models. Advances in neural information processing systems, 35:24824– 24837. Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Sti- ennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862. Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catan- zaro. 2023. Retrieval meets long context large lan- guage models. arXiv preprint arXiv:2310.03025. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025.Qwen3 technical report.arXiv preprint arXiv:2505.09388. An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, and 1 others. 2024. Qwen2. 5-math technical report: Toward mathemat- ical expert model via self-improvement.arXiv preprint arXiv:2409.12122. Mo Yu, Jiangnan Li, Shunyu Yao, Wenjie Pang, Xi- aochen Zhou, Zhou Xiao, Fandong Meng, and Jie Zhou. 2023. Personality understanding of fictional characters during book reading. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14784–14802, Toronto, Canada. Association for Computational Linguistics. Xinfeng Yuan, Siyu Yuan, Yuhan Cui, Tianhe Lin, Xin- tao Wang, Rui Xu, Jiangjie Chen, and Deqing Yang. 2024. Evaluating character understanding of large language models via character profiling from fictional works. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8015–8036, Miami, Florida, USA. Association for Computational Linguistics. Weiwei Zhang, Jackie Chi Kit Cheung, and Joel Oren. 2019. Generating character descriptions for auto- matic summarization of fiction. In Proceedings of the AAAI Conference on Artificial Intelligence, vol- ume 33, pages 7476–7483. Xuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine, and Dawn Song. 2025. Learning to rea- son without external rewards.arXiv preprint arXiv:2505.19590. Lixing Zhu, Runcong Zhao, Lin Gui, and Yulan He. 2023. Are NLP models good at tracing thoughts: An overview of narrative understanding. In Find- ings of the Association for Computational Linguis- tics: EMNLP 2023, pages 10098–10121, Singapore. Association for Computational Linguistics. A Implementation Details Training and Inference We use the Open- RLHF (Hu et al., 2024) library for GRPO training with the hyperparameters listed in Table 5. We also employ the Qwen-3-8b model to provide LLM-as- a-judge rewards during GRPO-training. For super- vised finetuning, we use LoRA (Hu et al., 2022) from the unsloth (Han et al., 2023) library. We report the supervised finetuning hyperparameters in Table 6. For inference, we use the vLLM en- gine (Kwon et al., 2023) with sample decoding and temperature0.4. We evaluate the statistical sig- nificance of the results using two-sided paired ap- proximate randomization test (10,000 permutations andα = 0.05) (Noreen, 1989), based on the mean value across 4 inference seeds. We used four H200 Nvidia GPUs for GPRO training and a single H100 or H200 GPU for all the inference experiments and SFT training. HyperparameterValue actor learning rate5× 10 −7 kL coefficient0.01 train batch size64 samples per prompt8 prompt max length17, 684 generation max length2, 048 Table 5: Training hyperparameters for Group Relative Policy Optimization (GRPO) training. HyperparameterValue learning rate10 −6 max input length32, 768 batch size1 gradient accumulation steps8 alpha256 rank128 lora dropout0.1 Table 6: Training hyperparameters for supervised fine- tuning (SFT) training with LoRA. Prompts We provide the prompts used for our experiments in Tables 7 and 8. Evaluation We use the VeriScore (Song et al., 2024) codebase to extract facts (prompts are ad- justed based on Papoudakis et al. (2024)) for both PRISMA and NLI metrics. Fact verification is per- formed using the MiniCheck (Tang et al., 2024) model in both metrics. We use rouge-score 4 imple- mentation for Rouge-L metric and NLTK 5 toolkit 4 https://github.com/google-research/ google-research/tree/master/rouge 5 https://w.nltk.org/ Context: context Describe charactercharacterfrom bookbook based on the given context. Return your output as a single paragraph (close to length words) including the important information. Table 7: Character description generation prompt tem- plate. We replace the variable length with the average output length of the corresponding dataset. Context: context Your task is to generate questions answer pairs about character:characterfrom book:bookgiven the above chunk of the book. You should focus on under- standing aspects of the character (e.g., role, relationships, personality, events) that are mentioned in the context. Each qa-pair should be labelled as role, relationship, per- sonality, event or other. We provide the definitions for these below. Definitions: Role: defines what part the character plays in the story, narrator, major/minor character. Relationship: connections the character has with others, such as friendships or family ties. Personality: character’s behavior, traits, and attributes. Event: actions and decisions the character is involved in throughout the story. Other: any other fact that doesn’t belong to the above categories. Output format: Q1: <question> E1: <explanation> A1: <answer> T1: <type> Q2: <question> E2: <explanation> A2: <answer> T2: <type> ... Generate an explanation, 1-2 sentences that fully justify your answer, do not simply repeat the answer. Type of qa has to be one of Role, Relationship, Personality, Event or Other. The answer should be short 1-4 words. Generate QA-pairs only related to character:character. Gen- erate QA-pairs only for information mentioned in the provided context. Do not include unanswered questions. The questions has to mention the name of the character: character. Do not generate repetitive QA-pairs with same answer. If the character is not mentioned, simply return None. Table 8: QA generation prompt template. We adopt the definitions for the different categories from Papoudakis et al. (2024) for named-entity extraction in entity-mention F1 calculation. B Additional Experiments and Statistics QA reasoning ablationWe run an ablation study to demonstrate the effect of the different compo- nents in the reasoning trace. We use the Qwen-3-8B QA ReasoningDescription Method#QAPRF1PRI QAEnt guided-QA8.60 15.07 17.68 16.2714.45 14.45 30.41 w/o expl.12.05 13.89 16.93 15.2614.34 15.01 29.94 w/o types11.44 12.46 15.87 13.9513.41 14.34 30.37 w/o expl., types15.46 13.90 12.16 13.6113.71 14.92 30.08 w/o expl., ans., types 13.44 —14.02 14.22 29.62 Table 9: Ablation of QA reasoning components using Qwen-3-8b with coreference retrieval on BookWorm validation set (zero-shot). P/R/F1 = reasoning preci- sion/recall/F1; PRI = PRISMA F1; Ent = EntMent F1. model as a backbone with a coreference-based re- trieval method to extract context and employ the guided reasoning method in zero-shot mode. Table 9 examines the contribution of each com- ponent in the QA reasoning trace: explanations, question types, and answers. The full Guided QA method achieves the best PRISMA (14.45) and EntMent (30.41), indicating that all components contribute to faithful and entity-rich descriptions. Removing explanations (w/o expl.) increases the number of generated QA pairs (8.60→12.05) and improves QA F1 (14.45→15.01), but reduces pre- cision (15.07→13.89) and recall (17.68→16.93), suggesting that explanations help maintain infor- mative and precise reasoning traces. Removing question types (w/o types) causes a drop in reasoning F1 (16.27→13.95) and PRISMA (14.45→13.41), indicating that typed questions encourage topical diversity and more in- formative traces. Removing both explanations and types (w/o expl., types) generates the most QA pairs (15.46) but with lower precision (13.90) and recall (12.16), yielding mixed downstream results. Finally, retaining only questions (w/o expl., ans., types) prevents reasoning evaluation (no answers to verify) and degrades PRISMA (14.02), confirm- ing that answers are essential for grounding the generation model. Overall, each component serves a distinct role: types encourage diversity, explana- tions improve precision, and answers provide the factual content that guides description generation. Reasoning Statistics Table 10 reports summary statistics for intermediate reasoning traces and the final character descriptions when conditioning on the retrieved context in BookWorm. Overall, the description outputs remain similar in length across settings (approximately 109–118 tokens) and ex- hibit comparable lexical diversity (about 71–73% unique unigrams) for the no-reasoning baseline and both guided-QA variants. This suggests that per- ReasoningDescription Method#QATokUniTokUni None—11171 built-in—2706010978 guided-QA (ZS)839136 11873 guided-QA (GRPO)115353111171 Table 10: Statistics for generated traces and their corre- sponding descriptions for different reasoning methods using retrieved context on BookWorm validation set. #QA = QA pairs; Tok = avg. tokens; Uni = unique uni- grams (%). formance differences in our main experiments are unlikely to be driven by longer or more lexically diverse descriptions, but instead by the content and structure of the intermediate reasoning signal. In contrast, the reasoning traces differ substan- tially. Built-in reasoning produces moderately long traces (270 tokens) with higher unigram diversity than guided-QA, whereas guided-QA traces are much longer (391–535 tokens) but show markedly lower unigram diversity (31–36%), which is largely attributable to the structured QA format (e.g., re- peated question templates and answer markers). Comparing the two guided-QA variants, GRPO yields longer traces (11 QA pairs; 535 tokens) than the zero-shot reasoner (8 QA pairs; 391 tokens), while maintaining similar diversity, suggesting that optimization encourages the model to select more, higher-yield QA pairs but without increasing ver- bosity. Finally, built-in reasoning leads to the high- est description unigram diversity (78%) despite lower faithfulness in our main results, consistent with the interpretation that unstructured delibera- tion may introduce additional (and potentially un- supported) details rather than improving ground- ing. Experiments with Gemma-3 We further evalu- ate our approach using the Gemma-3-12b-it (Team et al., 2025) model for coreference-based retrieval and Lead-128k methods on the BookWorm dataset. We report a reasoning experiment in which the model is prompted to explicitly reason, generating a chain-of-thought inside<think>· </think> tokens before producing the character description. Although the Gemma-3 model is not trained to explicitly reason, we found that it consistently follows the reasoning format in all the experi- ments. The results in Table 11 demonstrate that explicit reasoning harms performance, for exam- ple, both PRISMA (11.76→11.16) and QA-F1 MethodPRISMAQAEntMentR-L Coref-32k13.6512.4927.9316.20 + reasoning11.9011.1125.6015.49 + guided-QA13.1013.6028.8216.61 Lead-128k11.7611.7826.3915.99 + reasoning11.1610.6725.6615.16 + guided-QA13.1213.3229.3416.62 Table 11: Results on BookWorm dataset using Gemma- 12b-it model. Zero-shot experiments without any rea- soning trace, explicitly prompting the model to reason and QA-guided reasoning for coref-based retrieval and Lead-128k methods. (11.78→10.67) degrade for Lead-128k. In con- trast, zero-shot QA-guided reasoning in the coref- erence method improves performance over the corresponding approaches in all metrics except PRISMA, achieving also the highest QA-F1 score (13.60). Similarly, QA-guided reasoning for Lead- 128k improves over the corresponding baselines across all metrics, achieving the highest EntMent (29.34) and R-L (16.62). Overall, these results indi- cate that findings from the main experiments also apply to a model from a different family, which is not trained for explicit reasoning, and that the proposed QA-guided method can also be effective under this setting. QA generation methods We compare different question generation methods, evaluating the cor- responding reasoning traces. Precisely, we tested Savaal (Noorbakhsh et al., 2025), a multi-step topic-driven approach for QA generation in long documents. This approach first extracts topics from each chunk of the input and then uses each topic to retrieve relevant passages. The retrieved passages and the corresponding topic are finally used to gen- erate QA-pairs. Next, we evaluate a plan-based QA generation method based on PFQS (planning first, question second) (Li and Zhang, 2024). We prompt a model to extract sentences from the given input that serve as candidate evidence for answers. The extracted sentences are provided together with the input con- text to generate QA pairs. We evaluate both of these approaches using the same output format, adding an explanation and type to each QA-pair (see Ta- ble 8). We also split the input context into chunks of 16k tokens. Table 12 shows the evaluation of the different methods and their average number of questions. The guided-QA method yields the highest preci- Method# QAPrecisionRecallF1 guided-QA8.6017.6815.0716.27 Savaal12.0115.2313.9214.54 Plan-based QA11.9017.3411.7113.98 Table 12: Experiments with different QA generation methods using coref-based retrieval on BookWorm dataset. We report the average number of QA-pairs along with precision, recall and F1 score for reason- ing evaluation. sion, recall, and F1 (16.27) despite having the lowest number of QA pairs. Savaal method ob- tains a moderate F1 score but performs worse than guided-QA, while plan-based generation has com- petitive precision, but the recall is significantly lower (11.71). Overall, these results demonstrate that simple chunk-based question generation can outperform topic- and plan-based methods, yield- ing more precise and informative QA-pairs even with fewer generated questions. C Examples We show examples of generated outputs and rea- soning traces in Tables 13, 14, 15 and 16. Character: Demon CopperheadBook: Demon Copperhead Lead-128k: Demon Copperhead is a complex and resilient character shaped by a life of hardship, neglect, and survival. Born to a drug-addicted mother, he was abandoned in a trailer park in Lee County, Virginia, and nearly died at birth due to his mother’s drug use, which led to his being born in the caul—a rare occurrence that his mother claimed was a sign of her own superstitions. He was taken in by the Peggots, who were not his real family but became his de facto parents, giving him a sense of belonging and a nickname, "Demon," that stuck with him as he grew. His early life was marked by instability, including being sent to live with the brutal Crickson family on a farm, where he endured harsh treatment and emotional abuse. He also experienced foster care with Miss Barks, who became a maternal figure but eventually left for a better-paying job as a teacher. Demon’s life took a tragic turn when his mother died in an overdose, leaving him in foster care again and further alienated. He later ran away to Nashville, where he encountered a series of misfortunes, including theft, abuse, and being caught in a pawnshop with a hooker. Eventually, he found refuge with his grandmother Betsy Woodall, who had long been estranged from him but now took him in after learning of his mother’s death. Demon is marked by his red hair and blue eyes, traits he inherited from his father—a Melungeon man he never knew but who left an indelible mark on his identity. He is also defined by his survival instincts, having learned to navigate the harsh realities of poverty, neglect, and abuse. Lead-128k + guided-QA (GRPO): Demon Copperhead, the protagonist of the novel, is a resilient and determined young man who navigates a life marked by hardship, neglect, and a complex web of relationships. Born to an addicted mother, he faces a difficult start in life, including a traumatic birth and early experiences with foster care. His early years are shaped by the Peggots, who become a significant part of his life, though their relationship is strained and complicated. He develops a close friendship with Maggot, another foster child, and later forms a bond with Fast Forward, who becomes a mentor and friend. Demon’s nickname "Demon" is given to him by others, reflecting his troubled past and the stigma he carries. His relationship with Stoner, his stepfather, is fraught with tension and abuse, leading to a confrontation that results in Demon’s injury and eventual placement in foster care. He is placed with Mr. Crickson, a harsh farmer who runs a slave-boy farm, but eventually finds a more supportive home with Coach Winfield and his family in Jonesville. Demon’s personality is marked by his determination to prove his worth despite his circumstances, as well as his artistic talent for drawing superheroes and comic strips. He has a complex relationship with his biological mother, who dies from an overdose, and his biological father, whose death he never fully understands. His journey includes running away from home, stealing a tree for Christmas, and eventually finding love and support with Dori, who becomes his girlfriend and later his fiancée. Demon’s story is one of resilience, as he struggles with addiction but ultimately seeks redemption through his art and relationships. His character is defined by his search for belonging and self-worth in a world that often overlooks him. Goldstandard Description: Demon Copperhead is the eponymous, first-person narrator and protagonist of the novel. Because the novel’s narrative voice deals with both the events of Demon’s life and his reflection on those events, the novel is structured like a memoir. It is also a bildungsroman in which Demon endures several life-altering conflicts and eventually finds peace with himself. The world has been cruel to Demon, and his early life demonstrates The Failure of Society to Protect Its Children. He was born without a father, and his mother struggled with addiction throughout her short life. He navigates several different forms of false families, such as life with his abusive stepfather, placements in two abusive foster homes, and rejection from his father’s mother. The lack of love in Demon’s childhood and early adolescence makes him radically self-reliant, but it also makes him distrustful of the people who do provide him with security and care. Demon is independent because he must be to survive, but he looks for people to believe in—often misguidedly, as with Fast Forward and Dori. Similarly, he has an attachment to Lee County because it is a space he at least understands, even if it keeps bringing him down. Demon eventually learns that he must focus on himself to receive real love from other people. [truncated] FactLead-128kLead-128k guided-QA Demon Copperhead is the eponymous narrator of the novel.FalseTrue Demon Copperhead is the first-person narrator of the novel.FalseFalse Demon Copperhead is the protagonist of the novel.FalseTrue The narrative voice of the novel "Demon Copperhead" deals with the events of Demon’s life. FalseTrue The narrative voice of the novel "Demon Copperhead" includes Demon’s reflection on those events. FalseFalse The novel "Demon Copperhead" is structured like a memoir.FalseFalse "Demon Copperhead" is a bildungsroman.FalseTrue In "Demon Copperhead," Demon endures several life-altering conflicts.TrueTrue In "Demon Copperhead," Demon eventually finds peace with himself.FalseFalse The world has been cruel to Demon Copperhead.TrueTrue Demon’s early life demonstrates The Failure of Society to Protect Its Children.TrueTrue Demon was born without a father.TrueFalse Demon’s mother struggled with addiction throughout her short life.FalseTrue Demon navigates several different forms of false families.TrueTrue Demon lived with an abusive stepfather.FalseTrue Demon was placed in two abusive foster homes.TrueTrue Demon experienced rejection from his father’s mother.FalseFalse The lack of love in Demon’s childhood made him radically self-reliant.TrueTrue [Facts truncated] Table 13: Example of generated descriptions in CroSS dataset with QA-guided Lead-128k and the corresponding baseline. We demonstrate facts evaluated based on the PRISMA-recall metric. Character: Demon CopperheadBook: Demon Copperhead Reasoning for 1st Chunk: Q1: What is Demon Copperhead’s role in the story? E1: Demon Copperhead is the protagonist and narrator of the story, experiencing various challenges and relationships throughout his life. A1: Protagonist T1: Role Q2: Who is Demon Copperhead’s mother? E2: Demon Copperhead’s mother is a central character in the story, struggling with addiction and neglect. A2: Addicted mother T2: Relationship Q3: What is Demon Copperhead’s personality like? E3: Demon Copperhead is portrayed as a resilient and determined character, despite facing numerous hardships. A3: Resilient and determined T3: Personality Q4: What event leads to Demon Copperhead’s birth? E4: Demon Copperhead is born to a mother who is struggling with addiction, leading to a difficult start in life. A4: Difficult birth T4: Event Q5: Who is Demon Copperhead’s foster family? E5: Demon Copperhead is raised by the Peggots, who become a significant part of his life. A5: The Peggots T5: Relationship Q6: What is Demon Copperhead’s relationship with Maggot? E6: Demon Copperhead and Maggot are close friends who grow up together in the same community. A6: Close friend T6: Relationship Q7: What is Demon Copperhead’s nickname? E7: Demon Copperhead is given the nickname "Demon" by others, which becomes a part of his identity. A7: Demon T7: Personality Q8: What event leads to Demon Copperhead’s time with the Peggots? E8: Demon Copperhead spends time with the Peggots after his mother’s struggles with addiction and his own difficult circum- stances. A8: Time with Peggots T8: Event Q9: What is Demon Copperhead’s relationship with Stoner? E9: Demon Copperhead has a complex relationship with Stoner, who becomes his stepfather and has a significant impact on his life. A9: Stepfather T9: Relationship Q10: What event leads to Demon Copperhead’s conflict with Stoner? E10: Demon Copperhead conflicts with Stoner due to Stoner’s strict parenting style and his own struggles with identity and self-worth. A10: Conflict with Stoner T10: Event Reasoning for 2st Chunk: Q11: What event leads to Demon Copperhead being placed in foster care? E11: Demon Copperhead’s mother becomes addicted to drugs and is involved in a violent relationship, leading to his placement in foster care. A11: Mother’s addiction and abuse. T11: Event Q12: Who becomes Demon Copperhead’s foster father? E12: Demon Copperhead’s foster father is Mr. Crickson, a farmer who takes in foster children. A12: Mr. Crickson. T12: Relationship Q13: What event causes Demon Copperhead to run away from home? E13: Demon Copperhead runs away from home after a violent confrontation with his stepfather, Stoner. A13: Confrontation with Stoner. T13: Event Q14: What is Demon Copperhead’s attitude towards his foster family? E14: Demon Copperhead has a complex attitude towards his foster family, initially feeling out of place but gradually forming bonds with some members. A14: Complex attitude. T14: Personality Q15: What event leads to Demon Copperhead being placed in a temporary foster home? E15: Demon Copperhead is placed in a temporary foster home after his mother is hospitalized and unable to care for him. A15: Mother’s hospitalization. T15: Event [Trace truncated] Table 14: Example of generated reasoning trace in CroSS dataset. The trace is from the GRPO-trained Qwen-3-8b reasoning model using Lead-128k. Character: John MontravilleBook: Charlotte Temple Lead-128k: John Montraville is a young, impetuous soldier who becomes infatuated with Charlotte Temple, a pure and innocent girl, and secretly courts her despite his duty and honor. He is torn between his love for Charlotte and his sense of responsibility, but his passion leads him to elope with her, only to discover she was deceived by Mademoiselle La Rue. Montraville’s heart breaks upon learning of Charlotte’s betrayal and her subsequent suffering, ultimately leading to his tragic downfall and death, which he brings upon himself in a fit of vengeance against Belcour. Fact John Montraville is a young soldier.True John Montraville is an impetuous soldier.False John Montraville becomes infatuated with Charlotte Temple.True Charlotte Temple is a pure and innocent girl.False John Montraville secretly courts Charlotte Temple despite his duty.False John Montraville secretly courts Charlotte Temple despite his honor.False John Montraville is torn between his love for Charlotte Temple and his sense of responsi- bility. False John Montraville’s passion leads him to elope with Charlotte Temple.False John Montraville discovers that Charlotte Temple was deceived by Mademoiselle La Rue. False Montraville’s heart broke upon learning of Charlotte’s betrayal.False Montraville’s heart broke upon learning of Charlotte’s subsequent suffering.False Montraville’s heart breaking led to his tragic downfall.False Montraville’s tragic downfall led to his death.False Montraville brought his death upon himself in a fit of vengeance against Belcour.False Lead-128k + guided-QA: John Montraville is a Lieutenant in the army, an ardent and impetuous soldier who becomes infatuated with Charlotte Temple, a young, innocent girl he meets at a boarding school. He secretly courts her, leading her to elope with him, but later discovers her relationship with Belcour, his friend, which causes his heartbreak. Despite his initial honor and compassion, he succumbs to passion, abandons Charlotte, and marries Julia Franklin. His guilt over betraying Charlotte and the death of her child lead to a tragic end, as he kills Belcour in a fit of rage and dies shortly after, haunted by remorse. Fact John Montraville is a Lieutenant in the army.False John Montraville is an ardent soldier.False John Montraville is an impetuous soldier.False John Montraville became infatuated with Charlotte Temple.True Charlotte Temple is a young and innocent girl.False John Montraville met Charlotte Temple at a boarding school.False John Montraville secretly courted Charlotte Temple.False John Montraville led Charlotte Temple to elope with him.False John Montraville later discovered Charlotte Temple’s relationship with Belcour.False Belcour is a friend of John Montraville.True John Montraville’s discovery of Charlotte Temple’s relationship with Belcour caused him heartbreak. False John Montraville initially displayed honor and compassion.False John Montraville succumbed to passion.True John Montraville abandoned Charlotte Temple.True John Montraville married Julia Franklin.True John Montraville felt guilt over betraying Charlotte Temple.True John Montraville felt guilt over the death of Charlotte Temple’s child.False John Montraville killed Belcour in a fit of rage.False John Montraville died shortly after killing Belcour.False John Montraville was haunted by remorse after his actions.True Goldstandard Description: A charming, dashing young officer in the British Army, John Montraville falls in love with Charlotte Temple in Chichester and hopes to marry her. However, his father refuses to endorse a marriage to a poor woman, so Montraville makes Charlotte his mistress and brings her to America with him. He tries to treat her honorably but is easily deceived by his friend Belcour, who ruins his relationship with Charlotte so he can have her as his own mistress. Despite his inability to control his womanizing impulses, Montraville feels guilty about the misfortune he causes Charlotte and others. Table 15: Examples of generated descriptions in BookWorm dataset with QA-guided Lead-128k and the correspond- ing baseline. We demonstrate the facts evaluated based on PRISMA-precision metric. Character: John MontravilleBook: Charlotte Temple Reasoning for 1st Chunk: Q1: What role does John Montraville play in the story? E1: John Montraville is a major character in the story, a young soldier who becomes involved with Charlotte Temple. A1: Major character T1: Role Q2: Who is John Montraville related to in the story? E2: John Montraville is related to Charlotte Temple, as he becomes romantically involved with her. A2: Charlotte Temple T2: Relationship Q3: What is John Montraville’s personality like? E3: John Montraville is portrayed as a tender, eloquent, and ardent lover, but also impetuous and driven by his passions. A3: Ardent and impetuous T3: Personality Q4: What event leads John Montraville to meet Charlotte Temple? E4: John Montraville meets Charlotte Temple when he encounters her and her governess at a church in Chichester. A4: Meeting Charlotte Temple T4: Event Q5: What decision does John Montraville make regarding his relationship with Charlotte Temple? E5: John Montraville decides to elope with Charlotte Temple, despite knowing their families would not approve. A5: Decide to elope T5: Event Q6: What is John Montraville’s profession? E6: John Montraville is a soldier, as he is mentioned as a Lieutenant in the army. A6: Soldier T6: Role Q7: What is John Montraville’s relationship with Belcour? E7: John Montraville is friends with Belcour, who accompanies him on his visits to Charlotte Temple. A7: Friend T7: Relationship Q8: What is John Montraville’s attitude towards marriage? E8: John Montraville is initially cautious about marriage, as his father warns him against rushing into it. A8: Cautious T8: Personality Q9: What is one of the events involving John Montraville and Charlotte Temple? E9: One event involving John Montraville and Charlotte Temple is their clandestine meetings at a summer-house. A9: Clandestine meetings T9: Event Q10: What is John Montraville’s relationship with his father? E10: John Montraville has a respectful relationship with his father, who gives him advice about marriage and duty. A10: Respectful T10: Relationship Reasoning for 2st Chunk: Q11: Who is John Montraville romantically involved with? E11: John Montraville is romantically involved with Charlotte Temple and Julia Franklin. A11: Charlotte Temple T11: Relationship Q12: What event leads to John Montraville’s conflict with Charlotte Temple? E12: John Montraville’s conflict with Charlotte Temple arises from her elopement and his subsequent discovery of her relationship with Belcour. A12: Elopement T12: Event Q13: What event leads to John Montraville’s conflict with Julia Franklin? E13: John Montraville’s conflict with Julia Franklin arises from his indecision about marrying her while still being involved with Charlotte Temple. A13: Indecision T13: Event Q14: What event leads to John Montraville’s discovery of Charlotte’s infidelity? E14: John Montraville discovers Charlotte’s infidelity when he finds her in bed with Belcour. A14: Discovery T14: Event Q15: What is John Montraville’s role in the story’s resolution? E15: John Montraville plays a significant role in the story’s resolution by ultimately leaving Charlotte and pursuing a relationship with Julia Franklin. A15: Pursuing relationship T15: Event Q16: What is John Montraville’s personality trait regarding honor? E16: John Montraville is portrayed as having a strong sense of honor, which conflicts with his romantic entanglements. A16: Strong sense of honor T16: Personality [Trace truncated] Table 16: Example of generated reasoning trace in BookWorm dataset. The trace is from the GRPO-trained Qwen-3- 8b reasoning model with Lead-128k as input context.