Paper deep dive
Leveraging Weighted Syntactic and Semantic Context Assessment Summary (wSSAS) Towards Text Categorization Using LLMs
Shreeya Verma Kathuria, Nitin Mayande, Sharookh Daruwalla, Nitin Joglekar, Charles Weber
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/15/2026, 1:21:30 AM
Summary
The paper introduces the Weighted Syntactic and Semantic Context Assessment Summary (wSSAS), a deterministic framework designed to improve the reliability and accuracy of Large Language Models (LLMs) in text categorization. By organizing data into a hierarchy of Themes, Stories, and Clusters and applying a Signal-to-Noise Ratio (SNR) to filter input, wSSAS mitigates the stochastic nature of attention mechanisms, ensuring LLMs focus on high-value semantic features.
Entities (5)
Relation Signals (3)
wSSAS → buildsupon → SSAS
confidence 100% · The Weighted SSAS (wSSAS) methodology builds upon the foundational SSAS methodology
wSSAS → utilizes → Signal-to-Noise Ratio
confidence 95% · It then leverages a Signal-to-Noise Ratio (SNR) to prioritize high-value semantic features
wSSAS → improves → Gemini 2.0 Flash Lite
confidence 90% · Experimental results using Gemini 2.0 Flash Lite... demonstrate that wSSAS significantly improves clustering integrity
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The use of Large Language Models (LLMs) for reliable, enterprise-grade analytics such as text categorization is often hindered by the stochastic nature of attention mechanisms and sensitivity to noise that compromise their analytical precision and reproducibility. To address these technical frictions, this paper introduces the Weighted Syntactic and Semantic Context Assessment Summary (wSSAS), a deterministic framework designed to enforce data integrity on large-scale, chaotic datasets. We propose a two-phased validation framework that first organizes raw text into a hierarchical classification structure containing Themes, Stories, and Clusters. It then leverages a Signal-to-Noise Ratio (SNR) to prioritize high-value semantic features, ensuring the model's attention remains focused on the most representative data points. By incorporating this scoring mechanism into a Summary-of-Summaries (SoS) architecture, the framework effectively isolates essential information and mitigates background noise during data aggregation. Experimental results using Gemini 2.0 Flash Lite across diverse datasets - including Google Business reviews, Amazon Product reviews, and Goodreads Book reviews - demonstrate that wSSAS significantly improves clustering integrity and categorization accuracy. Our findings indicate that wSSAS reduces categorization entropy and provides a reproducible pathway for improving LLM based summaries based on a high-precision, deterministic process for large-scale text categorization.
Tags
Links
- Source: https://arxiv.org/abs/2604.12049v1
- Canonical: https://arxiv.org/abs/2604.12049v1
Trouble viewing inline? Open PDF directly →
Full Text
105,055 characters extracted from source content.
Expand or collapse full text
LEVERAGING WEIGHTED SYNTACTIC AND SEMANTIC CONTEXT ASSESSMENT SUMMARY (WSSAS) TOWARDS TEXT CATEGORIZATION USING LLMS Shreeya Verma Kathuria ∗1 , Nitin Mayande 1 , Sharookh Daruwalla 1 , Nitin Joglekar †2 , and Charles Weber ‡2 1 Tellagence Inc. 2 Villanova School of Business, Villanova University 3 Maseeh College of Engineering and Computer Science, Portland State University ABSTRACT The use of Large Language Models (LLMs) for reliable, enterprise-grade analytics such as text categorization is often hindered by the stochastic nature of attention mechanisms and sensitivity to noise that compromise their analytical precision and reproducibility. To address these technical frictions, this paper introduces the Weighted Syntactic and Semantic Context Assessment Summary (wSSAS), a deterministic framework designed to enforce data integrity on large-scale, chaotic datasets. We propose a two-phased validation framework that first organizes raw text into a hierarchical classification structure containing Themes, Stories, and Clusters. It then leverages a Signal-to-Noise Ratio (SNR) to prioritize high-value semantic features, ensuring the model’s attention remains focused on the most representative data points. By incorporating this scoring mechanism into a Summary-of-Summaries (SoS) architecture, the framework effectively isolates essential information and mitigates background noise during data aggregation. Experimental results using Gemini 2.0 Flash Lite across diverse datasets—including Google Busi- ness reviews, Amazon Product reviews, and Goodreads Book reviews—demonstrate that wSSAS significantly improves clustering integrity and categorization accuracy. Our findings indicate that wSSAS reduces categorization entropy and provides a reproducible pathway for improving LLM based summaries based on a high-precision, deterministic process for large-scale text categorization. Keywords Natural Language Processing (NLP)· Artificial Intelligence (AI)· Text Summarization· Categorization 1 Introduction The field of text categorization and summarization has fundamentally shifted, evolving from a complex engineering challenge—which historically necessitated extensive feature engineering, massive datasets, and prolonged training [1] —into a core capability powered by Large Language Models (LLMs). This transition is underpinned by the superior semantic understanding of LLMs, enabling zero-shot and few-shot learning—the ability to categorize text with little to no prior training [2]. By replacing rigid, bespoke classifiers with fluid foundational models, LLMs have unlocked applications across high-stakes sectors, ranging from real-time misinformation detection in social media [3] to the precise organization of patient records in clinical healthcare environments [4], [5]. Despite this versatility, the path to enterprise-grade reliability remains obstructed by several technical frictions. Cur- rent LLM performance is sensitive to prompt engineering; such that minor syntactic variations in instructions can yield drastically different classification outcomes [6]. Furthermore, the inherent constraints of In-Context Learning token limits restrict the volume of reference examples a model can process [7]. At a cognitive level, LLMs struggle ∗ shreeya, nitin, sharookh@tellagence.com † nitindra.joglekar@villanova.edu ‡ webercm@pdx.edu arXiv:2604.12049v1 [cs.CL] 13 Apr 2026 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs with nuanced linguistic phenomena such as irony, intensification, and latent bias [8]. For organizations, these lim- itations—compounded by a lack of model interpretability and the scarcity of high-quality annotated data for niche domains—create a significant gap between experimental capability and production-ready accuracy. 1.1 The Paradox of LLM Creativity: Why Generative AI underperforms in Data Science The primary obstacle to utilizing LLMs for rigorous data science lies in a fundamental architectural conflict: the paradox of generative creativity. LLMs are, by design, engines of probability. While their underlying mechanics are revolutionary for creative synthesis, they are inherently poorly suited for the rigid, invariant requirements of data analytics. The root of this instability is found in the attention mechanism. In a standard generative configuration, the attention mechanism dictates which tokens the model prioritizes during processing [9]. Because these models are optimized for novelty and fluency, the mechanism may assign disparate weights to the same input tokens across successive runs [10]. This stochasticity is an asset for creative tasks, but it represents a significant liability for data science tasks where latent space stability is required to ensure data integrity. In text categorization, this creative variance manifests as decreased accuracy and poor generalization. Research suggests that the inclusion of irrelevant information within the input context can be damaging to performance, as it forces the model to attend to inconsequential patterns [11]. This creates a signal-to-noise deficit that is particularly acute in modern marketing and commercial datasets, where the sheer volume of data often buries actionable insights under layers of technical friction [12] This paper addresses the necessity for a deterministic analytical framework capable of improving LLMs from creative assistants into precise instruments of categorization and summarization. By optimizing the Input Context, we seek to bridge the gap between AI potential and analytical execution [13]. Our inquiry is guided by two pivotal research questions: 1. Dynamic Context Improvement: Can categorization accuracy be measurably enhanced by replacing static prompts with dynamically generated, custom-tailored context information for specific requests? 2.Context-Quality Correlation: Is there a quantifiable relationship between the linguistic quality of provided contextual "hints" and the resulting precision of the categorization? By identifying, isolating, and removing background noise, we aim to ensure the attention mechanism remains focused exclusively on relevant context, thereby establishing improved LLM-driven data integrity for text categorization and summarization. 1.2 Hierarchical Contextual Framework for Analytical Integrity We propose Syntactic & Semantic Attention Summarization (SSAS) [14] [15], a hierarchical contextual framework that replaces the black box unpredictability of standard LLMs with a structured methodology designed to enforce integrity on chaotic datasets [13]. This approach aligns with the Information Bottleneck (IB) principle, which suggests that an optimal model should compress the input to retain only the information most relevant to the target output [16]. The philosophy is operationalized through a specialized two-phase framework: 1. Contextual Relevance: The process begins by evaluating the data within its specific context. By identifying information relevancy at a granular level, the system determines which data points are pertinent to the defined problem and which are extraneous. 2. Noise Reduction and Reliability Improvement: Using the derived context from Phase 1, the system systemati- cally reduces dataset noise. By feeding only refined, relevant context into the LLM, we reduce the variance and significantly improve the consistency and reliability of the output This methodology refines raw, chaotic data into a reliable and analytical relevant dataset. In parallel, by narrowing the model’s focus through derived context, this framework ensures that the refined input consistently yields the same results—addressing, to a large extent, the "stochasticity" problem inherent in generative architectures. This framework mitigates operational complexity for domain experts, enabling a shift in focus from algorithmic calibration toward high-level strategic analysis [17] The Weighted SSAS (wSSAS) methodology builds upon the foundational SSAS methodology by introducing a rigorous, data-driven prioritization layer. Technical details of SSAS and wSSAS approaches are elaborated in Section 3. 2 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs 2 Related Work The emergence of Large Language Models (LLMs) has fundamentally redefined text categorization, shifting the paradigm from supervised feature engineering toward zero-shot and few-shot learning [2] [18]. However, as these models move from creative synthesis to enterprise-grade analytics, their inherent instability presents significant challenges. Our work builds upon three primary areas of research: the sensitivity of in-context learning [19], the mechanics of attention-based noise [20] [11], and hierarchical data summarization. 2.1 In-Context Learning and Prompt Instability The efficacy of Large Language Models (LLMs) in zero-shot and few-shot regimes is largely governed by the paradigm of In-Context Learning (ICL). However, despite their sophisticated semantic latent spaces, LLMs exhibit a profound and "notorious" sensitivity to the specificities of the input context. Zhao et al. [6] characterized this as "prompt instability," demonstrating that stochastic variations—such as the permutation of few-shot examples or minor syntactic shifts in instruction templates—can induce significant fluctuations in classification accuracy. This volatility suggests that the standard attention mechanism often converges on "surface-level" patterns rather than underlying logical structures. Furthermore, the architectural constraints of the transformer’s context window present a dimensional bottleneck. As noted by Dong et al. [7], fixed token limits necessitate a zero-sum trade-off between the depth of individual examples and the breadth of the reference set. In enterprise analytics, where datasets are high-dimensional and noisy, this limitation often leads to "recency bias" or the inclusion of non-representative outliers that confound the model’s outcomes. The wSSAS framework departs from traditional ICL by replacing static, heuristically-derived prompts with a dynamically synthesized context [21]. By applying a precision-filtering pipeline to the input background, we ensure that the "hints" provided to the model are mathematically optimized for representative signals. This transforms the context from a variable, human-engineered instruction into a stable, feature-engineered instrument, effectively addressing the inherent stochasticity of the generative process. 2.2 Attention Mechanisms and the Signal-to-Noise Challenge The "LLM Paradox" identified in this study—wherein generative fluency inversely correlates with analytical preci- sion—is fundamentally rooted in the transformer’s attention mechanism [9]. While the attention layer excels at global dependency modeling, its probabilistic nature becomes a liability when processing chaotic, non-curated datasets. In these environments, the model often fails to distinguish high-value "signal" from background "noise," leading to a degradation of the latent space stability required for rigorous categorization. Empirical evidence by [11] suggests that the inclusion of irrelevant information within the input context is more detrimental to model performance than the omission of relevant data. This occurs because extraneous tokens force the attention mechanism to allocate significant weights to inconsequential patterns, effectively "diluting" the focus on salient features. This challenge aligns with the Information Bottleneck (IB) principle [16], which posits that an optimal learning system should maximize the compression of input data while retaining only the information most pertinent to the target output. The wSSAS method- ology operationalizes the IB principle by implementing a pre-inference filtering stage involving data refinement. By calculating a Signal-to-Noise Ratio (SNR), the framework systematically suppresses irrelevant data and outliers before they are ingested by the LLM model. This intervention enforces a deterministic focus, ensuring that the transformer’s limited attention budget is reserved exclusively for contextually dense, representative data points. Consequently, the methodology bridges the gap between the stochasticity of generative architectures and the invariance required for enterprise-grade analytics. 2.3 Hierarchical Information Compression and Semantic Alignment Historically, clustering and dimensionality reduction have been the standard tools for organizing large, chaotic datasets into meaningful structures. However, these traditional methods typically treat summarization as a simple, one- dimensional task, failing to account for the complex, multi-layered nature of enterprise data. Standard Retrieval- Augmented Generation (RAG) and recursive summarization techniques often suffer from "information dilution," where the specific nuances of data points are lost during mid-level aggregation [22]. While recent advancements in hierarchical information processing have improved document-level understanding, current models often struggle to reconcile strategic top-down intent [23] with bottom-up empirical evidence [24]. Our wSSAS framework addresses this by implementing a dual-flow logic that ensures narrative consistency across three distinct levels: Themes, Stories, and Clusters. The integration of syntactic alignment (structural hierarchy) [25] and semantic alignment (latent meaning) [26], [27] is a recognized frontier in NLP. While Named Entity Recognition (NER) [28] and Topic Modeling [29] [30] provide semantic labels, they do not inherently "weight" the importance of data points based on their representative power within a broader narrative. Our approach draws inspiration from Selective Attention mechanisms in cognitive 3 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Syntactic AlignmentSemantic Alignment Defines structural rules governing how words combine into grammatical sentences. Acts as a mechanism for models to understand the structural hierarchy of data. Delves into the meaning of words and sentences. Explores how syntactic structures map onto semantic roles to extract the “who, what, when, where, and why” of data. Examples:Examples: Coarse-to-Fine Retrieval, Spatial Reasoning Improvements [32], Efficient Tuning for Document Visual QA [33] Word Embeddings [34], Named Entity Recognition (NER) [28], Topic Modeling [29, 30] Table 1: Syntactic vs. Semantic Alignment modeling [31], where non-essential "noise" is suppressed prior to high-level cognitive processing. By utilizing the Summary-of-Summaries (SoS) architecture, wSSAS creates a bounded attention environment that forces the LLM to focus on distilled, sentiment-dense narratives rather than being distracted by the stochastic variance of raw, unweighted text. Finally, our work builds upon the Information Bottleneck (IB) principle [16], which suggests that an optimal analytical model must compress input to retain only the features most relevant to the target output. While generative AI is optimized for creative novelty, the wSSAS methodology enforces analytical integrity by treating the input context as a precision-engineered feature set. This transforms the LLM from a probabilistic generator into a deterministic instrument, providing a scalable solution to the "black box" unpredictability often cited in current enterprise AI research [17]. 3 Syntactic & Semantic Attention Summarization (SSAS) The strategic rationale behind our SSAS methodology [14] is the implementation of a bounded attention mechanism. By pre-processing raw text through synchronized syntactic and semantic filters, we constrain the LLM’s focus to high-signal tokens, effectively performing feature engineering at the prompt level. This ensures that the model recognizes the structural hierarchy of the data before the semantic layer interprets the underlying mood or sentiment following the Compositional Semantics principle [35]. Table 1 contrasts the two foundational pillars of our methodology i.e. Syntactic alignment and Semantic alignment. By combining these alignments, our methodology creates an accurate, distilled summary of the dataset. This summary functions as a specific input prompt that focuses the LLM’s attention mechanism [9] to focus on the provided essential information rather than being distracted by the surrounding noise [11]. This alignment is optimized when applied across a rigorous data hierarchy. 3.1 Hierarchical Data Classification: Themes, Stories, and Clusters To transform large-scale, chaotic datasets into actionable insights, a structured hierarchical classification is necessary. Our framework ensures that every data point is evaluated for its contribution to the macro-narrative, preventing the loss of signal in high-volume environments. The SSAS methodology implements three distinct levels found in natural language taxonomies: 1.Themes: The most general classification level, identifying the primary macro-topic across all data points within the set. 2.Stories: The intermediate level of classification, ensuring narrative consistency by identifying specific subtopics within a theme. 3.Clusters: The lowest level of classification, where the algorithm utilizes localized precision to identify similar data points. The architecture operates on a dual-flow logic that reconciles strategic intent with empirical evidence, shown in Figure 1: 1. Top-Down Taxonomy (Strategic Intent): The classification flow (Themes -> Stories -> Clusters) organizes data into increasingly granular, manageable segments, a standard approach in Recursive Hierarchy Decoding [23]. 4 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs 2.Bottom-Up Aggregation (Data Evidence): The insight flow (Cluster Contexts -> Stories Contexts -> Theme Context) aggregates data to build the Summary of Summaries (SoS), ensuring that high-level insights are grounded in the localized precision of the underlying clusters [24]. Figure 1: SSAS Architecture for Context Assessment 3.2 Summary-of-Summaries (SoS) The implementation culminates in the context localized Summary-of-Summaries (SoS) architecture. This approach on data pre-processing strategically bounds the LLM’s focus by providing a concise, iterative summary as the primary input prompt, effectively reducing the probability of stochastic drift —a phenomenon where the model loses its objective over long context windows [22]. The process follows a specific aggregation path: 1. Cluster Context: A syntactic summary of individual clusters. 2. Story Context: An aggregated summary based on the summary of cluster summaries within that story [24]. 3. Theme Context: An aggregated summary based on the summary of story summaries within that theme. The strategic value of SoS is its role as a feature engineering step. By distilling raw text into a sentiment-dense narrative, we force the LLM to align with core structures and factual content mitigating the risk of "distraction" from irrelevant tokens [36] 3.3 Signal-to-Noise Ratio (SNR) Maintaining data integrity requires a rigorous weighting algorithm to isolate high-value signal from noise. SSAS further uses a weighting logic to validate data points across the hierarchical strata. The primary metric is the Signal to Noise Ratio (SNR), which is the weighted aggregate of three distinct dimensions and is calculated using Equation (1). The signal-to-noise ratio (SNR i ) is calculated as follows: SNR i = X (S T heme + S Story + S Cluster )(1) where: • S T heme Theme Signal-to-Noise Ratio that measures global alignment i.e. whether the data point fits the macro-topic • S Story Story Signal-to-Noise Ratio that measures narrative consistency i.e. whether the data point fits the sub-topic • S Cluster Cluster Signal-to-Noise Ratio that measures localized precision, i.e., whether the data point fits the immediate group. In addition, the methodology incorporates Weighted Amplitude, where keywords are weighted by frequency to enhance the signal. The outcome is Precision Filtering, which suppresses data points that share keywords but lack the contextual depth required for stable attention [11]. This ensures that the LLM is prompted only with high-signal data that perfectly fits the hierarchy, preventing out-of-context noise. 5 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs 3.4 Noise Mitigation: Irrelevant Data and Outlier Management Reliable text categorization necessitates aggressive noise removal to prevent the dilution of the model’s focus. SSAS categorizes noise into ranked ordering in two steps: 1.Irrelevant Data: This comprises data that does not fit into any defined classification level. Our algorithm labels this as irrelevant data. 2. Outliers: These are data points within the classification levels that have a negligible impact on the whole level. Through this rank-ordering, the most representative and contextually dense data points are elevated, while outliers and irrelevant data are suppressed to the bottom of the dataset. This ensures the LLM is prompted only with high-signal data that fits the hierarchy perfectly, preventing out-of-context noise. By dramatically reducing noise in the input data, wSSAS accelerates the identification of core insights, leading to more accurate LLM categorization and superior business decision-making. 3.5 Comparison of SSAS and wSSAS Methodology The core difference between SSAS and wSSAS lies in the assignment of analytical value to the derived context summary: 1. SSAS (Syntactic & Semantic Attention Summarization): This unweighted framework provides a structural (syntactic) and meaning-based (semantic) summary, resulting in "Unweighted Context Summary." While it establishes the hierarchical relationships (Themes, Stories, Clusters), it treats all synthesized information as having equal descriptive value. The attention mechanism is bounded by the scope of the summary but is not directed toward the most critical features. 2.wSSAS (Weighted Syntactic & Semantic Context Assessment Summarization): This evolved framework applies the calculated Signal-to-Noise Ratio (SNR) to the SSAS-generated summaries, resulting in "Weighted Context Summary." The SNR mathematically prioritizes high-value semantic clusters and narratives, sup- pressing statistically insignificant data points. This weight layer transforms the context from a complete map of the data (SSAS) into a precision-filtered instrument that actively directs the LLM’s attention to the most representative and contextually dense features, effectively isolating "Signal" from "Noise." 4 Experimental Design: A Two-Phased Validation Framework The efficacy of Weighted Syntactic and Semantic Context Assessment Summarization (wSSAS) is validated through a two-phase experimental design. This framework isolates and measures two key components: the quality of the generated context summary and the accuracy of the final categorization, ensuring performance improvements are directly linked to the enhanced input. •Phase 1: Context Summary Quality Assessment: This phase focuses on the algorithmic transformation of raw data. The SSAS algorithm organizes the data into a hierarchy of Themes, Stories, and Clusters. We then generate and compare two distinct context summary types—Unweighted (SSAS) and Weighted (wSSAS)—to determine which offers the most accurate and rich representation of the underlying data signal. •Phase 2: Categorization Performance Measurement: This phase evaluates the business impact of the context summary on Large Language Model (LLM) performance.The LLM Gemini 2.0 Flash Lite [37] is used to identify primary and secondary topics for each data point. To simplify analysis and enhance interpretability, K-Means clustering is applied to the output, grouping related topics into cohesive category-clusters 1 . The experiment compares three scenarios to isolate the wSSAS impact: 1. Baseline: Direct LLM input with no context. 2. Unweighted Context (SSAS): Categorization using standard SSAS context summary. 3. Weighted Context (wSSAS): Categorization using the enhanced wSSAS context summary. As illustrated in Figure 2, the experimental pipeline moves from raw input data through hierarchical context generation to final categorization (refer Appendix B).The validity of these experiments relies on the rigor of metrics specifically designed to evaluate abstractive intelligence. 1 It is important to distinguish Category-Clusters from the Clusters (or Cluster context summary) produced by the SSAS algorithm. Specifically, Category-Clusters are the result of K-Means clustering applied to the Topics that the LLM generates during the categorization phase. 6 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs INPUT DATA Themes, Stories, Clusters CONTEXT: SSAS wSSAS LLM (Gemini 2.0 Flash Lite) Categorization Assessment Metrics: Question-Answer Generation(QAG), G-Eval Scores Assessment Metrics: Silhouette Score, Davies-Bouldin Index, Calinski-Harabasz Index Sankey plots Data Pre-processing Output: Category-Clusters Figure 2: Experimental Design and Assessment Metrics 4.1 Context Summary Evaluation Metrics: QAG and G-Eval Traditional metrics such as ROUGE are insufficient for abstractive summaries as they rely on simple n-gram overlap, failing to capture semantic nuance or factual consistency [38], [39]. We therefore adopt a "LLM-as-a-judge" framework using reference-free metrics. 1. QAG Mechanics and Embedding Engine: QAG acts as a reference-free "polygraph test" for factual consistency [40]. The system generates up to five factual, close-ended questions from the source text and verifies the summary’s ability to provide accurate answers. To calculate semantic similarity between true responses and extracted responses, we utilized the sentence-transformers/all-MiniLM-L6-v2 embedding model [41] (a)Triage and Encoding: QAG scores were encoded into a 0 (as good as), 1 (better than), or -1 (worse than) scale, comparing weighted vs. unweighted outputs. A critical triage process was applied to prioritize semantic similarity over verbatim alignment. This prevents the penalization of the LLM for utilizing sophisticated paraphrasing while ensuring that factual hallucinations—which an exact-match algorithm might miss—are identified and suppressed. 2. G-Eval Assessment: G-Eval complements QAG by leveraging the LLM to approximate human-like judgment across four qualitative dimensions [39]: Coherence: Logical structure and organizational flow. Fluency: Grammatical precision and linguistic naturalism. Relevance: The concentration of high-value information. Consistency: Factual alignment with the source manifold. This approach has been shown to outperform traditional metrics in correlation with human preference. 4.2 Categorization Quality Metrics The final evaluation phase focuses on the structural integrity of the generated category-clusters. Internal validation metrics allow us to assess cluster quality without the need for external, human-labeled ground truth [42]. Table 2 describes in detail three metrics used in this study to quantify category-cluster quality. Table 2: Clustering Evaluation Metrics and Interpretations MetricDescriptionInterpretationGoal Silhouette Score [42]Measures cohesion vs. separation for each sample. Range: [−1, +1] +1: Well-separated 0: Overlapping -1: Misassigned Maximize Davies-Bouldin Index [43] Calculates average similarity be- tween clusters. Lower score indicates better sepa- ration and compactness. 0 is the minimum. Minimize Calinski-Harabasz Index [44]Ratio of between-cluster dispersion to within-cluster dispersion. Higher score indicates dense and well-separated clusters. Maximize 7 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs These metrics mathematically confirm whether the wSSAS context summary enables the LLM to identify distinct, compact, and meaningful categories. To ensure generalizability, these metrics were applied across three diverse, industry-standard datasets. Table 3: Dataset Overview Dataset# ReviewsDate RangeQuarters Primary Entity Strategic Intent Amazon Product155,74501/01/2020 – 05/23/202314StoresProduct-related Google Business121,82603/01/2009 – 08/25/202145Book TitlesRestaurant-related Goodreads Book157,40712/07/2006 – 11/03/201751RestaurantsLiterary, subjective 4.3 Evaluation Datasets: Multi-Domain Selection and Characteristic Analysis To demonstrate the generalizability of the wSSAS methodology, we utilized three diverse, industry-standard datasets from the University of California, San Diego (UCSD) [45] (Table 3) 1.Google Business Reviews (American & Fast Food restaurants): 121K reviews from North Dakota, used for restaurant sentiment analysis. 2.Amazon Product Reviews (Health & Personal Care Products): 155K reviews, focused on product discourse over a 3.5-year window. 3.Goodreads Book Reviews (Spoilers): The full 157K dataset, testing the model’s ability to handle long-form narrative spoilers. The datasets showed significant variability in their timelines: Amazon provided the most compressed data (14 quarters), while Goodreads (45 quarters) and Google (51 quarters) offered longer-term data. We characterized the quarterly data using Normalized Volume (High/Low) and Review Distribution, a metric indicating signal stability by tracking the activity of specific sub-topics over time. (Details in Appendix A) 4.4 Hierarchical Dataset Analysis: Themes, Stories, and Clusters Understanding the data distribution across hierarchical strata (Themes -> Stories -> Clusters) is critical for identifying how noise removal impacts signal quality. By removing Theme -1 (Irrelevant data) and subsequent outliers, the wSSAS methodology refines the dataset for high-precision categorization. Table 4a, 4b, and 4c show the overall counts of themes, stories, clusters and data points within each of the three datasets. (See Appendix C for count of stories, clusters, data points within each theme in the dataset before and after removal of noisy data.) 5 Results The transition from a flat data structure to a hierarchical weighting framework is strategically necessary to ensure that the LLM focuses its finite attention mechanism on high-value information. The following results validate the efficacy of wSSAS in distinguishing meaningful semantic "Signal" from interference. 5.1 Comparative Performance of Weighted vs. Unweighted Context Summaries To evaluate context summary quality objectively, a reference-free Question-Answer Generation (QAG) framework was implemented using Gemini 2.0 Flash Lite. This method functions as a "polygraph test" for factual consistency, generating close-ended questions from the source data to determine if the generated contexts maintain narrative integrity. A rigorous triage process was applied to these scores to enhance reliability. Data indicates that QAG scores from weighted context summary showed a consistent relative improvement post-triage, and these superior QAG scores correlated directly with improved G-Eval metrics. Specifically, weighted context summaries demonstrated higher performance across majority of the G-eval metrics of coherence, fluency, relevance, and consistency, confirming that hierarchical weighting produces a more faithful representation of the information landscape. Table 5 shows an example of the evaluation of the two context summaries generated for a Story within the Google Business Reviews dataset. (See Appendix D for more examples). 8 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table 4: Data Processing Statistics across Datasets (a) Google Business Reviews Data StageThemes Stories Clusters Data points All Data151138,804121,826 Without Irrelevant & Outlier Data125431096,434 (b) Amazon Product Reviews Data StageThemes Stories Clusters Data points All Data1510322,791155,745 Without Irrelevant & Outlier Data14862,034116,102 (c) Goodreads Book Reviews Data StageThemes Stories Clusters Data points All Data128022,386157,407 Without Irrelevant & Outlier Data11661,842117,133 The wSSAS methodology has been robustly validated across three key industry datasets. Table 6 summarizes the overall QAG performance across the three datasets post-triage showing percentage of Stories and Themes where wSSAS context summary demonstrated equal or superior representation on QAG and G-Eval metrics as compared to SSAS context summary. Table 5: Comparative Analysis of Weighted vs. Unweighted Context Summaries (Google Business Reviews) Story ID: 76 Weighted Context Summary (wSSAS)Unweighted Context Summary (SSAS) An Unremarkable AssessmentA Mediocre Experience The text consistently uses "average" and its variations to describe a subject, emphasizing its lack of distinction. It expresses a desire for everything to be average, longing for the ordinary. Data indicates a slightly below-average performance for medical doctors, prompting further investigation. A negative assessment, described as "average at best," suggests disappointment. A neg- ative dining experience is detailed, with tasteless, overcooked food and unsatisfactory service, leading to a recommendation for alternative dining options The text describes a consistently average, and often disappoint- ing, experience. Medical doctors’ performance is assessed as slightly below average, prompting a desire for everything to be average. A negative dining experience is detailed, with taste- less, overcooked tortellini and inadequate service. The reviewer expresses dissatisfaction with the food quality and service, high- lighting the unremarkable nature of the subject being evaluated and suggesting alternatives for better experiences. The overall sentiment conveys a lack of enthusiasm and a longing for the ordinary. Quantitative Evaluation:Quantitative Evaluation: • QAG Pre-Triage: 1/4 • QAG Post-Triage: 2/4 • G-Eval Scores: – Coherence: 0.8 – Relevance: 1.0 – Fluency: 0.9 – Consistency: 0.5 • QAG Pre-Triage: 3/4 • QAG Post-Triage: 2/4 • G-Eval Scores: – Coherence: 0.4 – Relevance: 0.5 – Fluency: 1.0 – Consistency: 0.5 9 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs 0.0% 10.0% 20.0% 30.0% 40.0% 50.0% ScoreWSSAS > ScoreSSASScoreWSSAS = ScoreSSASScoreWSSAS < ScoreSSAS Pre-TriagePost-Triage Stories (Total: 112) (a) Stories 0.0% 10.0% 20.0% 30.0% 40.0% 50.0% ScoreWSSAS > ScoreSSASScoreWSSAS = ScoreSSASScoreWSSAS < ScoreSSAS Pre-TriagePost-Triage Themes (Total: 15) (b) Themes Figure 3: Overall QAG performance for Google Business Reviews 5.2 Quantitative Assessment of Categorization and Clustering Integrity To validate the quality of the categorization performed, we use three internal validation metrics: the Silhouette Score (measure of cohesion vs. separation), the Davies-Bouldin Index (measure of cluster similarity), and the Calinski-Harabasz (CH) Index (measure of dispersion ratio). Comparative performance across No Context (Baseline), Unweighted context summary (SSAS), and Weighted context summary(wSSAS) scenarios demonstrates the clear business impact of our contextual grounding approach (Table 7). The Weighted context approach consistently delivered superior and more actionable clustering across diverse datasets, making it the preferred method for strategic data analysis. 1.Google Business Reviews: The wSSAS context summary (CH Index: 8006.7) dramatically improved cluster definition and density compared to the "No context" scenario (CH Index: 3041.4), consolidating fragmented data into three strategic categories: "Customer Dissatisfaction & Service Failures," "Positive Dining Reviews," and "Restaurant Experience and Food Quality. 2.Amazon Product Reviews: While the SSAS context summary achieved a higher CH Index (4746.3), the superior Silhouette Score (0.049) and Davies-Bouldin Index (3.80) of wSSAS context summary indicate more effective cluster separation and internal cohesion. This suggests that the wSSAS approach provides a better qualitative definition of categories, even if the dispersion ratio is slightly lower than the unweighted model. This is crucial for distinguishing high-density generic feedback from specific, actionable issues like "Defective or Faulty Products." 3.Goodreads Book Reviews: The wSSAS context summary effectively consolidated complex review data into three highly defined clusters, delivering focused, high-density categories such as "Book Reviews and Criticism," "Book Series and Character Relationships," and "Romance and Suspense," preventing the fragmentation seen in the other two scenarios. Robustness Check: The structural integrity of the wSSAS approach was confirmed; removing irrelevant data points did not materially alter the core thematic architecture, proving the clusters are not easily disrupted by noise or outliers. (See Appendix E for details and Sankey Plot analysis showing the movement of data points between category-clusters for different experimental scenarios) Table 6: Overall QAG performance of the wSSAS context summary DatasetStories (%) Themes (%) Google Business Reviews79.5%86.7% Amazon Product Reviews87.3%80.0% Goodreads Book Reviews86.0%91.7% 10 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table 7: Comparative Clustering Performance and Data Distribution across Categories (a) Google Business Reviews ScenarioCount Category-Cluster Titles (% Vol)Silhouette Score Davies-Bouldin Index Calinski-Harabasz Index Weighted Context (wSSAS) 3 • Customer Dissatisfaction (17.7%) • Positive Dining Reviews (33.8%) • Restaurant Exp. and Food Quality (48.5%) 0.112.798006.7 Unweighted Context (SSAS) 3 • Restaurant Customer Satisfaction (26.1%) • Restaurant Exp. and Operations (59.9%) • Restaurant Service/Quality (14%) 0.072.922553.4 No context (Baseline) 4 • Customer Service/Quality (14.2%) • Positive Restaurant Experiences (19%) • Restaurant Exp. and Service (20%) • Restaurant Reviews/Dining (46.8%) 0.063.353041.4 (b) Amazon Product Reviews ScenarioCount Category-Cluster TitlesSilhouette Score Davies-Bouldin Index Calinski-Harabasz Index Weighted Context (wSSAS) 6 • Beauty and Grooming Products (15.3%) • Cleaning Products (9.7%) • Defective/Faulty Products (17.5%) • Digestive & Gut Health Supplements (10.8%) • Masks/Accessories (14.1%) • Product Reviews/Feedback (32.6%) 0.0493.804203.0 Unweighted Context (SSAS) 4 • Grooming & Personal Care (23.4%) • Pain Relief & Symptom Management (15.3%) • Product Defects & Dissatisfaction (21.8%) • Product Installation & User Experience (39.5%) 0.0474.134746.3 No context (Baseline) 7 • Assistive Devices (11.6%) • Cleaning & Maintenance (7.5%) • Pain & Symptom Relief (10.2%) • Personal Grooming & Hygiene (21.3%) • Positive Experiences & Reactions (12.6%) • Product Functionality & Performance (11.6%) • Quality and Performance Issues (25.2%) 0.0414.393264.0 (c) Goodreads Book Reviews ScenarioCount Category-Cluster TitlesSilhouette Score Davies-Bouldin Index Calinski-Harabasz Index Weighted Context (wSSAS) 3 • Book Reviews and Criticism (38.5%) • Book Series and Character Relationships (29.4%) • Romance and Suspense (32.1%) 0.0414.615887.4 Unweighted Context (SSAS) 5 • Book Review Criticism (41.0%) • Book Review Focus (28.3%) • Character Appreciation Focused Reviews (12.1%) • Content Evaluation and Reaction (7.0%) • Reader Disappointment/Enjoyment (11.7%) 0.0273.893453.8 No context (Baseline) 3 • Book Criticism and Appreciation (28.9%) • Book Review Themes and Tropes (54.4%) • Content Disappointment/Expectation (16.7%) 0.0214.735387.7 11 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs 6 Conclusion 6.1 Key Research Takeaways The Weighted Syntactic and Semantic Context Assessment Summary (wSSAS) methodology fundamentally reconfigures the data preprocessing landscape by moving beyond the limitations of unweighted architectures. While unweighted models—though providing a complete map of the information landscape—operate under a "flat value structure" that assumes all data points are equal, they ultimately lack the capacity to distinguish critical signals from low-relevance data, resulting in redundant category clusters. In contrast, the wSSAS methodology programmatically engineers a precision-filtered input by utilizing a Signal-to-Noise Ratio (SNR) that validates semantic integrity across three hierarchical strata: Cluster, Story, and Theme signals. This weighted approach ensures that the most representative data rises to the top while mathematically suppressing out-of-context outliers and statistical noise. As confirmed by internal validation metrics, including the Silhouette Score and the Calinski-Harabasz Index, this systematic isolation of high-value semantic signals produces superior, non-redundant category clusters. Ultimately, the primary value of this improved context is its direct contribution to focusing the attention mechanism of Large Language Models (LLMs), providing the high-quality foundation required for hyper-precise categorization. 6.2 Impact on Large-Scale Inference This methodology’s value proposition is its ability to convert heterogeneous data into refined, high-value informational resources, substantially increasing the velocity and precision of organizational decisional processes. By synthesizing hierarchical thematic outputs—comprising Themes, Stories, and Clusters—with automated categorization tools and raw metadata, the framework engineers a unified value proposition. These "derived segments" function as a precision-guided compass for stakeholders, allowing for accelerated diagnostic and growth activities across diverse sectors.For instance, restaurant owners can more effectively diagnose the underlying drivers of performance declines, while analysts at a consumer goods firm can leverage these segments to target and acquire new consumer populations with unprecedented accuracy. By converging these high-value components, organizations can extract actionable insights with significantly enhanced speed, ensuring that strategic resources are allocated with maximum efficiency (Figure 4) To realize these high-level advantages, however, organizations require a well-defined strategy for deploying the wSSAS methodology throughout the enterprise architecture. Themes/Stories/Clusters Categories Raw Meta Data BUSINESS VALUE Figure 4: True business value lies at the convergence of generated data-segments 6.3 Roadmap for Future Work Future research must expand the definition of context into a truly multi-dimensional construct, evolving toward even more granular weighting strata. The next generation of wSSAS will target currently unresolved linguistic 12 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs ambiguities—specifically complex phenomena such as irony, contrast, and intensification—where models traditionally struggle with reasoning. Beyond linguistic nuances, future iterations should integrate multi-dimensional contextual vectors that account for environmental factors, such as temporal shifts and the quarter-over-quarter trends observed in diverse review datasets. By refining these algorithmic dimensions, the framework will move beyond static summarization toward a dynamic contextual alignment architecture capable of navigating the most intricate intersections of human language and machine intelligence. 13 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs References [1] Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. Introduction to Information Retrieval. URL: https://nlp.stanford.edu/IR-book/. [2] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Nee- lakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam Mc- Candlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language Models are Few-Shot Learners, July 2020. arXiv:2005.14165 [cs]. URL: http://arxiv.org/abs/2005.14165, doi:10.48550/arXiv.2005.14165. [3] Canyu Chen and Kai Shu. Can LLM-Generated Misinformation Be Detected?, April 2024. arXiv:2309.13788 [cs]. URL: http://arxiv.org/abs/2309.13788, doi:10.48550/arXiv.2309.13788. [4]Elias Hossain, Rajib Rana, Niall Higgins, Jeffrey Soar, Prabal Datta Barua, Anthony R. Pisani, and Kathryn Turner. Natural Language Processing in Electronic Health Records in relation to healthcare decision-making: A systematic review. Computers in Biology and Medicine, 155:106649, March 2023. URL:https://w.sciencedirect. com/science/article/pii/S0010482523001142, doi:10.1016/j.compbiomed.2023.106649. [5]Monica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim, and David Sontag. Large language models are few-shot clinical information extractors. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1998– 2022, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL: https://aclanthology.org/2022.emnlp-main.130/, doi:10.18653/v1/2022.emnlp-main.130. [6]Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate Before Use: Improving Few-shot Performance of Language Models. In Proceedings of the 38th International Conference on Machine Learning, pages 12697–12706. PMLR, July 2021. URL: https://proceedings.mlr.press/v139/zhao21c.html. [7]Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, Baobao Chang, Xu Sun, Lei Li, and Zhifang Sui. A Survey on In-context Learning, October 2024. arXiv:2301.00234 [cs]. URL: http://arxiv.org/abs/2301.00234, doi:10.48550/arXiv.2301.00234. [8]Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Rahane, Anantharaman S. Iyer, Anders Andreassen, Andrea Madotto, Andrea Santilli, Andreas Stuhlmüller, Andrew Dai, Andrew La, Andrew Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karaka ̧s, B. Ryan Roberts, Bao Sheng Loe, Barret Zoph, Bartłomiej Bojanowski, Batuhan Özyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Howald, Bryan Orinion, Cameron Diao, Cameron Dour, Catherine Stinson, Cedrick Argueta, César Ferri Ramírez, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Chris Waites, Christian Voigt, Christopher D. Manning, Christopher Potts, Cindy Ramirez, Clara E. Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Dan Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel Moseguí González, Danielle Perszyk, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, David Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Dylan Schrader, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth Donoway, Ellie Pavlick, Emanuele Rodola, Emma Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A. Chi, Ethan Dyer, Ethan Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fanyue Xia, Fatemeh Siar, Fernando Martínez-Plumed, Francesca Happé, Francois Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, Germán Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Wang, Gonzalo Jaimovitch-López, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Hannah Kim, Hannah Rashkin, Hannaneh Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Shevlin, Hinrich Schütze, Hiromu Yakura, Hongming Zhang, Hugh Mee Wong, Ian Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, Jackson Kernion, Jacob Hilton, Jaehoon Lee, Jaime Fernández Fisac, James B. Simon, James Koppel, James Zheng, James Zou, Jan Koco ́ n, Jana Thompson, Janelle Wingfield, Jared Kaplan, Jarema Radom, Jascha Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jennifer Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Joan Waweru, John Burden, John Miller, John U. Balis, Jonathan Batchelder, Jonathan Berant, Jörg Frohberg, Jos 14 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Rozen, Jose Hernandez-Orallo, Joseph Boudeman, Joseph Guerr, Joseph Jones, Joshua B. Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakrishnan, Katerina Ignatyeva, Katja Markert, Kaustubh D. Dhole, Kevin Gimpel, Kevin Omondi, Kory Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras-Ochando, Louis-Philippe Morency, Luca Moschella, Lucas Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros Colón, Luke Metz, Lütfi Kerem ̧Senel, Maarten Bosma, Maarten Sap, Maartje ter Hoeve, Maheen Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, Maria Jose Ramírez Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew L. Leavitt, Matthias Hagen, Mátyás Schubert, Medina Orduna Baitemirova, Melody Arnaud, Melvin McElrath, Michael A. Yee, Michael Cohen, Michael Gu, Michael Ivanitskiy, Michael Starritt, Michael Strube, Michał Sw ̨edrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Mitch Walker, Mo Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, Mukund Varma T, Nanyun Peng, Nathan A. Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas Roberts, Nick Doiron, Nicole Martinez, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha S. Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Pu Liang, Paul Vicol, Pegah Alipoormolabashi, Peiyuan Liao, Percy Liang, Peter Chang, Peter Eckersley, Phu Mon Htut, Pinyu Hwang, Piotr Miłkowski, Piyush Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, Qing Lyu, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ramon Risco, Raphaël Millière, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan LeBras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Ruslan Salakhutdinov, Ryan Chi, Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib Singh, Saif M. Mohammad, Sajant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Samuel R. Bowman, Samuel S. Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Toshniwal, Shyam Upadhyay, Shyamolima, Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo-Hwan Lee, Spencer Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Biderman, Stephanie Lin, Stephen Prasad, Steven T. Piantadosi, Stuart M. Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq Ali, Tatsu Hashimoto, Te-Lin Wu, Théo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, Timofei Kornev, Titus Tunduny, Tobias Gerstenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Ramasesh, Vinay Uday Prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, Wout Vossen, Xiang Ren, Xiaoyu Tong, Xinran Zhao, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yangqiu Song, Yasaman Bahri, Yejin Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yufang Hou, Yuntao Bai, Zachary Seid, Zhuoye Zhao, Zijian Wang, Zijie J. Wang, Zirui Wang, and Ziyi Wu. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, June 2023. arXiv:2206.04615 [cs]. URL: http://arxiv.org/abs/2206.04615, doi:10.48550/arXiv.2206.04615. [9]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention Is All You Need, August 2023. arXiv:1706.03762 [cs]. URL:http://arxiv.org/ abs/1706.03762, doi:10.48550/arXiv.1706.03762. [10] Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, pages 610–623, New York, NY, USA, March 2021. Association for Computing Machinery. URL:https://dl.acm.org/doi/10.1145/3442188.3445922, doi:10.1145/3442188.3445922. [11]Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. Large Language Models Can Be Easily Distracted by Irrelevant Context. In Proceedings of the 40th International Conference on Machine Learning, pages 31210–31227. PMLR, July 2023. URL:https: //proceedings.mlr.press/v202/shi23a.html. [12]R. Andrew Kreek and Emilia Apostolova. Training and Prediction Data Discrepancies: Challenges of Text Classification with Noisy, Historical Data. In Wei Xu, Alan Ritter, Tim Baldwin, and Afshin Rahimi, editors, Proceedings of the 2018 EMNLP Workshop W-NUT: The 4th Workshop on Noisy User-generated Text, pages 104–109, Brussels, Belgium, November 2018. Association for Computational Linguistics. URL:https:// aclanthology.org/W18-6114/, doi:10.18653/v1/W18-6114. [13] Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv., 15 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs 55(9):195:1–195:35, January 2023. URL:https://dl.acm.org/doi/10.1145/3560815,doi:10.1145/ 3560815. [14]Nitin Mayande, Sharookh Daruwalla, Sumedh Khodke, Nitin Joglekar, and Charles Weber. Syntactic and Semantic Attention Summary (SSAS): An Approach to Improve LLM Summary Generation, October 2024. [15]Nitin Mayande, Sharookh Daruwalla, Shreeya Verma Kathuria, Nitin Joglekar, and Weber Charles. Leveraging Weighted Syntactic and Semantic Attention Summary (wSSAS) Towards Text Categorization Using LLMs, October 2025. [16]Naftali Tishby and Noga Zaslavsky. Deep Learning and the Information Bottleneck Principle, March 2015. arXiv:1503.02406 [cs]. URL: http://arxiv.org/abs/1503.02406, doi:10.48550/arXiv.1503.02406. [17]Thomas H. Davenport and Nitin Mittal. All-in On AI: How Smart Companies Win Big with Artificial Intelligence. January 2023. [18]Jiawei Ma, Yulei Niu, Jincheng Xu, Shiyuan Huang, Guangxing Han, and Shih-Fu Chang. DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection, March 2023. arXiv:2303.09674 [cs]. URL: http://arxiv.org/abs/2303.09674, doi:10.48550/arXiv.2303.09674. [19]Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. MetaICL: Learning to Learn In Context, May 2022. arXiv:2110.15943 [cs]. URL:http://arxiv.org/abs/2110.15943,doi:10.48550/arXiv. 2110.15943. [20]Zhaoyang Niu, Guoqiang Zhong, and Hui Yu. A review on the attention mechanism of deep learning. Neuro- computing, 452:48–62, September 2021. URL:https://w.sciencedirect.com/science/article/pii/ S092523122100477X, doi:10.1016/j.neucom.2021.03.091. [21]Tejpalsingh Siledar, Swaroop Nath, Sankara Sri Raghava Ravindra Muddu, Rupasai Rangaraju, Swaprava Nath, Pushpak Bhattacharyya, Suman Banerjee, Amey Patil, Sudhanshu Shekhar Singh, Muthusamy Chelliah, and Nikesh Garera. One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation, June 2024. arXiv:2402.11683 [cs]. URL: http://arxiv.org/abs/2402.11683, doi:10.48550/arXiv.2402.11683. [22]Qingyue Wang, Yanhe Fu, Yanan Cao, Shuai Wang, Zhiliang Tian, and Liang Ding. Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models, August 2025. arXiv:2308.15022 [cs]. URL: http://arxiv.org/abs/2308.15022, doi:10.48550/arXiv.2308.15022. [23] SangHun Im, GiBaeg Kim, Heung-Seon Oh, Seongung Jo, and Dong Hwan Kim. Hierarchical Text Classifi- cation as Sub-hierarchy Sequence Generation. Proceedings of the AAAI Conference on Artificial Intelligence, 37(11):12933–12941, June 2023. URL:https://ojs.aaai.org/index.php/AAAI/article/view/26520, doi:10.1609/aaai.v37i11.26520. [24]Janara Christensen, Stephen Soderland, Gagan Bansal, and Mausam. Hierarchical Summarization: Scaling Up Multi-Document Summarization. In Kristina Toutanova and Hua Wu, editors, Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 902–912, Baltimore, Maryland, June 2014. Association for Computational Linguistics. URL:https://aclanthology. org/P14-1085/, doi:10.3115/v1/P14-1085. [25]Weicheng Ma and Torsten Suel. Structural Sentence Similarity Estimation for Short Texts. URL:https: //aaai.org/papers/232-flairs-2016-12940/. [26]Daniel Jurafsky and James H. Martin. Speech and Language Processing. Pearson Education, December 2014. Google-Books-ID: Cq2gBwAAQBAJ. [27] Chengyu Nan. Semantic Map and HBV in English, Chinese and Korean—A Case Study of hand,Shou and Son. Journal of Language Teaching and Research, 7(6):1216, November 2016. URL:http://w. academypublication.com/issues2/jltr/vol07/06/21.pdf, doi:10.17507/jltr.0706.21. [28]Arya Roy. Recent Trends in Named Entity Recognition (NER), January 2021. arXiv:2101.11420 [cs]. URL: http://arxiv.org/abs/2101.11420, doi:10.48550/arXiv.2101.11420. [29]Hamed Jelodar, Yongli Wang, Chi Yuan, Xia Feng, Xiahui Jiang, Yanchao Li, and Liang Zhao. Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey. Multimedia Tools Appl., 78(11):15169– 15211, June 2019. doi:10.1007/s11042-018-6894-4. [30] Guangxu Xun, Vishrawas Gopalakrishnan, Fenglong Ma, Yaliang Li, Jing Gao, and Aidong Zhang. Topic Discovery for Short Texts Using Word Embeddings. pages 1299–1304, December 2016.doi:10.1109/ICDM. 2016.0176. [31]Robert Desimone and John Duncan. Neural Mechanisms of Selective Visual Attention. Annual Review of Neu- roscience, 18(Volume 18, 1995):193–222, March 1995. URL:https://w.annualreviews.org/content/ journals/10.1146/annurev.ne.18.030195.001205, doi:10.1146/annurev.ne.18.030195.001205. 16 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs [32]Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo, and Tsung-Yu Lin. Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs, April 2024. arXiv:2404.07449 [cs]. URL: http://arxiv.org/abs/2404.07449, doi:10.48550/arXiv.2404.07449. [33]Mike Lewis and Angela Fan. Generative Question Answering: Learning to Answer the Whole Question. September 2018. URL: https://openreview.net/forum?id=Bkx0RjA9tX. [34]Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient Estimation of Word Representations in Vector Space, September 2013. arXiv:1301.3781 [cs]. URL:http://arxiv.org/abs/1301.3781,doi: 10.48550/arXiv.1301.3781. [35]Barbara H. Partee. Lexical Semantics and Compositionality. 1995. URL:https://direct.mit.edu/books/ edited-volume/4671/chapter/214107/Lexical-Semantics-and-Compositionality,doi:10.7551/ mitpress/3964.001.0001. [36]Yucheng Li, Bo Dong, Chenghua Lin, and Frank Guerin. Compressing Context to Enhance Inference Efficiency of Large Language Models, October 2023. arXiv:2310.06201 [cs]. URL:http://arxiv.org/abs/2310.06201, doi:10.48550/arXiv.2310.06201. [37]Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalk- wyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael Isard, Paul R. Barham, Tom Hennigan, Benjamin Lee, Fabio Viola, Malcolm Reynolds, Yuanzhong Xu, Ryan Doherty, Eli Collins, Clemens Meyer, Eliza Rutherford, Erica Moreira, Kareem Ayoub, Megha Goel, Jack Krawczyk, Cosmo Du, Ed Chi, Heng-Tze Cheng, Eric Ni, Purvi Shah, Patrick Kane, Betty Chan, Manaal Faruqui, Aliaksei Severyn, Hanzhao Lin, YaGuang Li, Yong Cheng, Abe Ittycheriah, Mahdis Mahdieh, Mia Chen, Pei Sun, Dustin Tran, Sumit Bagri, Balaji Lakshminarayanan, Jeremiah Liu, Andras Orban, Fabian Güra, Hao Zhou, Xinying Song, Aurelien Boffy, Harish Ganapathy, Steven Zheng, HyunJeong Choe, Ágoston Weisz, Tao Zhu, Yifeng Lu, Siddharth Gopal, Jarrod Kahn, Maciej Kula, Jeff Pitman, Rushin Shah, Emanuel Taropa, Majd Al Merey, Martin Baeuml, Zhifeng Chen, Laurent El Shafey, Yujing Zhang, Olcan Sercinoglu, George Tucker, Enrique Piqueras, Maxim Krikun, Iain Barr, Nikolay Savinov, Ivo Danihelka, Becca Roelofs, Anaïs White, Anders Andreassen, Tamara von Glehn, Lakshman Yagati, Mehran Kazemi, Lucas Gonzalez, Misha Khalman, Jakub Sygnowski, Alexandre Frechette, Charlotte Smith, Laura Culp, Lev Proleev, Yi Luan, Xi Chen, James Lottes, Nathan Schucher, Federico Lebron, Alban Rrustemi, Natalie Clay, Phil Crone, Tomas Kocisky, Jeffrey Zhao, Bartek Perz, Dian Yu, Heidi Howard, Adam Bloniarz, Jack W. Rae, Han Lu, Laurent Sifre, Marcello Maggioni, Fred Alcober, Dan Garrette, Megan Barnes, Shantanu Thakoor, Jacob Austin, Gabriel Barth-Maron, William Wong, Rishabh Joshi, Rahma Chaabouni, Deeni Fatiha, Arun Ahuja, Gaurav Singh Tomar, Evan Senter, Martin Chadwick, Ilya Kornakov, Nithya Attaluri, Iñaki Iturrate, Ruibo Liu, Yunxuan Li, Sarah Cogan, Jeremy Chen, Chao Jia, Chenjie Gu, Qiao Zhang, Jordan Grimstad, Ale Jakse Hartman, Xavier Garcia, Thanumalayan Sankaranarayana Pillai, Jacob Devlin, Michael Laskin, Diego de Las Casas, Dasha Valter, Connie Tao, Lorenzo Blanco, Adrià Puigdomènech Badia, David Reitter, Mianna Chen, Jenny Brennan, Clara Rivera, Sergey Brin, Shariq Iqbal, Gabriela Surita, Jane Labanowski, Abhi Rao, Stephanie Winkler, Emilio Parisotto, Yiming Gu, Kate Olszewska, Ravi Addanki, Antoine Miech, Annie Louis, Denis Teplyashin, Geoff Brown, Elliot Catt, Jan Balaguer, Jackie Xiang, Pidong Wang, Zoe Ashwood, Anton Briukhov, Albert Webson, Sanjay Ganapa- thy, Smit Sanghavi, Ajay Kannan, Ming-Wei Chang, Axel Stjerngren, Josip Djolonga, Yuting Sun, Ankur Bapna, Matthew Aitchison, Pedram Pejman, Henryk Michalewski, Tianhe Yu, Cindy Wang, Juliette Love, Junwhan Ahn, Dawn Bloxwich, Kehang Han, Peter Humphreys, Thibault Sellam, James Bradbury, Varun Godbole, Sina Samangooei, Bogdan Damoc, Alex Kaskasoli, Sébastien M. R. Arnold, Vijay Vasudevan, Shubham Agrawal, Jason Riesa, Dmitry Lepikhin, Richard Tanburn, Srivatsan Srinivasan, Hyeontaek Lim, Sarah Hodkinson, Pranav Shyam, Johan Ferret, Steven Hand, Ankush Garg, Tom Le Paine, Jian Li, Yujia Li, Minh Giang, Alexander Neitz, Zaheer Abbas, Sarah York, Machel Reid, Elizabeth Cole, Aakanksha Chowdhery, Dipanjan Das, Dominika Rogozi ́ nska, Vitaliy Nikolaev, Pablo Sprechmann, Zachary Nado, Lukas Zilka, Flavien Prost, Luheng He, Mar- ianne Monteiro, Gaurav Mishra, Chris Welty, Josh Newlan, Dawei Jia, Miltiadis Allamanis, Clara Huiyi Hu, Raoul de Liedekerke, Justin Gilmer, Carl Saroufim, Shruti Rijhwani, Shaobo Hou, Disha Shrivastava, Anirudh Baddepudi, Alex Goldin, Adnan Ozturel, Albin Cassirer, Yunhan Xu, Daniel Sohn, Devendra Sachan, Reinald Kim Amplayo, Craig Swanson, Dessie Petrova, Shashi Narayan, Arthur Guez, Siddhartha Brahma, Jessica Landon, Miteyan Patel, Ruizhe Zhao, Kevin Villela, Luyu Wang, Wenhao Jia, Matthew Rahtz, Mai Giménez, Legg Yeung, James Keeling, Petko Georgiev, Diana Mincu, Boxi Wu, Salem Haykal, Rachel Saputro, Kiran Vodrahalli, James Qin, Zeynep Cankara, Abhanshu Sharma, Nick Fernando, Will Hawkins, Behnam Neyshabur, Solomon Kim, Adrian Hutter, Priyanka Agrawal, Alex Castro-Ros, George van den Driessche, Tao Wang, Fan Yang, Shuo-yiin Chang, Paul Komarek, Ross McIlroy, Mario Lu ˇ ci ́ c, Guodong Zhang, Wael Farhan, Michael Sharman, Paul Natsev, Paul Michel, Yamini Bansal, Siyuan Qiao, Kris Cao, Siamak Shakeri, Christina Butterfield, Justin Chung, 17 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Paul Kishan Rubenstein, Shivani Agrawal, Arthur Mensch, Kedar Soparkar, Karel Lenc, Timothy Chung, Aedan Pope, Loren Maggiore, Jackie Kay, Priya Jhakra, Shibo Wang, Joshua Maynez, Mary Phuong, Taylor Tobin, Andrea Tacchetti, Maja Trebacz, Kevin Robinson, Yash Katariya, Sebastian Riedel, Paige Bailey, Kefan Xiao, Nimesh Ghelani, Lora Aroyo, Ambrose Slone, Neil Houlsby, Xuehan Xiong, Zhen Yang, Elena Gribovskaya, Jonas Adler, Mateo Wirth, Lisa Lee, Music Li, Thais Kagohara, Jay Pavagadhi, Sophie Bridgers, Anna Bortsova, Sanjay Ghemawat, Zafarali Ahmed, Tianqi Liu, Richard Powell, Vijay Bolina, Mariko Iinuma, Polina Zablotskaia, James Besley, Da-Woon Chung, Timothy Dozat, Ramona Comanescu, Xiance Si, Jeremy Greer, Guolong Su, Martin Polacek, Raphaël Lopez Kaufman, Simon Tokumine, Hexiang Hu, Elena Buchatskaya, Yingjie Miao, Mohamed Elhawaty, Aditya Siddhant, Nenad Tomasev, Jinwei Xing, Christina Greer, Helen Miller, Shereen Ashraf, Aurko Roy, Zizhao Zhang, Ada Ma, Angelos Filos, Milos Besta, Rory Blevins, Ted Klimenko, Chih-Kuan Yeh, Soravit Changpinyo, Jiaqi Mu, Oscar Chang, Mantas Pajarskas, Carrie Muir, Vered Cohen, Charline Le Lan, Krishna Haridasan, Amit Marathe, Steven Hansen, Sholto Douglas, Rajkumar Samuel, Mingqiu Wang, Sophia Austin, Chang Lan, Jiepu Jiang, Justin Chiu, Jaime Alonso Lorenzo, Lars Lowe Sjösund, Sébastien Cevey, Zach Gleicher, Thi Avrahami, Anudhyan Boral, Hansa Srinivasan, Vittorio Selo, Rhys May, Konstantinos Aisopos, Léonard Hussenot, Livio Baldini Soares, Kate Baumli, Michael B. Chang, Adrià Recasens, Ben Caine, Alexander Pritzel, Filip Pavetic, Fabio Pardo, Anita Gergely, Justin Frye, Vinay Ramasesh, Dan Horgan, Kartikeya Badola, Nora Kassner, Subhrajit Roy, Ethan Dyer, Víctor Campos Campos, Alex Tomala, Yunhao Tang, Dalia El Badawy, Elspeth White, Basil Mustafa, Oran Lang, Abhishek Jindal, Sharad Vikram, Zhitao Gong, Sergi Caelles, Ross Hemsley, Gregory Thornton, Fangxiaoyu Feng, Wojciech Stokowiec, Ce Zheng, Phoebe Thacker, Ça ̆ glar Ünlü, Zhishuai Zhang, Mohammad Saleh, James Svensson, Max Bileschi, Piyush Patil, Ankesh Anand, Roman Ring, Katerina Tsihlas, Arpi Vezer, Marco Selvi, Toby Shevlane, Mikel Rodriguez, Tom Kwiatkowski, Samira Daruki, Keran Rong, Allan Dafoe, Nicholas FitzGerald, Keren Gu-Lemberg, Mina Khan, Lisa Anne Hendricks, Marie Pellat, Vladimir Feinberg, James Cobon-Kerr, Tara Sainath, Maribeth Rauh, Sayed Hadi Hashemi, Richard Ives, Yana Hasson, Eric Noland, Yuan Cao, Nathan Byrd, Le Hou, Qingze Wang, Thibault Sottiaux, Michela Paganini, Jean-Baptiste Lespiau, Alexandre Moufarek, Samer Hassan, Kaushik Shivakumar, Joost van Amersfoort, Amol Mandhane, Pratik Joshi, Anirudh Goyal, Matthew Tung, Andrew Brock, Hannah Sheahan, Vedant Misra, Cheng Li, Nemanja Raki ́ cevi ́ c, Mostafa Dehghani, Fangyu Liu, Sid Mittal, Junhyuk Oh, Seb Noury, Eren Sezener, Fantine Huot, Matthew Lamm, Nicola De Cao, Charlie Chen, Sidharth Mudgal, Romina Stella, Kevin Brooks, Gautam Vasudevan, Chenxi Liu, Mainak Chain, Nivedita Melinkeri, Aaron Cohen, Venus Wang, Kristie Seymore, Sergey Zubkov, Rahul Goel, Summer Yue, Sai Krishnakumaran, Brian Albert, Nate Hurley, Motoki Sano, Anhad Mohananey, Jonah Joughin, Egor Filonov, Tomasz K ̨epa, Yomna Eldawy, Jiawern Lim, Rahul Rishi, Shirin Badiezadegan, Taylor Bos, Jerry Chang, Sanil Jain, Sri Gayatri Sundara Padmanabhan, Subha Puttagunta, Kalpesh Krishna, Leslie Baker, Norbert Kalb, Vamsi Bedapudi, Adam Kurzrok, Shuntong Lei, Anthony Yu, Oren Litvin, Xiang Zhou, Zhichun Wu, Sam Sobell, Andrea Siciliano, Alan Papir, Robby Neale, Jonas Bragagnolo, Tej Toor, Tina Chen, Valentin Anklin, Feiran Wang, Richie Feng, Milad Gholami, Kevin Ling, Lijuan Liu, Jules Walter, Hamid Moghaddam, Arun Kishore, Jakub Adamek, Tyler Mercado, Jonathan Mallinson, Siddhinita Wandekar, Stephen Cagle, Eran Ofek, Guillermo Garrido, Clemens Lombriser, Maksim Mukha, Botu Sun, Hafeezul Rahman Mohammad, Josip Matak, Yadi Qian, Vikas Peswani, Pawel Janus, Quan Yuan, Leif Schelin, Oana David, Ankur Garg, Yifan He, Oleksii Duzhyi, Anton Älgmyr, Timothée Lottaz, Qi Li, Vikas Yadav, Luyao Xu, Alex Chinien, Rakesh Shivanna, Aleksandr Chuklin, Josie Li, Carrie Spadine, Travis Wolfe, Kareem Mohamed, Subhabrata Das, Zihang Dai, Kyle He, Daniel von Dincklage, Shyam Upadhyay, Akanksha Maurya, Luyan Chi, Sebastian Krause, Khalid Salama, Pam G Rabinovitch, Pavan Kumar Reddy M, Aarush Selvan, Mikhail Dektiarev, Golnaz Ghiasi, Erdem Guven, Himanshu Gupta, Boyi Liu, Deepak Sharma, Idan Heimlich Shtacher, Shachi Paul, Oscar Akerlund, François-Xavier Aubet, Terry Huang, Chen Zhu, Eric Zhu, Elico Teixeira, Matthew Fritze, Francesco Bertolini, Liana-Eleonora Marinescu, Martin Bölle, Dominik Paulus, Khyatti Gupta, Tejasi Latkar, Max Chang, Jason Sanders, Roopa Wilson, Xuewei Wu, Yi-Xuan Tan, Lam Nguyen Thiet, Tulsee Doshi, Sid Lall, Swaroop Mishra, Wanming Chen, Thang Luong, Seth Benjamin, Jasmine Lee, Ewa Andrejczuk, Dominik Rabiej, Vipul Ranjan, Krzysztof Styrc, Pengcheng Yin, Jon Simon, Malcolm Rose Harriott, Mudit Bansal, Alexei Robsky, Geoff Bacon, David Greene, Daniil Mirylenka, Chen Zhou, Obaid Sarvana, Abhimanyu Goyal, Samuel Andermatt, Patrick Siegler, Ben Horn, Assaf Israel, Francesco Pongetti, Chih-Wei "Louis" Chen, Marco Selvatici, Pedro Silva, Kathie Wang, Jackson Tolins, Kelvin Guu, Roey Yogev, Xiaochen Cai, Alessandro Agostini, Maulik Shah, Hung Nguyen, Noah Ó Donnaile, Sébastien Pereira, Linda Friso, Adam Stambler, Chenkai Kuang, Yan Romanikhin, Mark Geller, ZJ Yan, Kane Jang, Cheng-Chun Lee, Wojciech Fica, Eric Malmi, Qijun Tan, Dan Banica, Daniel Balle, Ryan Pham, Yanping Huang, Diana Avram, Hongzhi Shi, Jasjot Singh, Chris Hidey, Niharika Ahuja, Pranab Saxena, Dan Dooley, Srividya Pranavi Potharaju, Eileen O’Neill, Anand Gokulchandran, Ryan Foley, Kai Zhao, Mike Dusenberry, Yuan Liu, Pulkit Mehta, Ragha Kotikalapudi, Chalence Safranek-Shrader, Andrew Goodman, Joshua Kessinger, Eran Globen, Prateek Kolhar, Chris Gorgolewski, Ali Ibrahim, Yang Song, Ali Eichenbaum, Thomas Brovelli, Sahitya Potluri, Preethi Lahoti, Cip Baetu, Ali Ghorbani, Charles Chen, Andy Crawford, Shalini Pal, Mukund Sridhar, Petru Gurita, Asier Mujika, Igor Petrovski, Pierre-Louis Cedoz, Chenmei Li, Shiyuan 18 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Chen, Niccolò Dal Santo, Siddharth Goyal, Jitesh Punjabi, Karthik Kappaganthu, Chester Kwak, Pallavi LV, Sarmishta Velury, Himadri Choudhury, Jamie Hall, Premal Shah, Ricardo Figueira, Matt Thomas, Minjie Lu, Ting Zhou, Chintu Kumar, Thomas Jurdi, Sharat Chikkerur, Yenai Ma, Adams Yu, Soo Kwak, Victor Ähdel, Sujeevan Rajayogam, Travis Choma, Fei Liu, Aditya Barua, Colin Ji, Ji Ho Park, Vincent Hellendoorn, Alex Bailey, Taylan Bilal, Huanjie Zhou, Mehrdad Khatir, Charles Sutton, Wojciech Rzadkowski, Fiona Macintosh, Roopali Vij, Konstantin Shagin, Paul Medina, Chen Liang, Jinjing Zhou, Pararth Shah, Yingying Bi, Attila Dankovics, Shipra Banga, Sabine Lehmann, Marissa Bredesen, Zifan Lin, John Eric Hoffmann, Jonathan Lai, Raynald Chung, Kai Yang, Nihal Balani, Arthur Bražinskas, Andrei Sozanschi, Matthew Hayes, Héctor Fernández Alcalde, Peter Makarov, Will Chen, Antonio Stella, Liselotte Snijders, Michael Mandl, Ante Kärrman, Paweł Nowak, Xinyi Wu, Alex Dyck, Krishnan Vaidyanathan, Raghavender R, Jessica Mallet, Mitch Rudominer, Eric Johnston, Sushil Mittal, Akhil Udathu, Janara Christensen, Vishal Verma, Zach Irving, Andreas Santucci, Gamaleldin Elsayed, Elnaz Davoodi, Marin Georgiev, Ian Tenney, Nan Hua, Geoffrey Cideron, Edouard Leurent, Mahmoud Alnahlawi, Ionut Georgescu, Nan Wei, Ivy Zheng, Dylan Scandinaro, Heinrich Jiang, Jasper Snoek, Mukund Sundararajan, Xuezhi Wang, Zack Ontiveros, Itay Karo, Jeremy Cole, Vinu Rajashekhar, Lara Tumeh, Eyal Ben-David, Rishub Jain, Jonathan Uesato, Romina Datta, Oskar Bunyan, Shimu Wu, John Zhang, Piotr Stanczyk, Ye Zhang, David Steiner, Subhajit Naskar, Michael Azzam, Matthew Johnson, Adam Paszke, Chung-Cheng Chiu, Jaume Sanchez Elias, Afroz Mohiuddin, Faizan Muhammad, Jin Miao, Andrew Lee, Nino Vieillard, Jane Park, Jiageng Zhang, Jeff Stanway, Drew Garmon, Abhijit Karmarkar, Zhe Dong, Jong Lee, Aviral Kumar, Luowei Zhou, Jonathan Evens, William Isaac, Geoffrey Irving, Edward Loper, Michael Fink, Isha Arkatkar, Nanxin Chen, Izhak Shafran, Ivan Petrychenko, Zhe Chen, Johnson Jia, Anselm Levskaya, Zhenkai Zhu, Peter Grabowski, Yu Mao, Alberto Magni, Kaisheng Yao, Javier Snaider, Norman Casagrande, Evan Palmer, Paul Suganthan, Alfonso Castaño, Irene Giannoumis, Wooyeol Kim, Mikołaj Rybi ́ nski, Ashwin Sreevatsa, Jennifer Prendki, David Soergel, Adrian Goedeckemeyer, Willi Gierke, Mohsen Jafari, Meenu Gaba, Jeremy Wiesner, Diana Gage Wright, Yawen Wei, Harsha Vashisht, Yana Kulizhskaya, Jay Hoover, Maigo Le, Lu Li, Chimezie Iwuanyanwu, Lu Liu, Kevin Ramirez, Andrey Khorlin, Albert Cui, Tian LIN, Marcus Wu, Ricardo Aguilar, Keith Pallo, Abhishek Chakladar, Ginger Perng, Elena Allica Abellan, Mingyang Zhang, Ishita Dasgupta, Nate Kushman, Ivo Penchev, Alena Repina, Xihui Wu, Tom van der Weide, Priya Ponnapalli, Caroline Kaplan, Jiri Simsa, Shuangfeng Li, Olivier Dousse, Jeff Piper, Nathan Ie, Rama Pasumarthi, Nathan Lintz, Anitha Vijayakumar, Daniel Andor, Pedro Valenzuela, Minnie Lui, Cosmin Paduraru, Daiyi Peng, Katherine Lee, Shuyuan Zhang, Somer Greene, Duc Dung Nguyen, Paula Kurylowicz, Cassidy Hardin, Lucas Dixon, Lili Janzer, Kiam Choo, Ziqiang Feng, Biao Zhang, Achintya Singhal, Dayou Du, Dan McKinnon, Natasha Antropova, Tolga Bolukbasi, Orgad Keller, David Reid, Daniel Finchelstein, Maria Abi Raad, Remi Crocker, Peter Hawkins, Robert Dadashi, Colin Gaffney, Ken Franko, Anna Bulanova, Rémi Leblond, Shirley Chung, Harry Askham, Luis C. Cobo, Kelvin Xu, Felix Fischer, Jun Xu, Christina Sorokin, Chris Alberti, Chu-Cheng Lin, Colin Evans, Alek Dimitriev, Hannah Forbes, Dylan Banarse, Zora Tung, Mark Omernick, Colton Bishop, Rachel Sterneck, Rohan Jain, Jiawei Xia, Ehsan Amid, Francesco Piccinno, Xingyu Wang, Praseem Banzal, Daniel J. Mankowitz, Alex Polozov, Victoria Krakovna, Sasha Brown, MohammadHossein Bateni, Dennis Duan, Vlad Firoiu, Meghana Thotakuri, Tom Natan, Matthieu Geist, Ser tan Girgin, Hui Li, Jiayu Ye, Ofir Roval, Reiko Tojo, Michael Kwong, James Lee-Thorp, Christopher Yew, Danila Sinopalnikov, Sabela Ramos, John Mellor, Abhishek Sharma, Kathy Wu, David Miller, Nicolas Sonnerat, Denis Vnukov, Rory Greig, Jennifer Beattie, Emily Caveness, Libin Bai, Julian Eisenschlos, Alex Korchemniy, Tomy Tsai, Mimi Jasarevic, Weize Kong, Phuong Dao, Zeyu Zheng, Frederick Liu, Rui Zhu, Tian Huey Teh, Jason Sanmiya, Evgeny Gladchenko, Nejc Trdin, Daniel Toyama, Evan Rosen, Sasan Tavakkol, Linting Xue, Chen Elkind, Oliver Woodman, John Carpenter, George Papamakarios, Rupert Kemp, Sushant Kafle, Tanya Grunina, Rishika Sinha, Alice Talbert, Diane Wu, Denese Owusu-Afriyie, Chloe Thornton, Jordi Pont-Tuset, Pradyumna Narayana, Jing Li, Saaber Fatehi, John Wieting, Omar Ajmeri, Benigno Uria, Yeongil Ko, Laura Knight, Amélie Héliou, Ning Niu, Shane Gu, Chenxi Pang, Yeqing Li, Nir Levine, Ariel Stolovich, Rebeca Santamaria-Fernandez, Sonam Goenka, Wenny Yustalim, Robin Strudel, Ali Elqursh, Charlie Deck, Hyo Lee, Zonglin Li, Kyle Levin, Raphael Hoffmann, Dan Holtmann-Rice, Olivier Bachem, Sho Arora, Christy Koh, Soheil Hassas Yeganeh, Siim Põder, Mukarram Tariq, Yanhua Sun, Lucian Ionita, Mojtaba Seyedhosseini, Pouya Tafti, Zhiyu Liu, Anmol Gulati, Jasmine Liu, Xinyu Ye, Bart Chrzaszcz, Lily Wang, Nikhil Sethi, Tianrun Li, Ben Brown, Shreya Singh, Wei Fan, Aaron Parisi, Joe Stanton, Vinod Koverkathu, Christopher A. Choquette-Choo, Yunjie Li, TJ Lu, Prakash Shroff, Mani Varadarajan, Sanaz Bahargam, Rob Willoughby, David Gaddy, Guillaume Desjardins, Marco Cornero, Brona Robenek, Bhavishya Mittal, Ben Albrecht, Ashish Shenoy, Fedor Moiseev, Henrik Jacobsson, Alireza Ghaffarkhah, Morgane Rivière, Alanna Walton, Clément Crepy, Alicia Parrish, Zongwei Zhou, Clement Farabet, Carey Radebaugh, Praveen Srinivasan, Claudia van der Salm, Andreas Fidjeland, Salvatore Scellato, Eri Latorre-Chimoto, Hanna Klimczak-Pluci ́ nska, David Bridson, Dario de Cesare, Tom Hudson, Piermaria Mendolicchio, Lexi Walker, Alex Morris, Matthew Mauger, Alexey Guseynov, Alison Reid, Seth Odoom, Lucia Loher, Victor Cotruta, Madhavi Yenugula, Dominik Grewe, Anastasia Petrushkina, Tom Duerig, Antonio Sanchez, Steve Yadlowsky, Amy Shen, Amir Globerson, Lynette Webb, Sahil Dua, Dong Li, Surya Bhupatiraju, Dan Hurt, 19 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Haroon Qureshi, Ananth Agarwal, Tomer Shani, Matan Eyal, Anuj Khare, Shreyas Rammohan Belle, Lei Wang, Chetan Tekur, Mihir Sanjay Kale, Jinliang Wei, Ruoxin Sang, Brennan Saeta, Tyler Liechty, Yi Sun, Yao Zhao, Stephan Lee, Pandu Nayak, Doug Fritz, Manish Reddy Vuyyuru, John Aslanides, Nidhi Vyas, Martin Wicke, Xiao Ma, Evgenii Eltyshev, Nina Martin, Hardie Cate, James Manyika, Keyvan Amiri, Yelin Kim, Xi Xiong, Kai Kang, Florian Luisier, Nilesh Tripuraneni, David Madras, Mandy Guo, Austin Waters, Oliver Wang, Joshua Ainslie, Jason Baldridge, Han Zhang, Garima Pruthi, Jakob Bauer, Feng Yang, Riham Mansour, Jason Gelman, Yang Xu, George Polovets, Ji Liu, Honglong Cai, Warren Chen, XiangHai Sheng, Emily Xue, Sherjil Ozair, Christof Angermueller, Xiaowei Li, Anoop Sinha, Weiren Wang, Julia Wiesinger, Emmanouil Koukoumidis, Yuan Tian, Anand Iyer, Madhu Gurumurthy, Mark Goldenson, Parashar Shah, MK Blake, Hongkun Yu, Anthony Urbanowicz, Jennimaria Palomaki, Chrisantha Fernando, Ken Durden, Harsh Mehta, Nikola Momchev, Elahe Rahimtoroghi, Maria Georgaki, Amit Raul, Sebastian Ruder, Morgan Redshaw, Jinhyuk Lee, Denny Zhou, Komal Jalan, Dinghua Li, Blake Hechtman, Parker Schuh, Milad Nasr, Kieran Milan, Vladimir Mikulik, Juliana Franco, Tim Green, Nam Nguyen, Joe Kelley, Aroma Mahendru, Andrea Hu, Joshua Howland, Ben Vargas, Jeffrey Hui, Kshitij Bansal, Vikram Rao, Rakesh Ghiya, Emma Wang, Ke Ye, Jean Michel Sarr, Melanie Moranski Preston, Madeleine Elish, Steve Li, Aakash Kaku, Jigar Gupta, Ice Pasupat, Da-Cheng Juan, Milan Someswar, Tejvi M., Xinyun Chen, Aida Amini, Alex Fabrikant, Eric Chu, Xuanyi Dong, Amruta Muthal, Senaka Buthpitiya, Sarthak Jauhari, Urvashi Khandelwal, Ayal Hitron, Jie Ren, Larissa Rinaldi, Shahar Drath, Avigail Dabush, Nan-Jiang Jiang, Harshal Godhia, Uli Sachs, Anthony Chen, Yicheng Fan, Hagai Taitelbaum, Hila Noga, Zhuyun Dai, James Wang, Jenny Hamer, Chun-Sung Ferng, Chenel Elkind, Aviel Atias, Paulina Lee, Vít Listík, Mathias Carlen, Jan van de Kerkhof, Marcin Pikus, Krunoslav Zaher, Paul Müller, Sasha Zykova, Richard Stefanec, Vitaly Gatsko, Christoph Hirnschall, Ashwin Sethi, Xingyu Federico Xu, Chetan Ahuja, Beth Tsai, Anca Stefanoiu, Bo Feng, Keshav Dhandhania, Manish Katyal, Akshay Gupta, Atharva Parulekar, Divya Pitta, Jing Zhao, Vivaan Bhatia, Yashodha Bhavnani, Omar Alhadlaq, Xiaolin Li, Peter Danenberg, Dennis Tu, Alex Pine, Vera Filippova, Abhipso Ghosh, Ben Limonchik, Bhargava Urala, Chaitanya Krishna Lanka, Derik Clive, Edward Li, Hao Wu, Kevin Hongtongsak, Ianna Li, Kalind Thakkar, Kuanysh Omarov, Kushal Majmundar, Michael Alverson, Michael Kucharski, Mohak Patel, Mudit Jain, Maksim Zabelin, Paolo Pelagatti, Rohan Kohli, Saurabh Kumar, Joseph Kim, Swetha Sankar, Vineet Shah, Lakshmi Ramachandruni, Xiangkai Zeng, Ben Bariach, Laura Weidinger, Tu Vu, Alek Andreev, Antoine He, Kevin Hui, Sheleem Kashem, Amar Subramanya, Sissie Hsiao, Demis Hassabis, Koray Kavukcuoglu, Adam Sadovsky, Quoc Le, Trevor Strohman, Yonghui Wu, Slav Petrov, Jeffrey Dean, and Oriol Vinyals. Gemini: A Family of Highly Capable Multimodal Models, 2023. Version Number: 5. URL: https://arxiv.org/abs/2312.11805, doi:10.48550/ARXIV.2312.11805. [38] Chin-Yew Lin. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URL:https: //aclanthology.org/W04-1013/. [39]Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-Eval: NLG Eval- uation using Gpt-4 with Better Human Alignment. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore, December 2023. Association for Computational Linguistics. URL:https://aclanthology.org/ 2023.emnlp-main.153/, doi:10.18653/v1/2023.emnlp-main.153. [40] Potsawee Manakul, Adian Liusie, and Mark Gales. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9004–9017, Singapore, December 2023. Association for Computational Linguistics. URL:https://aclanthology.org/ 2023.emnlp-main.557/, doi:10.18653/v1/2023.emnlp-main.557. [41] Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China, November 2019. Association for Computational Linguistics. URL:https://aclanthology.org/D19-1410/,doi:10.18653/v1/D19-1410. [42]Peter J. Rousseeuw. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20:53–65, November 1987. URL:https://w.sciencedirect. com/science/article/pii/0377042787901257, doi:10.1016/0377-0427(87)90125-7. [43] David L. Davies and Donald W. Bouldin. A Cluster Separation Measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-1(2):224–227, April 1979. URL:https://ieeexplore.ieee.org/ document/4766909, doi:10.1109/TPAMI.1979.4766909. [44] T. Cali ́ nski and J Harabasz. A dendrite method for cluster analysis. Communications in Statistics, 3(1):1–27, January 1974. _eprint: https://doi.org/10.1080/03610927408827101. doi:10.1080/03610927408827101. 20 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs [45]Jianmo Ni, Jiacheng Li, and Julian McAuley. Justifying Recommendations using Distantly-Labeled Reviews and Fine-Grained Aspects. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 188–197, Hong Kong, China, November 2019. Association for Computational Linguistics. URL:https://aclanthology.org/D19-1018/, doi:10.18653/v1/D19-1018. 21 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Appendices Appendix A Data Characteristics (a) Google Business Reviews HIGHLOW Metric10051–990–5010051–990–50 # Business0 (0%)43 (17.1%)8 (3.2%)0 (0%)32 (12.7%)168 (66.9%) # Datapoints0 (0%)93,640 (77%)5,518 (5%) 0 (0%)8,734 (7%)13,934 (11%) (b) Amazon Product Reviews HIGHLOW Metric10051–990–5010051–990–50 # Stores87 (0.5%)1,053 (6.4%)1,983 (12.1%) 0 (0%)5 (0.0%)13,267 (80.9%) # Datapoints17,671 (11%)60,002 (39%)42,503 (27%)0 (0%)44 (0%)35,525 (23%) (c) Goodreads Book Reviews HIGHLOW Metric10051–990–5010051–990–50 # Books0 (0%)230 (1.2%)4,969 (25.2%) 0 (0%)0 (0%)14,526 (73.6%) # Datapoints0 (0%)18,720 (12%)93,213 (59%)0 (0%)0 (0%)45,474 (29%) Values are presented as Absolute Count (Percentage of Total). High and Low represent Businesses/Products/Books with Review Volume above the Mean or below the Mean respectively. • 100: Reviews present in every available quarter of dataset’s lifecycle. • 51-99: Reviews present in more than half, but not all, quarters • 0-50: Reviews present in half or fewer quarters 22 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Appendix B Design of Experiments Table B.2: Experimental Design: Scenarios, Datasets, and Design Rationale ScenarioDatasetOutputDesign Rationale Base (Themes, Stories, Clusters) AmazonDivided intoN Themes; for each theme, we report on # Stories, # Clusters, and # Datapoints Allows comparison of three industry-standard datasets (ALL, w/o Irrelevant, w/o Outliers) with varying context summaries Google Goodreads No context (Baseline) Amazon # Datapoints per category Categorization using direct LLM input with no context Google Goodreads Unweighted Context (SSAS) Amazon # Datapoints per categoryCategorization using standard SSAS contextGoogle Goodreads Weighted Context (wSSAS) Amazon # Datapoints per category Categorization using enhanced wSSAS context Google Goodreads 23 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Appendix C Themes, Stories, Clusters breakdown Table C.1: Detailed Data Distribution for Google Business Reviews across Noise Removal stages All Data W/o Irrelevant Data W/o Irrelevant & Outlier Data Th. #St. #Cl. #DP Th. #St. #Cl. #DP Th. #St. #Cl. #DP -1 1 147 273 0 11 3,978 88,107 0 10 3,978 84,789 0 10 190 74,451 1 10 1,146 10,668 1 9 1,146 10,400 1 5 40 7,631 2 10 919 9,299 2 9 919 9,059 2 7 27 7,066 3 10 564 3,257 3 9 564 3,184 3 6 9 1,917 4 10 435 2,506 4 9 435 2,471 4 4 14 1,386 5 10 286 2,145 5 9 286 2,113 5 7 10 1,563 6 10 284 1,556 6 9 284 1,531 6 4 5 950 7 9 311 992 7 8 311 965 7 2 3 230 8 2 182 842 8 2 182 842 8 1 3 444 9 11 196 804 9 10 196 803 9 4 4 200 10 5 70 650 10 5 70 650 10 2 3 486 11 5 129 370 11 5 129 370 11 2 2 110 12 5 105 198 12 4 105 196 13 4 52 159 13 3 52 155 15 113 8,804 121,826 14 101 8,657 117,528 12 54 310 96,434 Note: Th. refers to the Theme ID and #St, #Cl, #DP refer to the number of stories, clusters and data points in the theme. Theme -1 represents unclassified noise. The Subtotal row (Bottom) reflects the number of active themes, total stories, clusters, and datapoints retained in each stage. 24 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table C.2: Detailed Data Distribution for Amazon Product Reviews across Noise Removal stages All Data W/o Irrelevant Data W/o Irrelevant & Outlier Data Th. #St. #Cl. #DP Th. #St. #Cl. #DP Th. #St. #Cl. #DP -1 1 344 637 0 11 6,290 46,974 0 10 6,290 44,945 0 10 649 35,384 1 10 4,814 41,568 1 9 4,813 39,507 1 9 404 32,664 2 10 2,339 15,464 2 9 2,339 15,095 2 9 243 11,456 3 10 1,615 14,701 3 9 1,615 14,511 3 9 195 12,067 4 10 2,273 14,298 4 9 2,273 14,033 4 9 204 10,445 5 10 1,602 10,803 5 9 1,602 10,696 5 9 160 8,234 6 10 1,249 4,204 6 9 1,249 4,138 6 8 63 2,242 7 10 724 2,351 7 9 724 2,163 7 8 34 1,215 8 4 668 2,251 8 3 668 2,249 8 3 46 1,200 9 3 265 803 9 2 265 795 9 2 21 419 10 2 216 515 10 2 216 515 10 1 5 195 11 6 215 486 11 6 215 486 11 3 3 148 12 4 68 396 12 4 68 396 12 4 5 306 13 2 109 294 13 2 109 294 13 2 2 127 15 103 22,791 155,745 14 92 22,446 149,823 14 86 2,034 116,102 Note: Th. refers to the Theme ID and #St, #Cl, #DP refer to the number of stories, clusters and data points in the theme. Theme -1 represents unclassified noise. The Subtotal row (Bottom) reflects the number of active themes, total stories, clusters, and datapoints retained in each stage. 25 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table C.3: Detailed Data Distribution for Goodreads Book Reviews across Noise Removal stages All Data W/o Irrelevant Data W/o Irrelevant & Outlier Data Th. #St. #Cl. #DP Th. #St. #Cl. #DP Th. #St. #Cl. #DP -1 1 3,701 12,443 0 11 11,182 116,359 0 10 11,182 115,048 0 10 1,349 99,216 1 10 1,143 4,733 1 9 1,143 4,691 1 6 78 3,081 2 3 218 605 2 3 218 605 2 3 14 272 3 10 1,480 6,272 3 9 1,480 6,167 3 9 117 4,106 4 10 905 2,335 4 9 905 2,263 4 9 43 1,031 5 2 265 714 5 2 265 714 5 2 11 356 6 7 760 4,096 6 6 760 4,095 6 4 55 3,033 7 5 451 1,442 7 5 451 1,442 7 5 30 832 8 7 715 2,296 8 6 715 2,294 8 6 42 1,262 9 10 775 2,401 9 9 775 2,370 9 9 46 1,369 10 4 791 3,711 10 4 791 3,711 10 3 57 2,575 12 80 22,386 157,407 11 72 18,685 143,400 11 66 1,842 117,133 Note: Th. refers to the Theme ID and #St, #Cl, #DP refer to the number of stories, clusters and data points in the theme. Theme -1 indicates unclassified noise. The Subtotal row (Bottom) tracks the reduction in active themes and the refinement of cluster density through successive filtering stages. 26 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Appendix D Context Summary Evaluation using QAG / G-Eval Metrics Table D.1: Qualitative Comparison of LLM Summarization: Google Business Reviews Story ID: 85 Weighted Context Summary (wSSAS)Unweighted Context Summary (SSAS) Mixed Restaurant ExperiencesMixed Restaurant Experience The data presents a range of restaurant experiences. Some re- views are overwhelmingly positive, praising food quality and service, with one customer declaring a restaurant their new fa- vorite. Conversely, other reviews highlight negative experiences, including cold service, incorrect orders, and long wait times. One review focuses on the speed and efficiency of a drive-thru. Other reviews express contentment despite a lack of exceptional qual- ity, while another highlights the recurring problem of incomplete orders. One review describes an uncomfortable situation due to a coworker argument. The text also mentions a positive first visit and emphasizes the importance of consistently leaving situations feeling satisfied. Finally, the text describes an environment that is "ridiculously cold" and a persistent state of dissatisfaction. The text presents a mixed restaurant experience, highlighting both positive and negative aspects. The food, particularly caramel rolls, receives high praise, described as "amazing" and the "best in the world." However, the text also details issues with service, including a cold and inattentive waitress, incorrect orders, and long wait times. The data indicates a recurring problem with incomplete orders, suggesting operational inefficiencies. Despite the food’s quality, the inconsistent service and wait times detract from the overall dining experience, creating a trade-off between food quality and service efficiency. Quantitative Evaluation:Quantitative Evaluation: • QAG Pre-Triage: 1/5 • QAG Post-Triage: 2/5 • G-Eval Scores: ◦ Coherence: 0.7 ◦ Relevance: 1.0 ◦ Fluency: 1.0 ◦ Consistency: 0.5 • QAG Pre-Triage: 3/5 • QAG Post-Triage: 2/5 • G-Eval Scores: ◦ Coherence: 0.8 ◦ Relevance: 0.8 ◦ Fluency: 0.9 ◦ Consistency: 0.5 27 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table D.2: Qualitative Comparison of LLM Summarization: Amazon Product Reviews Story ID: 39 Weighted Context Summary (wSSAS)Unweighted Context Summary (SSAS) Product Review SummariesPain Relief Experiences: A Diverse Overview The text presents a collection of product reviews, primarily fo- cusing on pain relief. Many reviews are overwhelmingly positive, highlighting effective relief from various conditions like hand pain, arthritis, and back pain, with some users experiencing al- most immediate results. Products mentioned include pain relief patches, neck traction devices, and topical applications. Con- versely, some reviews report negative experiences, citing ineffec- tiveness, functional issues, and a lack of relief. The reviews also touch on the ease of use, convenience, and value of the products, with some users expressing a desire to repurchase. The data presents a wide array of experiences with pain relief products and methods. Positive outcomes are reported for IcyHot, neck traction devices, arthritis treatments, horse massages, red light belts, pain relief patches, and various sprays and creams, with many users experiencing significant and lasting relief. Con- versely, negative experiences are documented with ineffective products for finger pain, runner’s toe, neck issues, and back pain, along with issues like premature shoe failure and incorrect or- ders. Some products offer temporary relief, while others provide more sustained benefits. The effectiveness of each method varies, highlighting the diverse nature of pain management and the im- portance of finding the right solution. Quantitative Evaluation:Quantitative Evaluation: • QAG Pre-Triage: 2 / 5 • QAG Post-Triage: 3 / 5 • G-Eval Scores: ◦ Coherence: 0.8 ◦ Relevance: 1.0 ◦ Fluency: 0.9 ◦ Consistency: 0.5 • QAG Pre-Triage: 3 / 5 • QAG Post-Triage: 3 / 5 • G-Eval Scores: ◦ Coherence: 0.1 ◦ Relevance: 1.0 ◦ Fluency: 1.0 ◦ Consistency: 0.5 Story ID: 47 Weighted Context Summary (wSSAS)Unweighted Context Summary (SSAS) Shoe Horn Reviews: A Mixed BagShoe Horn Reviews: Functionality and Design The text presents a collection of reviews for shoe horns, revealing a wide range of experiences. Many users praise the telescop- ing shoe horns for their ease of use, especially for those with mobility issues, highlighting their sturdiness, adjustability, and travel-friendly design. However, other reviews detail significant design flaws, including flimsy construction, ineffective locking mechanisms, and handles that detach. Some users found the shoe horns helpful after surgery or for everyday use, while others considered them a waste of money due to their poor functionality. The text compiles numerous reviews of shoe horns, highlight- ing both positive and negative experiences. Many users praise telescoping shoe horns for their sturdiness, ease of use, and abil- ity to assist individuals with mobility issues, particularly those with back problems, arthritis, or recovering from surgery. These reviews often emphasize the product’s adjustability and portabil- ity. Conversely, several reviews criticize shoe horns for design flaws, such as flimsy materials, unstable handles, and ineffective spring mechanisms, rendering them difficult or impossible to use. Some users also report issues with durability, with parts breaking shortly after purchase. Quantitative Evaluation:Quantitative Evaluation: • QAG Pre-Triage: 1 / 5 • QAG Post-Triage: 4 / 5 • G-Eval Scores: ◦ Coherence: 0.9 ◦ Relevance: 1.0 ◦ Fluency: 1.0 ◦ Consistency: 0.5 • QAG Pre-Triage: 2 / 5 • QAG Post-Triage: 4 / 5 • G-Eval Scores: ◦ Coherence: 0.8 ◦ Relevance: 1.0 ◦ Fluency: 0.9 ◦ Consistency: 0.5 28 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table D.3: Qualitative Comparison of LLM Summarization: Goodreads Book Reviews Story ID: 31 Weighted Context Summary (wSSAS)Unweighted Context Summary (SSAS) Book Review RoundupBook Review Roundup The text summarizes a collection of book reviews, revealing diverse reader experiences. Some reviewers praised engaging plots, well-developed characters, and unexpected twists, while others expressed disappointment with pacing, predictability, or unmet expectations. Reviews covered various genres, with some readers eager for sequels and others expressing mixed feelings, highlighting the subjective nature of reading. The text summarizes a collection of book reviews, encompass- ing diverse genres and reader experiences. Reviews range from enthusiastic praise, highlighting engaging plots, well-developed characters, and unexpected twists, to critical assessments, citing predictable storylines, underdeveloped characters, and unmet expectations. The reviews reflect the subjective nature of read- ing, with some readers finding books deeply moving and others disappointed by various aspects of the narratives. The text also includes reviews of specific books, such as a negative review of a book about the Zodiac Killer, a positive review of a book about a family with secrets, and a mixed review of a book about a romance. Quantitative Evaluation:Quantitative Evaluation: • QAG Pre-Triage: 1 / 5 • QAG Post-Triage: 3 / 5 • G-Eval Scores: ◦ Coherence: 1.0 ◦ Relevance: 1.0 ◦ Fluency: 1.0 ◦ Consistency: 0.5 • QAG Pre-Triage: 2 / 5 • QAG Post-Triage: 2 / 5 • G-Eval Scores: ◦ Coherence: 1.0 ◦ Relevance: 0.5 ◦ Fluency: 1.0 ◦ Consistency: 0.5 Theme ID: 10 Weighted Context Summary(wSSAS)Unweighted Context Summary (SSAS) Diverse Reader Experiences in Book ReviewsDiverse Reader Reactions to Books The text summarizes a collection of book reviews, revealing a wide spectrum of reader opinions. Some reviews express strong enjoyment, praising engaging plots, relatable characters, and unique settings, while others express disappointment, citing un- interesting characters, predictable plots, and writing styles that failed to resonate. The reviews highlight the subjective nature of reading, with readers responding differently to similar elements like romance, world-building, and character dynamics, reflecting varied preferences and critiques across different genres. The text summarizes a collection of book reviews, revealing a wide spectrum of reader opinions. Some reviews express strong positive sentiments, praising engaging plots, well-developed char- acters, and unique premises, while others criticize pacing, charac- ter development, and unsatisfying endings. The reviews highlight the subjective nature of reading, with some readers finding books captivating and others disappointed by various elements. Over- all, the data reflects a range of opinions across different genres, indicating varied reader preferences and levels of satisfaction. Quantitative Evaluation:Quantitative Evaluation: • QAG Pre-Triage: 3 / 5 • QAG Post-Triage: 5 / 5 • G-Eval Scores: ◦ Coherence: 0.8 ◦ Relevance: 1.0 ◦ Fluency: 1.0 ◦ Consistency: 0.5 • QAG Pre-Triage: 5 / 5 • QAG Post-Triage: 5 / 5 • G-Eval Scores: ◦ Coherence: 0.8 ◦ Relevance: 0.8 ◦ Fluency: 0.8 ◦ Consistency: 0.5 29 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Appendix E Sankey Plots Google Business Reviews (All Data) (a) No Context (Baseline) vs Weighted Context Summary (wSSAS) (b) No Context (Baseline) vs Unweighted Context Summary (SSAS) (c) Unweighted Context Summary (SSAS) vs Weighted Context Summary (wSSAS) Figure 5: Detailed Sankey diagrams showing cluster transitions for Google Business Reviews. 30 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table E.1: Distribution of Review Categories across Experimental Scenarios (Google Business Reviews) ScenarioCategory-Cluster TitlesAll Data W/o Irrelevant W/o Irrelevant & Outliers No context (Baseline) Customer Service and Food Quality Complaints17,18916,56613,302 Restaurant Customer Satisfaction and Positive Feedback 31,81831,11627,991 Restaurant Experience and Service Quality24,39223,75718,922 Restaurant Reviews and Dining Experience57,05754,53943,972 Unweighted Context (SSAS) Customer Dissat. & Service Failures21,57820,78616,443 Positive Restaurant Experiences23,18822,66620,238 Restaurant Experience and Operations73,00270,07455,766 Restaurant Service and Quality Complaints17,00616,33812,677 Weighted Context (wSSAS) Positive Dining Reviews41,19740,23735,958 Restaurant Experience and Food Quality59,04456,49844,028 Customer Dissat. & Service Failures21,57820,78616,443 31 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Amazon Product Reviews (All Data) (a) No Context (Baseline) vs Weighted Context Summary (wSSAS) (b) No Context (Baseline) vs Unweighted Context Summary(SSAS) (c) Unweighted Context Summary (SSAS) vs Weighted Context Summary (wSSAS) Figure 6: Detailed Sankey diagrams showing cluster transitions for Amazon Product Reviews. 32 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table E.2: Distribution of Review Categories across Experimental Scenarios (Amazon Product Reviews) ScenarioCategory-Cluster TitlesAll Data W/o Irrelevant W/o Irrelevant & Outliers No context (Baseline) Assistive Devices and Aids18,10517,71213,821 Grooming and Personal Care Products36,45135,30927,846 Cleaning and Maintenance Products11,62611,2419,121 Pain and Symptom Relief15,78815,11010,226 Personal Grooming and Hygiene33,06632,52626,922 Product Installation and User Experience61,49158,69946,492 Product Quality and Performance Issues39,24438,10530,391 Unweighted Context (SSAS) Beauty and Grooming Products23,92723,49219,245 Pain Relief and Symptom Management23,76423,02216,749 Product Defects and Customer Dissat.33,85332,67024,893 Positive Experiences and Reactions19,57018,14314,116 Product Functionality and Performance18,06016,76311,342 Weighted Context (wSSAS) Cleaning Products and Supplies15,03514,50811,668 Defective or Faulty Products27,17726,13420,044 Digestive and Gut Health Supplements16,82016,27811846 Masks and Accessories22,03221,56516,417 Product Reviews and Feedback50,69447,80136,837 33 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Goodreads Book Reviews (All Data) (a) No Context (Baseline) vs Weighted Context Summary (wSSAS) (b) No Context (Baseline) vs Unweighted Context Summary (SSAS) (c) Unweighted Context Summary (SSAS) vs Weighted Context Summary (wSSAS) Figure 7: Detailed Sankey diagrams showing cluster transitions for Goodreads Book Reviews. 34 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Table E.3: Distribution of Review Categories across Experimental Scenarios (Goodreads Book Reviews) ScenarioCategory-Cluster TitlesAll Data W/o Irrelevant W/o Irrelevant & Outliers No context (Baseline) Book Criticism and Appreciation45,46040,41331,147 Book Review Themes and Tropes85,63482,70873,520 Content Disappointment and Expectation26,31120,27712,464 Unweighted Context (SSAS) Book Review Criticism64,38860,58652,847 Book Review Focus44,49543,41636,769 Character Appreciation Focused Reviews19,02316,65913,108 Content Evaluation and Reaction10,9467,8894,000 Reader Disappointment and Enjoyment18,34714,76110,350 Weighted Context (wSSAS) Book Reviews and Criticism60,57455,73446,506 Book Series and Character Relationships46,24638,51127,622 Romance and Suspense50,48349,05742,912 35 wSSAS: A Framework for Improved Text Categorization and Summarization using LLMs Appendix F Technical Stack and Implementation Environment This section details the software, models, and mathematical frameworks utilized to implement and validate the wSSAS methodology. F.1 Large Language Models (LLMs) •Primary Inference Engine:Gemini 2.0 Flash Litewas utilized for hierarchical text summarization and categorization across Themes, Stories, and Clusters. •LLM-as-a-Judge: The same model facilitated the reference-free evaluation framework, performing Question- Answer Generation (QAG) and qualitative G-Eval scoring. F.2 Embedding Models and Vectorization • High-Dimensional Vectorization:text-embedding-005was employed to generate high-fidelity vector representations of the raw text, leveraging its advanced semantic grasp for processing heterogeneous datasets. •Semantic Similarity Engine: Thesentence-transformers/all-MiniLM-L6-v2model was utilized within the QAG framework to calculate cosine similarity between true and extracted responses. F.3 Internal Validation Metrics Cluster integrity was mathematically confirmed using three primary metrics to ensure the cohesion of the generated themes: 1. Silhouette Score: Measures internal cohesion (a) versus cluster separation (b). 2.Davies-Bouldin Index: Measures the average similarity between clusters. Lower scores indicate better separation between thematic groups. 3.Calinski-Harabasz (CH) Index: Evaluates the ratio of between-cluster dispersion to within-cluster dispersion (the Variance Ratio Criterion). 36