Paper deep dive
DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi Text
Kunwar Arpit Singh, Ankush Prakash, Haroon R Lone
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/20/2026, 10:46:39 PM
Summary
The paper introduces DohaScript, a large-scale, multi-writer dataset for continuous handwritten Hindi text in the Devanagari script. Collected from 531 unique contributors, the dataset features a parallel stylistic corpus design where all writers transcribe the same six traditional Hindi dohas (couplets). This controlled lexical content allows for the isolation of writer-specific variation, supporting tasks like handwriting recognition, writer identification, and style analysis. The dataset includes demographic metadata, rigorous quality curation using CNN-based classification and Laplacian variance, and page-level difficulty annotations to facilitate stratified benchmarking in low-resource script settings.
Entities (10)
Relation Signals (12)
DohaScript → collectedfrom → 531 contributors
confidence 98% · collected from 531 unique contributors
DohaScript → contains → Hindi
confidence 95% · DohaScript, a large scale, multi writer dataset of handwritten Hindi text
DohaScript → contains → Devanagari
confidence 95% · DohaScript is a large-scale, multi-writer dataset of handwritten Hindi text... Devanagari script
Devanagari → hasfeature → shirorekha
confidence 90% · characters are connected through a shared shirorekha (horizontal headline)
DohaScript → supportstask → style analysis
confidence 90% · supports tasks such as handwriting recognition, writer identification, style analysis, and generative modeling
DohaScript → supportstask → handwriting recognition
confidence 90% · supports tasks such as handwriting recognition, writer identification, style analysis, and generative modeling
DohaScript → supportstask → writer identification
confidence 90% · supports tasks such as handwriting recognition, writer identification, style analysis, and generative modeling
DohaScript → supportstask → generative modeling
confidence 90% · supports tasks such as handwriting recognition, writer identification, style analysis, and generative modeling
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Despite having hundreds of millions of speakers, handwritten Devanagari text remains severely underrepresented in publicly available benchmark datasets. Existing resources are limited in scale, focus primarily on isolated characters or short words, and lack controlled lexical content and writer level diversity, which restricts their utility for modern data driven handwriting analysis. As a result, they fail to capture the continuous, fused, and structurally complex nature of Devanagari handwriting, where characters are connected through a shared shirorekha (horizontal headline) and exhibit rich ligature formations. We introduce DohaScript, a large scale, multi writer dataset of handwritten Hindi text collected from 531 unique contributors. The dataset is designed as a parallel stylistic corpus, in which all writers transcribe the same fixed set of six traditional Hindi dohas (couplets). This controlled design enables systematic analysis of writer specific variation independent of linguistic content, and supports tasks such as handwriting recognition, writer identification, style analysis, and generative modeling. The dataset is accompanied by non identifiable demographic metadata, rigorous quality curation based on objective sharpness and resolution criteria, and page level layout difficulty annotations that facilitate stratified benchmarking. Baseline experiments demonstrate clear quality separation and strong generalization to unseen writers, highlighting the dataset's reliability and practical value. DohaScript is intended to serve as a standardized and reproducible benchmark for advancing research on continuous handwritten Devanagari text in low resource script settings.
Tags
Links
- Source: https://arxiv.org/abs/2602.18089v1
- Canonical: https://arxiv.org/abs/2602.18089v1
Trouble viewing inline? Open PDF directly →
Full Text
46,823 characters extracted from source content.
Expand or collapse full text
DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi Text Kunwar Arpit Singh ∗ kunwar22@iiserb.ac.in IISER Bhopal Bhopal, India Ankush Prakash ∗ ankushs22@iiserb.ac.in IISER Bhopal Bhopal, India Haroon R. Lone haroon@iiserb.ac.in IISER Bhopal Bhopal, India Abstract Despite having hundreds of millions of speakers, handwritten De- vanagari text remains severely underrepresented in publicly avail- able benchmark datasets. Existing resources are limited in scale, focus primarily on isolated characters or short words, and lack con- trolled lexical content and writer-level diversity, which restricts their utility for modern data-driven handwriting analysis. As a result, they fail to capture the continuous, fused, and structurally complex nature of Devanagari handwriting, where characters are connected through a shared shirorekha (horizontal headline) and exhibit rich ligature formations. We introduceDohaScript, a large-scale, multi-writer dataset of handwritten Hindi text collected from 531 unique contributors. The dataset is designed as a parallel stylistic corpus, in which all writ- ers transcribe the same fixed set of six traditional Hindidohas(cou- plets). This controlled design enables systematic analysis of writer- specific variation independent of linguistic content, and supports tasks such as handwriting recognition, writer identification, style analysis, and generative modeling. The dataset is accompanied by non-identifiable demographic metadata, rigorous quality curation based on objective sharpness and resolution criteria, and page-level layout difficulty annotations that facilitate stratified benchmarking. Baseline experiments demon- strate clear quality separation and strong generalization to unseen writers, highlighting the dataset’s reliability and practical value. DohaScriptis intended to serve as a standardized and reproducible benchmark for advancing research on continuous handwritten De- vanagari text in low-resource script settings. The dataset is pub- licly available at Google Drive 1 . ACM Reference Format: Kunwar Arpit Singh, Ankush Prakash, and Haroon R. Lone. 2024. Do- haScript: A Large-Scale Multi-Writer Dataset for Continuous Handwrit- ten Hindi Text. In.ACM, New York, NY, USA, 11pages.https://doi.org/ X.X ∗ Both authors contributed equally to this research. 1 https://drive.google.com/drive/folders/1v2pjEE0MUkcLRn7YEfz3cro9OUZIXi0g Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full cita- tion on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy other- wise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conference’17, Washington, DC, USA © 2024 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/26/08 https://doi.org/X.X 1 Introduction Communication has played a central role in human evolution, and handwriting remains one of its most enduring manifestations. De- spite the widespread adoption of digital technologies, handwritten text continues to be extensively used in education, administration, and everyday documentation. Prior studies show that handwriting engages perceptual and motor processes that support cognitive de- velopment and literacy acquisition [13,35]. Beyond its functional role, handwriting is a personal and evolving form of expression that changes over time under the influence of age, gender, and socio-cultural context, resulting in distinctive writing styles [28]. This inherent variability makes handwriting a challenging object for systematic and computational analysis [ 20,36]. Motivated by these characteristics, recent years have witnessed growing interest in handwriting-related research areas, including optical character recognition (OCR) [31], handwritten text recognition (HTR) [1], handwriting generation [6], and writer identification [36]. How- ever, most existing work focuses on Latin-script languages, leav- ing many widely used writing systems insufficiently studied. This imbalance is particularly evident for Indic languages such as Hindi, Marathi, Nepali, Sanskrit, and Konkani, which are writ- ten in the Devanagari script and collectively serve hundreds of mil- lions of speakers [ 4,27]. Progress in handwritten Devanagari text analysis remains limited, largely due to the absence of large-scale, diverse, and well-annotated public datasets [26]. Unlike Latin scripts, where characters are typically discrete and visually separated, De- vanagari exhibits complex character composition involving con- sonant ligatures, combining vowel marks, and a continuousshi- rorekha(horizontal headline) that connects characters within a word. As a result, words often appear as single connected units rather than sequences of isolated symbols, complicating segmentation and recognition [ 5]. Methods developed for character-level pro- cessing in Latin scripts therefore transfer poorly to Devanagari handwriting. Data scarcity poses a major obstacle to modern data-driven ap- proaches, particularly deep learning models that require large and diverse datasets for effective generalization [33]. Limited train- ing data restricts the ability of models to learn robust represen- tations and leads to poor performance under variations in stroke width, slant, spacing, and writing style. These challenges are fur- ther amplified in Devanagari due to its large character inventory, numerous conjunct forms, and positional diacritics [ 18]. Captur- ing this structural and stylistic diversity reliably requires substan- tially larger and more heterogeneous datasets. Despite Hindi being the fourth most widely spoken language worldwide, with over 400 million native speakers and more than 600 million users of Devanagari-based languages [ 9,17], progress arXiv:2602.18089v1 [cs.CV] 20 Feb 2026 Conference’17, July 2017, Washington, DC, USAKunwar Arpit Singh, Ankush Prakash, and Haroon R. Lone in handwriting recognition and generation remains slow. The lack of large, publicly available datasets has fragmented research ef- forts, forcing individual studies to rely on private or institution- specific collections [ 8,24]. This fragmentation hinders fair compar- ison across methods and makes it difficult to assess true progress or establish reliable benchmarks. Existing public Devanagari handwriting datasets further limit progress. Resources such as IIIT-HW-Words, CALAM, and CHIPS primarily focus on isolated characters or short word-level units [ 12, 14,16]. While valuable for basic recognition tasks, these datasets do not adequately reflect realistic handwriting scenarios involv- ing continuous text and contextual dependencies across words and lines [11]. Moreover, most datasets provide limited writer cover- age and lack detailed demographic metadata, restricting the study of writer-specific variation and sociolinguistic patterns [36]. They also fail to meet the requirements of modern generative and style- transfer models, which rely on repeated lexical content across many writers. To address these limitations, we introduceDohaScript, a large- scale, multi-writer dataset of handwritten Hindi text collected from 531 unique contributors. In contrast to existing datasets that em- phasize lexical diversity,DohaScriptemphasizes lexical consistency across a large writer population. Each participant transcribed the same set of six traditional Hindidohas(couplets), yielding 531 sam- ples of a fixed 89-word corpus. This design forms aparallel stylis- tic corpusthat enables the isolation of handwriting style from lin- guistic content, supporting tasks such as writer identification, style analysis, and handwriting synthesis. DohaScriptfurther offers three key features. First, every sample is accompanied by non-identifiable demographic metadata, includ- ing age, gender, and regional state, enabling population-level and sociolinguistic analyses. Second, we employ a high-fidelity cura- tion pipeline based on Laplacian-variance assessment and CNN- driven quality classification to ensure sufficient stroke clarity and visual fidelity for modern learning architectures. Third, we provide page-level difficulty annotations (Easy, Medium, Complex) based on layout irregularity, facilitating stratified benchmarking and ro- bust evaluation. 2 Related datasets Offline handwritten text recognition has benefited from curated datasets across different scripts and annotation granularities. For Devanagari and other Indic scripts, most publicly available datasets have historically emphasized character- or word-level benchmarks rather than longer continuous handwritten documents. The IIIT-HW-Words dataset provides a large-scale benchmark of handwritten word images across multiple Indic scripts and has been widely used for word recognition and script-specific OCR [ 12]. However, its isolated-word formulation does not capture higher-level structural and contextual dependencies present in paragraph- or page-level handwriting, which are important for context-aware OCR and discourse-level analysis [ 2]. The CHIPS dataset, released within the PLATTER framework, represents an important step toward page-level Indic handwriting resources by providing synthetically generated page images con- structed through the aggregation and spatial arrangement of word- level samples from existing datasets, enabling layout-aware recog- nition research [ 14]. Similarly, datasets such as CALAM focus on isolated handwritten characters for Indic and Perso-Arabic scripts, supporting character-level recognition studies [16]. Beyond recognition-oriented benchmarks, writer-centric research directions such as writer identification, handwriting biometrics, style transfer, and writer-conditioned generation remain under- supported in existing Devanagari datasets due to the absence of standardized stylistic annotations and limited focus on generative modeling [29]. Moreover, generative handwriting synthesis mod- els typically require repeated lexical content across many writers, a requirement that current Indic resources only partially satisfy [36]. Additional datasets such as IIIT-HW-Dev, Parimal Hindi, and historical collections including AnciDev and Old Nepali further highlight the diversity of available resources, though they differ substantially in scope, annotation style, and intended research ap- plications (see Table 1). Recognition-focused datasets covering seg- mentation and classification tasks have also been introduced in re- cent years [21]. Overall, existing resources remain fragmented and insufficient for comprehensive analysis of continuous, stylistically diverse handwritten Devanagari text. 3 Data collection protocol 3.1 Study Design Doha:We selected six traditional Hindidohasthat are in De- vanagari script as the ideal text for handwriting collection.Dohas are rhyming couplets, which is very common in Hindi Literature. The selecteddohaswere composed by the renowned poets Kabir Das, Rahim Das, and Tulsidas, who represent traditional and classi- cal Hindi poetry that is culturally familiar to native speakers.These dohasare widely taught in Indian school, making them appropriate for broad participant demographics while avoiding copyright con- cerns. The combineddohasshow moderate linguistic complexity that is suitable for adult native speakers. The vocabulary includes common words that is used daily. The text corpus comprises of sixdohas, as shown in Figure 1. The English translations of thesedohasare provided in AppendixC. Thesedohaswere specially chosen to provide all the Devanagari Character. The combined text contains 89 words and 361 charac- ters (excluding spaces), with 55 distinct characters. It represents the diverse orthographic feature of Devanagari scripts. Figure 2summarizes the phonetic and orthographic coverage of the selecteddohas. The text includes 38 consonants spanning all five articulation categories (velar, palatal, retroflex, dental, and labial), along with semivowels and sibilants. It further covers all primary vowel classes (short, long, vocalic, and diphthongs), 13 vowel diacritics (mātrās), the halant for consonant clusters, nasal- ization marks, and the common conjunct characters. These combined consonants, called ligatures, are a special fea- ture of Brahmic scripts and are significant. DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi TextConference’17, July 2017, Washington, DC, USA Table 1: Quantitative comparison of existing Indic handwriting datasets with the proposed dataset. DatasetScript(s)Page-level Samples (# Pages) # WritersText TypeDemographics IIIT-HW-Words [12]Indic (multi)– (word-only dataset)135Isolated wordsNo IIIT-HW-Dev [8]Hindi (Devanagari)– (word-only dataset)12Continuous wordsLimited Numerals/Vowels Dataset [21]Devanagari– (character-only dataset)∼3800 subjectsIsolated charactersYes CHIPS (PLATTER) [14]Indic (multi)25,290 pages–ContinuousNo Parimal Hindi [7]Hindi (Devanagari)500 pages10Continuous paragraphsYes AnciDev [30]Ancient Devanagari500 pages–Continuous linesNo DohaScript (ours)Hindi (Devanagari)531 pages531Continuous coupletsYes Figure1: Textcorpuscomprisingsixdohasusedinthestudy. Figure 2: Comprehensive coverage of Devanagari charac- ters in the selecteddohas, including consonants by articu- lation category, vowel classes, vowel diacritics (mātrās), ha- lant, and common conjuncts. 3.2 Data Collection The dataset was collected from 531 writers, each contributing one handwritten A4 page containing the same fixed set ofdohas. Sam- ples were submitted through a Google Form as scanned images or mobile photographs, while a small number of physical sheets were digitized using a scanner. Both raw images and standardized versions were retained. All pages were resized to A4 dimensions (2480×3508 pixels at 300 DPI) to ensure uniform resolution and layout consistency. Figure 3: Age distribution of the 531 participants in the dataset. The dashed line indicates the median age. The de- tailed distribution is shown in Table8in the AppendixB. 3.2.1 Participants.Participants were recruited through the authors’ host institution in collaboration with five educational institutions across India, including one national-level institute in Madhya Pradesh, two state colleges in Bihar, and two elementary schools in Uttar Pradesh. Recruitment was conducted through direct contact with teachers and students, and participation was entirely voluntary. Only non-identifiable demographic information was collected, including age, gender, and city. The final cohort consists of 531 participants (135 female and 396 male). The dataset contains no names, signatures, or free-form personal text; all samples are re- stricted to the same fixed couplets, minimizing privacy and content leakage risks. The age distribution is shown in Figure 3. 3.2.2 Ethical Considerations.The study was approved by the In- stitutional Ethics Committee of the authors’ host institution. Par- ticipants were informed about the academic purpose and intended use of the dataset prior to submission. For minor participants, con- sent was obtained through classroom teachers under supervised collection. No personally identifiable information was collected, and all data were stored securely for research use only. 3.3 Dataset Description The dataset comprises handwritten Hindi text from 531 unique writers with diverse demographic and educational backgrounds. To capture natural variation in writing styles, no restrictions were Conference’17, July 2017, Washington, DC, USAKunwar Arpit Singh, Ankush Prakash, and Haroon R. Lone Figure 4: Frequency distribution of handwriting samples across states of India. The tabulated state wise distribution is show in Table7in AppendixA. imposed on handwriting proficiency. The dataset comprises con- tributors from states spanning the whole of India (Figure4), sup- porting evaluation under broad regional diversity in handwriting styles. Each participant transcribed the same set of six traditionaldo- hason an A4-sized plain white sheet. Each page contains 89 words, resulting in a total of 47,259 word instances across the dataset. This parallel design enables systematic analysis of stylistic variation in- dependent of linguistic content. Representative samples are shown in Figure 5. 4 Data Validation 4.1 Quality Requirements The quality of individual samples plays a foundational role for many learning algorithms. In this work, we developed an automated quality assessment system to evaluate and filter handwriting sam- ples based on visual fidelity. Rather than applying fixed manual thresholds, we trained convolutional neural network (CNN) classi- fiers to learn quality discrimination directly from image character- istics, which is enabling more robust and adaptive filtering. Image sharpness was quantified using the Laplacian variance method [ 19], which measures high-frequency edge content by com- puting the variance of the Laplacian operator—a second-order de- rivative filter responding strongly to rapid intensity changes. The blur score is defined as: Blur Score=Var(∇ 2 퐼)(1) where퐼denotes the grayscale image and∇ 2 is the Laplacian operator. Sharp images with well-defined strokes produce high variance values, whereas blurry images yield low variance due to smooth intensity transitions. 4.1.1 Dataset Quality Characterization and Labeling.Blur scores computed across all 531 collected samples revealed substantial vari- ation ranging from 9.7 to 13,865 (휇= 3268, median=4110,휎=2438), reflecting real-world handwriting data collection challenges includ- ing acquisition method variability (mobile camera vs. scanner), lighting conditions, ink quality, and paper texture. Rationale for CNN-Based Classification:While Laplacian variance provides an interpretable quality proxy, handwriting im- age quality is inherently multi-dimensional, encompassing ink den- sity, stroke continuity, background noise, and acquisition artifacts that single metrics cannot fully capture. We therefore employed blur scores for initial labeling, but trained CNNs to learn com- prehensive quality representations directly from raw pixel data. This approach enables the network to discover hierarchical fea- tures beyond edge variance that collectively determine image us- ability for handwriting recognition [ 34]. Furthermore, CNNs pro- vide learned, adaptive decision boundaries that generalize better than fixed thresholds across different acquisition conditions. Threshold Selection for Labeling:Quality class thresholds were determined through iterative analysis of the blur score distri- bution combined with visual validation. For binary classification, the threshold of 3000 (43rd percentile) creates balanced classes of 232 (43.7%) and 299 (56.3%) samples (see Figure6). Manual inspec- tion of 100 samples spanning the 2500-3500 range confirmed that 3000 effectively separates degraded images from those with suf- ficient stroke clarity for recognition tasks. For four-class stratifi- cation, thresholds at 1000, 3000, and 5000 were selected based on distributional quartiles and visual assessment. These boundaries maintain adequate class sizes (145, 87, 159, and 140 samples) while capturing meaningful quality gradations. Classification Formulations:We developed two CNN-based classifiers. Thefour-class classifierstratifies samples into Low (< 1000), Medium (1000–2999), Good (3000–4999), and Excellent (≥ 5000) levels, with 145 (27.3%), 87 (16.4%), 159 (29.9%), and 140 (26.4%) samples per class. Figure 7shows clear class separation with ex- pected statistical progression (Low:휇= 294; Medium:휇= 1432; Good:휇= 4269; Excellent:휇= 6129), though visible overlap be- tween Medium and Good classes indicates boundary ambiguity. Thebinary classifiersimplifies this into Low-to-Medium (< 3000) versus High-quality (≥ 3000) groups, reducing multi-class ambi- guity while maintaining effective discrimination for downstream filtering. The dataset was partitioned using stratified sampling into train- ing (339, 63.8%), validation (85, 16.0%), and test (107, 20.2%) sets, ensuring balanced class representation and preventing imbalance- induced bias. 4.1.2 Network Architecture and Training.The proposed CNN (Fig- ure8) follows a lightweight three-block convolutional backbone for feature extraction, followed by global average pooling and fully connected layers for classification into four quality categories. Training was conducted using PyTorch 2.0 on CUDA-enabled GPUs with the Adam optimizer and cross-entropy loss. 4.1.3 Results and Discussion. DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi TextConference’17, July 2017, Washington, DC, USA Figure 5: Sample images from the dataset. Each segment corresponds to one handwritten page instance. Figure 6: Blur score distribution with initial binary thresh- old labeling (before CNN refinement). Classification Performance:Tables2and3summarize test set performance. The binary classifier achieved 96.26% accuracy with balanced precision (0.94–0.99), recall (0.93–0.99), and F1-scores (0.96 for both classes), demonstrating robust discrimination. The four- class model achieved 85.98% accuracy with precision ranging from 0.75–0.96 and recall from 0.81–0.90. Medium-quality samples showed the weakest performance (precision=0.75, recall=0.83, F1=0.79) due to boundary ambiguity with adjacent Good class, as visualized in Figure 7. Low and Excellent extremes achieved stronger discrim- ination (F1=0.93 and 0.83), confirming that extreme quality levels are more reliably identifiable. Close validation-test alignment (bi- nary: 98.82% vs. 96.26%; four-class: 80.00% vs. 85.98%) indicates good generalization without overfitting. Quality-Based Dataset Filtering:The trained binary classifier was applied to all 531 images with a confidence threshold of 0.7 to con- struct a high-quality core subset while preserving the dataset’s Figure 7: Blur score distribution by quality class. Boxes show inter quartile ranges with medians (horizontal lines) and means (green triangles). Outliers appear as circles. Class separation validates stratification, while Medium- Good overlap explains reduced classification performance in boundary regions. Table 2: Overall quality classification performance ModelClasses Test Accuracy Val. Accuracy CNN (Binary)296.26%98.82% CNN (Four-class)485.98%80.00% real-world variability. This process retained 288 images (54.2%), which met acceptable legibility and acquisition quality standards, providing a reliable dataset for training handwriting recognition models. Importantly, this filtering does not imply that the remain- ing samples are irrelevant; rather, it reflects a structured separation Conference’17, July 2017, Washington, DC, USAKunwar Arpit Singh, Ankush Prakash, and Haroon R. Lone Figure 8: CNN architecture for handwriting quality classification. Three convolutional blocks extract hierarchical features, followed by adaptive pooling and fully connected layers for quality prediction. Table 3: Per-class performance on test set (107 samples) ModelClassPrec. Rec. F1 CNN (Binary) Low-Med. (<3000) 0.94 0.99 0.96 High (≥3000)0.99 0.93 0.96 CNN (Four-class) Low (<1000)0.96 0.90 0.93 Medium (1–3k)0.75 0.83 0.79 Good (3–5k)0.93 0.81 0.87 Excellent (≥5k)0.78 0.89 0.83 between clean reference data and naturally degraded real-world cases. Figure6reflects this initial threshold-based labeling, prior to the subsequent CNN refinement stage. Although 299 images ex- ceeded the blur threshold of 3000, the CNN-based filtering retained 288 images, indicating that the learned model captures broader quality cues beyond edge sharpness alone, such as ink consistency, illumination stability, and background interference. The retained subset shows substantially improved characteristics, with 157 Good (54.5%) and 131 Excellent (45.5%) samples and a mean blur score above 5000, representing a 53% increase over the original dataset mean (3268). This provides a strong benchmark set with sharper strokes and fewer acquisition artifacts while maintaining writer di- versity and character coverage. The remaining 243 samples (45.8%) form a challenging degraded subset, with mean blur scores below 1500, reflecting realistic issues such as defocus blur, poor lighting, low ink density, and background noise. Instead of being discarded, this subset demonstrates the dataset’s robustness under challeng- ing capture conditions. It can be useful for evaluating recognition performance in real-world settings, developing restoration or de- blurring methods, and testing model resilience to variations in im- age acquisition. 4.2 Data Variability 4.2.1 Line Segmentation Difficulty.The quality assessment frame- work described in the previous section4.1primarily filters sam- ples based on visual fidelity, targeting acquisition-related degra- dations such as defocus blur, poor illumination, and background noise. However, visual acceptability alone does not eliminate vari- ability at the structural level. Even among clear handwritten pages, substantial differences remain in inter-line spacing, baseline stabil- ity, stroke overlap, and shirorekha continuity, all of which directly affect the separability of text lines and complicate downstream doc- ument layout analysis. To quantify this remaining structural variability, we perform a page-level analysis of line segmentation difficulty over the en- tire dataset of 531 handwritten Hindi documents, using standard heuristic baselines as a practical proxy for characterizing layout complexity. This evaluation is deliberately carried out on the com- plete collection rather than restricting it to the quality-filtered sub- set. Although the quality assessment stage is designed to retain vi- sually reliable samples for downstream recognition model training, segmentation difficulty is shaped not only by acquisition-related factors such as blur or illumination, but also by inherent hand- writing properties. These include compressed inter-line spacing, irregular or drifting baselines, overlapping strokes, and interrup- tions in shirorekha continuity. Restricting the analysis to curated pages would therefore underrepresent the full spectrum of seg- mentation challenges present in real-world handwritten material, where uneven line gaps, skew variations, character overlap, and unpredictable spacing remain dominant sources of difficulty [ 23]. By profiling the entire dataset, we obtain an unbiased charac- terization of layout complexity under practical acquisition condi- tions. The objective is to isolate how naturally occurring handwrit- ing structure impacts line separability, independent of recognition performance. Such layout-level challenges remain central even in recent foundation models for full-page handwritten document understanding [ 10]. The resulting difficulty annotations support stratified evaluation of segmentation methods and help identify challenging layout regimes that persist even after quality-based filtering. 4.2.2 Experimental Setup.All experiments are conducted on the complete set of 531 raw handwritten page images, each containing 12 lines ofdohas. After standard preprocessing and binarization, we evaluate line segmentation difficulty at the page level to capture structural variability in the dataset. Given the diversity in spacing, baseline stability, and stroke in- teractions across writers, we employ a hybrid heuristic segmen- tation pipeline and assign difficulty labels based on the resulting DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi TextConference’17, July 2017, Washington, DC, USA Easy Medium Complex Figure 9: Representative handwritten pages from each segmentation difficulty class. Easy samples exhibit clear inter-line sep- aration, Medium pages contain moderate spacing variation, and Complex pages show overlap and irregular baselines leading to ambiguous line boundaries. line structures. Full implementation details of the segmentation heuristics and scoring procedure are provided in AppendixD. Difficulty Annotation Protocol.To enable dataset-level stratifica- tion, we assign each page a composite segmentation difficulty score 푆 ∈ [0, 100]based on structural consistency of the detected line re- gions. The score incorporates line count accuracy and additional layout regularity cues. Using푆and the absolute line count error, we assign each page to one of three segmentation difficulty levels. Pages are labeled asEasywhen segmentation is near-perfect (≤ 1line error) with a score of at least 65. Pages with moderate deviations (≤ 3lines) and scores above 45 are labeled asMedium. The remaining pages are assigned to theComplexcategory, corresponding to severe struc- tural irregularities and frequent segmentation failures. Thresholds are chosen empirically based on the observed score distribution. Representative samples from each segmentation difficulty category are presented in Figure 9. The segmented line boundaries are vi- sualized using horizontal markers, where green lines denote the upper boundary and red lines denote the lower boundary of each detected text line region. 4.2.3 Results and Discussion.We evaluate the performance of line segmentation on the full Hindi handwritten document collection (푁 = 531), where each page contains 12 expected text lines. Ta- ble 4summarizes the overall segmentation statistics. Perfect segmentation was achieved in 157 pages (29.57%). On average,11.22 ± 2.92lines were detected per page, indicating sub- stantial variability in line separability across handwriting styles. Segmentationdifficultydistribution.As shown in Table 5, only 110 pages (20.7%) are classified asEasy, while the majority (280 pages,52.7%) fall into theComplexcategory. Table 4: Line segmentation performance onDohaScript (N=531). MetricValue Total Documents531 Perfect Segmentation (12/12 lines) 157 (29.57%) Mean Segmentation Score43.46±22.55 95% CI[41.54, 45.39] Median Score42.91 Mean Lines Detected11.22±2.92 Table 5: Distribution of document segmentation difficulty. Difficulty Count Percentage Score Range Easy11020.7%65.1–91.8 Medium14126.6%45.9–66.4 Complex28052.7%7.6–46.7 Importantly, these difficult cases are not solely attributable to acquisition noise. Instead, complexity largely arises from intrin- sic handwriting characteristics, including dense inter-line spacing, overlapping strokes, baseline fluctuations, and shirorekha-induced continuity between adjacent lines. Error characteristics.Table 6illustrates the distribution of line detection errors, highlighting a non-trivial tail of highly challeng- ing pages. Overall, these findings demonstrate that segmentation-driven variability persists even after quality-based filtering. The result- ing Easy/Medium/Complex annotations provide a practical bench- mark for stratified evaluation and motivate adaptive segmentation models capable of handling the full spectrum of real-world Devana- gari handwriting. Conference’17, July 2017, Washington, DC, USAKunwar Arpit Singh, Ankush Prakash, and Haroon R. Lone Table 6: Distribution of line detection errors. Error (lines) Count Percentage 015729.6% 113826.0% 26913.0% 36412.1% 4285.3% 5285.3% 6101.9% 7183.4% 891.7% 971.3% 1020.4% 1110.2% 5 Dataset availability TheDohaScriptdataset is publicly available via a Google Drive 2 repository. All code for data preparation, preprocessing, quality as- sessment, segmentation, and experimentation is released through a public GitHub repository (https://github.com/KAS4453/DohaScript) to support full reproducibility and future research. 6 Potential Impact & Use Cases The dataset aims to fill a significant void in handwritten Devana- gari resources by offering a comprehensive, writer-diverse, and systematically regulated corpus of handwritten Hindi text. This dataset facilitates a wide array of study avenues that were previ- ously challenging or impractical for Indic scripts by integrating recurrent lexical elements from numerous authors with page-level document images and comprehensive quality assessments. We de- lineate critical domains in which this dataset can exert consider- able influence. Handwritten Text Recognition and OCR.The dataset provides page- level handwritten Devanagari text under realistic writing condi- tions, enabling evaluation of sequence-based HTR and OCR mod- els beyond isolated character or word settings. Writer Identification and Biometric Analysis.Each document is written by a unique writer with identical textual content, support- ing controlled studies of writer identification and handwriting-based biometrics. Handwriting Style Analysis and Clustering.The dataset supports analysis of handwriting style variation across writers, aided by de- mographic metadata and layout difficulty annotations for stratified evaluation. Handwriting Synthesis and Generative Modeling.Repeated lex- ical content and stylistic diversity make the dataset suitable for handwriting synthesis, style-conditioned generation, and OCR data augmentation. Document Layout Analysis and Segmentation.The page-level struc- ture and segmentation difficulty profiling enable evaluation of line 2 https://drive.google.com/drive/folders/1v2pjEE0MUkcLRn7YEfz3cro9OUZIXi0g segmentation and document layout analysis methods for Devana- gari handwriting. Benchmarking and Reproducible Research.The dataset is released with documented protocols and validation procedures, supporting fair benchmarking and reproducible research. 7 Conclusion We presentedDohaScript, a large-scale multi-writer dataset of con- tinuous handwritten Hindi text, addressing the limited availabil- ity of page-level Devanagari handwriting resources. The dataset comprises identicaldohascollected from 531 writers, enabling con- trolled investigation of handwriting style variation independent of linguistic content. This design supports a range of downstream tasks, including handwritten text recognition, writer-centric anal- ysis, and generative modeling. Beyond extensive quality curation, we introduced page-level seg- mentation difficulty annotations to capture intrinsic structural vari- ability in handwritten layouts. Our findings indicate that chal- lenges such as irregular inter-word spacing, baseline drift, and shi- rorekha discontinuities persist even in visually clean samples, high- lighting the need for structure-aware approaches. Overall,DohaScript establishes a standardized and challenging benchmark for advanc- ing research in continuous handwritten Devanagari text under- standing, particularly in low-resource settings. References [1]AfKaRi-FahandaRi, A., Shabaninia, E., Asadi-Zeydabadi, F., and Nezamabadi-PouR, H. A comprehensive survey of transformers in text recognition: Techniques, challenges, and future directions.ACM Computing Surveys 58, 5 (2025), 1–42. [2]AlKendi, W., GechteR, F., HeybeRgeR, L., and Guyeux, C. Advancements and challenges in handwritten text recognition: A comprehensive survey.Journal of Imaging 10, 1 (2024), 18. [3]ARivazhagan, M., SRinivasan, H., and SRihaRi, S. A statistical approach to line segmentation in handwritten documents. InProceedings of SPIE Document Recognition and Retrieval XIV(2007). [4]ARoRa, S., MaliK, L., Goyal, S., BhattachaRjee, D., NasipuRi, M., and KRej- caR, O. Devanagari character recognition: A comprehensive literature review. IEEE Access 13(2025), 1249–1263. [5]Babu, S., and Jangid, M. Touching character segmentation of devanagari script. InProceedings of the 7th International Conference on Computing, Communication and Networking Technologies(Dallas, TX, USA, 2016), ICCCNT ’16, ACM, p. 1– 6. [6]Dai, G., Zhang, Y., Wang, Q., Du, Q., Yu, Z., Liu, Z., and Huang, S. Dis- entangling writer and character styles for handwriting generation. InProceed- ings of the IEEE/CVF conference on computer vision and pattern recognition(2023), p. 5977–5986. [7]Dey, S., Alaei, A., and Roy, P. P. Handwritten text recognition for low resource languages.arXiv abs/2512.01348(2025). [8]Dutta, K., KRishnan, P., Mathew, M., and JawahaR, C. V. Offline handwrit- ing recognition on devanagari using a new benchmark dataset. InInternational Workshop on Document Analysis Systems (DAS)(2018). IIIT-HW-Dev dataset; discusses lack of publicly available word/line-level corpora. [9]EbeRhaRd, D. M., Simons, G. F., and Fennig, C. D., Eds.Ethnologue: Languages of the World, 28 ed. SIL International, Dallas, TX, USA, 2025. [10]Fadeeva, A., CoRiou, V., Antognini, D., Musat, C., and MaKsai, A. Inkfm: A foundational model for full-page online handwritten note understanding.arXiv preprint arXiv:2503.23081(2025). [11]GaRRido-MuÑoz, C., Rios-Vila, A., and Calvo-ZaRagoza, J. Handwritten text recognition: A survey.arXiv preprint arXiv:2502.08417v1 [cs.CV](2025). [12]Gongidi, S., and JawahaR, C. V. iiit-indic-hw-words: A dataset for indic hand- written text recognition. InProceedings of the International Conference on Docu- ment Analysis and Recognition (ICDAR)(2021), IEEE. [13]James, K. H. The importance of handwriting experience on the development of the literate brain.Current Directions in Psychological Science 26, 6 (2017), 502– 508. DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi TextConference’17, July 2017, Washington, DC, USA [14]Kasuba, B. V., Kudale, D., SubRamanian, V., ChaudhuRi, P., and RamaKR- ishnan, G. Platter: A page-level handwritten text recognition system for indic scripts.arXiv abs/2502.06172(2025). [15]LiKfoRman‐Sulem, L., et al. Text line segmentation of historical documents: a survey. InInternational Conference on Document Analysis and Recognition(2007). [16]NehRa, M. S., Nain, N., and Ahmed, M. Handwritten devnagari script database development for off-line hindi character with matra (modifiers). InProceedings of the International Conference on Information and Communication Technology for Competitive Strategies(2016), Springer. [17]Office of the RegistRaR GeneRal and Census CommissioneR, India.Lan- guage Atlas of India: Census of India 2011. Government of India, Ministry of Home Affairs, New Delhi, 2022. [18]Pal, U., and ChaudhuRi, B. B. Indian script character recognition: A survey. Pattern Recognition 37, 9 (2004), 1887–1899. [19]Pech-Pacheco, J. L., CRistÓbal, G., ChamoRRo-MaRtinez, J., and FeRnÁndez- Valdivia, J. Diatom autofocusing in brightfield microscopy: A comparative study. InProceedings of the 15th International Conference on Pattern Recognition (ICPR)(2000), vol. 3, p. 314–317. [20]Plamondon, R., and SRihaRi, S. N. Online and off-line handwriting recognition: a comprehensive survey.IEEE Transactions on pattern analysis and machine in- telligence 22, 1 (2002), 63–84. [21]PRashanth, D. S., Mehta, R. V. K., and Challa, N. P. A multi-purpose dataset of devanagari script comprising isolated numerals and vowels.Data in Brief 41 (2022), 107723. [22]PtaK, R., Zygadło, B., and Unold, O. Projection‐based text line segmentation with a variable threshold.International Journal of Applied Mathematics and Com- puter Science 27, 1 (2017), 195–206. [23]RomeRo, V., Sanchez, J. A., Bosch, V., Depuydt, K., and de Does, J. Influence of text line segmentation in handwritten text recognition. InProceedings of the 13th International Conference on Document Analysis and Recognition (ICDAR)(2015), IEEE, p. 1040–1044. Accessed: 2026-02-06. [24]Samanta, P. K., and Biswas, S. Aio-hb: A handwritten text image dataset of hindi and bengali indian scripts for handwritten text recognition. InPattern Recognition. ICPR 2024. Lecture Notes in Computer Science, vol. 15319 ofLecture Notes in Computer Science. Springer, Cham, 2024, p. 316–332. First online: 04 December 2024. [25]SchneideR, P. Combining morphological and histogram based text line segmen- tation in the ocr context.Journal of Data Mining & Digital Humanities(2021). arXiv:2103.08922. [26]Senthil, G., K, N., and SubRahmanyam, G. R. K. S. Handwritten hindi word generation to enable few instance learning of hindi documents. InProceedings of the IEEE International Conference on Signal Processing and Communications (SPCOM)(Bangalore, India, 2020), IEEE. [27]ShaKya, S., Sainju, S., ShRestha, S. K., Dawadi, P., and Khatiwada, S. De- vanagari script classification using cbow embeddings with attention-enhanced bilstm. InProceedings of the First Workshop on Challenges in Processing South Asian Languages (CHiPSAL)(Abu Dhabi, United Arab Emirates, Jan. 2025), In- ternational Committee on Computational Linguistics, p. 339–343. [28]ShaRma, B. K., Dhillon, D., Rana, V., Ishant, and Singh, B. Approximation of person’s age and gender from its handwriting characteristics: A review.SSRN Electronic Journal(2022). [29]ShaRma, R. A survey on offline recognition of handwritten indic scripts.Pattern Recognition Letters 133(2020), 41–57. [30]ShaRma, V., VeRma, R., and Saluja, R. Ancidev: A dataset for high-accuracy handwritten text recognition of ancient devanagari manuscripts. InProceedings of the 1st Workshop on Benchmarks, Harmonization, Annotation, and Standardiza- tion for Human-Centric AI in Indian Languages (BHASHA 2025)(Mumbai, India, 2025), Association for Computational Linguistics, p. 91–101. [31]Shi, B., Bai, X., and Yao, C. An end-to-end trainable neural network for image- based sequence recognition and its application to scene text recognition.IEEE transactions on pattern analysis and machine intelligence 39, 11 (2016), 2298–2304. [32]Shinde, A. B., and Dandawate, Y. H. Shirorekha extraction in character seg- mentation for printed devanagari text in document image processing. In2014 Annual IEEE India Conference (INDICON)(2014), p. 1–7. [33]Sun, C., ShRivastava, A., Singh, S., and Gupta, A. Revisiting unreasonable effectiveness of data in deep learning era.IEEE International Conference on Com- puter Vision (ICCV)(2017), 843–852. [34]Szandała, T. Convolutional neural network for blur images detection as an al- ternative for laplacian method. In2020 IEEE Symposium Series on Computational Intelligence (SSCI)(2020), p. 2901–2904. [35]Wiley, R. W., and Rapp, B. The effects of handwriting experience on literacy learning.Psychological Science 32, 7 (2021), 1086–1103. [36]Xing, L., and Qiao, Y. Deepwriter: A multi-stream deep cnn for text- independent writer identification. In2016 15th international conference on fron- tiers in handwriting recognition (ICFHR)(2016), IEEE, p. 584–589. Conference’17, July 2017, Washington, DC, USAKunwar Arpit Singh, Ankush Prakash, and Haroon R. Lone A State-wise Participant Distribution Table 7: State-wise distribution of participants. S.No. StateParticipants 1 Andhra Pradesh15 2 Arunachal Pradesh1 3 Assam3 4 Bihar150 5 Chandigarh2 6 Chhattisgarh3 7 Delhi3 8 Gujarat5 9 Haryana10 10 Himachal Pradesh2 11 Jammu & Kashmir4 12 Jharkhand11 13 Karnataka9 14 Kerala7 15 Madhya Pradesh28 16 Maharashtra36 17 Odisha8 18 Rajasthan25 19 Tamil Nadu8 20 Telangana14 21 Tripura2 22 Uttar Pradesh155 23 Uttarakhand6 24 West Bengal23 B Age-wise Participant Distribution Table 8: Distribution of participants across different age groups. Age ParticipantsAge Participants 741999 84 2076 914 2170 102 2247 117236 129243 1322283 1426291 157311 163401 1743 411 187842–50*4 C Dohas 1. When both the Guru and God stand before the seeker, whom should one revere first? I bow to the Guru, for it is through the Guru’s guidance that one attains knowledge of God. 2. The mind is urged to remain patient, as everything happens gradually and in its proper time. Even if a gardener waters a plant a hundred times, fruit appears only when the proper season arrives. 3. Compassion is the root of righteousness, and arrogance is the root of sin. Compassion should not be abandoned for as long as life remains in the body. 4. The world has spent its life reading books, yet no one became truly learned by that alone. One who understands the two and a half letters that form the word “love” (in Hindi) becomes truly learned. 5. There is no austerity equal to truth, and no sin equal to falsehood. One whose heart holds truth has God dwelling within their heart. 6. The Guru acts as the guardian of knowledge and keeps one’s thoughts pure. Even if all six philosophical systems are known, the true Guru alone is the foundation. D Segmentation Pipeline Details This appendix provides full implementation details of the heuris- tic line segmentation strategies and scoring procedure used for dataset-level difficulty annotation. D.1 Projection Profiles We use horizontal projection profiles as a structural cue for text- line organization [ 15,22]. Given a binarized page image퐵(푥, 푦), the horizontal projection is defined as: 푃 ℎ (푦) = ∑ 푥 퐵(푥, 푦),(2) DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi TextConference’17, July 2017, Washington, DC, USA which measures foreground ink density along the vertical axis. Dis- tinct peaks typically correspond to individual text lines, whereas compressed spacing or overlapping strokes lead to merged responses and ambiguous boundaries [ 3]. D.2 Hybrid Line Segmentation Strategy To handle the wide range of handwriting layouts present in the dataset, we employ a hybrid line segmentation pipeline that com- bines three complementary classical heuristics. A projection-based method estimates candidate line boundaries by identifying valleys in smoothed horizontal projection profiles, where Gaussian filtering (휎 = 2) reduces spurious fluctuations [ 3, 15]. In parallel, a contour-based strategy groups text regions through horizontal morphological dilation using a50 × 1kernel, followed by contour extraction and bounding-box merging [25]. Finally, a morphology-driven approach suppresses shirorekha- like horizontal strokes and applies connected component grouping to better isolate line structures in Devanagari handwriting [32]. For each page, all three strategies are run independently, and the segmentation whose detected line count is closest to the expected 12 lines is retained. D.3 Composite Difficulty Score Each page is assigned a composite heuristic score푆 ∈ [0, 100] based on structural cues derived from the detected line regions: 푆 = 푆 count + 푆 height + 푆 spacing + 푆 coverage + 푆 straight ,(3) where the terms capture line count accuracy, uniformity of line heights, spacing consistency, coverage ratio, and boundary regu- larity. The full scoring implementation is released with the dataset code- base for reproducibility.