Paper deep dive
Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts
Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/20/2026, 4:57:25 AM
Summary
This paper introduces an iterative fine-tuning pipeline for digitizing complex historical Sanskrit manuscripts using traditional OCR methods. The pipeline addresses layout and appearance distribution shifts by iteratively fine-tuning a Graph Neural Network (GNN) for layout analysis and a CNN-BiLSTM-CTC model for text recognition. The authors present a curated dataset of three Sanskrit manuscripts with granular annotations and demonstrate that iterative fine-tuning reduces human annotation effort and improves transcription accuracy compared to baseline methods and multi-modal large language models.
Entities (17)
Relation Signals (15)
Kaushik Gopalan â affiliatedwith â FLAME University
confidence 95% · Kaushik Gopalan Affiliation: FLAME University
Kartik Chincholikar â affiliatedwith â Centre for Inter-disciplinary Artificial Intelligence
confidence 95% · Kartik Chincholikar Affiliation: Centre for Inter-disciplinary Artificial Intelligence
Dataset â contains â Yajnavalakyasmritih
confidence 95% · The dataset consists of three historical Sanskrit manuscripts: Yajnavalakyasmritih
Dataset â contains â Tantra Raj With Yantra And Mantra Uddhara
confidence 95% · The dataset consists of three historical Sanskrit manuscripts: ... and Tantra Raj With Yantra And Mantra Uddhara.
Dataset â contains â Muhurta Martanda
confidence 95% · The dataset consists of three historical Sanskrit manuscripts: ... Muhurta Martanda
Yajnavalakyasmritih â haslayouttype â Moderate Layout
confidence 95% · Moderate Layout Manuscript: Yajnavalakyasmritih
Muhurta Martanda â haslayouttype â Dense Layout
confidence 95% · Dense Layout Manuscript: Muhurta Martanda
Tantra Raj With Yantra And Mantra Uddhara â haslayouttype â Circular Layout
confidence 95% · Circular Layout Manuscript: Tantra Raj With Yantra And Mantra Uddhara
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historical scholars to study them in new ways. Historical manuscripts, however, often exhibit complex heterogeneous layouts and non-standard appearance due to period-specific writing styles, page textures, camera noise, and other nuisance factors, making them difficult to perform OCR on. To tackle this challenge, we introduce a local traditional OCR pipeline, which can be iteratively fine-tuned on the target manuscript at the layout-level and the appearance-level. By adapting to the target manuscript distribution, the proposed Traditional OCR pipeline makes better predictions on subsequent pages, causing iterative reduction in human annotation effort, which is expensive and time-consuming as it requires historical domain expertise. Using this pipeline, we digitize text from three complex historical Sanskrit manuscripts and introduce a dataset with granular layout-level annotations, along with Unicode annotations in the standard PAGE-XML format. We demonstrate quantitative gains due to iterative fine-tuning of the proposed traditional OCR pipeline, and also benchmark the performance of leading Multi-Modal Large Language Models on the introduced Dataset. Code and dataset are available at: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.18696v1
- Canonical: https://arxiv.org/abs/2608.18696v1
Trouble viewing inline? Open PDF directly â
Full Text
59,392 characters extracted from source content.
Expand or collapse full text
Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts Kartik Chincholikar Affiliation: Centre for Inter-disciplinary Artificial Intelligence (CAI), Kaushik Gopalan Affiliation: FLAME University, Pune, India Mihir Hasabnis Affiliation: E-mail kartik.chincholikar,kaushik.gopalan@flame.edu.in Abstract Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historical scholars to study them in new ways. Historical manuscripts, however, often exhibit complex heterogeneous layouts and non-standard appearance due to period-specific writing styles, page textures, camera noise, and other nuisance factors, making them difficult to perform OCR on. To tackle this challenge, we introduce a local traditional OCR pipeline, which can be iteratively fine-tuned on the target manuscript at the layout-level and the appearance-level. By adapting to the target manuscript distribution, the proposed Traditional OCR pipeline makes better predictions on subsequent pages, causing iterative reduction in human annotation effort, which is expensive and time-consuming as it requires historical domain expertise. Using this pipeline, we digitize text from three complex historical Sanskrit manuscripts and introduce a dataset with granular layout-level annotations, along with Unicode annotations in the standard PAGE-XML format. We demonstrate quantitative gains due to iterative fine-tuning of the proposed traditional OCR pipeline, and also benchmark the performance of leading Multi-Modal Large Language Models on the introduced Dataset. Code and dataset are available at: https://github.com/flame-cai/gnn-synthetic-layout-historical/. Keywords: OCR Handwritten Text Recognition Obscure Domains 1 Introduction Digitization of historical manuscripts in Unicode Format makes them easily accessible to researchers and scholars in a digital format, avoiding the risk of damaging the original manuscripts, which are often in a fragile condition. Such digitization also allows historical scholars to search through the manuscripts more quickly, study changes in word usage over time, and track the frequency with which certain ideas appear. Digitization efforts of historical manuscripts often require specialized historical language and script expertise and are thus expensive and time-consuming to collect, causing the scarcity of annotated data. Furthermore, historical manuscript pages differ from modern digital or printed documents at the Layout-level and at the Appearance-level. Layout-level distribution shifts arise from high heterogeneity in historical page layoutsâincluding marginalia, interlinear glosses, footnotes, and irregular, curved text-lines. Appearance-level distribution shifts occur due to variation in scribal styles, period-specific writing conventions, physical degradation artifacts (e.g., ink bleed-through, fading, staining) and nuisance factors (e.g., camera noise, image compression, paper or palm-leaf textures, uneven illumination, darkened scans). In this work, we thus digitize three historical Sanskrit manuscripts with complex layouts and non-standard appearances, as illustrated in Fig. 1, using a traditional two-step OCR pipeline: first, we perform layout analysis and segment individual text-line images from manuscript pages, and second, we transcribe those text-line images into machine-readable Unicode text. The proposed traditional pipeline allows us to iteratively fine-tune and adapt the pipeline to the target manuscript - at the layout-level and at the appearance-level, thus progressively reducing the burden of human annotation on subsequent pages. This is important, as historical data is scarce, time-consuming, and expensive to annotate. With this context, our work makes two primary contributions: 1) Fine-tunable Traditional OCR pipeline. We introduce an open-source digitization pipeline, which can be iteratively fine-tuned from human supervision at the layout-level and the appearance-level, allowing progressive reduction in the burden of human annotation of subsequent pages. 2) Curated Dataset. We introduce a high-quality dataset, which is richly annotated at the layout-level and the appearance-level, and is available in the standard PAGE-XML format [53] 11 1 https://github.com/PRImA-Research-Lab/PAGE-XML. (a) Moderate Layout Manuscript: Yajnavalakyasmritih (Acharadhyayah)[7] (b) Dense Layout Manuscript: Muhurta Martanda[5] (c) Circular Layout Manuscript:Tantra Raj With Yantra And Mantra Uddhara[6] Figure 1: The manuscripts. Samples of pages from the three manuscripts digitized in this case study. The manuscripts exhibit non-standard layout variations, including dense marginalia, circular text, interlinear commentary, irregular text-line orientation, as well as non-standard appearance variations such as ink bleed-through, faded ink, paper degradation, scribe-specific handwriting styles, and period-specific orthographic conventions. 2 Literature Review Modern Multi-Modal LLMs have made significant progress in their OCR capabilities [70, 67, 21, 54, 45, 69]. However, the benchmark datasets [50, 47, 65] mostly consider documents in a printed format or in a standard digital format. Therefore, at the date of writing, performing OCR on historical documents with complex, dense layouts with non-standard appearance using Modern Multi-Modal LLMs remains inconsistent [20, 31, 19], because of a distribution shift of the target distribution from the pre-training distribution, and due to scarcity of annotated data from the target distribution. To efficiently learn from scarce data, traditional OCR pipelines are hence employed to digitize historical manuscripts [38, 36, 55, 56, 1, 22, 66, 33, 52, 68, 18, 46, 51, 61]. Traditional pipelines broadly consist of two steps: (a) Text-line Segmentation, where the layout of the page is analyzed and individual text-line images are segmented from manuscript pages and (b) text recognition OCR, where the text-line images are transcribed into machine-readable Unicode text. Early text-line segmentation methods follow a projection profile approach successfully; however, they are limited to applications where the manuscript layouts are known a-priori [12, 48, 15]. More recent methods perform end-to-end text-line segmentation (sometimes with subsequent algorithmic post-processing) by predicting bounding polygons or by predicting pixel level masks [58, 63, 35, 27, 9, 63, 39, 49, 30, 10]. Performing such dense end-to-end pixel-level predictions is, however, not robust to distribution shifts commonly encountered in historical manuscripts [22, 4, 3, 34, 16]. To alleviate this, instead of performing dense end-to-end pixel-level prediction, recent methods LineTR [2] and CurT [40] use deep learning to extract information from the manuscript images and use it to predict the parameters defining the geometry of piecewise line segments or cubic BĂ©zier curves representing the text lines, thus making effective use of inductive priors and the geometric structure of text-lines. To handle complex unconstrained historical layouts, systems like Kodym and HradiĆĄ [42] jointly estimate baselines, line heights, and pixel-wise text orientations to extract clean line regions. Challenges in text-line segmentation motivated the FEST Competition 2025 (Few-Shot Text-Line Segmentation) [71], with the aim of stimulating research in the design of methods that generalize to target manuscripts when trained on scarce data. Once the text-line images are segmented, the text content from the text-line images can be recognized in a Unicode format using methods which use a Convolutional Neural Network(CNN) for extracting images features, followed by a BiLSTM or an RNN for sequential modelling, with a CTC loss and decoder [29, 57, 37, 60, 23]. More recent methods use the Transformer architecture [64], where an image Transformer extracts the visual features and a text Transformer performs the language modeling and decoding [44]. Beyond architectural designs, recent line recognition methods address domain-specific challenges in historical HTR: AT-ST [41] introduces self-training adaptation to train line recognizers when target domain transcripts are scarce, while TS-Net [43] enables a single recognizer to switch dynamically between different transcription styles (e.g., diplomatic versus modernized output). In traditional pipelines, layout analysis and text-line segmentation are thus prerequisites for performing line level text recognition OCR, and errors in layout analysis and text-line segmentation can negatively impact the downstream text recognition OCR task. Hence, manual work and human supervision might still be required at the layout analysis stage [26]. For manuscripts with non-standard layouts and unconventional reading order, character segmentation (instead of text-line segmentation) is a promising task decomposition [59, 17], offering increased flexibility to perform layout analysis. 3 Dataset Table 1: Dataset statistics Manuscript Pages Text-lines per page Grapheme clusters per page Min. Max. Mean Min. Max. Mean Moderate Layout 15 12 28 21.07 285 525 392.27 Dense Layout 7 15 78 44.43 285 1098 691.86 Circular Layout 9 10 65 28.33 164 434 258.89 The dataset consists of three historical Sanskrit manuscripts: Yajnavalakyasmritih (Acharadhyayah), Muhurta Martanda, and Tantra Raj With Yantra And Mantra Uddhara. We will henceforth refer to the manuscripts as Moderate Layout Manuscript, Dense Layout Manuscript, and Circular Layout Manuscript respectively, based on their page layouts as illustrated in Fig. 1. Table 1 shows the number of pages, the number of text-lines, and grapheme clusters [62, 24] in each manuscript. A grapheme cluster corresponds to a visually identifiable unit in the script, but it is made up of two or more Unicode points. Annotation Methodology. As illustrated in Figure Fig. 2, we annotate the manuscripts at the Layout-level and Appearance-level. For layout-level annotations, we consider each character (or grapheme cluster) of the manuscript page as a node, with edges connecting nodes with their neighbours in the text-line together. Thus, all nodes belonging to the same text-line have the same label, as shown in Fig. 2(c). Similarly, all nodes belonging to the same text-region have the same label, as shown in Fig. 2(b). The user can hover over the predicted graph while pressing and holding keys "a" or "d" to add or delete edges, respectively. A "right-click" or "left-click" adds or deletes nodes, respectively. Similarly, text-regions can be annotated by pressing and holding "e" and hovering over the text-lines in the text-region. Each graph-based text-line is explicitly linked to itâs corresponding Unicode text content, as illustrated in Fig. 2(d). While annotating the dataset, we made the following assumptions: (a) We annotate text-regions, such that the reading order of the text-lines inside the text-region must be unambiguous. (b) Given the nature of the marginalia and commentary, the reading order of the text-regions is ambiguous. (c) The reference annotation symbols and numbers that link the main text to the commentary are not annotated. (d) When a watermark overlaps with the handwritten text-content, we give precedence to the handwritten text-content. (e) If the text-line segmentation incorrectly excludes a diacritic mark in the segmentation, we still annotate it in the Ground-Truth Unicode annotation. (f) If a character or a grapheme cluster is scratched out, causing the character to be illegible, we do not annotate it. Selection Criteria. The proposed Traditional Pipeline can be used to digitize rare historical manuscripts that fit the following selection criteria: (a) The frozen character segmentation model CRAFT [8] should perform satisfactorily, and (b) A pre-trained text recognition OCR model is available for the manuscriptâs script. We believe that these selection criteria should apply to most of the Sanskrit manuscript images archived in culture preservation projects like the eGangotri project22 2 https://egangotri.org/, and GyanBharatam33 3 https://gyanbharatam.com/. (a) Original image (b) Text-region annotation (c) Text-line annotation (d) Unicode annotation Figure 2: Rich granular annotations at the Layout-level and Appearance-level. The dataset is also exported in the standard PAGE-XML format. 4 Method (a) Original Image (b) Heatmap (c) Layout Graph (d) Unwrapped text-line image (e) Prediction (Pred) and Ground-Truth (GT) Text Frozen U-Net GNN Unwrapping and processing CNNâBiLSTMâCTC Figure 3: Iterative Fine-tuning Pipeline. The frozen U-Net CRAFT [8] detects character locations in the original image (a) as a heatmap (b). Next, the GNN predicts the layout graph (c) where characters (or grapheme clusters) belonging to the same text-line are connected. In the next step, text-line images are prepared using the GNN predictions, and curved text (if any) is unwrapped and straightened (d). Finally, the CNNâBiLSTMâCTC takes in the text-line images as input, and predicts its text contents as Unicode text (e). In both (c) and (e), the colors denote differences between the predictions and the ground-truth, and hence illustrate the scope of improvement from iterative fine-tuning of the GNN and CNNâBiLSTMâCTC. Orange denotes missing nodes, edges, or Unicode characters. Blue denotes extra nodes, edges, or Unicode characters. Maroon denotes modified text, and Black denotes correct nodes, edges, and unicode text. Historical manuscripts can vary at the Layout-level, concerning where text is placed: dense pages, marginalia, interlinear text, and curved or circular text-lines, and at the Appearance-level, concerning how text looks: handwriting style, period-specific writing conventions, paper texture, and uneven scans. Fig. 3 shows the traditional pipeline, which allows fine-tuning at the Layout-level and at the Appearance-level. 4.1 Layout-Level Fine-Tuning Layout Annotation. The first step of the pipeline is to detect character (or grapheme cluster) locations. The character detector is a frozen pre-trained U-Net, CRAFT [8]. It produces the heatmap in Fig. 3(b). Next, the 2D character locations we get from the heatmap in Fig. 3(b) are used by a Graph Neural Network based layout analysis backbone [16] to perform text-line segmentation. This graph-based problem formulation considers each character (or grapheme cluster) of the manuscript as a node, with edges connecting nodes of the same text-line to their neighbours in the same text-line. First, a preprocessing step is performed to get information-rich node and edge features using geometric inductive priors, after which the GNN performs binary edge classification to predict whether an edge exists between two nodes, giving us the predicted graph seen in Fig. 3(c). This problem formulation decomposes the text-line detection problem into two sub-problems: (i) character detection, and (i) binary edge classification (connecting characters belonging to the same text-line together). This task decomposition enables the GNN to be pre-trained on large scale, diverse, synthetic layout data in the geometric domain (only consisting of node locations and edge connections), while the character detection task is delegated to the pre-trained frozen CRAFT model [8]. (As illustrated in Fig. 3 (b) to (c)). Iterative GNN Fine-tuning. The predictions of the GNN can be (optionally) corrected manually by the human user for two reasons: (a) to create a supervised training data point which is used to fine-tune the GNN, and (b) for manual correction at inference time. Using the annotated training data points of the target manuscript, we fine-tune a Graph Neural Network with the SplineCNN architecture [25], which is pre-trained on diverse synthetic layouts in the geometric domain to perform a binary edge classification task, where an edge between two nodes is classified as 1 if the nodes are neighbours in the same text-line, and 0 otherwise. If doing manual correction at inference time, the user can also manually add/delete incorrect nodes (in addition to adding/deleting incorrect edges) to account for mistakes in character detection by CRAFT; however, these node manual corrections are not currently used to fine-tune the GNN, as it currently performs only a binary edge classification task. Text-Region annotations, as shown in Fig. 2(b) are also currently done manually. 4.2 Appearance-Level Fine-Tuning Once the layout graph has been corrected, the pipeline prepares one text-line image per text-line, as shown in Fig. 3(d). Curved Text Unwrapping. In this step, if the detected text is curved, we unwrap it and convert the text from the graph-based format into a rectangular text-line image format, which the downstream text recognition OCR model (CNNâBiLSTMâCTC) requires. For each node along the text-line, we construct a local coordinate system: the tangent direction becomes the horizontal direction of the crop, and the normal direction becomes its vertical direction. Using information from the heatmap, we then sample the manuscript image in these local coordinates to form the rectangular image in Fig. 3(d). In this sense, we consider a circular text-line as a straight line locally; walking along its circular baseline unwraps it into a conventional left-to-right strip for OCR. This step is heuristic, as the circular text is cut at the topmost point in the global page coordinate system to define the start and end of the unwrapped line. It must also prevent diacritics from adjacent text-lines from entering the crop, and choose a reading direction. OCR Fine-tuning. The text recognition OCR model CNNâBiLSTMâCTC takes the (optionally unwrapped) and processed text-line image as input and predicts its text content as Unicode text, as shown in Fig. 3(e). The predictions can then be corrected by human experts to get ground-truth input-label pairs(text-line image, corrected Unicode text). These input-label pairs are used to fine-tune the CNNâBiLSTMâCTC to adapt to the appearance-level distribution shift of the target manuscript, such as the scribeâs writing style, period-specific writing conventions, page texture, and other nuisance factors. This adaption when done iteratively, allows the CNNâBiLSTMâCTC to make better predictions on subsequent pages and thus also reduces the burden of human annotation iteratively. 4.3 Fine-tuning Configuration System Requirements. All inference, annotation, and fine-tuning were done locally on a laptop with an NVIDIA GeForce RTX 4050 Laptop GPU, an AMD Ryzen processor, and 16 GB RAM. The setup comprises a Vue.js front end and connects with the traditional pipeline Flask back end, which orchestrates the iterative fine-tuning and inference. After the user corrects a new page, the layout-level GNN and the appearance-level CNNâBiLSTMâCTC text recognizer are both fine-tuned sequentially using the newly corrected data. The resulting adapted checkpoints are then used to predict the subsequent pages of the same manuscript. On average, adapting the annotation tool to each newly annotated page required 41.10 s for GNN fine-tuning and 27.51 s for CNNâBiLSTMâCTC fine-tuning. Layout-level GNN fine-tuning. We fine-tune the pre-trained GNN using the human-corrected layout graph of the target manuscript page. As the layout-level data is in the 2D geometric domain, where each page is represented by character-node locations and candidate edges rather than manuscript pixels, we can easily augment the corrected page layout 50 times using transformations such as warping, skewing, node drop-out, and node jitter [16]. We fine-tune all GNN parameters for 10 epochs using Adam with a learning rate of 10â310^-3 and a batch size of 4. To address the imbalance between positive and negative candidate edges, we use focal loss with α=0.9α=0.9 and Îł=2.0Îł=2.0. The learning rate is linearly warmed up during the first five epochs. Checkpoint selection maximises text-line F1 (IoU â„0.5â„ 0.5) on the unaugmented fine-tuning pages themselves rather than a held-out split; the foldâs test pages are reserved exclusively for the reported metrics. Appearance-level text recognition OCR fine-tuning. The pre-trained CNNâBiLSTMâCTC recognizer [15], which is based on the code provided by Clova AIâs deep-text-recognition-benchmark repository44 4 https://github.com/clovaai/deep-text-recognition-benchmark and the EasyOCR python package, is fine-tuned on the corrected input-label pairs (text-line image, corrected Unicode text) from the newly annotated page. We use Adadelta with learning rate 0.2, Ï=0.95Ï=0.95, and Ï”=10â8Δ=10^-8 for 60 iterations, with batch size 1 and gradient clipping at 5. Text-line images are resized to a height of 50 pixels and padded to the maximum width within each batch. Checkpoint selection is in-sample: the accuracy and edit-distance snapshots are scored on unaugmented text-line imageâlabel pairs from the training data, and the two are decided between by page CER on the fine-tuning pages; the foldâs test pages are reserved for the reported metrics. Annotation Cost. The GNN and the OCR Recognition model are not supervised by the same annotations. The GNN learns from the corrected graph alone and never reads text, whereas the CNNâBiLSTMâCTC needs its input text-line images cut from an already corrected layout, which are then transcribed to get itâs (text-line image, corrected Unicode text) input-label training pairs. Hence, the annotation cost of fine-tuning the OCR Recognition model on one page subsumes that of fine-tuning the GNN on one page. 5 Experiment Setup Data Splits. For each manuscript, we create five folds using a fixed random seed. In each fold, three pages are reserved for fine-tuning the GNN and the CNNâBiLSTMâCTC, and all remaining pages form the held-out test set. We quantify the gains due to this fine-tuning on 1, 2, and 3 pages by using the held-out test data pages of the respective target manuscript, across 5 folds. Pages may reappear across folds, but the fine-tuning and test sets are disjoint within each fold. Each fine-tuned pipeline checkpoint is evaluated on the same test pages. Multi-modal large language model pipeline. We also evaluate four off-the-shelf multi-modal OCR systems: Gemini-3.5-Flash, OpenAI GPT-5.6-Terra, Claude-Sonnet-5, and Sarvam Vision Document Digitization. Gemini, OpenAI, and Claude receive the same resized manuscript-page image and the same end-to-end prompt (see Supplementary Material), which requests a diplomatic Unicode Devanagari transcription together with one polygon per visual text-line in a normalized 00â10001000 coordinate system following the conventions established in Pix2Seq [13] and PaLI [14], which may favour Gemini 3.5 Flash, as the respective research which set the convention, was done by researchers associated with Google. The JSON outputs are parsed into the standard PAGE-XML representation containing locations of the text-lines in a bounding polygon format, and the corresponding Unicode text content. As Sarvam uses its own Sanskrit document-digitization API and returns HTML without bounding polygons at the text-line level, we parse the HTML into the same PAGE-XML text-line representation, preserving their emitted order and Unicode text but leaving the geometric information of the text-line location empty. Due to this, Sarvamâs Page-CER metric is evaluated by treating the emitted HTML text-line order as reading order. In the event that API failures or parsing failures occur, we retry 3 more times before considering the prediction of the page as a full failure. Multi-modal large language model predictions are evaluated for the exact same held-out test pages, across the exact five folds used for the traditional pipeline evaluation. The exact MMLLM model identifier requested and the HTTP endpoint it is requested from can be accessed in the Supplementary Material. The raw model responses are provided with the dataset released. Metrics. We quantify the transcription accuracy using standard metrics TextEdit [50] and Page-CER [11, 28, 32]. The TextEdit metric ignores line order and geometry, as it pairs each ground-truth line with at most one predicted line and counts unmatched lines as missing or extra text. To calculate TextEdit we use the exact same OmniDocBench simple_match score55 5 Code (v1.5): https://github.com/opendatalab/OmniDocBench/tree/v1_5, pinned at commit 59b103c.. Each non-empty PAGE-XML TextLine is parsed as an atomic item using its direct TextEquiv/Unicode transcription and supplied as an individual text block to the matcher. TextRegion membership is ignored. To calculate Page,-CER we sort the ground-truth and predicted lines from top to bottom and left to right, join the text strings in that order, and then calculate the Character Edit Distance(CER). 6 Results Table 2: Benchmarking the performance of Multi-Modal Large Language Models using the metrics TextEdit and Page-CER. For Sarvam Vision, * marks Page-CER computed by treating the modelâs emitted HTML text-line order as reading order, as Sarvam does not provide text-line level bounding polygons. Each entry reports the metric value followed by its 95% confidence interval. Lower values are better. Manuscript Multi-Modal LLM TextEdit â Page-CER â Moderate Layout Claude Sonnet 5 0.60â[0.45, 0.76]0.60\,[0.45,\,0.76] 0.64â[0.50, 0.80]0.64\,[0.50,\,0.80] Gemini 3.5 Flash 0.26 [0.18, 0.37] 0.35 [0.29, 0.44] GPT 5.6 Terra 0.71â[0.70, 0.73]0.71\,[0.70,\,0.73] 0.74â[0.73, 0.76]0.74\,[0.73,\,0.76] Sarvam Vision 0.33â[0.31, 0.36]0.33\,[0.31,\,0.36] 0.39â[0.36, 0.41]0.39\,[0.36,\,0.41]* Dense Layout Claude Sonnet 5 0.80â[0.60, 1.00]0.80\,[0.60,\,1.00] 0.85â[0.65, 1.00]0.85\,[0.65,\,1.00] Gemini 3.5 Flash 0.30 [0.27, 0.34] 0.44 [0.38, 0.47] GPT 5.6 Terra 0.79â[0.77, 0.80]0.79\,[0.77,\,0.80] 0.79â[0.78, 0.79]0.79\,[0.78,\,0.79] Sarvam Vision 0.43â[0.25, 0.66]0.43\,[0.25,\,0.66] 0.64â[0.47, 0.79]0.64\,[0.47,\,0.79]* Circular Layout Claude Sonnet 5 0.94â[0.83, 1.00]0.94\,[0.83,\,1.00] 0.97â[0.90, 1.00]0.97\,[0.90,\,1.00] Gemini 3.5 Flash 0.51 [0.43, 0.59] 0.51 [0.44, 0.58] GPT 5.6 Terra 0.81â[0.78, 0.85]0.81\,[0.78,\,0.85] 0.79â[0.76, 0.82]0.79\,[0.76,\,0.82] Sarvam Vision 0.63â[0.49, 0.77]0.63\,[0.49,\,0.77] 1.44â[0.57, 2.69]1.44\,[0.57,\,2.69]* Table 3: Quantifying the gains due to fine-tuning the Traditional OCR pipeline across manuscripts using the metrics TextEdit and Page-CER. For each metric, Iterative Fine-tuning (Fully Automatic) runs the fine-tuned pipeline inference fully automatically on the pages in the test data, whereas Iterative Fine-tuning (With Manual Layout Correction) denotes performing manual correction of predicted layouts before the downstream text recognition OCR. Metric values are followed by their 95% confidence intervals. Lower values are better. Manuscript Pages Fine-tuned TextEdit â Page-CER â Iterative Fine-tuning (Fully Automatic) Iterative Fine-tuning (With Manual Layout Correction) Iterative Fine-tuning (Fully Automatic) Iterative Fine-tuning (With Manual Layout Correction) Moderate Layout 0 0.34â[0.32, 0.38]0.34\,[0.32,\,0.38] 0.30â[0.28, 0.32]0.30\,[0.28,\,0.32] 0.32â[0.30, 0.36]0.32\,[0.30,\,0.36] 0.30â[0.28, 0.32]0.30\,[0.28,\,0.32] 1 0.28â[0.26, 0.30]0.28\,[0.26,\,0.30] 0.24â[0.22, 0.25]0.24\,[0.22,\,0.25] 0.25â[0.24, 0.26]0.25\,[0.24,\,0.26] 0.23â[0.22, 0.24]0.23\,[0.22,\,0.24] 2 0.24â[0.22, 0.26]0.24\,[0.22,\,0.26] 0.20â[0.19, 0.21]0.20\,[0.19,\,0.21] 0.22â[0.20, 0.23]0.22\,[0.20,\,0.23] 0.20â[0.19, 0.21]0.20\,[0.19,\,0.21] 3 0.23 [0.21, 0.24] 0.19â[0.17, 0.20]0.19\,[0.17,\,0.20] 0.20 [0.19, 0.22] 0.18â[0.17, 0.19]0.18\,[0.17,\,0.19] Dense Layout 0 0.33â[0.29, 0.38]0.33\,[0.29,\,0.38] 0.21â[0.20, 0.23]0.21\,[0.20,\,0.23] 0.34â[0.30, 0.36]0.34\,[0.30,\,0.36] 0.25â[0.23, 0.27]0.25\,[0.23,\,0.27] 1 0.32â[0.27, 0.37]0.32\,[0.27,\,0.37] 0.19â[0.19, 0.21]0.19\,[0.19,\,0.21] 0.33â[0.30, 0.36]0.33\,[0.30,\,0.36] 0.23â[0.22, 0.24]0.23\,[0.22,\,0.24] 2 0.28â[0.24, 0.34]0.28\,[0.24,\,0.34] 0.18â[0.17, 0.19]0.18\,[0.17,\,0.19] 0.29â[0.27, 0.30]0.29\,[0.27,\,0.30] 0.22â[0.21, 0.23]0.22\,[0.21,\,0.23] 3 0.26 [0.21, 0.32] 0.16â[0.16, 0.17]0.16\,[0.16,\,0.17] 0.26 [0.25, 0.27] 0.20â[0.19, 0.21]0.20\,[0.19,\,0.21] Circular Layout 0 0.48â[0.44, 0.52]0.48\,[0.44,\,0.52] 0.35â[0.33, 0.38]0.35\,[0.33,\,0.38] 0.52â[0.48, 0.56]0.52\,[0.48,\,0.56] 0.40â[0.37, 0.44]0.40\,[0.37,\,0.44] 1 0.46â[0.40, 0.51]0.46\,[0.40,\,0.51] 0.33â[0.30, 0.36]0.33\,[0.30,\,0.36] 0.51â[0.47, 0.55]0.51\,[0.47,\,0.55] 0.38â[0.35, 0.42]0.38\,[0.35,\,0.42] 2 0.44â[0.39, 0.49]0.44\,[0.39,\,0.49] 0.32â[0.29, 0.35]0.32\,[0.29,\,0.35] 0.50â[0.45, 0.54]0.50\,[0.45,\,0.54] 0.37â[0.33, 0.40]0.37\,[0.33,\,0.40] 3 0.40 [0.35, 0.45] 0.30â[0.27, 0.33]0.30\,[0.27,\,0.33] 0.48 [0.43, 0.53] 0.36â[0.33, 0.40]0.36\,[0.33,\,0.40] Off-the-Shelf Multi-modal LLMs. Table 3 reports the end-to-end performance of the four off-the-shelf multi-modal LLMs on the held-out test pages, aggregated over five folds as described in Section 5. Gemini 3.5 Flash performs best on every manuscript and on both metrics, with Sarvam Vision being the second in every case. Every multi-modal LLM ranks the difficulty of the three manuscripts in the same order: Moderate << Dense << Circular on both metrics. The Page-CER of Sarvam Vision on the Circular Layout Manuscript reaches 1.44â[0.57, 2.69]1.44\,[0.57,\,2.69], which is caused due to "runaway generation" type of hallucinations in which the model over-generates random tokens, and thus causes the character edit distance to exceed the number of ground-truth characters. Off-the-Shelf Traditional Pipeline. The traditional pipeline, when used off-the-shelf (with no fine-tuning on the target manuscript and no manual layout correction), performs comparably with Gemini 3.5 Flash as seen the 0-page fine-tuned rows of the "Iterative Fine-tuning (Fully Automatic)" subcolumns of Table 3, and the Gemini 3.5 Flash rows in Table 3. For the PAGE-CER metric, the off-the-shelf traditional pipeline outperforms Gemini 3.5 Flash on Moderate and Dense Layout Manuscripts but underperforms on the Circular Layout Manuscript. For the TextEdit metric, the traditional pipeline performs worse than Gemini on Moderate and Dense Layout Manuscripts but better on the Circular Layout Manuscript. Iteratively Fine-tuned Traditional Pipeline (Fully Automatic). While off-the-shelf results are important when performing OCR in bulk quantities, the benefits of iterative fine-tuning become more apparent when performing Historical OCR, where the traditional pipeline needs to adapt to target manuscript heterogeneity. In this setting, fine-tuning the GNN and the CNNâBiLSTMâCTC of the traditional pipeline on pages of the target manuscript helps it adapt to the distribution of the target manuscript at the layout-level and appearance-level, respectively, and improves prediction quality monotonically with each new page fine-tuned, as seen in "Iterative Fine-tuning (Fully Automatic)" subcolumns of Table 3. Fine-tuning both models on three pages reduces TextEdit by 34.1%34.1\% on the Moderate Layout Manuscript, 22.9%22.9\% on the Dense Layout Manuscript, and 15.5%15.5\% on the Circular Layout Manuscript. The corresponding Page-CER reductions are 37.5%37.5\%, 22.9%22.9\%, and 8.1%8.1\%. Once fine-tuned on 3 pages of each target manuscript, the traditional pipeline performs substantially better than Gemini 3.5 Flash on all three manuscripts on both metrics. Iteratively Fine-tuned Traditional Pipeline (with Manual Layout Correction). In the above Iterative Fine-tuning (Fully Automatic) setting, the fine-tuned pipeline does inference fully automatically, without any manual layout correction. However, enabling historians and scholars to perform Manual layout correction at inference time (on the test set in this experiment) can ensure that there are no costly layout analysis mistakes that can cause disastrous consequences for the downstream line level text recognition OCR task [26]. In other words, for the "Iterative Fine-tuning (Fully Automatic)" subcolumn, both the GNN and the CNNâBiLSTMâCTC perform inference automatically when fine-tuned on up to three pages. In comparison, in the "Iterative Fine-tuning (With Manual Layout Correction)" setting, the GNNâs predictions are manually corrected at inference time, providing error-free text-line detection, on which the fine-tuned CNNâBiLSTMâCTC performs inference automatically. The performance gap between the columns "Iterative Fine-tuning (Fully Automatic)" and "Iterative Fine-tuning (With Manual Layout Correction)" quantifies the effects of object-level layout-detection errorsâand, equivalently, quantifies the benefits of correcting them. We observe that the benefit of manual layout correction is minimal for Moderate Layout Manuscript (12.6%12.6\% TextEdit and 17.8%17.8\% Page-CER), and is the greatest for Dense Layout Manuscript (36.2%36.2\% TextEdit and 36.9%36.9\% Page-CER). Across all three manuscripts, the traditional pipeline achieves the best performance when fine-tuned on three pages with manual layout corrections enabled. In this experiment, 38.038.0, 165.5165.5, and 109.0109.0 seconds per page were required on average to perform manual layout correction on the Moderate, Dense, and Circular Layout Manuscripts, respectively. 0.20.20.30.30.40.40.50.5Page CER â LayoutDense LayoutCircular Layout0.20.20.30.30.40.40.50.5TextEdit â GNN only fine-tuning OCR only fine-tuning GNN and OCR fine-tuning OCR only fine-tuning, Perfect Layout Figure 4: We fine-tune the traditional pipeline up to three pages (across the same 5-fold train-test split as described in Section 5) using three ablations: one where only the GNN is fine-tuned at layout-level, one where only the CNN-BiLSTM-CTC OCR recognition model is fine-tuned at the appearance-level, and one where both are fine-tuned. We also report a fourth ablation where a human annotator manually corrects the layout of the test set pages at inference time, and an iteratively fine-tuned CNN-BiLSTM-CTC OCR recognition model is used to predict the text content from the text-line images. This fourth ablation is meant to mimic the expected working conditions of the Traditional Pipeline Annotation Tool, where manual layout correction takes a few minutes at most. Fine-tuning Ablations. Fig. 4 compares the gains in downstream OCR accuracy due to GNN-only fine-tuning, OCR-only fine-tuning, and combined fine-tuning. For the Moderate Layout manuscript, we observe that GNN-only layout-level fine-tuning does not help as much as the OCR-only fine-tuning because the pre-trained GNNâs layout predictions are already close to ground-truth, whereas fine-tuning the OCR Recognition model displays rapid adaptation to the target manuscript at the appearance-level. For Dense Layout and Circular Layout manuscripts, GNN-only and OCR-only gains due to fine-tuning are comparable on the TextEdit metric. The combined fine-tuning of GNN and OCR gives compounded improvement. On Circular, fine-tuning both models beats the sum of the two single-model gains by 19% (TextEdit). Considering the annotation cost in terms of annotation time, this compounding is free: a page supervising the OCR recognition model must have its layout corrected before its text-line images can be cut, so the GNN and OCR combined fine-tuning (black, open circles) costs no more than OCR-only fine-tuning (blue, filled squares), while GNN-only fine-tuning (grey dotted, triangles) requires manual layout correction alone. 7 Discussion When used without adaptation to the target manuscript, the traditional pipeline achieves comparable accuracy with the best-performing Multi-Modal LLM Gemini 3.5 Flash. However, fine-tuning the GNN and the CNNâBiLSTMâCTC of the Traditional pipeline on three corrected pages reduces TextEdit by up to 34.1%34.1\% and Page-CER by up to 37.5%37.5\%, thus achieving better results than Gemini 3.5 Flash on all three manuscripts on both metrics. Fine-tuning on each new page thus effectively reduces the human annotation effort required for the next page. This is desirable especially when the data is scarce, and annotation is costly and time-consuming. In addition to fine-tuning, performing manual layout correction at inference time removes the pipelineâs object-level layout-detection errors, such as incorrectly merged or split text-lines by the GNN, or incorrectly predicted missing or extra nodes by CRAFT. The errors that remain are attributable to the text-line unwrapping and processing, and the fine-tuned iterations of the text-line image recognition model CNNâBiLSTMâCTC. When combined with iterative fine-tuning, inference time Manual layout correction further reduced the TextEdit by up to 51.3%51.3\%, and roughly required tens of seconds to a few minutes per page. We thus conclude that the specialized Traditional pipeline, which leverages domain knowledge in various ways, is well suited to digitize Sanskrit manuscripts where the data is scarce, and annotation is time-consuming and expensive. However, a limitation of the pipeline is that it is brittle [32], because of itâs step by step nature, use of heuristics in unwrapping and processing the text-line images, and itâs dependence on the frozen pre-trained character detector CRAFT. We use this brittle but task-specific and locally fine-tunable traditional pipeline to bootstrap the creation of a richly annotated dataset containing layout-level annotations and downstream Unicode transcriptions, represented in both a graph-based format and the standard PAGE-XML format. Notably, the final outputs of the digitization taskânamely, text-line locations and Unicode text contentâcan be externally verified, making them suitable for post-training and fine-tuning modern multimodal large language models. This direction is promising because multimodal large language models are less susceptible to the brittleness of task-specific traditional pipelines. However, they are more data-intensive and require an initial curated dataset, which the proposed traditional pipeline can effectively bootstrap. Acknowledgements The authors wish to express their thanks to Lalchand Research Library, DAV College, Chandigarh, India, the eGangotri Project, and the Gyan Bharatam Project, for making manuscript data publicly available for educational and research purposes. The authors also wish to express their gratitude to the anonymous reviewers, Dr. Petar VeliÄkoviÄ, Dr. Dhaval Patel, Dr. Oliver Hellwig and Dr. Tarinee Awasthi for their invaluable support and feedback. The authors also wish to thank their colleagues Shagun Dwivedi, Janhavi Vaishampayan and Ansh Kushwaha for reviewing this work, and for their insightful suggestions. References [1] D. Adiga, R. Saluja, V. Agrawal, G. Ramakrishnan, P. Chaudhuri, K. Ramasubramanian, and M. Kulkarni (2018) Improving the learnability of classifiers for Sanskrit OCR corrections. In The 17th World Sanskrit Conference, Vancouver, Canada. IASS, p. 143â161. Cited by: §2. [2] V. Agrawal, N. Vadlamudi, M. Waseem, A. Joseph, S. Chitluri, and R. K. Sarvadevabhatla (2024) LineTR:unified text line segmentation for challenging palm leaf manuscripts. ICPR. Cited by: §2. [3] V. Agrawal, N. Vadlamudi, M. Waseem, A. Joseph, S. Chitluri, and R. K. Sarvadevabhatla (2025) LineTR: unified text line segmentation for challenging palm leaf manuscripts. In International Conference on Pattern Recognition, p. 217â233. Cited by: §2. [4] M. Aubreville, C. Bertram, M. Veta, R. Klopfleisch, N. Stathonikos, K. Breininger, N. ter Hoeve, F. Ciompi, and A. Maier (2021) Quantifying the scanner-induced domain gap in mitosis detection. arXiv preprint arXiv:2103.16515. Cited by: §2. [5] U. Author (Unknown Year) Muhurta Martanda (841 Gha Alm 4 Shlf 5 Devanagari Jyotish). Note: eGangotri Digital Preservation Trust. Accessed online at: https://archive.org/details/MuhurtaMartanda841GhaAlm4Shlf5DevanagariJyotish/mode/2upAccessed 16-Aug-2026 Cited by: 1(b), 1(b). [6] U. Author (Unknown Year) Tantra Raj With Yantra And Mantra Uddhara 5890 1430 Ka Almira 26 Shlf 3 Devanagari Stotr. Note: eGangotri Digital Preservation Trust. Accessed online at: https://archive.org/details/TantraRajWithYantraAndMantraUddhara58901430KaAlmira26Shlf3DevanagariStotr/mode/2upAccessed 16-Aug-2026 Cited by: 1(c), 1(c). [7] U. Author (Unknown Year) Yajnavalakyasmritih (Acharadhyayah). Note: Lalchand Research Library, DAV College, Chandigarh, India. Accessed online at: https://dav.splrarebooks.com/collection/view/yajnavalakyasmritih-acharadhyayahAccessed: 16-Aug-2026 Cited by: 1(a), 1(a). [8] Y. Baek, B. Lee, D. Han, S. Yun, and H. Lee (2019) Character region awareness for text detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9365â9374. External Links: Document Cited by: §3, Figure 3, Figure 3, §4.1. [9] M. Boillet, C. Kermorvant, and T. Paquet (2021) Multiple document datasets pre-training improves text line detection with deep neural networks. In 2020 25th International Conference on Pattern Recognition (ICPR), p. 2134â2141. Cited by: §2. [10] M. Boillet, C. Kermorvant, and T. Paquet (2021) Multiple document datasets pre-training improves text line detection with deep neural networks. In 2020 25th International Conference on Pattern Recognition (ICPR), p. 2134â2141. External Links: Link, Document Cited by: §2. [11] M. Boillet, C. Kermorvant, and T. Paquet (2022) Robust text line detection in historical documents: learning and evaluation methods. International Journal on Document Analysis and Recognition (IJDAR) 25 (2), p. 95â114. Cited by: §5. [12] R. Chamchong and C. C. Fung (2012) Text Line Extraction Using Adaptive Partial Projection for Palm Leaf Manuscripts from Thailand. In 2012 International Conference on Frontiers in Handwriting Recognition, Bari, Italy, p. 588â593 (en). External Links: ISBN 978-1-4673-2262-1, Link, Document Cited by: §2. [13] T. Chen, S. Saxena, L. Li, D. J. Fleet, and G. Hinton (2021) Pix2seq: a language modeling framework for object detection. arXiv preprint arXiv:2109.10852. Cited by: §5, §8.1. [14] X. Chen, J. Djolonga, P. Padlewski, B. Mustafa, S. Changpinyo, J. Wu, C. R. Ruiz, S. Goodman, X. Wang, Y. Tay, et al. (2023) Pali-x: on scaling up a multilingual vision and language model. arXiv preprint arXiv:2305.18565. Cited by: §5, §8.1. [15] K. Chincholikar, S. Dwivedi, K. Gopalan, and T. Awasthi (2025) A case study of handwritten text recognition from pre-colonial era sanskrit manuscripts. In Computational Sanskrit and Digital Humanities-World Sanskrit Conference 2025, p. 52â69. Cited by: §2, §4.3. [16] K. Chincholikar, K. Gopalan, and M. Hasabnis (2026) Towards text-line segmentation of historical documents using graph neural networks. In ICLR 2026 Workshop on Geometry-grounded Representation Learning and Generative Modeling, External Links: Link Cited by: §2, §4.1, §4.3. [17] T. Clanuwat, A. Lamb, and A. Kitamoto (2019) Kuronet: pre-modern japanese kuzushiji character recognition with deep learning. In 2019 International Conference on Document Analysis and Recognition (ICDAR), p. 607â614. Cited by: §2. [18] P. Community (n.d.) PaddleOCR. Note: https://paddlepaddle.github.io/PaddleOCR/main/en/index.htmlAccessed: 23-Nov-2024 Cited by: §2. [19] D. Coquenet, C. Chatelain, and T. Paquet (2023) Dan: a segmentation-free document attention network for handwritten document recognition. IEEE transactions on pattern analysis and machine intelligence 45 (7), p. 8227â8243. Cited by: §2. [20] G. Crosilla, L. Klic, and G. Colavizza (2025) Benchmarking large language models for handwritten text recognition. Journal of Documentation 81 (7), p. 334â354. Cited by: §2. [21] C. Cui, T. Sun, S. Liang, T. Gao, Z. Zhang, J. Liu, X. Wang, C. Zhou, H. Liu, M. Lin, et al. (2025) PaddleOCR-vl: boosting multilingual document parsing via a 0.9 b ultra-compact vision-language model. arXiv preprint arXiv:2510.14528. Cited by: §2. [22] D. Das (2021) Enhancing ocr performance with low supervision. Ph.D. Thesis, International Institute of Information Technology Hyderabad. Cited by: §2, §2. [23] A. Dwivedi, R. Saluja, and R. K. Sarvadevabhatla (2020) An ocr for classical indic documents containing arbitrarily long words. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, p. 560â561. Cited by: §2. [24] S. Dwivedi and K. Gopalan (2026) Comparative analysis of the intrinsic metrics for tokenizers and their effect on downstream tasks for Hindi and Marathi. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 22652â22663. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §3. [25] M. Fey, J. E. Lenssen, F. Weichert, and H. MĂŒller (2018) Splinecnn: fast geometric deep learning with continuous b-spline kernels. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 869â877. Cited by: §4.1. [26] N. Fischer, A. Hartelt, and F. Puppe (2023) Line-level layout recognition of historical documents with background knowledge. Algorithms 16 (3). External Links: Link, ISSN 1999-4893, Document Cited by: §2, §6. [27] F. C. Fizaine, P. Bard, M. Paindavoine, C. Robin, E. BouyĂ©, R. LefĂšvre, and A. Vinter (2024) Historical text line segmentation using deep learning algorithms: mask-rcnn against u-net networks. Journal of Imaging 10 (3), p. 65. Cited by: §2. [28] F. C. Fizaine, P. Bard, M. Paindavoine, C. Robin, E. BouyĂ©, R. LefĂšvre, and A. Vinter (2024) Historical text line segmentation using deep learning algorithms: mask-rcnn against u-net networks. Journal of Imaging 10 (3). External Links: Link, ISSN 2313-433X, Document Cited by: §5. [29] A. Graves, S. FernĂĄndez, and J. Schmidhuber (2007) Multi-dimensional recurrent neural networks. In International conference on artificial neural networks, p. 549â558. Cited by: §2. [30] T. GrĂŒning, R. Labahn, M. Diem, F. Kleber, and S. Fiel (2018) Read-bad: a new dataset and evaluation scheme for baseline detection in archival documents. In 2018 13th IAPR International Workshop on Document Analysis Systems (DAS), p. 351â356. Cited by: §2. [31] Z. He, C. Zhang, Z. Wu, Z. Chen, Y. Zhan, Y. Li, Z. Zhang, X. Wang, and M. Qiu (2025) Seeing is believing? mitigating OCR hallucinations in multimodal large language models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2. [32] H. Heidenreich, B. Elliott, O. Dinica, and Y. Getachew (2026) GutenOCR: a grounded vision-language front-end for documents. arXiv preprint arXiv:2601.14490. Cited by: §5, §7. [33] O. Hellwig OCR and digitization software for hindi and sanskrit - ind.senz. Note: Accessed 16-Aug-2026 External Links: Link Cited by: §2. [34] D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song (2021) Natural adversarial examples. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 15262â15271. Cited by: §2. [35] A. Jindal and R. Ghosh (2023) Text line segmentation in indian ancient handwritten documents using faster r-cnn. Multimedia Tools and Applications 82 (7), p. 10703â10722. External Links: Document Cited by: §2. [36] P. Kahle, S. Colutto, G. Hackl, and G. MĂŒhlberger (2017) Transkribus-a service platform for transcription, recognition and retrieval of historical documents. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), Vol. 4, p. 19â24. Cited by: §2. [37] T. Karayil, A. Ul-Hasan, and T. M. Breuel (2015) A segmentation-free approach for printed devanagari script recognition. In 2015 13th International Conference on Document Analysis and Recognition (ICDAR), p. 946â950. Cited by: §2. [38] B. Kiessling (2019) Kraken-an universal text recognizer for the humanities. In ADHO, Ăd., Actes de Digital Humanities Conference, Cited by: §2. [39] B. Kiessling (2020) A Modular Region and Text Line Layout Analysis System. In 2020 17th International Conference on Frontiers in Handwriting Recognition (ICFHR), Dortmund, Germany, p. 313â318 (en). External Links: ISBN 978-1-72819-966-5, Link, Document Cited by: §2. [40] B. Kiessling (2022) CurT: end-to-end text line detection in historical documents with transformers. In International Conference on Frontiers in Handwriting Recognition, p. 34â48. Cited by: §2. [41] M. KiĆĄ, K. BeneĆĄ, and M. HradiĆĄ (2021) AT-st: self-training adaptation strategy for ocr in domains with limited transcriptions. In Document Analysis and Recognition â ICDAR 2021, p. 463â477. External Links: ISBN 9783030863371, ISSN 1611-3349, Link, Document Cited by: §2. [42] O. Kodym and M. HradiĆĄ (2021) Page layout analysis system for unconstrained historic documents. External Links: 2102.11838, Link Cited by: §2. [43] J. KohĂșt and M. HradiĆĄ (2021) TS-net: ocr trained to switch between text transcription styles. In Document Analysis and Recognition â ICDAR 2021, p. 478â493. External Links: ISBN 9783030863371, ISSN 1611-3349, Link, Document Cited by: §2. [44] M. Li, T. Lv, J. Chen, L. Cui, Y. Lu, D. Florencio, C. Zhang, Z. Li, and F. Wei (2022) TrOCR: transformer-based optical character recognition with pre-trained models. External Links: 2109.10282, Link Cited by: §2. [45] H. Liao, A. RoyChowdhury, W. Li, A. Bansal, Y. Zhang, Z. Tu, R. K. Satzoda, R. Manmatha, and V. Mahadevan (2023) Doctr: document transformer for structured information extraction in documents. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 19584â19594. Cited by: §2. [46] M. Liwicki (2014) 3.8 divadia & hisdoc 2.0 approaches at the university of fribourg to digital paleography. Digital Palaeography: New Machines and Old Texts, p. 123. Cited by: §2. [47] O. Nath, S. Kukkala, M. Khapra, and R. K. Sarvadevabhatla (2026) IndicDLP: a foundational dataset for multi-lingual and multi-domain document layout parsing. In Document Analysis and Recognition â ICDAR 2025, X. Yin, D. Karatzas, and D. Lopresti (Eds.), Cham, p. 23â39. External Links: ISBN 978-3-032-04614-7 Cited by: §2. [48] T. Nguyen, J. Burie, T. Le, and A. Schweyer (2022) An effective method for text line segmentation in historical document images. In 2022 26th International Conference on Pattern Recognition (ICPR), p. 1593â1599. Cited by: §2. [49] S. A. Oliveira, B. Seguin, and F. Kaplan (2018) DhSegment: a generic deep-learning approach for document segmentation. In 2018 16th International Conference on Frontiers in Handwriting Recognition (ICFHR), p. 7â12. Cited by: §2. [50] L. Ouyang, Y. Qu, H. Zhou, J. Zhu, R. Zhang, Q. Lin, B. Wang, Z. Zhao, M. Jiang, X. Zhao, J. Shi, F. Wu, P. Chu, M. Liu, Z. Li, C. Xu, B. Zhang, B. Shi, Z. Tu, and C. He (2025) OmniDocBench: benchmarking diverse pdf document parsing with comprehensive annotations. External Links: 2412.07626, Link Cited by: §2, §5. [51] C. Papadopoulos, S. Pletschacher, C. Clausner, and A. Antonacopoulos (2013) The impact dataset of historical document images. In Proceedings of the 2Nd international workshop on historical document imaging and processing, p. 123â130. Cited by: §2. [52] V. Paruchuri (n.d.) Surya: A Sanskrit OCR tool. Note: https://github.com/VikParuchuri/suryaAccessed: 23-Nov-2024 Cited by: §2. [53] S. Pletschacher and A. Antonacopoulos (2010) The page (page analysis and ground-truth elements) format framework. In 2010 20th International Conference on Pattern Recognition, p. 257â260. Cited by: §1. [54] J. Poznanski, A. Rangapur, J. Borchardt, J. Dunkelberger, R. Huff, D. Lin, C. Wilhelm, K. Lo, and L. Soldaini (2025) Olmocr: unlocking trillions of tokens in pdfs with vision language models. arXiv preprint arXiv:2502.18443. Cited by: §2. [55] R. Saluja, D. Adiga, P. Chaudhuri, G. Ramakrishnan, and M. Carman (2017) Error Detection and Corrections in Indic OCR using LSTMs. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), Vol. 1, p. 17â22. External Links: Document Cited by: §2. [56] R. Saluja, D. Adiga, G. Ramakrishnan, P. Chaudhuri, and M. Carman (2017) A framework for document specific error detection and corrections in indic ocr. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), Vol. 4, p. 25â30. External Links: Document Cited by: §2. [57] N. Sankaran, A. Neelappa, and C. Jawahar (2013) Devanagari text recognition: a transcription based formulation. In 2013 12th International Conference on Document Analysis and Recognition, p. 678â682. Cited by: §2. [58] S. Sharan, S. Aitha, A. Kumar, A. Trivedi, A. Augustine, and R. K. Sarvadevabhatla (2021) Palmira: a deep deformable network for instance segmentation of dense and uneven layouts in handwritten manuscripts. In International Conference on Document Analysis and Recognition, p. 477â491. External Links: Document Cited by: §2. [59] A. Sharma, P. Jena, A. Joseph, and R. K. Sarvadevabhatla (2026) EpiSAM: character segmentation in challenging stone inscriptions. External Links: 2606.28859, Link Cited by: §2. [60] B. Shi, X. Bai, and C. Yao (2016) An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE transactions on pattern analysis and machine intelligence 39 (11), p. 2298â2304. Cited by: §2. [61] A. Trivedi and R. K. Sarvadevabhatla (2019) Hindola: a unified cloud-based platform for annotation, visualization and machine learning-based layout analysis of historical manuscripts. In 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW), Vol. 2, p. 31â35. External Links: Document Cited by: §2. [62] Unicode Consortium (2025) Unicode text segmentation. Unicode Standard Annex Technical Report 29, Unicode Consortium. Note: Revision 47, Unicode 17.0.0; edited by Josh Hadley External Links: Link Cited by: §3. [63] N. Vadlamudi, R. Krishna, and R. K. Sarvadevabhatla (2023) SeamFormer: high precision text line segmentation for handwritten documents. In International Conference on Document Analysis and Recognition, p. 313â331. External Links: Document Cited by: §2. [64] A. Vaswani (2017) Attention is all you need. Advances in Neural Information Processing Systems. Cited by: §2. [65] V. K. Venna, S. M. Gunda, J. S. Jinka, H. S. Rachakonda, A. Srinivasan, and R. K. Sarvadevabhatla (2026) M3Grounder: mask-based multi-span and multi-granular grounding for document qa. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 23685â23695. Cited by: §2. [66] V. Vinitha and C. Jawahar (2016) Error detection in indic ocrs. In 2016 12th IAPR Workshop on Document Analysis Systems (DAS), p. 180â185. External Links: Document Cited by: §2. [67] H. Wei, Y. Sun, and Y. Li (2025) DeepSeek-ocr: contexts optical compression. arXiv preprint arXiv:2510.18234. Cited by: §2. [68] S. Weil, R. Smith, and Z. Podobny (n.d.) Tesseract. Note: https://github.com/tesseract-ocr/tesseractAccessed: 23-Nov-2024 Cited by: §2. [69] Y. Xu, M. Li, L. Cui, S. Huang, F. Wei, and M. Zhou (2020) Layoutlm: pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, p. 1192â1200. Cited by: §2. [70] Y. Yin, H. Liu, Y, Q. Xie, C. Liu, S. Yang, S. Wang, Z. Liu, H. Zou, J. Chen, S. Wei, J. Wu, M. Huang, Z. Wu, G. Wang, T. Du, and L. Jia (2026) Unlimited ocr works. External Links: 2606.23050, Link Cited by: §2. [71] S. Zottin, A. De Nardin, G. Branca, C. Piciarelli, and G. L. Foresti (2025) ICDAR 2025 competition on few-shot text line segmentation of ancient handwritten documents (fest). In International Conference on Document Analysis and Recognition, p. 586â602. Cited by: §2. 8 Supplementary Material 8.1 Transcription using Off-the-shelf Multi-Modal LLMs Gemini-3.5-Flash, OpenAI GPT-5.6-Terra, Claude-Sonnet-5 received the exact same manuscript-page image and the same end-to-end prompt shown in Section 8.1 below, which requests a diplomatic Unicode Devanagari transcription together with one polygon per visual text-line in a normalized 00â10001000 coordinate system following the conventions established in Pix2Seq [13] and PaLI [14]. This convention might favour Gemini 3.5 Flash, as the research papers which set the convention were authored by researchers associated with Google. ⏠prompt_text = """ You are an expert Indologist and Paleographer specializing in handwritten Sanskrit manuscripts. Your Task: Perform a diplomatic transcription (OCR) of the manuscript image and provide text-line geometry. CRITICAL INSTRUCTIONS: 1. Output Format: Output ONLY raw valid JSON. No Markdown. 2. Coordinates: Coordinates are normalized from 0 to 1000, where [0,0] is top-left and [1000,1000] is bottom-right. 3. Geometry: For every visual text-line, output polygon_2d as [[y,x], ...]. Use a tight polygon following the visible line. If the line is straight and rectangular, box_2d [ymin,xmin,ymax,xmax] is also acceptable. For curved or circular lines, polygon_2d is mandatory. 4. Granularity: Transcribe at the visual text-line level. 5. Script: Unicode Devanagari. JSON SCHEMA: "status": "success", "regions": [ "id": "region_0", "type": "main_text", "polygon_2d": [[y,x], [y,x], [y,x]], "box_2d": [ymin, xmin, ymax, xmax], "lines": [ "id": "line_0", "polygon_2d": [[y,x], [y,x], [y,x]], "box_2d": [ymin, xmin, ymax, xmax], "text": "Transcribed text here" ] ] """ 8.2 Multi-Modal LLM Provider Information The exact MMLLM model identifier requested, and the HTTP endpoint it is requested from is shown in Table 4. The exact model responses are documented in the dataset released with this paper. Table 4: The four end-to-end VLM baselines: the exact model identifier requested and the HTTP endpoint it is requested from. Gemini, OpenAI, and Claude receive the resized page image followed by the same VLM_END_TO_END_PROMPT and return JSON; Sarvam receives the image alone and returns HTML, hence its different call shape. Method Model API endpoint gemini_e2e gemini-3.5-flash POST https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash:generateContent openai_e2e gpt-5.6-terra POST https://api.openai.com/v1/responses claude_e2e claude-sonnet-5 POST https://api.anthropic.com/v1/messages sarvam_e2e sarvam-visionâ POST https://api.sarvam.ai/doc-digitization/job/v1 ⥠and, relative to that base: POST /upload-files POST /job_id/start GETT /job_id/status POST /job_id/download-files â Sarvam Vision Document Digitization API accepts no model parameter: the request body carries only language=sa-IN and output_format=html. sarvam-vision is the identifier recorded in VLM_PROVIDER_SPECS and pinned into the cache fingerprint, not a value sent over the wire. ⥠Sarvam Vision is asynchronous and job-based: one page is one job, so a single page costs five calls in the order shown (create, upload, start, poll status until terminal, download). upload-files and download-files return presigned object-storage URLs; the page image and the result ZIP transfer over those, not over api.sarvam.ai.