Paper deep dive
Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment
Jahangir Alam SM, Md Khalid Syfullah, Saad Ahmed, Munira Akter Mou, A K Z Rasel Rahman, A. K. M. Masudur Rahman, Mohammed Sowket Ali
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This data article describes a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE) grading metadata. The collection contains 485 answer submissions from 415 consenting students at four academic institutions. Eight faculty contributors supplied examination materials covering nine subjects and 12 distinct question templates. Each answer-level item links a scanned PDF to a randomized identifier, subject label, question, model answer, criterion definitions, performance-level descriptions, criterion marks, and a total mark. The 12 rubrics contain 47 criteria in total. The scans retain realistic academic content, including handwriting, printed text, equations, tables, code, figures, sketches, and diagrams. CamScanner, Adobe Scan, and conventional scanners contributed variation in illumination, contrast, orientation, compression, and resolution. Diverse handwriting, crossed-out work, revised calculations, and inserted corrections add further visual variability for robustness and generalization studies. Preparation involved heterogeneous-source consolidation, label and text standardization, score validation, identifier randomization, filename randomization, and JSON-to-PDF integrity checks. An answer-level audit confirmed 485 unique identifiers, 485 unique PDF filenames, agreement between each total mark and its criterion-mark sum, and scores within the applicable rubric maximum. The data can support rubric-aware automated evaluation, multimodal document understanding, criterion-level feedback, score prediction, and privacy-aware OBE assessment research. Access is restricted to research use and is available from the corresponding author upon reasonable request.
Tags
Links
- Source: https://arxiv.org/abs/2608.22346v1
- Canonical: https://arxiv.org/abs/2608.22346v1
Trouble viewing inline? Open PDF directly →
Full Text
31,498 characters extracted from source content.
Expand or collapse full text
Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment Jahangir Alam SM a,∗ , Md Khalid Syfullah a,1 , Saad Ahmed a,1 , Munira Akter Mou a,1 , A K Z Rasel Rahman a,1 , A.K.M. Masudur Rahman a , Mohammed Sowket Ali a a Department of Computer Science and Engineering (CSE), Bangladesh Army University of Science and Technology (BAUST), Saidpur, Bangladesh Abstract This data article describes a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE) grad- ing metadata. The collection contains 485 answer submissions from 415 consenting students at four academic institutions. Eight faculty contribu- tors supplied examination materials covering nine subjects and 12 distinct question templates. Each answer-level item links a scanned PDF to a ran- domized identifier, subject label, question, model answer, criterion definitions, performance-level descriptions, criterion marks, and a total mark. The 12 rubrics contain 47 criteria in total. The scans retain realistic academic con- tent, including handwriting, printed text, equations, tables, code, figures, sketches, and diagrams. CamScanner, Adobe Scan, and conventional scanners contributed variation in illumination, contrast, orientation, compression, and resolution. Diverse handwriting, crossed-out work, revised calculations, and inserted corrections add further visual variability for robustness and gener- alization studies. Preparation involved heterogeneous-source consolidation, label and text standardization, score validation, identifier randomization, filename randomization, and JSON-to-PDF integrity checks. An answer-level audit confirmed 485 unique identifiers, 485 unique PDF filenames, agreement ∗ Corresponding author Email address: jahangir@baust.edu.bd (Jahangir Alam SM) 1 These authors contributed equally to this work. arXiv:2608.22346v1 [cs.CV] 23 Aug 2026 1. Value of the Data •The collection links visual examination evidence with question context, model answers, expert-authored OBE criteria, performance descriptions, criterion marks, and total marks at the answer level. •The nine-subject coverage and mixed visual content support research on handwritten and printed document understanding, optical character recog- nition, vision-language modeling, and cross-subject generalization. •Variation in lighting, handwriting, corrections, page geometry, and capture 2 pipelines spanning CamScanner, Adobe Scan, and conventional scanners supports robustness studies under realistic acquisition conditions. •Criterion-level labels permit development of systems that report which parts of an answer satisfy examiner expectations, extending automated grading beyond a single aggregate score. • The paired PDF-JSON structure supports controlled experiments in score prediction, rubric selection, evidence grounding, feedback generation, cali- bration, and auditability. •The preparation workflow provides a reproducible pattern for consolidat- ing heterogeneous examiner files, validating scores, randomizing record identifiers, and preserving file-to-record mappings. •Education researchers, document-AI researchers, assessment specialists, and institutions developing OBE workflows can reuse the data under approved research-only access conditions. 2. Background Outcome-Based Education organizes curriculum, instruction, and assess- ment around explicit capabilities that learners are expected to demonstrate [1,2]. Assessment within this framework requires a visible connection among a question, the intended learning evidence, the criteria used by an examiner, and the marks assigned to the response. An overall score records final attain- ment in compact form. It does not preserve which required components were demonstrated, omitted, or expressed with partial quality. Analytic rubrics address this limitation by dividing a task into criteria and describing performance at multiple quality levels. Clear, focused, task-specific rubrics can strengthen the consistency and interpretability of performance assessment [3,4]. A rubric-aware dataset needs more than images and total marks. It must retain criterion definitions, performance descriptions, maximum allocations, criterion-level marks, and the link to the relevant question and reference answer. Automatic short-answer grading has traditionally emphasized typed re- sponses and aggregate labels. Prior reviews describe a progression from rule-based systems to statistical and neural approaches, with substantial vari- ation in tasks and evaluation settings [5]. Real examination scripts introduce a broader document-understanding problem. Student work may contain hand- writing, equations, tables, code, graphs, or diagrams. Page layout and visual evidence can influence interpretation. Modern document-AI models combine 3 text, layout, and image information [6], while optical-character-recognition- free architectures learn direct mappings from document images to structured outputs [7]. Such systems require data that preserve both page appearance and structured assessment context. The present collection was assembled for this purpose. Its unit of orga- nization is one complete answer submission. Each submission retains the scanned PDF and an answer-level JSON record. The record supplies the subject, question, model answer, total mark, and a list of rubrics with crite- rion text, gained marks, and performance-level marking rules. This structure supports transparent experiments in which predicted scores can be traced to criterion-specific evidence. The data article focuses on the data, its organization, and its preparation. It does not report model comparisons or operational grading performance. The dataset is intended for controlled research and is not intended to determine official student grades. 3. Data Description 3.1. Collection overview The collection [8] contains 485 answer submissions associated with 415 participating students from four institutions. The difference between the two counts reflects answer-level organization: a participant can contribute more than one answer submission. Eight faculty contributors supplied the scanned scripts and assessment metadata. The coverage includes Machine Learning, Digital Image Processing, Database Systems, Computer Networks, Data Mining, Algorithms, Fisheries, Object Oriented Programming, and E-commerce. Each scan retains the original page appearance. The visual content includes plain text, mathematical notation, tables, figures, diagrams, code, and mixed layouts. Individual files vary in length. Some scripts contain one response, while others contain several questions or subquestions. Figure 1 presents representative page samples from the collection. The scans were acquired with CamScanner, Adobe Scan, and conventional document scanners. Mobile capture and flatbed or sheet-fed scanning intro- duced different lighting conditions, brightness and contrast levels, shadows, page orientation, cropping, compression, and resolution. The scripts also preserve distinct handwriting styles, crossed-out text, overwritten symbols, corrected calculations, arrows, insertions, and other authentic revision marks. 4 Sample 1 (p.1)Sample 1 (p.2)Sample 2 (p.1)Sample 2 (p.2)Sample 2 (p.3)Sample 2 (p.4)Sample 2 (p.5) Sample 3 (p.1)Sample 4 (p.1)Sample 4 (p.2)Sample 4 (p.3)Sample 5 (p.1)Sample 6 (p.1)Sample 6 (p.2) Sample 7 (p.1)Sample 7 (p.2)Sample 8 (p.1)Sample 8 (p.2)Sample 8 (p.3)Sample 8 (p.4) Figure 1: Representative scanned answer pages from the dataset. The overview shows vari- ation in handwriting, page count, equations, tables, diagrams, and institutional document formats. This combination broadens the visual domain represented by the dataset and supports evaluation of generalization across writing styles, document content, and acquisition conditions. External validation remains necessary before conclusions are extended beyond the participating institutions and subjects. 3.2. Answer-level files and metadata The final organization uses paired visual and structured records. One randomized PDF filename identifies the scanned response, and one JSON object contains its metadata. Each object storesanswer_id,pdf_file, Subject,question,model_answer, andtotal_marks. Therubricsfield is a variable-length array whose length equals the number of criteria defined for the associated question. Each rubric element stores the criterion text, itsgained_marks, and the performance descriptions undermarking_rules. The four performance keys are Excellent, Good, Average, and Poor. The JSON structure keeps the grading context next to the file reference, as illustrated in Figure 2. It supports answer-level loading without joining separate question, rubric, and mark tables. The underlying PDF remains unchanged as visual evidence. This arrangement is suitable for pipelines 5 "answer_id": "...", "pdf_file": "...", "Subject": "...", "question": "...", "model_answer": "...", "total_marks": "...", "rubrics": [ "criteria": "...", "gained_marks": "...", "marking_rules": "Excellent": "...", "Good": "...", "Average": "...", "Poor": "..." ] Figure 2: Answer-level JSON representation linking a randomized scan with its question, reference material, rubric criteria, obtained marks, and performance-level rules. that render pages, extract text, encode document images, or combine both modalities. 3.3. Subject distribution Table 2 summarizes the subject distribution derived from the workbook and verified against the final JSON. Dataset shares are calculated from 485 answers. Mean marks are normalized by the maximum of each answer’s question rubric prior to subject-level averaging. This normalization makes questions with 5, 6, 9, 10, or 16 available marks comparable within a subject. 6 Table 2: Subject-level composition. Mean normalized mark is descriptive metadata rather than a model evaluation. SubjectAnswers QuestionsShare (%) Mean normal- ized mark (%) Visual answer content Machine Learning 88218.1477.90Primarily text Digital Image Processing 75215.4677.20Primarily text Database Systems 80216.4962.91Text, tables, equations Computer Networks 70114.4364.57Text, optional diagrams Data Mining 4218.6656.40Primarily text Algorithms83117.1182.98Text, equations, diagrams Fisheries2314.7491.49Primarily text Object Oriented Program- ming 1913.9253.45Text and code E- commerce 511.0348.89Primarily text Total48512100.00 3.4. Question and rubric structure The 12 question templates contain 47 criteria. Nine questions use four criteria, two use three criteria, and one uses five criteria. Question maxima range from 5 to 16 marks. Table 3 reports the answer type, number of answers, criterion count, rubric maximum, mean total mark, and observed mark range for each question. The verified answer-level JSON supplies the descriptive values, while the workbook supplies the answer-type categories. Rubrics were defined at the question level. Each criterion names an assess- able component and associates it with textual descriptions of performance quality. The supplied template uses four performance columns. Numeric allocations remain question-specific, and the final mark for a criterion is stored ingained_marks. Performance labels provide descriptive guidance; they are 7 Model Answer: CriteriaExcellent (4)Good (3)Average (2)Poor (1) [Criteria 1] [Statement for excellent on Criteria 1] [Statement for good on Criteria 1] [Statement for average on Criteria 1] [Statement for poor on Criteria 1] [Criteria 2] [Statement for excellent on Criteria 2] [Statement for good on Criteria 2] [Statement for average on Criteria 2] [Statement for poor on Criteria 2] [Criteria 3] [Statement for excellent on Criteria 3] [Statement for good on Criteria 3] [Statement for average on Criteria 3] [Statement for poor on Criteria 3] [Criteria 4] [Statement for excellent on Criteria 4] [Statement for good on Criteria 4] [Statement for average on Criteria 4] [Statement for poor on Criteria 4] MARKS ARE IN THE SECOND SHEET OF THIS EXCEL Question: [All Questions Here] [Model Answer 1 Here] [Model Answer 2 Here (If Applicable)] [Model Answer 3 Here (If Applicable)] Rubrics (a) Illustrative blank rubric tem- plate. Model Answer: CriteriaExcellent (4)Good (3)Average (2)Poor (1) Understanding of Concept Fully correct and clearly demonstrates the concept Mostly correct with minor issues Partial understanding Shows little understanding; answer mostly incorrect Step-by-Step Procedure All steps clearly shown and well structuredMostly correct stepsFew correct stepsMissing or wrong steps Accuracy of Final Answer Fully accurate and justifiedMostly correctPartially correctWrong result Clarity & Presentation Very clear, neat diagrams/tables, strong justificationClear explanation Understandable but not neatPoorly organized MARKS ARE IN THE SECOND SHEET OF THIS EXCEL Question 1. You are given the following values: 12, -5.4, 10000, 3.14, 0, 87, -200, 45.6, 999, 1.2. Choose the most efficient among Counting Sort, Radix Sort, or Bucket Sort for this dataset. Justify your selection and state the time and space complexity of all three algorithms. 2. A directed weighted graph is given below. Determine whether the graph contains a negative cycle using a step-by-step solution. Graph (text form): Vertices = A, B, C, D; Edges = A→B (4), B→C (-6), C→D (2), D→B (1), A→D (5). 3. You are given coin denominations: 1, 3, 4, 5. Find the minimum number of coins required to make 10 taka. Show the DP table and derive the final answer. Answer 1: Bucket Sort is the most efficient for this dataset because it contains floating-point numbers and a wide range of values including negatives. Counting Sort is not suitable due to large range and non-integers, and Radix Sort is inefficient for floating-point values. Time Complexity: Counting Sort = O(n + k), Space = O(k); Radix Sort = O(d(n + k)), Space = O(n + k); Bucket Sort = O(n + k) average, worst O(n²), Space = O(n + k). Therefore, Bucket Sort is preferred. Answer 2: Using Bellman-Ford algorithm: Step 1 initialize distances (A=0, others=∞). Step 2 relax edges repeatedly: A→B (4), B→C (-6), C→D (2), D→B (1). After V-1 iterations, distances keep decreasing due to cycle B→C→D→B with total weight (-6 + 2 + 1 = -3). Step 3 one more relaxation still reduces distance, confirming a negative cycle exists. Therefore, the graph contains a negative cycle. Answer 3: Using Dynamic Programming: Let dp[x] = minimum coins to make x. Initialize dp[0]=0. Fill table up to 10: dp[1]=1, dp[2]=2, dp[3]=1, dp[4]=1, dp[5]=1, dp[6]=2, dp[7]=2, dp[8]=2, dp[9]=2, dp[10]=2. Minimum coins for 10 = 2 (5+5 or 4+3+3 is larger). Final answer: 2 coins. Rubrics (b) Illustrative completed ques- tion, model answers, rubric, and criterion-level marks. SLAnswer ID Understanding of Concept Step-by-Step Procedure Accuracy of Final Answer Clarity & Presentation Total Marks 1PseudoID1.pdf332311 2PseudoID2.pdf342413 3PseudoID3.pdf332311 4PseudoID4.pdf343414 5PseudoID5.pdf443415 6PseudoID6.pdf443415 7PseudoID7.pdf443415 8PseudoID8.pdf343414 9PseudoID9.pdf443415 10PseudoID10.pdf443415 11PseudoID11.pdf342312 12PseudoID12.pdf342312 13PseudoID13.pdf443415 14PseudoID14.pdf332311 15PseudoID15.pdf332311 16PseudoID16.pdf443415 17PseudoID17.pdf443415 18PseudoID18.pdf443415 19PseudoID19.pdf443415 20PseudoID20.pdf342413 21PseudoID21.pdf433313 22PseudoID22.pdf443314 23PseudoID23.pdf334212 24PseudoID24.pdf433313 25PseudoID25.pdf444315 26PseudoID26.pdf433313 27PseudoID27.pdf333413 28PseudoID28.pdf434415 29PseudoID29.pdf433313 30PseudoID30.pdf443314 (c) Criterion marks and total marks. Figure 3: Outcome-Based Education rubric materials used during data acquisition. The three source PDFs are included directly and displayed as subfigures. not a substitute for the criterion maximum. Figure 3 presents the blank template and the two pages of one completed rubric example as requested. 4. Experimental Design, Materials and Methods 4.1. Data acquisition and expert rubric design The source data were collected through participating faculty members at four academic institutions. A total of eight faculty contributors supplied scanned examination scripts and examiner-prepared assessment spreadsheets. Student consent was obtained for preparation and research use of the scripts. The source pool represented 415 students and produced 485 answer submis- sions. Faculty contributors provided the question, model or reference answer, rubric, criterion descriptions, criterion-level marks, and total mark for each applicable answer. Rubrics were designed for specific questions with involve- ment from an OBE expert. Criteria reflected the knowledge or performance components expected in each response, including conceptual correctness, com- pleteness, procedural accuracy, examples, technical accuracy, interpretation, and clarity. Different questions retained different numbers of criteria and mark allocations. 8 The scanned PDFs were preserved rather than transcribed into text- only responses. This decision retained handwriting, page layout, equations, tables, sketches, diagrams, code, and other evidence that may be relevant to multimodal assessment. The original scripts also retained realistic differences in page count and question organization. The answer scripts were digitized through multiple acquisition pathways. Mobile scanning applications, including CamScanner and Adobe Scan, were used alongside conventional document scanners. The resulting files retain variations in illumination, brightness, contrast, shadows, orientation, crop- ping, compression, and resolution. Student responses also exhibit diverse handwriting, overwritten text, crossed-out work, inserted corrections, and revised calculations. Preserving these traits broadens the visual domain of the collection and supports evaluation across realistic acquisition and writing conditions. 4.2. Consolidation of heterogeneous examiner data The examination materials were collected from different faculty mem- bers and institutions, and the raw files were not initially available in one standardized representation. Differences existed in file naming, spreadsheet organization, question formatting, criterion formatting, and mark representa- tion. The first processing operation combined these independently submitted sources into a common answer-level data structure. For each student answer, the corresponding PDF was linked with: • subject information; • examination question; • model or reference answer; • OBE rubric criteria; • rubric-level performance descriptions; • criterion-level obtained marks; and • total obtained marks. The merged representation was organized with one record for each com- plete answer submission. Figure 2 shows the resulting JSON pattern. This organization preserves the relationship among the visual answer, academic prompt, reference material, rubric definitions, and examiner-assigned marks. The consolidation process established a one-to-one mapping between each metadata object and one scanned PDF. Rubric objects were stored as arrays to preserve question-specific criterion counts. No fixed four-criterion 9 structure was imposed. This choice retained the three-criterion Fisheries and E-commerce rubrics and the five-criterion Computer Networks rubric. 4.3. Metadata standardization and completion Subject labels from examiner files were mapped to the nine finalized names reported in Table 2. Capitalization, spacing, and common abbreviations were standardized. Question strings were cleaned to remove redundant score annotations where they did not form part of the academic prompt. Question numbers and ordered subquestions were retained. Criterion text received typographic and spacing corrections while its educational meaning and original allocation were preserved. 4.4. Score validation Every criterion mark was checked against the applicable criterion max- imum. Any source value above its valid maximum was corrected during preparation, and the total mark was updated where the correction changed the aggregate. The final answer-level JSON was audited across all 485 records. Each stored total equals the sum of its criterion marks. No total exceeds its question-level rubric maximum. The verified range for Question 1 is 0 to 5, with a mean of 2.82 marks. Question maxima vary across the collection. Normalized descriptive marks in Table 2 were calculated for each answer as s i = m i M q(i) ,(1) wherem i is the total mark for answeriandM q(i) is the maximum for its question. A subject mean is the arithmetic mean ofs i over all answers in that subject. The question-level table retains the original mark units. 4.5. De-identification and referential integrity The 485 records were assigned a randomly shuffled set of unique integers from 0 to 484. Original filenames were replaced by randomly generated unique PDF filenames. Thepdf_filevalue was updated at the same time as the physical filename to preserve the correct mapping. Integrity checks verified answer-ID uniqueness, filename uniqueness, one PDF reference per record, and preservation of the pre-randomization record- to-file mapping. The final metadata contain 485 unique IDs and 485 unique 10 randomized filenames. The file preparation process checked that each refer- enced PDF existed in the assembled file directory. Record-level de-identification does not guarantee removal of text or logos embedded within every scan. Visible identifiers that remain in image content are addressed through restricted research access and the privacy conditions described in Section 6. 4.6. Preparation workflow Figure 4 summarizes the four-stage process. Raw questions and rubric designs were combined with contributions from 415 students. Data acquisition linked scans, examiner metadata, and OBE rubrics. Standardization merged sources, cleaned metadata, validated marks, randomized IDs and filenames, and checked integrity. The final organization contains 485 answer-level PDF- JSON pairs. Stage 1: Raw Data Question & Rubric Design 415 Participating Students 4 institutions 9 Subjects 1. Machine Learning 2. Digital Image Processing 3. Database Systems (DBMS) 4. Computer Networks 5. Data Mining 6. Algorithms 7. Fisheries 8. Object Oriented Programming 9. E-commerce Stage 2: Data Acquisition Stage 3: Standardization and De- identification Stage 4: Final Dataset 485 name anonymized PDF-JSON pairs OBE Rubrics Contributors Faculty exam materials from multiple institutions Scanned Exam Answers PDFs with text, equations, figures, and diagrams Examiner Metadata Questions, model answers, marks, and criterion scores OBE Rubrics Criterion-wise grading rules and expected performance levels Coverage 8 contributors 9 subjects Complete Examiner Package PDFs + questions + model answers + rubrics + criterion scores + obtained marks 1. Merge Sources Combine PDFs, metadata, and rubrics into one answer-level dataset 2. Metadata Cleaning Standardize names and formatting; remove empty criterias 3. Score Validation Check criterion scores, correct over-allocation and total marks 4. De-identification Assign random answer IDs and random unique PDF names 5. Integrity Check Verify JSON-to-PDF mapping and confirm every file exists Figure 4: Dataset preparation workflow from raw examination materials to the standardized answer-level collection. The supplied process diagram is included directly. 4.7. Suggested research use A typical research pipeline can load a JSON object, retrieve its paired PDF, render one or more pages, and construct a target from the total mark or criterion marks. Models may consume page images, optical-character- recognition text, layout features, or multimodal combinations. Training and evaluation partitions should be created at the question level or with grouped controls when the research objective involves transfer to unseen questions. Random answer-level splitting can place visually related responses to the same prompt in every partition and may overstate cross-question generalization. Criterion marks can be represented in their original units or divided by criterion maxima. Question and subject labels should remain available for 11 stratification. The smallest subject contains five answers, while the largest contains 88. Sampling plans and uncertainty estimates should reflect this imbalance. Any generated feedback should be treated as a research output and must not be presented as an official grade. The paired files and machine-readable metadata support principles of interoperability and reuse [9]. The restricted access condition reflects the presence of educational records and residual identifiers in some source images. The naturally occurring image variation also supports generalization- oriented evaluation. Robustness studies can group pages by observable characteristics such as illumination, contrast, skew, compression, handwriting style, correction density, and document complexity. The combination of mobile scanning applications and conventional scanners creates acquisition shifts that are useful for assessing whether a model remains stable outside a single capture pipeline. These variations strengthen the dataset as a research resource for evaluating multimodal document understanding under realistic examination conditions. 5. Limitations The collection has several constraints relevant to reuse. It covers nine sub- jects and only 12 question templates, and its subject distribution is uneven. E-commerce has five answers, while Machine Learning has 88. Question- specific rubrics use different maxima, criterion counts, and textual scales. Comparisons across questions require normalization or question-aware model- ing. 6. Ethics and Privacy Statement Student consent was obtained for preparation and research use of the examination scripts. The dataset and associated research activities were separated from official grading decisions. The data were not used to determine, alter, or replace the direct grades assigned to participating students. Answer IDs and PDF filenames were randomized during preparation. Some scanned pages retain student names, student IDs, university names, institutional logos, or related identifiers within the page image. These elements are retained only within the controlled research context. Dataset access is limited to approved research purposes. Recipients must protect the files, 12 avoid re-identification, avoid public redistribution, and report only aggregate or de-identified information. Any future use involving automated assessment must retain human moni- toring. Outputs from research models must not be treated as official grades or used as the sole basis for a decision affecting a student. Acknowledgments This project was funded by the Bangladesh Accreditation Council (BAC) under the project AI-Powered Automated Exam Evaluation and OBE-Based University Assessment System. The authors acknowledge the participating stu- dents, faculty contributors, institutions, and the OBE expert who supported rubric preparation and data collection. Data and Code Availability The dataset is deposited on Zenodo (DOI: 10.5281/zenodo.22058761) under restricted access and is available for research use upon reasonable request via the repository, subject to privacy review and the safeguards stated in this article. No source code is distributed with the manuscript package; supporting preparation code may be shared by the corresponding author under the same research-only conditions. Declaration of Generative AI and AI-Assisted Technologies in the Manuscript Preparation Process Generative artificial intelligence was used to assist with preparation of the manuscript. The human authors monitored and reviewed the manuscript, numerical statements, tables, figures, and references. The authors take full responsibility for the final content. Declaration of Competing Interest The authors declare that there are no known competing financial interests or personal relationships that could have appeared to influence the work reported in this article. 13 References [1]W. G. Spady, Outcome-Based Education: Critical Issues and Answers, American Association of School Administrators, Arlington, VA, 1994. [2]J. Biggs, Enhancing teaching through constructive alignment, Higher Education 32 (3) (1996) 347–364. doi:10.1007/BF00138871. [3]A. Jönsson, G. Svingby, The use of scoring rubrics: Reliability, validity and educational consequences, Educational Research Review 2 (2) (2007) 130–144. doi:10.1016/j.edurev.2007.05.002. [4]S. M. Brookhart, F. Chen, The quality and effectiveness of descriptive rubrics, Educational Review 67 (3) (2015) 343–368.doi:10.1080/0013 1911.2014.929565. [5] S. Burrows, I. Gurevych, B. Stein, The eras and trends of automatic short answer grading, International Journal of Artificial Intelligence in Education 25 (1) (2015) 60–117. doi:10.1007/s40593-014-0026-8. [6]Y. Huang, T. Lv, L. Cui, Y. Lu, F. Wei, LayoutLMv3: Pre-training for document AI with unified text and image masking, in: Proceedings of the 30th ACM International Conference on Multimedia, 2022, p. 4083–4091. doi:10.1145/3503161.3548112. [7]G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, S. Park, OCR-free document understanding transformer, in: Computer Vision – ECCV 2022, Vol. 13688 of Lecture Notes in Computer Science, Springer, 2022, p. 498–517.doi:10.1007/978-3-031-19815-1 _29. [8] J. Alam SM, S. Ahmed, M. A. Mou, A. K. Z. R. Rahman, M. K. Sy- fullah, A. M. Rahman, Multimodal examination answer dataset with expert-designed outcome-based education (obe) rubrics for criterion-level assessment (2026). doi:10.5281/zenodo.22058761. [9] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, et al., The FAIR guid- ing principles for scientific data management and stewardship, Scientific Data 3 (2016) 160018. doi:10.1038/sdata.2016.18. 14 Table 3: Question-level grading structure, answer type, and mark coverage. QSubject and focus Question summary Answer typeAnswers CriteriaMean mark Range 1Machine Learning, R 2 Importance, mathematical interpretation, limitations, and model comparison Plain text2542.82/5 0 to 5 2 Machine Learning, paradigms Supervised, unsupervised, and reinforcement learning Plain text634 8.64/100 to 10 3 Digital Image Processing, quantization Purpose of quantization and the Max-Lloyd algorithm Plain text4048.53/102 to 10 4 Digital Image Processing, histograms Histogram equalization and histogram matching Plain text3546.80/104 to 10 5 Database Systems, normalization Redundancy, update anomalies, normal forms, and examples Text, tables, equations 504 6.25/102 to 9.5 6 Database Systems, keys Schema, instance, database keys, and relation-specific identification Text, tables, equations 304 6.37/103 to 10 7 Computer Networks, OSI model Seven layers, technical accuracy, protocols, and presentation Text, optional diagrams 705 6.46/103 to 10 8 Data Mining, KDD Definition, KDD steps, importance, and application Plain text4249.02/160 to 16 9 AlgorithmsSorting choice, negative-cycle detection, and dynamic programming Text, equations, diagrams 834 13.28/167 to 16 10 FisheriesClimate, climate variables, and temperature- related changes in rohu Plain text2335.49/63.75 to 6 11 Object Oriented Programming Method overloading, overriding, comparison, and examples Text and code194 8.55/164 to 15 12 E-commerceEncryption, firewall, and symmetric-key encryption Plain text534.40/9 0 to 8 15