Paper deep dive
A Standardized Framework for Machine Learning in Power System Protection
Julian Oelhaf, Georg Kordowich, Paula Andrea Pérez-Toro, Christian Bergler, Johann JÀger, Andreas Maier, Siming Bayer
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/21/2026, 4:15:07 AM
Summary
This paper proposes a standardized evaluation framework for machine learning-based power system protection to address the lack of reproducibility and comparability in current research. The framework defines seven critical dimensions: protection objective, physical scope, observability, timing, targets, validation protocol, and evaluation outputs. It is instantiated in a case study using the PROTECT-90 benchmark, demonstrating that an MLP achieves high performance in fault classification and localization under specific conditions, while highlighting the impact of observability and timing on robustness.
Entities (7)
Relation Signals (5)
Julian Oelhaf â affiliatedwith â Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg
confidence 99% · Julian Oelhaf julian.oelhaf@fau.de ... organization=Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg
PROTECT-90 â usedin â Case Study
confidence 95% · The framework is instantiated in a bounded case study on the public PROTECT-90 electromagnetic-transient benchmark
Multi-Layer Perceptron â achievedperformanceon â Fault Classification
confidence 93% · a multi-layer perceptron (MLP) achieved a five-fold mean macro-averaged F1 score of 0.991 +/- 0.001 for classification
Multi-Layer Perceptron â achievedperformanceon â Fault Localization
confidence 93% · a localization mean absolute error of 10.20 +/- 0.25% of line length
Proposed Framework â addresses â Reproducibility Crisis
confidence 90% · This echoes the broader reproducibility crisis in machine-learning-based science... The framework turns evaluation assumptions into explicit, reproducible evidence
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Studies of machine-learning-based power-system protection increasingly report near-perfect scores, yet the meaning of those scores depends strongly on the evaluation setting. Protection task, physical scope, measurements, timing, targets, preprocessing, and validation often vary jointly and remain incompletely specified. This paper proposes a standardization-oriented framework that treats evaluation design as part of the scientific contribution. It defines seven required study dimensions: protection objective, physical scope, observability, timing and decision windows, targets and sample validity, validation protocol, and evaluation outputs. The framework is instantiated in a bounded case study on the public PROTECT-90 electromagnetic-transient benchmark, comprising 9022 simulated episodes from a 90 kV double-line topology, for onset-conditioned fault classification and localization. Under centralized sensing, simulation-metadata-aligned 20 ms windows, and episode-grouped validation, a multi-layer perceptron (MLP) achieved a five-fold mean macro-averaged F1 score of 0.991 +/- 0.001 for classification and a localization mean absolute error of 10.20 +/- 0.25% of line length (mean +/- std across episode-grouped folds). Extending the decision horizon to 50 ms preserved this task-dependent performance asymmetry, while reduced observability approximately doubled the MLP localization error but had little effect on classification. A synchronized two-ended conventional locator outperformed the learning locators under its richer clean information set, and measurement degradation showed that clean predictive performance did not determine robustness. The framework turns evaluation assumptions into explicit, reproducible evidence and provides a basis for more comparable, auditable evaluation and future certification-oriented assessment of machine-learning protection functions.
Tags
Links
- Source: https://arxiv.org/abs/2608.20181v1
- Canonical: https://arxiv.org/abs/2608.20181v1
Trouble viewing inline? Open PDF directly â
Full Text
141,805 characters extracted from source content.
Expand or collapse full text
[orcid=0009-0008-8204-589X] [orcid=0000-0003-2225-7926] [orcid=0000-0002-2727-2116] [orcid=0000-0002-9550-5284] [orcid=0000-0003-2874-4805] A Standardized Framework for Machine Learning in Power System Protection Julian Oelhaf julian.oelhaf@fau.de https://lme.tf.fau.de Georg Kordowich georg.kordowich@fau.de https://ees.tf.fau.de Paula Andrea PĂ©rez-Toro paula.andrea.perez@fau.de Christian Bergler c.bergler@oth-aw.de https://oth-aw.de Johann JĂ€ger johann.jaeger@fau.de Andreas Maier andreas.maier@fau.de Siming Bayer siming.bayer@fau.de organization=Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg, addressline=Martensstr. 3, postcode=91058, city=Erlangen, country=Germany organization=Institute of Electrical Energy Systems, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg, addressline=Cauerstr. 4, postcode=91058, city=Erlangen, country=Germany organization=Department of Electrical Engineering, Media and Computer Science, Ostbayerische Technische Hochschule Amberg-Weiden, addressline=Kaiser-Wilhelm-Ring 23, postcode=92224, city=Amberg, country=Germany Abstract Studies of machine-learning-based power-system protection increasingly report near-perfect scores, yet the meaning of those scores depends strongly on the evaluation setting. Protection task, physical scope, measurements, timing, targets, preprocessing, and validation often vary jointly and remain incompletely specified. This paper proposes a standardization-oriented framework that treats evaluation design as part of the scientific contribution. It defines seven required study dimensions: protection objective, physical system scope, observability and measurements, timing and decision windows, targets and valid samples, training and validation protocol, and evaluation outputs. The framework is instantiated in a bounded case study on the public PROTECT-90 electromagnetic-transient benchmark, comprising 9022 simulated episodes from one 90 kV double-line topology for onset-conditioned fault classification and fault localization. Under centralized sensing, simulation-metadata-aligned 20 ms windows, and episode-grouped validation, a multi-layer perceptron (MLP) achieved a five-fold mean macro-averaged F1 score of 0.991±0.0010.991± 0.001 for classification and a localization mean absolute error of 10.20±0.25%10.20± 0.25\% of line length, where the standard deviations describe variation across the episode-grouped folds. Extending the decision horizon to 50 ms preserved this task-dependent performance asymmetry, while reduced observability approximately doubled the MLP localization error but had little effect on classification. A synchronized two-ended conventional locator outperformed the learning locators under its richer clean information set, and measurement degradation showed that clean predictive performance did not determine robustness. The framework turns evaluation assumptions into explicit and reproducible evidence and provides a basis for more comparable, auditable research evaluation and future certification-oriented assessment of machine-learning protection functions. keywords Power system protection ,Machine learning ,Evaluation framework ,Fault classification ,Fault localization ,Electromagnetic transient simulation â credit: Conceptualization, Methodology, Investigation, Writing â original draftâ credit: Conceptualization, Writing - review & editingâ credit: Conceptualization, Writing - review & editingâ credit: Writing - review & editing, Supervisionâ credit: Supervision, Project administration, Funding acquisitionâ credit: Supervision, Project administration, Funding acquisitionâ credit: Supervision, Project administration, Funding acquisition, Writing â review and editingâ corresponding: Corresponding author. 1 Introduction Power system protection is a safety-critical function: protective relays must decide correctly within milliseconds, and in practice they are trusted only after standardized type-testing and certification against explicitly specified behavior. Two simultaneous shifts now strain this basis of trust â the changing physics of the grid, and a change in the methods proposed to protect it. The transition toward decentralized power systems, driven by growing renewable energy sources and distributed energy resources, is reshaping grid operation. Rising shares of inverter-based generation and the adoption of hybrid ac (ac)â dc (dc) architectures [53] expand the range of operating and fault scenarios encountered in practice [64]. Meshed topologies, multi-terminal configurations, dynamic redispatch, and varying grid-connected or islanded operating modes increase protection-relevant uncertainty, while inverter-based resources contribute fault currents that differ from those of synchronous machines and challenge established protection functions [58, 7, 14, 54]. Consequently, conventional protection schemes based on deterministic logic, fixed thresholds, and static system models are increasingly stressed under variable operating conditions [5]. This has motivated growing interest in data-driven and ml (ml)-based approaches that can exploit nonlinear structure in protection-relevant measurements. Yet, unlike the conventional functions they aim to augment or replace, these ml-based approaches enter the field with no established basis for testing, auditing, or certifying them â precisely the discipline that makes conventional protection trustworthy. In protection engineering, performance claims matter only insofar as they can be verified under defined operating, sensing, timing, and failure conditions. A protection function is not trusted because it performs well in one isolated experiment, but because its behavior can be inspected, repeated, and tested against explicit assumptions. In conventional practice, this trust is established through standardized type-testing and certification: protection equipment is evaluated against common requirements such as iec (iec) 60255 [20] using dedicated relay-test tooling and third-party conformity assessment before deployment. Safety-critical domains are now extending analogous assurance and certification regimes to machine learning â aviation learning-assurance guidance [9], cross-sector ai (ai) risk management [59], auditable ai-management-system standards [22, 21], and safety-case standards for autonomous products [62] â alongside a growing literature on assuring the machine-learning lifecycle [4]. No widely adopted protection-specific framework currently integrates the requirements needed to evaluate, report, and audit such ml-based systems. This requirement is not yet reflected in much of the ml-based protection literature. Public datasets, released code, and high reported scores are useful first steps, but they do not by themselves create deployment-relevant evidence. A model result becomes interpretable only when the protection task, information available at decision time, data-processing pipeline, and validation protocol are specified together. Otherwise, reported performance remains inseparable from hidden study assumptions, and individual papers cannot accumulate into reliable engineering knowledge. Despite this growing research activity, reported results in ml-based protection remain difficult to interpret and compare. Existing studies often differ simultaneously in physical system scope, sensing assumptions, temporal representation, target construction, and validation protocol. These choices are not merely implementation details; they define the effective inference problem by determining which information is available, which timing constraints apply, and what the model is asked to predict. Reported performance therefore reflects not only algorithmic design, but also differences in task formulation and information availability. When these factors vary jointly with model choice, empirical comparisons conflate task difficulty, data-generation choices, and learning effects. A reported accuracy of 99% in one study may correspond to a comparatively simple task on a radial feeder, whereas 90% in another may reflect a more demanding setting. Without a common evaluation structure, the community cannot distinguish algorithmic progress from differences in the underlying benchmark. This paper places evaluation design at the center of the methodological contribution rather than as a neutral experimental backdrop. The central question is not only whether ml-based protection can work, but under which physical, observational, and temporal assumptions reported success or failure should be interpreted. This question is especially important because most studies rely on simulated datasets whose physical fidelity, scenario coverage, and documentation vary substantially across works. Without a framework that fixes and reports the core study assumptions, these factors cannot be separated consistently across studies. This work therefore does not propose a new protection method, nor a ranking of models on one benchmark. Its contribution is a standardized evaluation and reporting framework that makes the effective inference problem explicit before performance is interpreted [37]. 1.1 Limitations of Current Evaluation Practice and Related Work A substantial body of literature has applied ml-based methods to protection tasks such as fd (fd), fli (fli), fc (fc), and fl (fl). For transmission-line protection, wavelet-derived features and artificial neural networks have been used for ultrafast fault detection [1], while artificial neural networks and convolutional neural networks have been compared for fault-type identification from voltage and current measurements [35]. Controlled comparative work has also evaluated multiple learning methods for fault detection and line identification under a shared experimental setting [46]. A later controlled comparison on the benchmark used in the present case study examined fault classification and localization under shared sensing, timing, and validation assumptions [43]. Combined classification and localization have been studied for three-terminal transmission circuits [36], while support-vector-machine models embedded in distribution relays have been used to classify faults, estimate their region, and support switching decisions on a simulated feeder [24]. Other studies have addressed joint fault detection, classification, and localization using optimized tree-based and neuro-fuzzy methods [40], as well as machine-learning-based electrical fault detection and localization in broader grid settings [65]. Application-specific protection schemes have further included a modwt (modwt)â xgboost (xgboost) pipeline for fault detection and classification across changing microgrid topologies and operating modes [48], and a hybrid protection scheme based on deep reinforcement learning [28]. Graph-neural-network-based protection has likewise been explored as a topology-aware learning approach [31]. Collectively, these studies have shown that ml methods can capture temporal patterns and nonlinear relationships in protection-relevant voltage and current signals. At the same time, they have been evaluated under substantially different assumptions regarding topology, fault space, measurement access, preprocessing, timing, and targets. A central limitation of current practice is its predominantly model-centric perspective. Learning algorithms are usually assessed inside study-specific setups. In each work, the physical system model, data-generation process, measurement configuration, preprocessing pipeline, window construction, target definitions, and validation protocol are specified independently. These choices define the information available to the model, the temporal context available at decision time, and therefore the protection problem actually being solved. As a result, studies that nominally address the same task may in fact evaluate different inference problems. Reported performance is therefore shaped not only by model capability, but also by implicit assumptions about observability, timing, data realism, and system complexity. Recent survey and scoping studies have confirmed that this is not an isolated issue, but a structural limitation of the field. Reviews of ml in power system protection have documented substantial heterogeneity in simulation setups, preprocessing pipelines, feature representations, and evaluation metrics, which limits reproducibility and meaningful cross-study comparison [63, 52, 32, 45]. This echoes the broader reproducibility crisis in machine-learning-based science, where incomplete documentation of data, methods, and evaluation has repeatedly prevented reported results from being reproduced [13]. In a prior scoping review by the authors covering 119 studies, only 16.0% used real-world data, 82.5% provided no data access, and only 1.7% publicly released code or models; even basic metadata such as sampling frequency were omitted in 52.1% of studies [45]. These omissions are not secondary reporting defects: they determine whether a result can be reproduced, compared, or interpreted under protection-relevant constraints. Practical challenges such as noisy or incomplete measurements, class imbalance, shifts between training and deployment conditions, and limited interpretability further complicate the use of ml in safety-critical settings [52, 44]. These findings indicate that the central limitation is not only a lack of stronger models, but a lack of shared evaluation discipline for turning model results into reproducible protection evidence. This ambiguity is amplified by the widespread reliance on simulated data [45]. Domain-specific studies targeting hvdc (hvdc) systems, wind integration, or hybrid ac/dc networks provide valuable insights for particular applications [8, 61, 69], but they also introduce additional variation in system modeling, measurement conditions, and operating assumptions. Application-specific studies have further illustrated this dependence on the physical setting, including decision-tree-based fault detection on a tcsc (tcsc)-compensated line during power swing [34] and an enhanced relaying scheme for a compensated line connected to a dfig (dfig)-based wind farm [39]. When simulation fidelity, scenario coverage, and documentation of the data-generation process are only partly specified, it remains unclear whether reported success or failure is driven by model design, information availability, task formulation, or the realism of the underlying data [45]. The difficulty is partly interdisciplinary. In protection engineering, the validity of a result depends on physical scope, measurement access, synchronization, timing, operating conditions, and failure behavior. In machine learning, validity depends on data construction, preprocessing, split design, leakage prevention, model selection, and evaluation metrics [26, 25]. ml-based protection studies require both perspectives simultaneously, but current reporting practices often leave one of them implicit. The resulting gap is procedural: the field needs an evaluation structure that connects protection assumptions with ml validity requirements. General-purpose ai-assurance standards establish the governance scaffolding but do not supply these protection-specific dimensions; conversely, protection studies rarely adopt ml validity safeguards. The framework proposed here is intended to occupy exactly this intersection. Table 1 provides an illustrative structured mapping of how representative protection, machine-learning, and ai guidance treats the seven dimensions formalized in the proposed framework. Table 1: Comparison of representative guidance against the seven framework dimensions. Source family Obj. Scope Obs. Time Targets Validation Outputs Dataset documentation [11] â â â â â â â ML reporting and validation guidance [25, 26, 55] â â â â â â â Protection reviews and scoping studies [63, 52, 45] â â â â â â â Protection-testing standards [20, 18] â â â â â â â General ai risk and assurance guidance [59] â â â â â â â Proposed framework â â â â â â â Note: This table is an illustrative structured mapping rather than a quantitative score. âdenotes an explicit prescriptive requirement accompanied by an operational procedure, â denotes acknowledgement without full operationalization, and â denotes no identified treatment in the cited sources. Prior-work rows aggregate closely related sources by guidance family. The comparison identifies an integration gap rather than an absence of prior guidance: existing source families address complementary subsets of the evaluation problem, whereas the proposed framework operationalizes all seven dimensions jointly. In particular, it treats sample-validity rules, including onset conditioning, as part of the evaluated protection problem rather than as an implicit preprocessing choice; where prior guidance addresses onset at all â as in the fault-inception-angle definition of iec 60255-121 â it does so as a controlled test parameter, not a sample-validity rule. Unlike survey and scoping contributions, including prior work by the authors [45], this paper does not primarily catalog the literature or summarize reported methods. Its contribution is prescriptive rather than descriptive. It defines a standardized evaluation structure, a reporting logic, and a bounded reference instantiation that make study assumptions explicit before model performance is compared. Overall, the literature has provided substantial evidence of the technical feasibility of ml for protection tasks, but it does not yet provide a consistent basis for interpreting why results differ across studies. Addressing this gap requires evaluation frameworks that make task formulation, observability, timing, data provenance, and validation design explicit before model comparison. 1.2 Objective and Contributions The objective of this paper is to develop a standardized evaluation framework for ml-based power system protection that makes reported results interpretable as conditional and auditable evidence rather than isolated model scores. The framework therefore requires explicit specification of the physical scope, observability, timing constraints, target construction, data provenance, validation design, and reporting outputs before model performance is interpreted or compared. In doing so, it makes the effective inference problem an explicit part of the scientific contribution. The contributions of this paper are threefold. First, it proposes a standardized, protection-specific evaluation framework that structures the complete path from protection objective and information availability to validation and reporting. The framework defines what must be specified, reported, and examined for an ml-based protection result to be reproducible, comparable, and open to independent audit. It thereby provides a first methodological step toward certification-oriented assessment without claiming to constitute a formal certification procedure. Second, it formalizes an interpretation logic in which reported performance is treated as conditional evidence tied to task formulation, physical scope, observability, timing, target construction, and validation design. This separates apparent model capability from differences in the underlying inference problem and prevents aggregate scores from being interpreted independently of the assumptions under which they were obtained. Third, it instantiates the framework in a bounded and reproducible case study on fc and fl. Under shared sensing, timing, sample-validity, and validation assumptions, the case study demonstrates how the framework exposes task-dependent behavior, observability effects, diagnostic patterns, conventional-reference comparisons, robustness differences, and runtime trade-offs beyond aggregate performance measures. The case study serves as a worked example of the framework; it does not claim benchmark completeness or practical superiority over conventional protection. The proposed framework complements established protection-testing practice [20] and emerging ai-assurance frameworks [59, 22, 21] by supplying protection-specific evaluation dimensions, including observability and synchronization, decision horizons, and onset-conditioned sample validity. 1.3 Paper Organization The remainder of this paper is organized as follows. Section 2 presents the proposed standardized evaluation framework and its seven-step protocol. Section 3 shows how the framework is instantiated in a bounded case study on a public emt (emt) benchmark for fc and fl. Section 4 reports the resulting findings and illustrates what the framework reveals beyond aggregate performance measures. Section 5 discusses generalizability, reuse, limitations of the present instantiation, and directions for future work. Section 6 concludes the paper. 2 Proposed Standardized Evaluation Framework The central contribution of this paper is a standardized evaluation framework for ml-based power system protection. The framework does not prescribe a particular model, dataset, or grid architecture. Rather, it specifies the minimum structure a study must define before its results can be interpreted, reproduced, or compared. The motivation is straightforward: reported performance depends not only on model choice, but also on task definition, physical scope, observability, timing, target construction, data provenance, and validation design. When these elements vary across studies, similar performance numbers may correspond to different underlying inference problems and therefore lack direct comparability. The proposed framework addresses this issue by making the evaluation design explicit, inspectable, and reusable. Accordingly, the framework structures a study into seven steps: protection objective, physical system scope, observability and measurements, timing and decision windows, targets and valid samples, training and validation protocol, and evaluation outputs and diagnostics. Together, these steps define a standardized reporting structure that makes assumptions explicit and evaluation settings inspectable. The structure is informed by established protection-equipment testing and performance-documentation practice, general ai test, evaluation, verification, and validation guidance, and machine-learning reproducibility and temporal-benchmark literature [17, 18, 59, 51, 26, 68, 27]. Its contribution is not that every dimension is individually unprecedented, but that these previously separate requirements are integrated into one protection-specific and operational evaluation workflow. DEFINE THE PROTECTION OBJECTIVE12345 Physical system scope Observability & measurements Timing & decision windows Targets & valid samples 6TRAINING AND VALIDATION PROTOCOLTest PhaseData AcquisitionPreprocessingFeature ExtractionClassificationClass PredictionData AcquisitionPreprocessingFeature ExtractionModel LearningTraining Phase7EVALUATION EVIDENCE Task performance Diagnostics Robustness Deployment STANDARDIZED REPORTING PACKAGE Figure 1: Seven-step framework for evaluating machine-learning-based power system protection studies. Step 1 defines the protection objective, Steps 2â5 specify the effective inference setting, Step 6 places the classical training and test pipeline within an explicit validation protocol, and Step 7 structures the resulting evaluation evidence. The final reporting package links conclusions to the complete study definition. The pattern-recognition pipeline is adapted from Niemann [41]. 2.1 Design Goals and Guiding Principles The proposed framework is guided by a small set of principles that define what a rigorous ml-based protection study must achieve. First, it must support comparability: studies should specify the task, sensing assumptions, timing constraints, targets, and outputs in a form that permits meaningful comparison. Second, it must support reproducibility: the full path from raw signals to reported results should be reconstructable, including preprocessing, sample construction, splitting, and reporting. Third, it must remain physically grounded: reported performance should be interpretable in the context of topology, operating conditions, fault space, and information availability. Fourth, it must remain protection-relevant: evaluation should reflect operational constraints such as short decision times, realistic sensing assumptions, and deployment-oriented decision settings. In addition, the framework requires leakage-aware validation, since related samples can otherwise inflate reported performance. It requires diagnostic interpretability, so that evaluation extends beyond aggregate metrics and reveals failure modes across classes, locations, sensing regimes, or operating conditions. Finally, it requires deployment awareness, so that predictive quality is assessed together with runtime, sensing requirements, communication assumptions, and robustness under degraded conditions. Together, these principles define the framework as a reusable evaluation standard for ml-based power system protection rather than a benchmark recipe. 2.2 Framework Overview Figure 1 summarizes the framework as a seven-step workflow. The sequence runs from problem definition to result interpretation. It first fixes what is to be learned, under which physical and observational conditions, and within which timing constraints. It then specifies how targets are constructed, how training and validation are performed, and which outputs are required to support interpretable conclusions. The seven steps do not prescribe a particular dataset, topology, model family, or signal domain. Instead, they define the dimensions that must be made explicit in any rigorous protection study. This keeps the framework applicable across different tasks and grid settings while preserving comparability at the level of study design. The framework output is not only a trained model or a set of performance numbers. It is a standardized reporting package that documents the evaluation setting together with predictive, diagnostic, and deployment-relevant outputs. In this way, the framework supports both reproducible experimentation and interpretation of reported ml results relative to explicit assumptions. The following subsections define each framework step in turn. Problem Specification. Steps 1â5 define the task, system, inputs, timing, and prediction target. 2.3 Step 1: Protection Objective The first step is to define the protection objective precisely. A study must state which protection function is being addressed, for example, fd, fc, fli, or fl, and must specify the corresponding learning formulation, such as binary classification, multi-class classification, or regression. Relevance. The protection objective determines what the model is expected to predict, which errors are operationally relevant, and which evaluation outputs are meaningful. The objective must therefore be stated in operational rather than generic ml terms. This is consistent with guidance requiring an ai systemâs intended purpose, context of use, expected impacts, and evaluation criteria to be documented before assessment [59]. Protection standards similarly specify functions through operational characteristics such as starting, directionality, and time-delay behavior [18]. For example, âfault analysisâ is too broad to support comparison unless it is decomposed into explicit tasks such as fault detection, fault type classification, or distance estimation. Fixing the protection objective at the outset prevents ambiguity and establishes the basis for later choices such as target construction, metric selection, and diagnostic analysis. 2.4 Step 2: Physical System Scope Step 2 defines the physical system in which the protection task is posed. A study must specify the network type, topology class, voltage level, operating range, fault and disturbance space, scenario-generation method, coverage of critical operating conditions, data source, and known representativeness limitations. The data source may consist of simulation, hardware-in-the-loop testing, digital-twin-based generation, or field recordings. Together, these choices define both the physical environment from which the data arise and the limits of the protection problem represented by the study. Relevance. Task difficulty depends strongly on physical scope. A result obtained on a radial distribution feeder is not directly comparable to a result obtained on a meshed transmission system, even if both are reported for the same nominal task. The same applies to differences in fault and disturbance coverage, scenario-generation assumptions, critical operating conditions, and data origin. Making the physical system scope explicit ensures that reported performance is interpreted relative to the grid conditions under which it was obtained, rather than attributed to the model alone. For simulation-derived studies, established verification and validation practice further distinguishes conceptual-model validity, model verification, operational validity, and data validity [57]. 2.5 Step 3: Observability and Measurements The third step specifies what information from the physical system is available to the model. A study must specify which measurements are provided, where they are taken, in which signal domain they are represented, how they are synchronized, and whether access is centralized, local, or distributed. This includes, for example, whether the model operates on emt waveforms, sampled values, rms (rms) quantities, phasors, or pmu (pmu)-based features, as well as which voltage, current, frequency, or derived channels are used. These representations are not interchangeable: sampled-value standards define the communication of waveform samples, whereas synchrophasor standards impose explicit time-tagging, synchronization, and static- and dynamic-performance requirements [19, 17]. A study must also state whether the model receives auxiliary non-waveform information beyond the primary measurements, such as topology encodings, line parameters, bus or line identifiers, operating-point variables, load information, switching states, or other metadata. If channels or auxiliary inputs are missing, delayed, noisy, or otherwise degraded, these conditions must also be reported. Relevance. Observability is a primary determinant of task difficulty. Two studies may evaluate the same physical system but still solve different inference problems if one model receives only local measurements while another also receives synchronized multi-location signals, topology information, or operating metadata. Making the available information explicit separates model capability from information availability and turns observability into a comparable evaluation axis rather than a hidden implementation detail. In fault prediction, feature-selection studies provide a complementary way to identify which inputs carry task-relevant information [29]. 2.6 Step 4: Timing and Decision Windows The fourth step defines the decision-time setting under which the protection task is evaluated. A study must specify the temporal reference used for sample construction, the decision horizon, the window length, the step size or stride between windows, and any additional real-time constraints that limit what information is available at inference time. This includes the event reference used for alignment, the amount of pre-event and post-event context, the window extraction rule, and whether overlapping windows are used. Together, these choices determine how much pre-fault, fault-onset, and post-fault information the model can use and therefore shape the protection problem being solved. Relevance. Protection tasks are time-critical by definition. Protection standards accordingly treat operating-time and time-delay characteristics as explicit quantities to be evaluated under defined test conditions [18]. A result obtained from short windows near fault inception is not directly comparable to one based on longer windows or broader post-event context. Changing the stride also changes the number, overlap, and temporal diversity of the available samples, which can affect both dataset composition and evaluation outcomes. Defining timing and decision windows explicitly ties reported performance to a clear protection-time budget and prevents it from being interpreted independently of the available decision time. 2.7 Step 5: Targets and Sample Validity The fifth step fixes what the model must predict and which samples are valid for that prediction. A study must specify how labels or regression targets are constructed, which inclusion and exclusion rules are applied, and how boundary cases are handled. This includes, for example, whether samples are labeled by event presence, fault type, faulted line, or fault position, and whether the dataset retains all extracted windows, only windows whose time span overlaps with a fault, or only windows in which the fault begins inside the window. If samples are filtered, merged, discarded, or reassigned, these rules must be stated explicitly. Relevance. The task name alone does not fully define the prediction problem. Two studies may both report fc or fl results while using different class definitions, regression targets, or sample-validity rules. The same applies to boundary handling, such as windows near fault inception, multiple events, or ambiguous cases. Defining targets and sample validity explicitly ensures that reported performance is tied to a fully specified labeling and filtering scheme rather than to labels that are only partly described. More generally, time-series benchmark research has shown that flawed anomaly placement, ambiguous event boundaries, and permissive temporal scoring can materially distort apparent progress and method rankings [68, 27]. Evaluation Protocol. Steps 6â7 define how models are compared and what evidence a study must report. 2.8 Step 6: Training and Validation Protocol The sixth step specifies how models are trained, validated, and compared. A study must specify the split design, the unit of independence used for splitting, the treatment of related samples, the preprocessing pipeline, and the conditions under which different models are evaluated. This includes, for example, whether splitting is performed by event, episode, line, feeder, or recording, whether folds are grouped to prevent leakage, and whether preprocessing is fitted on training data only [55, 26]. Relevance. Protection datasets often contain strong dependencies between samples. Multiple windows may originate from the same event, share the same operating point, or overlap in time. If such samples are split across training and test sets, reported performance may reflect leakage rather than generalization. The protocol must therefore state which samples are considered dependent and how this dependency is handled during validation. It must also define the boundary between training and evaluation. Any normalization, feature extraction, target transformation, dimensionality reduction, or parameter selection must be fitted on training data only and then applied to held-out data without re-estimation. If hyperparameters are tuned, the tuning procedure must be reported explicitly. Competing models must be compared under the same data partitions, target definitions, preprocessing assumptions, and metric definitions so that reported differences are attributable to model behavior rather than to inconsistent evaluation conditions. Reproducibility guidance further emphasizes transparent reporting of data, code, experimental conditions, and model-selection procedures, while benchmark studies show that data sampling, initialization, and hyperparameter choices can materially affect comparative conclusions [51, 6]. 2.9 Step 7: Evaluation Outputs and Diagnostics The seventh step defines what outputs a study must report for its results to support clear conclusions. A study must report more than a single headline number. At minimum, the evaluation output should include task-appropriate aggregate metrics, structured diagnostics, and implementation-relevant indicators. Reported uncertainty must identify its source explicitly, for example, variation across held-out groups, model initializations, data samples, or perturbation realizations, because these quantities are not interchangeable [59, 6]. Where repeated evaluations are used, their unit and aggregation procedure must therefore be reported. Depending on the task, this may include classification scores, regression errors, class-resolved confusion patterns, location-resolved error profiles, observability-conditioned comparisons, robustness analyses, and runtime measurements. Relevance. Aggregate performance alone does not show why a model succeeds or fails. Two models may achieve similar overall scores while exhibiting different failure modes across fault classes, locations, sensing regimes, or operating conditions. A meaningful evaluation must therefore expose not only average performance, but also the structure and stability of that performance. The reported outputs must also match the protection objective. Classification tasks require metrics that reflect class-wise behavior and not only overall accuracy. Localization or other regression tasks require error measures with direct physical interpretation and, where relevant, stratified analyses across faulted elements or operating regions. If robustness is assessed, the study must state which factors are varied. If runtime is reported, the measurement setup must be stated clearly enough to support comparison. Recent streaming evaluations further show that short decision windows or early model decisions do not necessarily imply correspondingly short end-to-end protection latency, because the complete signal-processing and inference pipeline contributes additional delay [3]. Defining evaluation outputs in this way ensures that a study produces evidence rather than only scores. 2.10 Standardized Reporting Package A study that follows the proposed framework should report its evaluation setting in a compact and standardized form, following the broader principle that datasets and evaluation artifacts should be documented for reuse and inspection [25, 11]. Just as datasheets document datasets and model cards document trained models for reuse and scrutiny [38], the proposed reporting package documents the evaluation setting of a protection study, so that a result can be inspected, reproduced, and reused under its stated conditions. This reporting package makes the core assumptions of the study explicit and allows readers to judge what is comparable, reproducible, and interpretable. Table 2 summarizes the minimum information that should be reported. Table 2: Minimum reporting package for ml-based power system protection studies under the proposed framework. Framework step Minimum reported information 1. Protection objective Protection task, operational interpretation, learning formulation, and predicted output 2. Physical system scope Network type, topology class, voltage level, operating range, fault and disturbance space, scenario-generation method, critical-condition coverage, data source, and known representativeness limitations 3. Observability and measurements Measurement locations, signal domain, channel set, synchronization assumptions, access regime, and auxiliary inputs 4. Timing and decision windows Event reference, pre-/post-event context, decision horizon, window length, stride, and overlap rule 5. Targets and sample validity Label or regression-target construction, inclusion/exclusion rules, boundary handling, and retained sample set 6. Training and validation protocol Split design, unit of independence, leakage-prevention strategy, preprocessing boundary, and model-comparison conditions 7. Evaluation outputs and diagnostics Aggregate metrics, structured diagnostics, robustness analyses, observability-conditioned results, and runtime setup Reporting artifacts Data/code availability, configuration details, and any restrictions affecting reproducibility or comparison 3 Case Study: Instantiation on a Public EMT Benchmark This section instantiates the evaluation framework introduced in Section 2 on the public PROTECT-90 emt benchmark [42, 30]. The case study addresses fc and fl under shared assumptions on physical system scope, observability, timing, target construction, and validation, as summarized in Table 3. The evaluated methods span conventional protection, classical ml, task-specific deep learning, and a pre-trained time-series foundation model, as specified in Sections 3.3 and 3.4. Controlled analyses then vary timing, observability, measurement fidelity, and fault-resistance distribution while keeping the remaining evaluation assumptions fixed. Additional stride and hyperparameter checks assess the stability of the resulting interpretation. Together, these experiments demonstrate how the framework separates model effects from changes in information availability, decision time, data realism, and evaluation configuration. Table 3: Case-study instantiation of the proposed standardized evaluation framework. Framework step Reference instantiation 1. Protection objective fc and fl under shared sensing, timing, and validation assumptions 2. Physical system scope PROTECT-90; 90 kV double-line topology; 9022 emt episodes; randomized fault and operating conditions 3. Observability Synchronized three-phase voltage/current measurements from 8 relays; centralized full-observability reference setting 4. Timing ± 80 ms around fault inception; 20 ms sliding windows; fixed 5 ms stride 5. Targets Onset-conditioned 11-class fc with one non-onset class and ten fault-type classes; fl as normalized fault position; fault-onset-containing windows for fl 6. Validation 5-fold grouped cross-validation by simulation episode; shared preprocessing and split design across tasks and models 7. Outputs Aggregate metrics, structured diagnostics, conventional baselines, runtime assessment, and controlled analyses of timing, observability, measurement fidelity, fault-resistance distribution shift, stride, and hyperparameters 3.1 Physical System and Observability The case study is based on the public PROTECT-90 dataset [42]. The benchmark is derived from a 90 kV transmission-system âDouble Lineâ topology with multiple relay locations and diverse short-circuit events. All episodes are generated in DIgSILENT PowerFactory using the emt simulation module. This simulation-based design reflects the current scarcity of publicly accessible real-world protection datasets with sufficiently detailed waveform recordings and labels for reproducible ml evaluation [45, 67, 10, 12]. These public resources provide valuable event signatures, real-world oscillograms, or synthetic transmission-grid data, but differ in signal representation, measurement layout, available metadata, and supported protection targets and therefore do not constitute directly interchangeable benchmarks for the present fc/fl protocol. Scenario diversity is introduced through domain randomization over fault and operating parameters, including variations in fault resistance, inception time, line characteristics, loading conditions, and external grid properties within predefined physically plausible ranges [42, 66]. Figure 2 summarizes the corresponding data-generation process. Such simulation-based benchmarks require explicit documentation of the data-generation assumptions because synthetic power-system datasets can otherwise be difficult to interpret or validate across studies [33]. The resulting benchmark contains 9022 simulation episodes. Each episode spans T=1âsT=1\,s and is sampled at fs=6400âHzf_s=6400\,Hz, yielding 6400 time steps per episode. The simulated events cover the main short-circuit categories considered in this study, namely slg (slg), l (l), llg (llg), and l (l). The measurement model is based on primary three-phase voltage and current waveforms extracted at protection-relevant locations. In total, each episode provides signals from NPR=8N_PR=8 locations corresponding to the relay positions in the studied topology. At each location, three-phase voltage and three-phase current signals are available, resulting in six channels per location. Under the centralized reference configuration, all available measurements are jointly provided to the learning model, yielding the episode-level representation XââNPRĂ6ĂTs, X ^N_PRĂ 6Ă T_s,\@add@centering (1) where Ts=6400T_s=6400 denotes the number of discrete time steps per episode. No auxiliary non-waveform metadata are used in the reference configuration. All measurements are treated as fully synchronized, noise-free, and continuously available, with no communication delay, jitter, or packet loss. This centralized full-observability configuration serves as an upper-bound reference sensing setting for the case study. Later analyses vary the available measurements by restricting the input either to a single relay location or to both terminals of the same line, while preserving the remaining study design. These reduced-observability experiments therefore act as controlled sensitivity analyses of the observability dimension rather than as separate case-study definitions. The clean reference configuration excludes non-fault disturbances, ct (ct)/ vt (vt) nonidealities, missing channels, synchronization errors, and communication effects, all of which are relevant for protection-equipment assessment and practical relay evaluation [20]. Additive noise, a simplified current-transformer saturation proxy, and synchronization jitter are subsequently evaluated as controlled measurement-fidelity axes in Section 4.5. Sensor-availability degradation is addressed in the companion study [47], whereas communication latency and broader non-fault disturbances remain outside the present scope. Grid Topology Double-line 90 kV system Scenario Sampling Fault and operating conditions EMT Simulation Time-domain fault simulation Data Extraction Waveforms, labels, timing metadata PROTECT-90 Dataset Synchronized waveform episodes, labels, and timing metadata Figure 2: Case-study data generation flow for the PROTECT-90 dataset [42]. A fixed grid topology is combined with randomized fault and operating scenarios, simulated in the emt domain, and processed into synchronized waveform episodes with labels and metadata. 3.2 Timing, Windowing, and Targets Sample construction is defined relative to the fault inception time tft_f. For each episode, the raw tensor in Eq. 1 is restricted to a ± 80 ms interval around tft_f, retaining a short pre-fault baseline together with the early post-fault transient. This bounded crop is used as a reference timing setting that preserves protection-relevant onset information without introducing substantially longer post-event context that would change the effective decision problem. Importantly, this cropping uses the ground-truth tft_f from simulation metadata, which is unavailable pre-detection in a deployed relay. The reference configuration therefore isolates fc and fl from the upstream detection problem; a deployment-oriented instantiation would require either a coupled detection stage or a trigger derived from observable signal features alone. From this cropped segment, overlapping sliding windows are extracted. Let L denote the window length in samples and S the step size between consecutive windows. The i-th window is xi=X[:,:,iS:iS+L],x_i=X[:,:,iS:iS+L], (2) with start time ti=iâS/fst_i=iS/f_s. The reference timing configuration uses a fixed step size of S=32S=32 samples (5âms5\,ms) and a fixed window length of L=128L=128 samples, corresponding to a decision horizon of 20âms20\,ms. At a nominal system frequency of 50 Hz, this window covers approximately one fundamental cycle and therefore provides a physically interpretable reference horizon for protection-oriented evaluation. The 5 ms stride provides overlapping temporal views of the same transient while retaining bounded sample density; it therefore affects both sample overlap and the effective composition of the resulting dataset. For this reason, stride is treated as part of the evaluation setting rather than as a neutral implementation detail. Additional window lengths are examined later as a controlled sensitivity analysis of the timing dimension. Both tasks are defined on this shared windowed representation. For fc, each window receives a discrete label yFCâc0,c1,âŠ,c10,y_FCâ\c_0,c_1,âŠ,c_10\, (3) where c1c_1âc10c_10 correspond to the short-circuit classes AG, BG, CG, AB, BC, CA, ABG, BCG, CAG, and ABC. A window receives one of these fault-type labels only when the fault inception lies within the window; all other windows, including pre-fault windows and windows in which the fault is already active, receive c0c_0. Thus, c0c_0 denotes the absence of a fault inception within the current window rather than a physically fault-free state. For fl, the target is the normalized fault position along the affected line segment, yFL=dfaultâlineâ 100,y_FL= d_fault _line· 100, (4) where dfaultd_fault is the distance from a fixed reference terminal and âline _line is the line length. Reporting the target in percent of line length makes localization results comparable across line segments. For fl, only windows containing the fault onset are retained. A window xix_i is used only if ti+Ï”<tf<ti+LfsâÏ”,t_i+Δ<t_f<t_i+ Lf_s-Δ, (5) where Ï”=2/fsΔ=2/f_s is a two-sample margin. This excludes boundary cases and ensures that each retained regression sample contains both pre-fault information and the initial post-fault transient. This onset-conditioned definition is an intentional task-construction choice that focuses localization on early transient evidence rather than on broader post-fault context. It should therefore be interpreted as one bounded formulation of fl, not as the only valid definition of the localization problem. 3.3 Validation Protocol and Model Families The reference, timing, observability, conventional-baseline, measurement-fidelity, stride, and hyperparameter analyses use five-fold cross-validation grouped by simulation episode. The fault-resistance distribution-shift analysis instead uses the disjoint episode-level training and test partitions defined in Appendix C. All windows originating from the same episode are assigned to the same fold. This grouping is necessary because sliding windows from the same episode are strongly dependent through shared waveform structure, fault metadata, and operating conditions. Grouped splitting therefore prevents leakage that would otherwise inflate reported performance [55, 26]. Within each investigation, all compared models use the same deterministic episode-grouped split, task definition, timing configuration, observability setting, and metric definitions. The split is also reused across investigations wherever the underlying sample construction is unchanged. Unless stated otherwise, reported values are the mean and standard deviation across the five held-out fold scores. These fold-wise standard deviations describe variation across episode groups and do not quantify variation across model initializations, confidence intervals, or standard errors. For the classical ml estimators, each extracted window is converted into a fixed-length feature vector by reshaping the windowed tensor to xiââLĂFâŠx~iââLâ F,x_i ^LĂ F\; \; x_i ^L· F, (6) where L denotes the number of time steps in the window and F=NPRâ 6F=N_PR· 6 the number of input features per time step. In implementation terms, this corresponds to reshape(n_samples, n_timesteps * n_features), that is, a time-major flattening in which all features at one time step are kept together before proceeding to the next. For each fold, preprocessing consists of standardizing features to zero mean and unit variance, with the transformation fitted on the training partition only and then applied unchanged to the corresponding held-out fold. Hyperparameters are fixed across folds and selected without access to test-fold results; no fold-specific retuning is performed. The cnn1d (cnn1d) receives the same fold-wise standardized waveform windows as the classical models, but the standardized arrays are reshaped to channels-by-time rather than flattened for prediction. For each outer fold, the standardizer is fitted on the complete outer training partition only and then applied unchanged to the held-out fold. The fixed, untuned architecture comprises three one-dimensional convolution blocks with 64, 128, and 128 output channels; each block uses a kernel width of 5, batch normalization, rectified-linear activation, and max pooling by a factor of 2. Global average pooling and a linear output layer produce the task prediction. Models are trained with Adam at a learning rate of 10â310^-3 and batch size 256 for at most 50 epochs. Early stopping with patience 8 uses a 10% episode-grouped validation subset of the outer training fold. Cross-entropy loss is used for fc and mean-squared-error loss for fl. MOMENT-1-large is used as a frozen pre-trained time-series encoder, with its representations read out by either a linear probe or a two-layer mlp (mlp) prediction head with approximately 25.4 million trainable parameters. The input-length adaptation, channel-wise encoding and embedding aggregation, linear-probe solver and regularization, mlp-head dimensions and activation, optimizer, learning rate, batch size, epoch limit, task-specific loss, validation split, and stopping criterion are fixed in the released experiment configuration and reused unchanged across folds. For MOMENT, the fold-local protocol is applied end to end: within every outer fold, the raw-window standardizer is fitted on the outer-training windows only, the training and held-out embeddings are regenerated independently using that foldâs standardizer and the frozen encoder, and the linear probe and mlp head are fitted without any access to the held-out fold. These models use the same episode-grouped outer folds, task definitions, decision horizons, and task-specific metrics as the classical learning models; the cnn1d results are reported in Section 4.10 and the MOMENT results in Appendix E. The learned-model evaluation uses three complementary model panels. The broad classical panel comprises ridge (ridge), knn (knn), gb (gb), and mlp estimators, together with their corresponding regression variants for fl, and is used for the reference and timing analyses. The focused sensitivity panel comprises the neural mlp and histogram-based gb, which provide contrasting nonlinear estimators for the observability, measurement-fidelity, fault-resistance-shift, runtime, stride, and hyperparameter analyses. The cross-family panel comprises the mlp, a task-specific cnn1d, and MOMENT-1-large, with the mlp providing continuity with the preceding analyses. The classical models are implemented in scikit-learn 1.8.0 [49]. Their fixed configurations are summarized in Table G.1, while the separate single-parameter sensitivity analyses are reported in Appendix G. Conventional protection methods are evaluated separately in Section 3.4. 3.4 Conventional Protection Baselines Conventional protection methods are evaluated under the same episode-grouped five-fold splits and task-specific metrics as the learning models. The per-window localization comparison uses the onset-conditioned windows defined in Section 3.2. Because the phasor-based locators require settled post-onset quantities, separate per-episode settled estimates are additionally reported and treated as a distinct sample-validity setting rather than as validity-matched learning-model results. All methods use phasors obtained from a full-cycle discrete Fourier transform at 50 Hz, following standard digital-relaying and distance-protection practice [50, 70]. No additional signal pre-filtering is applied before phasor estimation, which should be considered when interpreting the conventional results relative to production relay implementations. At fs=6400f_s=6400 Hz, a 20 ms window contains exactly 128 samples and therefore one nominal cycle. Conventional baselines are consequently evaluated at the 20 ms and 50 ms horizons; the 10 ms horizon is marked phasor-invalid. For fc, a rule-based symmetrical-component phase selector determines ground involvement from the residual-current ratio |3âI0|/|I1||3I_0|/|I_1|, where I0I_0 and I1I_1 denote the zero- and positive-sequence currents, and identifies faulted phases from their current increase relative to a pre-fault reference, following established residual- and phase-current selection principles [5, 50]. The resulting phase set and ground flag are mapped to the same eleven classes used for the learning models. The phase-pickup threshold Ïp _p and ground-involvement threshold Ïg _g are selected by grid search on the training folds using macro- f1 (f1) and are frozen for the held-out fold, thereby placing conventional parameter selection within the Step-6 training boundary. For fl, an uncompensated single-ended reactance-based locator, grounded in established one-terminal fault-location formulations [60, 56], and a synchronized two-ended positive-sequence locator, following established two-terminal fault-location principles [23, 15], represent single-relay and relay-pair observability, respectively. Both receive the ground-truth faulted line and fault type for loop selection, matching the isolated fl formulation used for the learning models. They additionally use the episode-specific line length and sequence impedances from the benchmark metadata and therefore operate under best-case parameter observability. By contrast, the centralized learning models use waveform inputs only. Estimated distance is expressed as a percentage of line length from the lower-index bus. Results are reported per window and, for the conventional locators, per episode using the settled post-onset estimate. 3.5 Measurement-Fidelity Non-Idealities To address the realism of the reference configuration, three measurement-fidelity non-idealities are introduced as controlled, one-axis-at-a-time sensitivity analyses, mirroring the existing timing and observability sweeps. Additive white Gaussian noise is applied per channel and per window at a target signal-to-noise ratio; a proxy for current-transformer saturation magnitude-clips the current channels at a fraction of their clean peak (a deliberately simple, symmetric proxy; a faithful flux-driven current-transformer model [16] is designated future work); and synchronization jitter shifts each relayâs channels by an independent integer sample offset. Sensor-availability degradation (missing channels, downsampling, communication dropout) is not reproduced here, as it is the subject of a separate published study on the same data [47]; the present analysis adds the complementary measurement-fidelity axes that study does not cover. Perturbations are injected on the raw waveform windows before standardization. The primary protocol is train-clean and test-perturbed: the clean fold models are trained once and evaluated on perturbed held-out folds, with the standardizer kept fit on clean training statistics, mirroring a system calibrated under nominal conditions that then meets degraded inputs. Each stochastic level is repeated over five realizations seeded deterministically from a single global seed. For each fold and perturbation level, the metric is first averaged across the five realizations; reported values are then computed as the mean and standard deviation across the five fold-level means. The clean level of every axis is the identity operator and therefore reproduces the corresponding reference means. The clean rows report the corresponding five-fold reference statistics; perturbed levels are aggregated over folds and realizations as stated in the table captions. 3.6 Reproducibility Except for the explicitly multi-seed fault-resistance-shift experiment, all experiments use a global pseudorandom seed of 42, with deterministic derivation of perturbation seeds for each combination of fold, evaluation axis, level, and realization. The episode-grouped five-fold split is deterministic and reused across tasks, models, decision horizons, and observability settings. Learning-model hyperparameters are fixed a priori and are not selected using held-out benchmark results. The ablations reported in Appendix G are diagnostic sensitivity analyses rather than part of the model-selection procedure. For the conventional phase selector, thresholds are selected independently within each training fold and then frozen for evaluation on the corresponding held-out fold. The classical-model environment is pinned to scikit-learn 1.8.0, and the released configuration records all explicitly set estimator parameters. The cnn1d runs used PyTorch 2.8.0 with CUDA 12.8; their logs record the software version, pseudorandom seed, fold-wise stopping epoch, and GPU model. NumPy and PyTorch pseudorandom seeds are fixed to 42, although bitwise-identical execution across different CUDA devices is not claimed. The released code also records the configuration and hardware used for runtime measurements. The fc reference, timing, observability, runtime, and hyperparameter-ablation experiments, including the 10 ms configurations, are regenerated under this pinned environment. The knn and ridge baselines remain unchanged, while the regenerated mlp and histogram-based gb results differ from the earlier reported values by at most 0.0080.008 in macro-f1. All regenerated estimators use the fixed pseudorandom seed recorded in the released configuration. The released code and configurations support regeneration of the reported experiments subject to the documented data and dependency-access requirements. 3.7 Reported Outputs Consistent with the reporting logic defined in Section 2, the case study reports aggregate predictive performance, structured diagnostics, conventional-method context, runtime measurements, and controlled analyses of timing, observability, measurement fidelity, fault-resistance distribution shift, stride, and hyperparameter sensitivity. For fc, the primary aggregate metric is macro-f1, complemented by confusion-based diagnostics that make class-resolved behavior explicit, including confusions between the non-onset class and the onset-conditioned fault-type classes. The classification target is strongly imbalanced â the non-onset class accounts for 89.7%89.7\% of windows and the rarest fault-type class for 0.83%0.83\% (an â108:1â108:1 ratio; see Table D.1) â so macro-averaged f1 is reported rather than accuracy, since a trivial non-onset-only classifier attains 89.7%89.7\% accuracy but only â0.09â0.09 macro-f1. For fl, the primary error measure is mae (mae). Runtime measurements refer to model-side inference on pre-windowed inputs under the stated reference assumptions. Compact robustness analyses are included to assess whether the main conclusions remain stable under reasonable parameter variation. Together, these outputs define the evidence reported for the instantiated study. Section 4 uses them to analyze the empirical findings of the reference configuration and to show how controlled variations of selected evaluation dimensions affect the resulting conclusions. 4 Results and Framework-Guided Interpretation This section reports the empirical findings of the bounded reference case study introduced in Section 3. The purpose is not to present benchmark scores in isolation, but to show what becomes interpretable once task definition, physical scope, observability, timing, targets, and validation are fixed explicitly. The results are therefore read as evidence for the framework rather than as a model-ranking exercise. The section begins with the fixed 20 ms reference configuration and then examines timing and observability as controlled problem-definition axes. Conventional methods are subsequently placed within the same reporting structure to expose the role of sensing and auxiliary information. Measurement-fidelity degradation and fault-resistance distribution shift then test robustness beyond the clean in-distribution setting, while structured diagnostics, runtime measurements, stride sensitivity, and hyperparameter analyses examine the stability and practical interpretation of the reported results. Finally, a cross-family comparison of the mlp and cnn1d (with a pre-trained foundation-model baseline in Appendix E) tests whether the task-level interpretation persists beyond the classical model panel. Table 4: Reference performance across representative case-study models for the fixed 20 ms timing configuration. fc is reported by macro-f1; fl by mae in percent of normalized line length. Values are mean ± std over 5 folds. Task Metric mlp gb knn Ridge fc Macro-f1 â 0.991 ± 0.001 0.745 ± 0.017 0.799 ± 0.006 0.095 ± 0.003 fl mae [%] â 10.20 ± 0.25 14.78 ± 0.25 20.10 ± 0.16 26.12 ± 0.36 4.1 Reference Performance Under the Fixed 20 ms Configuration Table 4 summarizes the reference results for the fixed 20 ms configuration defined in Section 3. Because fc and fl are evaluated under the same physical scope, observability setting, temporal budget, and validation protocol, the comparison isolates task-dependent differences rather than differences induced by the evaluation setup. For fc, the reference results indicate that short-window fault classification is already highly effective under the present benchmark and sensing assumptions. The mlp attains a macro-f1 of 0.991±0.0010.991± 0.001, whereas the remaining representative models perform substantially worse. The weak performance of the linear ridge baseline further suggests that the classification task is strongly non-linear in the present representation. For fl, the pattern is different. Although the mlp again yields the strongest result, the lowest error remains an mae of 10.20 ± 0.25 % of normalized line length, with gb, knn, and ridge producing progressively larger errors. Under the same benchmark, observability setting, and temporal budget, classification approaches its metric ceiling, whereas localization retains physically material error. The fixed 20 ms reference configuration therefore reveals a task-dependent performance asymmetry under otherwise shared evaluation assumptions. Table 5: Controlled timing sensitivity for fault classification. Macro-f1 across decision horizons for the representative case-study models. Mean ± std over 5 folds; higher is better. The 20 ms row corresponds to the fixed reference timing configuration. Window mlp gb knn ridge 10 ms 0.985 ± 0.011 0.418 ± 0.009 0.738 ± 0.012 0.097 ± 0.002 20 ms (ref.) 0.991 ± 0.001 0.745 ± 0.017 0.799 ± 0.006 0.095 ± 0.003 30 ms 0.988 ± 0.005 0.977 ± 0.001 0.831 ± 0.004 0.093 ± 0.002 40 ms 0.988 ± 0.007 0.982 ± 0.002 0.851 ± 0.003 0.091 ± 0.002 50 ms 0.990 ± 0.005 0.982 ± 0.001 0.863 ± 0.003 0.089 ± 0.002 Table 6: Timing sensitivity for fl. mae is reported as percent of normalized line length across decision horizons. Values are mean ± std over 5 folds; lower is better. The 20 ms setting is the reference configuration. Window mlp gb knn ridge 10 ms 10.64 ± 0.36 14.65 ± 0.19 20.38 ± 0.32 26.14 ± 0.32 20 ms (ref.) 10.20 ± 0.25 14.78 ± 0.25 20.10 ± 0.16 26.12 ± 0.36 30 ms 10.18 ± 0.35 14.64 ± 0.18 19.65 ± 0.19 26.11 ± 0.35 40 ms 10.46 ± 0.52 14.59 ± 0.19 19.33 ± 0.16 26.12 ± 0.32 50 ms 9.92 ± 0.27 14.66 ± 0.20 19.12 ± 0.14 26.13 ± 0.31 4.2 Controlled Timing Sensitivity Tables 5 and 6 summarize the effect of varying the decision horizon while keeping the remaining case-study design fixed. In framework terms, this analysis isolates Step 4 and examines how the two protection objectives depend on available decision time under otherwise identical assumptions. For fc, Table 5 shows that performance is most sensitive at the shortest horizons and then approaches saturation. The mlp already attains a macro-f1 of 0.985±0.0110.985± 0.011 at 10 ms and remains strong across all horizons, whereas histogram-based gb benefits markedly from longer windows and knn improves more gradually. The linear ridge baseline remains ineffective throughout. Under the present benchmark and sensing assumptions, extending the horizon beyond 20 ms therefore provides little additional benefit for the strongest classifier, while weaker nonlinear models still benefit from additional temporal context. For fl, the timing dependence in Table 6 is much weaker. The best mae decreases only modestly from 10.64 ± 0.36 % at 10 ms to 9.92 ± 0.27 % at 50 ms, and the model ranking remains stable across horizons. Histogram-based gb and ridge change little, while knn improves only moderately. Longer windows therefore provide some benefit, but they do not remove the substantial localization error observed in this benchmark setting. Taken together, the timing analysis shows that decision horizon is part of the effective task definition rather than a secondary implementation choice: under the present evaluation setting, fc reaches very high performance at short horizons, whereas localization retains physically material error across the same timing range. Table 7: Controlled observability sensitivity for fault classification. Macro-f1 under full, relay-pair, and single-relay sensing for mlp and histogram-based gb models at 20 ms and 50 ms. Full-observability values are the mean ± standard deviation over the five episode-grouped folds; the relay-pair and single-relay values are the mean over the four same-line relay pairs and the eight individual relays, respectively, with the minâmax range across those configurations in brackets. Higher is better. Observability mlp (20 ms) gb (20 ms) mlp (50 ms) gb (50 ms) Full 0.991 ± 0.001 0.745 ± 0.017 0.990 ± 0.005 0.982 ± 0.001 Relay pair 0.989 (0.988â0.990) 0.750 (0.717â0.777) 0.987 (0.987â0.989) 0.971 (0.968â0.975) Single relay 0.986 (0.975â0.993) 0.734 (0.629â0.811) 0.986 (0.966â0.991) 0.952 (0.900â0.973) Table 8: Controlled observability sensitivity for fault localization. mae in percent of normalized line length under full, relay-pair, and single-relay sensing for mlp and histogram-based gb models at 20 ms and 50 ms. Full-observability values are the mean ± standard deviation over the five episode-grouped folds; the relay-pair and single-relay values are the mean over the four same-line relay pairs and the eight individual relays, respectively, with the minâmax range across those configurations in brackets. Lower is better. Observability mlp (20 ms) gb (20 ms) mlp (50 ms) gb (50 ms) Full 10.20 ± 0.25 14.78 ± 0.25 9.92 ± 0.27 14.66 ± 0.20 Relay pair 19.65 (17.40â21.67) 21.43 (20.30â22.42) 19.62 (17.23â21.63) 21.37 (20.12â22.38) Single relay 22.22 (20.17â23.69) 23.20 (22.46â23.94) 22.17 (20.05â23.62) 23.07 (22.19â23.77) 4.3 Controlled Observability Sensitivity Tables 7 and 8 examine observability as a controlled framework dimension. Reduced-observability results are reported for the focused sensitivity panel, comprising the neural mlp and histogram-based gb, at the fixed 20 ms reference horizon and at 50 ms as a longer-window comparison. In all cases, the physical system, target construction, and validation protocol remain unchanged, so the observed differences can be attributed directly to information availability. For fc, the effect of reduced observability is limited. As shown in Table 7, the mlp remains close to its full-observability baseline under both relay-pair and single-relay sensing, while histogram-based gb shows only moderate degradation overall, aside from one small relay-pair improvement at 20 ms. For fl, the pattern is markedly different. Table 8 shows substantial error increases under reduced observability for both models and at both horizons. Moving from full observability to relay-pair sensing already raises the mae by roughly 6.7â9.7 percentage points, and single-relay sensing increases it further to roughly 8.4â12.3 points above the full-observability baseline. Taken together, these results show that, under the present benchmark and sensing assumptions, observability has only a limited effect on fault classification but is a dominant determinant of fault localization performance. In the language of the proposed framework, information availability is therefore part of the effective task definition rather than a secondary implementation detail. 4.4 Conventional Protection Baselines Under a Shared Evaluation Structure Table 9 places conventional impedance-based fault location alongside the learning models under the shared data partitions, targets, and task-specific evaluation structure. On clean, synchronized data the synchronized two-ended method is the strongest locator. Its per-window mae is 5.42%5.42\% at 20 ms, below the best classical-learning result in Table 4 (10.20%10.20\%), and, when a single estimate is latched per fault at settled current, it reaches 2.74%2.74\% at 20 ms and 1.57%1.57\% at 50 ms. This outcome is expected and useful for the framework: a physics-based estimator that optimally combines both synchronized terminals, and that additionally receives the true line parameters and the ground-truth loop selection denied to the reference learning models, should beat a general learner on clean data. The comparison is therefore not equal-information, and the conventional methodâs advantage is reported explicitly rather than presented as an equal-footing win. The single-ended reactance method is far weaker at 20 ms (41.9%41.9\%) but improves sharply to 10.9%10.9\% at 50 ms, because its reactance estimate requires a settled post-fault phasor that only the longer window provides. The symmetrical-component phase selector attains a fault-only macro-f1 of 0.830.83 at 20 ms and 0.840.84 at 50 ms. These values provide a descriptive threshold-based reference, but they are not compared numerically with the 11-class learning-model macro-f1 because the latter additionally includes the non-onset class and is evaluated over both onset-containing and non-onset windows. Two framework points follow directly. First, the observability axis is exactly the axis that separates the conventional locators: two-ended location uses a relay pair, and single-ended location uses a single relay. Table A.2 shows that under matched observability, though with the parameter advantage noted above, the conventional two-ended method (2.74%2.74\% at relay-pair sensing) far outperforms the learning models restricted to the same relay pair (19.619.6â21.4%21.4\%), while single-ended location only becomes competitive with single-relay learning at the longer horizon. Second, the decision-horizon axis constrains the conventional methods too: single-ended location improves markedly, while the phase selector improves slightly, from 20 ms to 50 ms as the fundamental-cycle phasor settles, and the 10 ms horizon is phasor-invalid. A further framework observation is that the onset-conditioned sample-validity rule, designed for and validated on the learning models, does not transfer unchanged to single-ended impedance location, which needs a settled phasor; the per-window and per-episode-settled columns of Table 9 quantify this gap. Sample validity is therefore method-class-dependent, a distinction the framework makes explicit rather than conceals. Table 9: fl performance of learning and conventional impedance-based methods under shared episode partitions and task-specific metrics. mae [% line length], mean ± standard deviation over five folds. Per-window conventional results use the onset-conditioned evaluation windows, whereas settled results use a separately reported per-episode post-onset validity rule. Conventional locators additionally receive per-episode line parameters and ground-truth loop selection. Method Observability 20 ms 50 ms mlp full 10.20 ± 0.25 9.92 ± 0.27 gb full 14.78 ± 0.25 14.66 ± 0.20 knn full 20.10 ± 0.16 19.12 ± 0.14 Ridge full 26.12 ± 0.36 26.13 ± 0.31 Two-ended (per window) relay pair 5.42 ± 0.09 3.62 ± 0.23 Two-ended (settled) relay pair 2.74 ± 0.05 1.57 ± 0.03 One-ended (per window) single relay 156.2 ± 4.9 77.1 ± 5.6 One-ended (settled) single relay 41.9 ± 1.2 10.9 ± 0.6 4.5 Measurement-Fidelity Degradation Table 10 summarizes measurement-fidelity robustness for the focused mlpâgb sensitivity panel at the 20 ms horizon, reporting the clean value and the worst-case level across the three axes; the full per-level tables are given in Appendix B (Tables B.1 and B.2). The clean columns remain consistent with the corresponding reference results, confirming that the perturbation harness leaves the reference configuration effectively unchanged. Two findings stand out. First, degradation is monotonic within each evaluated axis, but the most damaging non-ideality is task- and model-dependent. Current-transformer saturation produces the largest localization degradation, whereas severe additive noise is particularly damaging for histogram-based gb classification. Second, and more consequentially for the framework, robustness does not track clean predictive performance, and the robustness ranking of the models is itself task-dependent. For localization, the clean-best model is the most fragile: the mlp, which attains the lowest clean error, inflates by approximately 1.21.2 to 2.4Ă2.4Ă across the severe perturbation levels at the evaluated 20 ms horizon, whereas histogram-based gb degrades far more gracefully (1.01.0 to 1.4Ă1.4Ă). For classification the ranking reverses (Table B.2): the mlp retains a macro-f1 above 0.950.95 under every perturbation, while histogram-based gb collapses under additive noise, falling from 0.740.74 to 0.270.27 at 10 dB. Aggregate clean performance is therefore insufficient to rank models for deployment; robustness must be measured rather than inferred from nominal performance or model family. Robustness is a distinct evaluation dimension, and the framework makes it comparable across the evaluated learning models and tasks. Applying the same perturbation protocol to the conventional methods remains necessary for a direct cross-method-class robustness comparison. Table 10: Compact measurement-fidelity summary at 20 ms for the focused sensitivity panel: clean value and worst-case level across the three degradation axes (additive noise, current-transformer saturation, synchronization jitter). Full per-level results in Appendix B (Tables B.1 and B.2). Task (metric) Model Clean Worst case fl (mae [%], â ) mlp 10.20 24.36 fl (mae [%], â ) gb 14.78 20.56 fc (macro-f1, â ) mlp 0.991 0.959 fc (macro-f1, â ) gb 0.745 0.274 Figure 3: Row-normalized confusion matrix for fc on fault-only cases under the 20 ms mlp configuration. Rows denote the true fault class and columns the predicted fault class. Diagonal entries indicate per-class recall, while off-diagonal cells represent misclassifications. Figure 4: fl error as a function of the true fault position for the reference 20 ms mlp configuration (out-of-fold evaluation). Results are aggregated in 10% line-length bins; the solid curve shows the median absolute error, and the shaded band spans the median to the 95th percentile. 4.6 Structured Diagnostics Aggregate metrics summarize overall performance but do not show where models fail. The task-specific diagnostics therefore expose the structure behind the reported scores for the 20 ms mlp configuration. For fc, Fig. 3 shows that the errors are sparse and structured for the 20 ms mlp. The diagonal remains close to one across all fault classes, indicating that the strong aggregate macro-f1 is not driven by only a small subset of easy cases. Per-class fault-only f1 values remain uniformly high, ranging from 0.987 to 0.996. Because Fig. 3 is restricted to onset-containing fault cases, complementary pre-fault behavior is reported separately. On the held-out folds, the 20 ms mlp assigns a fault-type class to 3 of 117 286 genuine pre-fault windows, corresponding to a pooled pre-fault-window false-positive rate of 0.0026 %. This is a bounded window-level diagnostic on the pre-fault portions of fault episodes. It should not be interpreted as an episode-level false-trip probability because no relay pickup, persistence, latching, or trip logic is modeled and the benchmark contains no dedicated non-fault disturbances. For fl, Fig. 4 shows that the residual error is not spatially uniform: median error is lowest in the mid-line region and increases toward the line ends, while the 95th-percentile envelope widens substantially near the boundaries. A coarse line-level summary shows the same tendency, with lower mae for Line 1â2 A/B (9.08 and 9.18) than for Line 2â3 A/B (10.58 and 11.05). Taken together, these diagnostics indicate that fc remains uniformly strong across classes, whereas fl is shaped by topology- and location-dependent difficulty even in the strongest evaluated configuration. Table 11: Compact runtime summary for fault classification under the reference assumptions. Macro-f1 is shown for context. Model Window f1 Train [s] Infer. [ÎŒ ] Thr. [k/s] mlp 20 ms 0.991 603.7 7.32 136.7 gb 20 ms 0.745 132.6 24.47 41.1 mlp 50 ms 0.990 1289.5 16.37 61.1 gb 50 ms 0.982 1968.8 82.91 12.1 Table 12: Compact runtime summary for fault localization under the reference assumptions. mae [%] is shown for context. Model Window mae Train [s] Infer. [ÎŒ ] Thr. [k/s] mlp 20 ms 10.20 346.3 6.45 155.0 gb 20 ms 14.78 27.2 22.49 44.5 mlp 50 ms 9.92 2258.0 15.91 62.9 gb 50 ms 14.66 140.4 49.36 20.3 4.7 Runtime Under the Reference Assumptions Tables 11 and 12 summarize training time and model-side inference time under the reference assumptions. Runtime is measured on pre-windowed inputs. The reported inference times include feature scaling and model prediction, but exclude window extraction, disk I/O, communication overhead, and synchronization overhead. The values should therefore be interpreted as relative computational indicators under a common protocol, not as end-to-end relay latencies. For context, real-time processing at the 5 ms stride requires approximately 200 windows per second per episode; all reported throughput figures exceed this threshold by more than one order of magnitude. At the 20 ms reference horizon, the mlp combines the stronger predictive result with the lower inference time within the profiled mlpâgb subset, reaching a macro-f1 of 0.991 at 7.32 ÎŒ per window for fc and an mae of 10.20 at 6.45 ÎŒ for fl. At 50 ms, inference cost increases for both model families, but substantially more for histogram-based gb, where fc inference rises from 24.47 to 82.91 ÎŒ per window. Within the profiled mlpâgb subset, the mlp is therefore the more favorable option on both predictive and computational grounds. 4.8 Stride Sensitivity Analysis A compact stride-sensitivity check for the representative mlp and histogram-based gb models (Appendix F, Tables F.1 and F.2) confirms that this secondary windowing choice does not overturn the task-level interpretation: stride can shift absolute scores and some model-family gaps â notably, histogram-based gb improves by more than 0.220.22 macro-f1 at 20 ms with larger strides â but classification continues to approach its metric ceiling while localization retains material error, and the dominant effects remain task definition, timing, and observability. 4.9 Hyperparameter Analysis Single-parameter hyperparameter ablations for the representative mlp and histogram-based gb models (Appendix G, Table G.4) are diagnostic sensitivity checks rather than a hyperparameter search. The mlp is highly stable for fc while histogram-based gb is more configuration-sensitive at the 20 ms horizon, and both vary more for fl; favorable settings improve absolute error but do not change the qualitative interpretation â classification continues to approach its metric ceiling while localization retains material error, and the dominant effects remain task definition, decision horizon, and observability. 4.10 Cross-Family Learned Baselines To examine whether the task-level conclusions persist across model families, the comparison includes a classical mlp and a compact cnn1d trained from scratch. All models use the same episode-grouped five-fold protocol, sample-validity rules, decision horizons, and task-specific metrics, while retaining the model-specific input representation and preprocessing described in Section 3.3. A pre-trained time-series foundation-model baseline (MOMENT-1-large, a frozen transformer with approximately 341 million parameters) is reported separately in Appendix E. Table 13 summarizes the main-text results. At the 50 ms horizon, the fl mae is 9.92%9.92\% for the mlp and 9.09%9.09\% for the cnn1d, while the pre-trained foundation-model baseline reaches 8.90%8.90\% (Appendix E). For fc, the cnn1d achieves the highest macro-f1 score at 0.9990.999, followed by the mlp at 0.9900.990, with the foundation-model baseline at 0.9890.989. Despite the differences in absolute performance, the broader interpretation remains unchanged: under the shared evaluation setting, classification approaches its metric ceiling while localization retains material error across classical ml, task-specific deep learning, and transfer from a pre-trained foundation model. Table 13: Main-text learned-baseline comparison under the shared episode-grouped five-fold protocol. fc is evaluated using macro-f1 (â ), and fl using mae in percent of line length (â ). Values are mean ± standard deviation across the five held-out folds; the best mean in each column is shown in bold. The pre-trained MOMENT-1-large results are reported separately in Appendix E (Table E.1). Model fc 20 ms fc 50 ms fl 20 ms fl 50 ms mlp (reference) 0.991 ± 0.001 0.990 ± 0.005 10.20 ± 0.25 09.92 ± 0.27 cnn1d (from scratch) 0.999 ± 0.001 0.999 ± 0.001 09.65 ± 0.67 09.09 ± 0.28 4.11 Framework-Level Synthesis Across the shared evaluation configuration and the controlled analyses, seven findings emerge. First, fc and fl exhibit fundamentally different behavior despite identical sensing, timing, and validation assumptions. Second, decision horizon and observability materially affect the apparent difficulty of both tasks and must therefore be treated as explicit evaluation dimensions. Third, the classificationâlocalization asymmetry persists across classical ml, task-specific deep learning, and foundation-model transfer, indicating that it is not specific to a single model family. Fourth, the synchronized two-ended conventional locator outperforms the learned locators on clean data when measurements from both terminals, line parameters, and ground-truth loop selection are available, demonstrating that information-set differences must be reported explicitly. Fifth, measurement degradation shows that clean-data performance does not determine robustness and that robustness rankings can differ between tasks. Sixth, the directional fault-resistance holdout reveals model- and direction-dependent performance differences (Appendix C). Seventh, stride and hyperparameter choices can affect differences between model families, while structured diagnostics and runtime measurements reveal limitations that aggregate metrics alone do not expose. Together, these findings support the central premise of the framework: performance in machine-learning-based protection can be interpreted only relative to the complete evaluation setting. 5 Discussion This paper proposes a standardized evaluation framework for ml-based power system protection and demonstrates it through one bounded case study. The main contribution is not a claim about one universally best model or one definitive benchmark result, but a reusable way to specify, evaluate, and interpret protection studies before model performance is compared. The framework defines what must be fixed and reported; the case study shows what becomes visible when those requirements are met. 5.1 Main Contributions of the Framework The reusable contribution is the seven-step evaluation logic and the associated reporting structure. Its central claim is that reported results are only scientifically interpretable when the effective inference problem â task, observability, timing, and validation â is made explicit before model performance is compared. Under this perspective, the value of the case study is not that it adds another benchmark leaderboard, but that it shows what becomes visible once these dimensions are fixed under a common protocol: fc and fl behave differently even under shared assumptions, and observability materially changes the apparent difficulty of the task. The sensitivity analyses further show that even model-family comparisons are not invariant to seemingly secondary configuration choices. Histogram-based gb improves substantially for fc at the 20 ms horizon under alternative stride and hyperparameter settings, while the mlp remains comparatively stable. Thus, part of the observed mlpâgb gap in the default configuration reflects configuration sensitivity rather than only model-family capability. This illustrates why window construction, preprocessing, model configuration, and validation protocol must be reported as part of the evaluated protection problem, rather than treated as incidental implementation details. The framework also accommodates multi-function protection schemes in which fd, fc, and fl are required jointly rather than in isolation. In such settings, Step 1 must enumerate each protection function and its learning formulation explicitly, while Steps 4 and 5 must state whether timing constraints and target definitions are shared or task-specific. This avoids the interpretive ambiguity that arises when distinct protection objectives share a single aggregate metric, masking task-specific failure modes. The frameworkâs novelty therefore lies in integrating complementary protection and machine-learning requirements into one operational evaluation workflow, not in claiming that its individual dimensions are unprecedented. Applied retrospectively, the reporting package also exposes which assumptions would need to be reconstructed from supplementary material or clarified by the original authors. The exact macro-f1 values, mae values, model rankings, and the persistent localization error are specific to PROTECT-90, the Double Line topology, the chosen sensing regimes, the representative model set, and the stated configuration choices used here. These findings should therefore be read as benchmark-conditioned evidence, not as universal statements about all protection settings. What the community can reuse is not the specific performance profile of this benchmark, but the evaluation structure, reporting package, and interpretation logic that tie empirical results to explicit protection assumptions. 5.2 Cross-Method-Class Comparison, Robustness, and Consistency The conventional comparisons show that the framework applies beyond learning models, because their performance likewise depends on observability, decision horizon, auxiliary inputs, and sample-validity rules. The observability axis is the axis that separates one- from two-terminal fault location, and under matched observability the synchronized two-ended locator outperforms the learning models on clean data; the decision-horizon axis is the budget that governs phasor estimability, and single-ended location and the phase selector both improve as the fundamental cycle settles. Recent work comparing dynamic-state-estimation-based protection with Transformer-based diagnosis illustrates a complementary division of functions: the physics-based method provides rapid anomaly detection, whereas the Transformer differentiates physical faults from measurement attacks and supplies measurement-level diagnostic information [2]. Such combinations should therefore be evaluated using function-specific timing budgets, targets, and outputs rather than through a single aggregate performance ranking. Within the mlpâgb fidelity comparison, robustness does not track clean predictive performance and the ranking is task-dependent: for localization, the lower-error mlp degrades more strongly than gb, whereas for classification the ranking reverses. Aggregate clean performance therefore does not by itself rank models for deployment. More broadly, realism and robustness, together with cross-method-class comparability, are themselves evaluation dimensions that the framework exposes. These results are consistent with the surrounding literature on the same benchmark. The reference and timing errors agree with the authorsâ controlled model-comparison study on this benchmark [43], which likewise fixes default hyperparameters; the fc and fl reference and timing values reported here are shared with that controlled-comparison study, and this overlap is disclosed explicitly, as it is for the companion robustness study. The sensor-availability robustness that the present study deliberately does not repeat is characterized in a companion paper on the same data [47]; the lower localization error reported there under performance-tuned hyperparameters is consistent with the present studyâs diagnostic ablation, which shows the favorable-hyperparameter error moving toward that value. This is consistent with the framework rather than an inconsistency, and it underscores the requirement to report the hyperparameter procedure. Finally, the framework fixes and reports the effective inference problem so that comparisons are legible; it does not by itself guarantee that a benchmark is representative of the physical-operational problem, and representativeness and critical-condition coverage remain open problems that the framework highlights but does not solve. A utility or relay manufacturer could use the framework as a study-design and reporting checklist for internal and third-party evaluations, not as a deployment certification. Operationally, adoption can proceed in three stages. First, the protection task, physical scope, information access, and timing budget are frozen before model development. Second, nominal, degraded, and critical-condition test matrices are defined together with task-specific acceptance criteria. Third, a versioned reporting package records data provenance, preprocessing boundaries, validation groups, model configuration, diagnostics, and runtime conditions. This supports auditable internal studies, supplier comparisons, and procurement-oriented evaluation without replacing formal product qualification or protection certification. 5.3 Comparability, Reproducibility, and Boundaries The framework improves comparability by shifting the scientific point of comparison from models alone to the evaluation setting that defines the problem those models are solving. In much of the existing literature, grid assumptions, sensing access, decision windows, target definitions, and split protocols vary simultaneously, so similar headline scores may correspond to different protection problems. Here, these dimensions are fixed and reported explicitly before model performance is interpreted, making it easier to distinguish task-dependent conclusions from differences induced by information availability, window construction, model configuration, or validation choices. The same logic improves reproducibility: reproducibility is not only a matter of releasing code, but of making the evaluation path inspectable enough that another group can reconstruct, critique, and reuse it under the same stated conditions. Moving from reproducible research evaluation toward deployment requires additional verification and validation procedures for neural-network-based protection schemes [37]. Openness of the underlying data strengthens reproducibility, but full data release is not always feasible: protection recordings and grid models are frequently subject to confidentiality, cybersecurity, and intellectual-property constraints at utilities and equipment vendors. The framework remains applicable under these limits, because it preserves comparability at the level of study design. Even when raw waveforms cannot be shared, a complete reporting package â the declared physical scope, observability, timing, target and sample-validity definitions, and validation protocol â allows a result to be interpreted, critiqued, and reused under its stated conditions. Data openness and its institutional constraints are thus reported as an explicit study dimension rather than treated as an all-or-nothing requirement. That said, the present instantiation is bounded in several respects. Empirically, it covers one public emt benchmark derived from a single 90 kV Double Line topology. The observed task-dependent performance asymmetry between fc and fl may not generalize to meshed or ring-bus topologies, where overlapping current paths can make fault type discrimination substantially harder while simultaneously providing richer spatial information for localization. In terms of realism, the reference setting assumes synchronized, noise-free, continuously available measurements. Measurement-fidelity non-idealities (additive noise, current-transformer saturation, and synchronization jitter) are characterized as controlled sensitivity axes in Section 4.5, and sensor-availability degradation on the same data is characterized in the companion study [47]; communication latency and broader non-fault disturbances remain outside the present scope. In particular, security-critical non-fault events â such as transformer inrush, capacitor-bank switching, motor starting, and power swings â are not represented in the present benchmark. The pre-fault-window false-positive rate reported among the structured diagnostics is therefore a bounded diagnostic rather than an episode-level false-trip probability; a systematic false-trip and failure-to-trip security assessment under such conditions is left to future work. The model coverage is investigation-specific: the broad classical panel is used for the reference and timing analyses, the mlpâgb subset for focused sensitivity analyses, and the mlpâcnn1dâMOMENT panel for the cross-family comparison; conventional protection methods are evaluated separately in Section 4.4. Finally, the reported runtime values are model-side indicators on pre-windowed inputs, not end-to-end relay latencies; the cnn1d and MOMENT baselines were not profiled for runtime or memory and therefore do not support conclusions regarding deployment suitability. The framework applies equally to field-recording settings, where ground-truth metadata is unavailable or uncertain. In such cases, Step 2 must flag the data source and its label provenance explicitly; Step 5 must document the labeling procedure and its associated uncertainty; and Step 4 must adapt the temporal reference to a proxy such as relay trip time if fault inception is unknown. Label uncertainty then becomes a declared study dimension that propagates into the interpretation of all downstream metrics, rather than a hidden assumption. 5.4 Future Extensions of the Framework The next step is not simply to add more models, but to extend the same evaluation logic to broader and more realistic protection settings. This includes additional topologies and operating regimes, broader conventional protection methods, and additional recurrent, topology-aware, and end-to-end adapted foundation-model baselines under the same protocol, including graph-based protection models [31]. It also includes extension of the disturbance space toward higher-resistance and high-impedance fault regimes beyond the present benchmark range, as well as security-critical non-fault events. Sensing studies should extend the present measurement-fidelity analyses toward physically detailed ct/vt models, correlated or nonstationary noise, missing channels, and communication latency, jitter, or loss. In this context, feature- and channel-selection methods provide one route for studying which measurements are most relevant under reduced or degraded observability [29]. For studies working with smaller datasets, the grouped cross-validation protocol requires careful adaptation. When the number of independent groups per fold is low, fold-level estimates become sensitive to partition randomness. In such settings, leave-one-group-out cross-validation is preferable; results should additionally verify class coverage across folds before aggregate metrics are interpreted. At the reporting level, future work should complement aggregate predictive metrics with outputs more directly tied to protection security, such as false-positive behavior under extreme non-fault conditions, and move toward end-to-end latency accounting where deployment relevance is the goal. The feasible extension of the framework will depend on the availability of suitable public datasets and on access to information that is often vendor-locked, especially with respect to hardware behavior, instrument nonidealities, and relay implementation details. Pursued systematically, these extensions would move the framework from a controlled research standard toward an evaluation basis for deployment-relevant protection assessment. 6 Conclusion Reported performance in machine-learning-based power system protection is not a property of the model alone. It depends on the protection objective, physical system, available measurements, decision time, target construction, validation protocol, and reported evidence. The proposed seven-step framework makes these conditions explicit and treats evaluation design as part of the scientific contribution. The unit of comparison therefore becomes the complete inference problem, not an isolated algorithm or headline score. The PROTECT-90 case study shows why this matters. Under a fixed physical scope, sensing configuration, timing definition, and episode-grouped validation protocol, onset-conditioned fault classification approached its metric ceiling while fault localization retained physically material error. Reduced observability had little effect on classification but approximately doubled localization error. The synchronized two-ended locator outperformed the waveform-based learners when given measurements from both terminals, line parameters, and ground-truth loop selection. Measurement degradation also showed that the best clean-data model need not be the most robust. These results are specific to the evaluated benchmark and do not establish universal rankings. Their value lies in showing which assumptions produce each result. The framework is not yet a deployment certificate; it defines the evidence needed for future qualification and certification. Reaching that stage requires evaluation across additional topologies and operating regimes, broader conventional and learning-based method families, uncertain event timing, field and hardware-in-the-loop recordings, security-critical non-fault disturbances, instrument-transformer nonidealities, communication degradation, and end-to-end latency. Near-perfect scores on isolated benchmarks are insufficient. Protection models need evidence that remains interpretable and comparable as tasks, sensors, and operating conditions change. By making those conditions explicit and testable, the proposed framework turns separate studies into cumulative engineering evidence and provides a path toward dependable, auditable, and certifiable machine-learning-based protection. Data and Code Availability The case study is based on the publicly available PROTECT-90 dataset. The accompanying paper documents the simulation design, waveform structure, metadata, and intended benchmark use [42]; dataset version 1.0.0 is available through Zenodo at 10.5281/zenodo.21109169 [30]. The corresponding framework implementation, including code for preprocessing, windowing, task construction, leakage-aware validation, and result reporting, is available at https://github.com/julianoelhaf/protection-eval-framework. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Acknowledgment This project was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 535389056. References [1] A. Abdullah (2018) Ultrafast Transmission Line Fault Detection Using a DWT-Based ANN. IEEE Transactions on Industry Applications 54 (2), p. 1182â1193. External Links: ISSN 0093-9994, 1939-9367, Link, Document Cited by: §1.1. [2] E. Abukhousa, S. Zonouz, and A. P. S. Meliopoulos (2026) Transformer is All You Need: Attention-Based Anomaly Detection and Classification in Inverter-Rich Power Systems. In 2026 IEEE International Conference on Communications, Control, and Computing Technologies for Smart Grids (SmartGridComm), College Station, TX, USA. Note: Accepted at IEEE SmartGridComm 2026. External Links: Link Cited by: §5.2. [3] E. Abukhousa, S. Zonouz, and A.P. S. Meliopoulos (2026) Latency-Aware Deep Learning Benchmark for Real-Time Cyber-Physical Attack and Fault Classification in Inverter-Dominated Power Grids. In 2026 IEEE/PES Transmission and Distribution Conference and Exposition (T&D), Chicago, IL, USA, p. 1â5. External Links: ISBN 979-8-3315-5569-6, Link, Document Cited by: §2.9. [4] R. Ashmore, R. Calinescu, and C. Paterson (2022) Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges. ACM Computing Surveys 54 (5), p. 1â39 (en). External Links: ISSN 0360-0300, 1557-7341, Link, Document Cited by: §1. [5] J. L. Blackburn and T. J. Domin (2014) Protective Relaying: Principles and Applications, Fourth Edition. 4 edition, Taylor & Francis Group, Baton Rouge (eng). External Links: ISBN 978-1-4398-8812-4 Cited by: §1, §3.4. [6] X. Bouthillier, P. Delaunay, M. Bronzi, A. Trofimov, B. Nichyporuk, J. Szeto, N. Mohammadi Sepahvand, E. Raff, K. Madan, V. Voleti, et al. (2021) Accounting for variance in machine learning benchmarks. Proceedings of machine learning and systems 3, p. 747â769. Cited by: §2.8, §2.9. [7] W. Chen (2005) The electrical engineering handbook. Elsevier Academic Press, Boston (eng). Note: OCLC: 57371415 External Links: ISBN 978-1-4175-5266-5 Cited by: §1. [8] A. S. Da Silva, R. C. Dos Santos, and G. T. De Alencar (2025) An Intelligent Time-Domain ANN-Based Method for Fault Identification in CSC-HVDC Systems. Smart Grids and Sustainable Energy 10 (2), p. 50 (en). External Links: ISSN 2731-8087, Link, Document Cited by: §1.1. [9] European Union Aviation Safety Agency (2024) Artificial Intelligence Concept Paper Issue 2: Guidance for Level 1 & 2 Machine Learning Applications. Technical Report European Union Aviation Safety Agency (EASA). External Links: Link Cited by: §1. [10] A. Evdakov, G. Filatova, A. Yablokov, A. Kovalenko, E. Skachkov, and I. Makarov (2026) A dataset of real-world oscillograms from electrical power grids. Scientific Data 13 (1), p. 262 (en). External Links: ISSN 2052-4463, Link, Document Cited by: §3.1. [11] T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. Iii, and K. Crawford (2021) Datasheets for datasets. Communications of the ACM 64 (12), p. 86â92 (en). External Links: ISSN 0001-0782, 1557-7317, Link, Document Cited by: Table 1, §2.10. [12] M. Gillioz, G. Dubuis, and P. Jacquod (2025) A large synthetic dataset for machine learning applications in power transmission grids. Scientific Data 12 (1), p. 168 (en). External Links: ISSN 2052-4463, Link, Document Cited by: §3.1. [13] O. E. Gundersen and S. Kjensmo (2018) State of the Art: Reproducibility in Artificial Intelligence. Proceedings of the AAAI Conference on Artificial Intelligence 32 (1). External Links: ISSN 2374-3468, 2159-5399, Link, Document Cited by: §1.1. [14] A. Haddadi, E. Farantatos, I. Kocar, and U. Karaagac (2021) Impact of Inverter Based Resources on System Protection. Energies 14 (4), p. 1050 (en). External Links: ISSN 1996-1073, Link, Document Cited by: §1. [15] IEEE Power and Energy Society (2015) IEEE Guide for Determining Fault Location on AC Transmission and Distribution Lines. Standard, IEEE. External Links: Link, Document Cited by: §3.4. [16] IEEE (2023) IEEE Guide for the Application of Current Transformers Used for Protective Relaying Purposes. IEEE. External Links: ISBN 978-1-5044-9541-7, Link, Document Cited by: §3.5. [17] International Electrotechnical Commission and IEEE (2018) Measuring relays and protection equipment- Part 118-1: Synchrophasor for power systems -Measurements. International Standard, IEEE, New York, NY, USA. External Links: Link Cited by: §2.5, §2. [18] International Electrotechnical Commission (2014) Measuring relays and protection equipment â Part 121: Functional requirements for distance protection. International Standard, 1.0 edition, International Electrotechnical Commission, Geneva, Switzerland (English/French). Note: Issue: IEC 60255-121:2014 External Links: ISBN 978-2-8322-1399-5 Cited by: Table 1, §2.3, §2.6, §2. [19] International Electrotechnical Commission (2020) Communication networks and systems for power utility automation - Part 9-2: Specific communication service mapping (SCSM) - Sampled values over ISO/IEC 8802-3. International Standard, IEC, Geneva, Switzerland (English). Note: Consolidated version incorporating IEC 61850-9-2:2011 and Amendment 1:2020. External Links: ISBN 978-2-8322-7886-4 Cited by: §2.5. [20] International Electrotechnical Commission (2022) Measuring relays and protection equipment - Part 1: Common requirements. Standard, International Electrotechnical Commission (IEC), Geneva, Switzerland. External Links: Link Cited by: §1.2, Table 1, §1, §3.1. [21] ISO/IEC (2022) Framework for Artificial Intelligence (AI) Systems Using Machine Learning (ML). Standard, International Organization for Standardization. External Links: Link Cited by: §1.2, §1. [22] ISO/IEC (2023) Information Technology - Artificial Intelligence - Management System. Standard, 1 edition, International Organization for Standardization, Geneva, Switzerland. External Links: Link Cited by: §1.2, §1. [23] A.T. Johns and S. Jamali (1990) Accurate fault location technique for power transmission lines. IEE Proceedings C Generation, Transmission and Distribution 137 (6), p. 395 (en). External Links: ISSN 01437046, Link, Document Cited by: §3.4. [24] C. B. Jones, A. Summers, and M. J. Reno (2021) Machine Learning Embedded in Distribution Network Relays to Classify and Locate Faults - 2021 IEEE Power & Energy Society Innovative Smart Grid Technologies Conference (ISGT). In 2021 IEEE Power & Energy Society Innovative Smart Grid Technologies Conference (ISGT), Washington, DC, USA, p. 1â5. External Links: Document Cited by: §1.1. [25] S. Kapoor, E. M. Cantrell, K. Peng, T. H. Pham, C. A. Bail, O. E. Gundersen, J. M. Hofman, J. Hullman, M. A. Lones, M. M. Malik, P. Nanayakkara, R. A. Poldrack, I. D. Raji, M. Roberts, M. J. Salganik, M. Serra-Garcia, B. M. Stewart, G. Vandewiele, and A. Narayanan (2024) REFORMS: Consensus-based Recommendations for Machine-learning-based Science. Science Advances 10 (18) (en). External Links: ISSN 2375-2548, Link, Document Cited by: §1.1, Table 1, §2.10. [26] S. Kapoor and A. Narayanan (2023) Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4 (9), p. 100804 (en). External Links: ISSN 26663899, Link, Document Cited by: §1.1, Table 1, §2.8, §2, §3.3. [27] S. Kim, K. Choi, H. Choi, B. Lee, and S. Yoon (2022) Towards a Rigorous Evaluation of Time-Series Anomaly Detection. Proceedings of the AAAI Conference on Artificial Intelligence 36 (7), p. 7194â7201. External Links: ISSN 2374-3468, 2159-5399, Link, Document Cited by: §2.7, §2. [28] G. Kordowich, M. Jaworski, T. Lorz, C. Scheibe, and J. Jaeger (2022) A hybrid Protection Scheme based on Deep Reinforcement Learning. In 2022 IEEE PES Innovative Smart Grid Technologies Conference Europe (ISGT-Europe), Novi Sad, Serbia, p. 1â6. External Links: ISBN 978-1-6654-8032-1, Link, Document Cited by: §1.1. [29] G. Kordowich, J. Oelhaf, S. Bayer, A. Maier, M. Kereit, and J. Jaeger (2026) Feature selection for fault prediction in distribution systems. Electric Power Systems Research 261, p. 113498 (en). External Links: ISSN 03787796, Link, Document Cited by: §2.5, §5.4. [30] G. Kordowich, J. Oelhaf, C. Bergler, A. Maier, S. Bayer, and J. JĂ€ger (2026) PROTECT-90: A Fault Dataset for Power System Protection: Open, Standardized Voltage and Current Waveforms for Reproducible Protection and Transient Analysis. Zenodo (en). Note: [Dataset]. Zenodo, version 1.0.0https://doi.org/10.5281/zenodo.21109169 External Links: Link, Document Cited by: §3, Data and Code Availability. [31] G. Kordowich, J. Oelhaf, A. Maier, S. Bayer, and J. Jaeger (2025) A Graph Neural Network-Based Approach for Power System Protection. In 2025 IEEE Kiel PowerTech, Kiel, Germany, p. 1â6. External Links: ISBN 979-8-3315-4397-6, Link, Document Cited by: §1.1, §5.4. [32] M. Kouraichi, M. Mansouri, M. Trabelsi, A. MâHalla, A. S. Abdel-Khalik, and A. Sakly (2025) Deep Learning for Fault Diagnosis in Power Transmission Lines: Current Trends, Limitations, and Future Directions. IEEE Access 13, p. 192105â192142. External Links: ISSN 2169-3536, Link, Document Cited by: §1.1. [33] V. Krishnan, B. Bugbee, T. Elgindy, C. Mateo, P. Duenas, F. Postigo, J. Lacroix, T. G. S. Roman, and B. Palmintier (2020) Validation of Synthetic U.S. Electric Power Distribution System Data Sets. IEEE Transactions on Smart Grid 11 (5), p. 4477â4489. External Links: ISSN 1949-3053, 1949-3061, Link, Document Cited by: §3.1. [34] S. Kumar Mohanty, A. Swetapadma, P. Kumar Nayak, and O. P. Malik (2023) Decision tree approach for fault detection in a TCSC compensated line during power swing. International Journal of Electrical Power & Energy Systems 146, p. 108758 (en). External Links: ISSN 01420615, Link, Document Cited by: §1.1. [35] T. S. Kumar, R. Meena, P. Mani, S. Ramya, K. L. Khandan, A. Mohammed, and M. S. Ramkumar (2022) Deep Learning based Fault Detection in Power Transmission Lines. In 2022 4th International Conference on Inventive Research in Computing Applications (ICIRCA), Coimbatore, India, p. 861â867. Note: Section: 0 External Links: Document Cited by: §1.1. [36] H. Livani and C. Y. Evrenosoglu (2013) A Fault Classification and Localization Method for Three-Terminal Circuits Using Machine Learning. IEEE Transactions on Power Delivery 28 (4), p. 2282â2290. External Links: ISSN 0885-8977, 1937-4208, Link, Document Cited by: §1.1. [37] C. Mederer, G. Kordowich, J. Oelhaf, A. Maier, S. Bayer, and J. Jaeger (2025) Verification of neural network based power system protection schemes. IET Conference Proceedings 2025 (5), p. 155â159 (en). External Links: ISSN 2732-4494, Link, Document Cited by: §1, §5.3. [38] M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru (2019) Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, Atlanta GA USA, p. 220â229 (en). External Links: ISBN 978-1-4503-6125-5, Link, Document Cited by: §2.10. [39] S. K. Mohanty, P. K. Nayak, P. K. Bera, and H. H. Alhelou (2024) An Enhanced Protective Relaying Scheme for TCSC Compensated Line Connecting DFIG-Based Wind Farm. IEEE Transactions on Industrial Informatics 20 (3), p. 3425â3435. External Links: ISSN 1551-3203, 1941-0050, Link, Document Cited by: §1.1. [40] M. Najafzadeh, J. Pouladi, A. Daghigh, J. Beiza, and T. Abedinzade (2024) Fault Detection, Classification and Localization Along the Power Grid Line Using Optimized Machine Learning Algorithms. International Journal of Computational Intelligence Systems 17 (1), p. 49 (en). External Links: ISSN 1875-6883, Link, Document Cited by: §1.1. [41] H. Niemann (1983) Klassifikation von Mustern. Springer Berlin Heidelberg, Berlin, Heidelberg (ger). External Links: ISBN 978-3-540-12642-3 978-3-642-47517-7, Document Cited by: Figure 1. [42] J. Oelhaf, G. Kordowich, C. Bergler, A. Maier, J. JĂ€ger, and S. Bayer (2026) PROTECT-90: A Fault Dataset for Power System Protection. In 2026 IEEE PES Innovative Smart Grid Technologies Conference Europe (ISGT-Europe), Note: Accepted for publication External Links: Link Cited by: Figure 2, §3.1, §3, Data and Code Availability. [43] J. Oelhaf, G. Kordowich, C. Kim, P. A. PĂ©rez-Toro, C. Bergler, A. Maier, J. JĂ€ger, and S. Bayer (2026) Controlled Comparison of Machine Learning Models for Fault Classification and Localization in Power System Protection. In 2026 IEEE PES Innovative Smart Grid Technologies Conference Europe (ISGT-Europe), Note: Accepted for publication External Links: Link Cited by: §1.1, §5.2. [44] J. Oelhaf, G. Kordowich, C. Kim, P. A. PĂ©rez-Toro, A. Maier, J. JĂ€ger, and S. Bayer (2025) Impact of Data Sparsity on Machine Learning for Fault Detection in Power System Protection. In 2025 33rd European Signal Processing Conference (EUSIPCO), Palermo, Italy, p. 1997â2001. External Links: ISBN 978-94-645936-2-4, Link, Document Cited by: §1.1. [45] J. Oelhaf, G. Kordowich, M. Pashaei, C. Bergler, A. Maier, J. JĂ€ger, and S. Bayer (2025) A Scoping Review of Machine Learning Applications in Power System Protection and Disturbance Management. International Journal of Electrical Power & Energy Systems 172, p. 111257 (en). External Links: ISSN 01420615, Link, Document Cited by: §1.1, §1.1, §1.1, Table 1, §3.1. [46] J. Oelhaf, G. Kordowich, P. A. PĂ©rez-Toro, T. Arias-Vergara, A. Maier, J. JĂ€ger, and S. Bayer (2025) A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power Grids. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, p. 1â5. External Links: ISBN 979-8-3503-6874-1, Link, Document Cited by: §1.1. [47] J. Oelhaf, M. Pashaei, G. Kordowich, C. Bergler, A. Maier, J. JĂ€ger, and S. Bayer (2026) Robustness evaluation of machine learning models for fault classification and localization in power system protection. IET Conference Proceedings 2026 (3), p. 188â193 (en). External Links: ISSN 2732-4494, Link, Document Cited by: Appendix G, §3.1, §3.5, §5.2, §5.3. [48] B. Patnaik, M. Mishra, R.C. Bansal, and R.K. Jena (2021) MODWT-XGBoost Based Smart Energy Solution for Fault Detection and Classification in a Smart Microgrid. Applied Energy 285, p. 116457. External Links: Link, Document Cited by: §1.1. [49] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay (2011) Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 12, p. 2825â2830. Cited by: §3.3. [50] A. G. Phadke and J. S. Thorp (2009) Computer Relaying for Power Systems. 2 edition, Wiley (en). External Links: ISBN 978-0-470-05713-1 978-0-470-74972-2, Link, Document Cited by: §3.4, §3.4. [51] J. Pineau, P. Vincent-Lamarre, K. Sinha, V. LariviĂšre, A. Beygelzimer, F. dâAlchĂ©-Buc, E. Fox, and H. Larochelle (2021) Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program). J. Mach. Learn. Res. 22 (1). External Links: ISSN 1532-4435 Cited by: §2.8, §2. [52] G. Porawagamage, K. Dharmapala, J. S. Chaves, D. Villegas, and A. Rajapakse (2024) A review of machine learning applications in power system protection and emergency control: opportunities, challenges, and future directions. Frontiers in Smart Grids 3, p. 1371153. External Links: ISSN 2813-4311, Link, Document Cited by: §1.1, Table 1. [53] Protection and automation (B5) and A. distribution systems and distributed energy resources (C6) (2015) Protection of distribution systems with distributed energy resources. Technical Brochure CIGRE (en-GB). External Links: Link Cited by: §1. [54] J. C. Quispe and E. Orduña (2022) Transmission line protection challenges influenced by inverter-based resources: a review. Protection and Control of Modern Power Systems 7 (1), p. 28 (en). External Links: ISSN 2367-2617, 2367-0983, Link, Document Cited by: §1. [55] D. R. Roberts, V. Bahn, S. Ciuti, M. S. Boyce, J. Elith, G. GuilleraâArroita, S. Hauenstein, J. J. LahozâMonfort, B. Schröder, W. Thuiller, D. I. Warton, B. A. Wintle, F. Hartig, and C. F. Dormann (2017) Crossâvalidation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 40 (8), p. 913â929 (en). External Links: ISSN 0906-7590, 1600-0587, Link, Document Cited by: Table 1, §2.8, §3.3. [56] M. M. Saha, J. Izykowski, and E. Rosolowski (2010) Fault Location on Power Networks. Power Systems, Springer London, London. External Links: ISBN 978-1-84882-885-8 978-1-84882-886-5, Link, Document Cited by: §3.4. [57] R. G. Sargent (2013) Verification and validation of simulation models. Journal of Simulation 7 (1), p. 12â24 (en). External Links: ISSN 1747-7778, 1747-7786, Link, Document Cited by: §2.4. [58] J. Schindler, J. Prommetta, and J. JĂ€ger (2020) Secure and Dependable Protection Relay Behaviour in Extremely High Loaded Transmission Systems. In 15th International Conference on Developments in Power System Protection (DPSP 2020), Liverpool, UK, p. 6 p.â6 p.. External Links: Link, Document Cited by: §1. [59] E. Tabassi (2023) Artificial Intelligence Risk Management Framework (AI RMF 1.0). Technical report Technical Report NIST AI 100-1, National Institute of Standards and Technology (U.S.), Gaithersburg, MD. External Links: Link, Document Cited by: §1.2, Table 1, §1, §2.3, §2.9, §2. [60] T. Takagi, Y. Yamakoshi, M. Yamaura, R. Kondow, and T. Matsushima (1982) Development of a New Type Fault Locator Using the One-Terminal Voltage and Current Data. IEEE Transactions on Power Apparatus and Systems PAS-101 (8), p. 2892â2898. External Links: ISSN 0018-9510, Link, Document Cited by: §3.4. [61] M.N. Uddin, N. Rezaei, and M.S. Arifin (2022) Hybrid Machine Learning-based Intelligent Distance Protection and Control Schemes with Fault and Zonal Classification Capabilities for Grid-connected Wind Farms. In Conference Record - IAS Annual Meeting (IEEE Industry Applications Society), Vol. 2022, Detroit, MI, USA, p. 1â8. External Links: Link, Document Cited by: §1.1. [62] UL Standards & Engagement (2023) Standard for Safety for the Evaluation of Autonomous Products. Standard, UL. External Links: Link Cited by: §1. [63] R. Vaish, U.D. Dwivedi, S. Tewari, and S.M. Tripathi (2021) Machine learning applications in power system fault diagnosis: Research advancements and perspectives. Engineering Applications of Artificial Intelligence 106, p. 104504 (en). External Links: ISSN 09521976, Link, Document Cited by: §1.1, Table 1. [64] VDE (2015) Der Zellulare Ansatz - VDE Studie. Fachinformation VDE Verband der Elektrotechnik Elektronik Informationstechnik e.V. (German). External Links: Link Cited by: §1. [65] B. Vivek, B. Teja, B. Mallala, and G. Srinitha (2024) Electrical Fault Detection And Localization Using Machine Learning. In 2024 International Conference on Expert Clouds and Applications (ICOECA), Bengaluru, India, p. 820â825. External Links: ISBN 979-8-3503-8579-3, Link, Document Cited by: §1.1. [66] M. Wang, G. Kordowich, and J. JĂ€ger (2022) A Generic Data Generation Framework for Short Circuit Detection Training of Neural Networks. In PESS + PELSS 2022; Power and Energy Student Summit, Kassel, Germany, p. 49â54 (English). External Links: ISBN 978-3-8007-6013-8, Link Cited by: §3.1. [67] A. J. Wilson, A. Riza Ekti, J. Follum, S. Biswas, C. Annalicia, J. Joo, O. Aziz, and J. Lian (2024) The Grid Event Signature Library: An Open-Access Repository of Power System Measurement Signatures. IEEE Access 12, p. 76207â76218. External Links: ISSN 2169-3536, Link, Document Cited by: §3.1. [68] R. Wu and E. J. Keogh (2022) Current Time Series Anomaly Detection Benchmarks are Flawed and are Creating the Illusion of Progress (Extended Abstract). In 2022 IEEE 38th International Conference on Data Engineering (ICDE), Kuala Lumpur, Malaysia, p. 1479â1480. External Links: ISBN 978-1-6654-0883-7, Link, Document Cited by: §2.7, §2. [69] G. K. Yadav, M. K. Kirar, S.C. Gupta, and J. Rajender (2025) Integrating ANN and ANFIS for effective fault detection and location in modern power grid. Science and Technology for Energy Transition 80, p. 34. External Links: ISSN 2804-7699, Link, Document Cited by: §1.1. [70] G. Ziegler (2011) Numerical distance protection: principles and applications. 4 edition, Publicis Publishing, Erlangen. External Links: ISBN 978-3-89578-381-4 978-3-89578-667-9 Cited by: §3.4. Appendix A Detailed Reduced-Observability Results This appendix reports relay-level reduced-observability results for the representative 50 ms mlp fl configuration. Table A.1 complements the aggregate observability summary in Table 8 by comparing single-terminal and same-line two-terminal sensing for each protected line. The results show that two-terminal same-line sensing consistently reduces localization error relative to either individual terminal, while the remaining variation across lines indicates that measurement usefulness is topology-dependent rather than uniform across the grid. Table A.1: Relay-level fl error for the 50 ms mlp regressor under reduced observability. The improvement is computed relative to the better of the two single-terminal settings for each line. Line Single-terminal mae [%] Two-terminal Improvement Terminal 1 Terminal 2 mae [%] [% points] 01â02A 21.948 20.055 17.229 2.826 01â02B 22.316 22.712 20.814 1.502 02â03A 20.510 23.202 18.797 1.712 02â03B 22.997 23.620 21.630 1.367 fl performance under observability matched between conventional and learning methods is reported in Table A.2; these values also appear, differently organized, in Table 9. Table A.2: fl performance by observability at 20 ms for sensing-matched conventional and learning methods. Conventional values use the separately reported per-episode settled estimates and additionally receive episode-specific line parameters and ground-truth loop selection; the comparison is therefore sensing-matched but not input- or sample-validity-matched. mae [% line length]; lower is better. Observability Conventional Conv. mlp gb Full (n/a) â 10.20 14.78 Relay pair two-ended synchr. 2.74 19.65 21.43 Single relay one-ended reactance 41.9 22.22 23.20 Appendix B Detailed Measurement-Fidelity Results This appendix reports the full per-level measurement-fidelity results summarized compactly in Table 10. Tables B.1 and B.2 give localization and classification degradation, respectively, across all evaluated levels of additive noise, current-transformer saturation, and synchronization jitter, for the focused mlpâgb sensitivity panel at the 20 ms horizon. Table B.1: fl robustness to measurement degradation at 20 ms. The clean row reproduces the five-fold reference statistic. For each perturbed level, the five realizations are first averaged within each fold; values are then reported as mean ± standard deviation across the five fold-level means. Lower is better. Axis Level mlp gb Clean â 10.20 ± 0.25 14.78 ± 0.25 Noise 20 dB 10.48 ± 0.26 14.91 ± 0.24 Noise 10 dB 12.55 ± 0.35 15.75 ± 0.22 ct saturation c=0.5c=0.5 15.91 ± 0.24 17.44 ± 0.35 ct saturation c=0.3c=0.3 24.36 ± 0.67 20.56 ± 0.37 Jitter 2 samples 12.30 ± 1.35 15.05 ± 0.26 Jitter 4 samples 15.06 ± 2.94 15.28 ± 0.31 Table B.2: fc robustness to measurement degradation at 20 ms. The clean row reproduces the five-fold reference statistic. For each perturbed level, the five realizations are first averaged within each fold; values are then reported as mean ± standard deviation across the five fold-level means. Higher is better. Axis Level mlp gb Clean â 0.991 ± 0.001 0.745 ± 0.017 Noise 20 dB 0.989 ± 0.003 0.592 ± 0.024 Noise 10 dB 0.982 ± 0.011 0.274 ± 0.028 ct saturation c=0.5c=0.5 0.982 ± 0.002 0.682 ± 0.021 ct saturation c=0.3c=0.3 0.959 ± 0.003 0.534 ± 0.021 Jitter 4 samples 0.983 ± 0.006 0.682 ± 0.034 Appendix C Directional Holdout Across Fault-Resistance Ranges The reference and preceding sensitivity analyses preserve the same fault-resistance distribution between training and test data. To examine performance under a directional fault-condition holdout, the episode set is partitioned into disjoint training and test ranges according to fault resistance RfR_f. Two directional splits are evaluated at the 20 ms reference horizon: training on the lower 80% of the RfR_f range and testing on the upper 20%, and training on the upper 80% and testing on the lower 20%. The split is performed by episode, so no windows from the same episode occur in both partitions. Table C.1 reports the focused mlpâgb sensitivity panel under these directional holdouts. The ordinary episode-grouped five-fold result is included for context but is not a training-size-matched control. Table C.1: Directional fault-resistance holdout at 20 ms. Models are trained on the lower 80% and tested on the upper quintile of RfR_f, or vice versa. fc: macro-f1; fl: mae [% line length]. The episode-grouped five-fold result is shown for context and is not a training-size-matched control. Shifted mlp columns are mean ± standard deviation across five explicitly varied model seeds (0â4); the histogram-based gb columns report one run with seed 42. Task Model Five-fold reference High-RfR_f test Low-RfR_f test fc (macro-f1) mlp 0.991 0.975±0.0040.975± 0.004 0.979±0.0070.979± 0.007 fc (macro-f1) gb 0.745 0.731 0.725 fl (mae [%]) mlp 10.20 10.33±0.2210.33± 0.22 12.34±0.6312.34± 0.63 fl (mae [%]) gb 14.78 16.79 14.30 Under these directional fault-resistance holdouts, the representative models retain broadly similar performance, although the comparison with the five-fold reference is descriptive rather than training-size matched. The mlp classifier remains within 0.020.02 macro-f1 of its baseline, while the mlp locator changes only slightly on the high-resistance range (10.33±0.22%10.33± 0.22\%, compared with the 10.20%10.20\% five-fold in-distribution reference) and degrades moderately to 12.34±0.63%12.34± 0.63\% on the low-resistance range. Histogram-based gb localization is more sensitive to the shift toward higher resistance, with the mae increasing to 16.79%16.79\%, whereas its error decreases slightly on the low-resistance test range. The observed differences are direction- and model-dependent, showing that fault-resistance coverage is a distinct robustness axis that should be reported explicitly. This within-benchmark experiment is illustrative rather than exhaustive; shifts in topology, loading, fault families, and data source remain future instantiations under the same evaluation framework. Appendix D Class Distribution and Test-Set Sizes Table D.1 reports the class distribution of the 20 ms onset-conditioned classification windows. The target is dominated by the non-onset class (89.66%89.66\% of windows), while the rarest fault-type class accounts for 0.83%0.83\%, an imbalance of roughly 108:1108:1; this motivates macro-averaged f1 rather than accuracy as the classification metric. Each grouped five-fold test partition contains â52.3â52.3k windows over 1,8041,804â1,8051,805 episodes for fc (all windows), and â5.4â5.4k fault-onset windows over the same episodes for fl (fault-only). Table D.1: Class distribution of the 20 ms fault-classification windows (event_type; N=261,638N=261,638 windows over 9,0229,022 episodes). Class Type Windows Share Non-onset â 234 572 89.66% AG / BG / CG 1phâg 2295 / 2259 / 2283 0.88 / 0.86 / 0.87% AB / BC / CA 2ph 2175 / 2262 / 2316 0.83 / 0.86 / 0.89% ABG / BCG / CAG 2phâg 2175 / 2205 / 2277 0.83 / 0.84 / 0.87% ABC 3ph 6819 2.61% Total 261 638 100% Appendix E Pre-Trained Foundation-Model Baseline This appendix reports the pre-trained foundation-model baseline from the cross-family comparison in Section 4.10. MOMENT-1-large, a frozen time-series transformer with approximately 341 million parameters, is evaluated using either a linear probe or a two-layer mlp head on its per-channel embeddings. The fixed input-adaptation, embedding-aggregation, readout, optimization, and stopping settings are documented in the released experiment configuration. Both readouts follow the same episode-grouped five-fold protocol, sample-validity rules, decision horizons, and task-specific metrics as the main-text learned models, with fold-local standardization and independent regeneration of training and held-out embeddings in every outer fold. Table E.1 reports their results alongside the reference mlp and cnn1d. The nonlinear head outperforms the linear probe on both tasks, indicating that the frozen embeddings contain task-relevant information that is not fully linearly accessible. At 50 ms, the mlp head achieves the lowest mean fl error in the fixed cross-family comparison (8.90%8.90\%), while the main interpretation remains unchanged: classification approaches its metric ceiling, whereas localization retains material error across classical ml, task-specific deep learning, and foundation-model transfer. Table E.1: Pre-trained foundation-model (MOMENT-1-large) baseline under the shared episode-grouped five-fold protocol, with the main-text mlp and cnn1d repeated for reference. fc: macro-f1 (â ); fl: mae [% line length] (â ). Mean ± standard deviation over 5 folds; best mean per column in bold. Model fc 20 ms fc 50 ms fl 20 ms fl 50 ms mlp (reference) 0.991 ± 0.001 0.990 ± 0.005 10.20 ± 0.25 09.92 ± 0.27 cnn1d (from scratch) 0.999 ± 0.001 0.999 ± 0.001 09.65 ± 0.67 09.09 ± 0.28 MOMENT â linear probe 0.953 ± 0.005 0.975 ± 0.002 14.82 ± 0.10 14.99 ± 0.09 MOMENT â mlp head 0.985 ± 0.002 0.989 ± 0.001 10.59 ± 0.22 08.90 ± 0.12 Appendix F Detailed Stride Sensitivity Analysis Tables F.1 and F.2 report a compact stride-sensitivity check for the representative mlp and histogram-based gb models. The purpose is not stride optimization, but to assess whether a secondary windowing choice changes the interpretation of the reference evaluation setting. For fc, the mlp remains largely stable across the tested strides, whereas histogram-based gb is more sensitive at short and intermediate windows. In particular, its 20 ms performance improves by more than 0.22 macro-f1 at larger strides, showing that some apparent model differences can reflect configuration choices rather than only model-family capability. At 50 ms, both models remain close to baseline. For fl, the mlp is stable at 10 ms but degrades at larger strides for the 20 ms and 50 ms windows, while histogram-based gb shows mixed changes with no consistent gain from increasing stride. Overall, stride can affect absolute scores and relative model gaps in selected settings, and the large improvement of histogram-based gb at 20 ms shows that model-family comparisons can be configuration-dependent. However, this sensitivity does not overturn the task-level interpretation of the case study: classification continues to approach its metric ceiling while localization retains material error, and the broader conclusions remain shaped primarily by task definition, timing, and observability. Table F.1: Compact stride sensitivity for fault classification. Baseline performance is reported at the default 5 ms stride; additional columns show Î macro-f1 relative to that baseline for larger strides. Positive values indicate improvement. Model Window Baseline f1 Î @ 10 ms Î @ 20 ms Î @ 50 ms mlp 10 ms 0.985 +0.003 â â mlp 20 ms 0.991 -0.008 -0.010 â mlp 50 ms 0.990 -0.004 -0.002 -0.003 gb 10 ms 0.418 +0.100 â â gb 20 ms 0.745 +0.221 +0.225 â gb 50 ms 0.982 -0.005 -0.003 +0.004 Table F.2: Compact stride sensitivity for fault localization. Baseline performance is reported at the default 5 ms stride; additional columns show Î relative to that baseline for larger strides. Negative values indicate improvement. Model Window Baseline mae Î @ 10 ms Î @ 20 ms Î @ 50 ms mlp 10 ms 10.64 +0.00 â â mlp 20 ms 10.20 +0.93 +0.13 â mlp 50 ms 9.92 +1.10 +1.24 +1.44 gb 10 ms 14.65 +0.00 â â gb 20 ms 14.78 +0.28 -0.36 â gb 50 ms 14.66 +0.49 +0.65 -0.20 Appendix G Detailed Hyperparameter Analysis Results Table G.1: Fixed configurations of the classical learning models used in the reference evaluation. Parameters not listed retain the scikit-learn 1.8.0 defaults. The same settings are used for the corresponding classifier and regressor unless stated otherwise. Model Fixed configuration Ridge Regularization strength α=1.0α=1.0; automatic solver selection; pseudorandom seed 42. RidgeClassifier is used for fc and Ridge for fl. knn Five neighbors; uniform weighting; Minkowski distance with p=2p=2; automatic neighbor-search algorithm; parallel prediction enabled. Histogram-based gb 100 boosting iterations; learning rate 0.1; unrestricted tree depth; minimum 20 samples per leaf; no L2 regularization; pseudorandom seed 42. mlp One hidden layer with 100 rectified-linear units; Adam optimizer; L2 penalty 10â410^-4; initial learning rate 10â310^-3; automatic batch size; maximum 200 iterations; no early stopping; pseudorandom seed 42. This appendix reports the single-parameter hyperparameter ablations underlying the robustness summary in the main text. The benchmark results reported in the main analysis use the tagged default configurations; the ablations are diagnostic sensitivity checks and are not used as a hyperparameter-selection procedure for the reported models. In each ablation, one hyperparameter group is varied while all remaining settings are held fixed at the tagged default baseline. Table G.2 defines the evaluated ablation design, and Table G.3 summarizes the resulting task- and window-specific sensitivity. The compact robustness analyses reported here are intended as diagnostic sensitivity checks; broader robustness stress testing of ML-based protection models is treated separately in [47]. Table G.2: Single-parameter hyperparameter ablation design. Model Hyperparameter Default Tested values gb Learning rate 0.1 0.03, 0.2 gb Maximum tree depth unlimited 3, 5, 10 gb Boosting iterations 100 50, 300 gb Minimum samples per leaf 20 5, 50 gb L2 regularization 0 10â410^-4, 10â210^-2 mlp Hidden layers (100) (50), (100, 50), (256, 128) mlp L2 penalty 10â410^-4 10â510^-5, 10â310^-3 mlp Initial learning rate 10â310^-3 10â510^-5, 10â410^-4, 10â210^-2 mlp Batch size auto 64, 128, 256 mlp Training iterations 200 100, 300, 400 Table G.3: Condensed hyperparameter robustness summary across single-parameter ablations. For fc, Îbest _best denotes the increase in macro-f1; for fl, it denotes the reduction in mae. Positive values therefore indicate improvement in both tasks. The maximum spread reports the largest bestâworst difference within any single-parameter ablation group. The default cells are independent results from the ablation harness and can differ slightly from the Table 4 reference; for example, fl mlp at 20 ms is 9.9989.998 here versus 10.2010.20 in the reference evaluation. Comparisons within each ablation group therefore use the corresponding default generated by the same harness. This difference is not a repeated-seed uncertainty estimate. Task Model Window Default Best metric Îbest _best Max. spread fc gb 20 ms 0.745 0.967 +0.222 (min leaf) 0.327 (learn. rate) fc gb 50 ms 0.982 0.988 +0.006 (iters) 0.116 (learn. rate) fc mlp 20 ms 0.990 0.996 +0.006 (init. lr) 0.015 (init. lr) fc mlp 50 ms 0.991 0.994 +0.003 (init. lr) 0.011 (init. lr) fl gb 20 ms 14.782 13.232 +1.550 (iters) 5.530 (depth) fl gb 50 ms 14.662 12.942 +1.720 (iters) 5.671 (depth) fl mlp 20 ms 9.998 9.322 +0.675 (hidden) 6.209 (init. lr) fl mlp 50 ms 9.955 8.296 +1.659 (hidden) 3.694 (init. lr) Table G.4: Compact hyperparameter robustness summary for the representative mlp and histogram-based gb models based on the single-parameter ablations reported in this appendix. Ablation default refers to the best tagged default result within the independent ablation campaign across the reported decision horizons. Best denotes the strongest result observed in the corresponding ablation setting. For fc, Îbest _best denotes the increase in macro-f1; for fl, it denotes the reduction in mae. Positive values therefore indicate improvement in both tasks. Model Task Ablation default Best Îbest _best Max. spread mlp fc 0.991 0.996 +0.005 0.015 gb fc 0.982 0.988 +0.006 0.327 mlp fl 9.955 8.296 +1.659 6.209 gb fl 14.662 12.942 +1.720 5.671