Paper deep dive
Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift
Jayakumar Manoharan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/16/2026, 3:57:50 AM
Summary
The paper introduces Audited Selective Verification (ASV-N1), a risk-budgeted screening method for real-time N-1 thermal contingency analysis in energy management systems. It addresses the computational and assurance trade-off by combining a cheap linear surrogate for triage, an online audit via full AC power flow on a random sample, and a Learn-Then-Test calibrated threshold to certify violation-rate bounds. The method provides distribution-free guarantees robust to arbitrary deployment shifts caused by advanced controllers, outperforming deterministic screens on IEEE and PEGASE test systems while reducing full power-flow computations by 29–75%.
Entities (8)
Relation Signals (8)
Audited Selective Verification (ASV-N1) → evaluatedon → PEGASE 1354-bus System
confidence 96% · fails on IEEE 118-bus and the 1354-bus PEGASE system
Audited Selective Verification (ASV-N1) → evaluatedon → IEEE 300-bus System
confidence 96% · On three public transmission systems up to 1354 buses... IEEE 300-bus
Audited Selective Verification (ASV-N1) → solves → N-1 Thermal Contingency Screening
confidence 95% · This paper introduces Audited Selective Verification, a risk-budgeted screening and triage layer for any controller's output
Audited Selective Verification (ASV-N1) → robustto → Deployment Shift
confidence 94% · Validity rests on real verification and the audit rather than on surrogate accuracy, so it holds under arbitrary deployment shift.
Audited Selective Verification (ASV-N1) → uses → Full AC Power Flow
confidence 93% · an online audit runs full power flow on a small random sample each window
Linear Sensitivity Screening → failsunder → Deployment Shift
confidence 92% · fast linear-sensitivity screening gives no statistical guarantee and can silently pass unsafe operating points, especially when a controller drives the system into unfamiliar regimes.
Audited Selective Verification (ASV-N1) → uses → Linear Sensitivity Screening
confidence 91% · A cheap surrogate proposes which outages to skip
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Real-time N-1 contingency screening in an energy management system trades assurance against cost: verifying every credible outage with full power flow is too slow, while fast linear-sensitivity screening gives no statistical guarantee and can silently pass unsafe operating points, especially when a controller drives the system into unfamiliar regimes. This paper introduces Audited Selective Verification, a risk-budgeted screening and triage layer for any controller's output (optimization, model-predictive, or learned). A cheap surrogate proposes which outages to skip; an online audit runs full power flow on a small random sample each window; and a calibrated threshold certifies a thermal-violation-rate bound for the skipped set at a chosen budget and confidence, with a corresponding bound for the unverified trusted subset. Validity rests on real verification and the audit rather than on surrogate accuracy, so it holds under arbitrary deployment shift. It is a risk-budgeted screen, not a replacement for deterministic verification when policy requires checking every credible contingency. On three public transmission systems up to 1354 buses, the realized violation rate stays within budget, standard deterministic and calibrated screens become unsafe under shift, and the method cuts full power-flow studies by 29 to 75 percent per real-time operating point.
Tags
Links
- Source: https://arxiv.org/abs/2607.13221v1
- Canonical: https://arxiv.org/abs/2607.13221v1
Trouble viewing inline? Open PDF directly →
Full Text
63,623 characters extracted from source content.
Expand or collapse full text
Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift Jayakumar Manoharan J. Manoharan is with the Electric Power Research Institute (EPRI), Charlotte, NC, USA (e-mail: jmanoharan@epri.com; ORCID: 0009-0009-7765-9165). Abstract Real-time N-1 contingency screening in an energy management system trades assurance against cost: verifying every credible outage with full power flow is too slow, while fast linear-sensitivity screening gives no statistical guarantee and can silently pass unsafe operating points, especially when a controller drives the system into unfamiliar regimes. This paper introduces Audited Selective Verification, a risk-budgeted screening and triage layer for any controller’s output (optimization, model-predictive, or learned). A cheap surrogate proposes which outages to skip; an online audit runs full power flow on a small random sample each window; and a calibrated threshold certifies a thermal-violation-rate bound for the skipped set at a chosen budget and confidence, with a corresponding bound for the unverified trusted subset. Validity rests on real verification and the audit rather than on surrogate accuracy, so it holds under arbitrary deployment shift. It is a risk-budgeted screen, not a replacement for deterministic verification when policy requires checking every credible contingency. On three public transmission systems up to 1354 buses, the realized violation rate stays within budget, standard deterministic and calibrated screens become unsafe under shift, and the method cuts full power-flow studies by 29 to 75 percent per real-time operating point. I Introduction Keeping the transmission system secure against any single credible outage, the N-1 criterion, is a core task of real-time operation, and an energy management system must reassess it continuously as operating points change [1]. Verifying it exactly means a full AC power flow for every credible contingency, too slow to repeat after every candidate action, so operators rely on fast screening. This assurance requirement also gates adoption of the faster, more adaptive controllers now entering the control room, reinforcement-learning agents for topology control and model-predictive schemes for redispatch [2, 3, 4]: no operator will hand control to a policy without evidence that its actions respect N-1 security. Providing that evidence cheaply is hard. The standard fast alternative is deterministic screening with linear sensitivities (power transfer and line outage distribution factors), which ranks contingencies and skips those it predicts to be safe [1]. Screening is fast, but it offers no statistical guarantee: it trusts a linear approximation, and when that approximation is wrong it silently passes unsafe actions. The problem is sharpest precisely when a learned controller is deployed, because the controller changes the distribution of operating points away from the historical operator, so any safety estimate calibrated on past data can become invalid at deployment. This is a feedback-induced distribution shift, and it breaks the exchangeability that distribution-free calibration normally relies on. This paper asks: can we provide a distribution-free finite-sample guarantee that risk-controls a black-box controller’s N-1 thermal exposure, cheaply, and robustly to the controller’s own shift? We answer yes, with a method we call Audited Selective Verification (ASV-N1). The key idea is to stop trusting the cheap surrogate for safety and instead use it only to save work. In each control window, ASV-N1 (i) scores every credible contingency with a cheap linear surrogate, (i) draws a small random audit and runs full AC power flow on it, and (i) uses the audited (score, violation) pairs to calibrate, by a Learn-Then-Test risk-control procedure [5, 6], the largest surrogate threshold below which the skipped contingencies are certified to violate with probability at most a budget. Contingencies above the threshold are verified by full AC; those below it, outside the audit, are trusted. Because the certificate rests on real verification and an audit drawn from the deployment distribution, its validity does not depend on whether the surrogate is accurate, and it holds under arbitrary shift. A poor surrogate does not make the certificate unsafe; it only forces more contingencies to be verified, raising compute cost. We first show why simpler designs do not suffice. A static threshold certificate, and even an adaptive one based on adaptive conformal inference [7], control the violation rate only when the surrogate separates safe from unsafe contingencies in the operating regime. On three public systems this holds on IEEE 300-bus but fails on IEEE 118-bus and the 1354-bus PEGASE system, where under stress the surrogate labels as safe many contingencies that in fact violate. This negative result motivates verification-backed validity. Contributions. 1. We introduce a distribution-free, per-window risk-controlled N-1 thermal screening certificate for arbitrary black-box grid controllers, formulated as a transductive instance of Learn-Then-Test risk control [5, 6], whose validity does not depend on surrogate accuracy or on distribution stationarity, and is therefore robust to controller-induced feedback shift (Sections V and VI). It controls a violation-rate budget over the screened contingencies, not worst-case N-1 security for every individual skipped outage. To our knowledge this is the first such certificate for N-1 thermal screening, although the underlying risk-control machinery is standard. 2. We give an online-audited selective-verification mechanism that combines cheap-surrogate triage with Learn-Then-Test calibration of the skip threshold, plus an adaptive audit-sizing rule that minimizes total AC-solve cost (Section V). 3. We establish a negative result that motivates the design: static and adaptive threshold certificates on a cheap surrogate are valid only under a surrogate-stability condition that fails on real systems, and deterministic screening is unsafe under shift (Sections IV and VII). 4. We provide an open, reproducible evaluation on public systems (IEEE 118-bus and 300-bus, PEGASE 1354-bus, with the pandapower solver [8]) reporting the safety-versus-cost tradeoff under both load-induced and genuine controller-induced shift (Section VII). The method certifies the action of any controller, learned or not; learned controllers are our motivation, and we demonstrate robustness to a genuine controller-induced shift using an optimization-based redispatch policy, leaving an end-to-end study with a trained learning-based controller to future work. In operational terms, ASV-N1 is an online transmission security-assessment and risk-based contingency-screening layer for bulk power systems: it complements security-constrained optimal power flow and online contingency analysis and does not replace deterministic verification where policy mandates it. The strongest single result is that ASV-N1 stays within budget on all three systems while a deterministic screen is unsafe under shift, while skipping the majority of full power-flow solves where the surrogate is reliable. I Related Work Contingency screening and security assessment Fast N-1 screening is a mature part of operations. Linear sensitivity factors (power transfer distribution factors and line outage distribution factors) estimate post-contingency flows in closed form, and screening rules rank and skip contingencies whose estimated loading is well below limits, reserving full AC analysis for the rest [1, 9, 10]. Security-constrained optimal power flow (SCOPF) embeds N-1 constraints directly into the dispatch optimization [11, 12], and recent work studies adversarially robust learning of such constraints [13]. Unlike SCOPF, which re-optimizes the dispatch to satisfy worst-case N-1 constraints, ASV-N1 does not change the dispatch: it gives an explicit finite-sample statistical bound on the violation rate among the contingencies it skips, per control window, for whatever action the controller proposes, and can therefore wrap an SCOPF, an RL policy, or any other controller as a verification-backed overlay. These methods are fast and widely used, but they are deterministic: they assume a fixed operating point and provide no probabilistic guarantee that a skipped contingency is actually safe. We show that under deployment shift this assumption fails and a deterministic screen can skip contingencies that violate. ASV-N1 keeps the same linear surrogate for efficiency but replaces the deterministic skip rule with an audited, distribution-free certificate. Beyond deterministic screening, risk-based security assessment quantifies operational risk across contingencies for control-room decisions [14, 15], but these methods estimate expected risk under assumed contingency probabilities rather than bounding the screened violation rate. Unlike prior risk-based security assessment, ASV-N1 provides finite-sample, per-window distribution-free control of the screened violation rate using online AC-backed auditing, valid under arbitrary deployment shift. Distribution-free uncertainty and risk control Conformal prediction and its risk-control extensions provide finite-sample, distribution-free guarantees for black-box predictors [16]. Risk-controlling prediction sets and the Learn-Then-Test framework calibrate a decision threshold so that a monotone risk is controlled with high probability [6, 5], and conformal risk control extends this to general monotone losses [17]. Standard guarantees assume exchangeability; weighted conformal prediction handles covariate shift when the likelihood ratio is known [18], adaptive conformal inference maintains long-run coverage under arbitrary shift by online threshold updates [7], and analyses beyond exchangeability bound the coverage loss under distribution drift [19]. Crucially, these guarantees either assume exchangeability between calibration and deployment, or they remain valid only under a shift that is explicitly modeled (for example a known or estimated likelihood ratio) [18, 7, 19]. Recent work brings conformal ideas closer to control and to power systems: conformal policy control calibrates how far a deployed policy may deviate from a reference [20], decision-calibrated prediction sets calibrate forecast uncertainty so a robust optimization stays feasible [21], and, in robotics, adaptive conformal prediction has been combined with control-barrier functions to enforce state constraints during safe reinforcement learning [22]. All of these either presume exchangeability or a modeled shift, or they trust the cheap predictor whose error they calibrate. ASV-N1 differs on exactly this point: it is a transductive instance of Learn-Then-Test in which the calibration set is an online audit drawn from the deployment distribution itself, so its validity is robust to arbitrary controller-induced feedback shift and does not rely on surrogate accuracy. We further show, as a negative result, that directly applying a static or adaptive threshold certificate to a cheap surrogate score is insufficient for N-1 safety when the surrogate is unreliable under stress. Safe learning for power systems Safe reinforcement learning for power systems has grown rapidly [4]. Shielding constrains a learning agent to a precomputed safe action set [23], and decision-calibrated prediction sets calibrate forecast uncertainty so that a downstream robust optimization remains feasible [21]. These approaches either modify the learning process, assume an exact shield, or certify uncertainty in inputs rather than the security of a chosen action. ASV-N1 differs in three ways: it certifies the action of any frozen black-box controller after the fact, its guarantee is backed by real AC verification rather than by an assumed-exact model, and it targets per-action N-1 security under controller-induced shift. Benchmarks for learning-based grid control are also emerging [24, 25, 26]; our open evaluation of the safety-versus-cost tradeoff complements them by scoring safety certificates rather than control performance. I Problem Setup I-A N-1 thermal security of a control action Consider a transmission network operated in discrete control windows. A window is one control decision: the controller commits to an operating point, and that point holds until the next decision. The natural window is therefore a single operating point. As we show in Section VII, a single operating point yields a safe but conservative certificate, because the audit is small; batching the credible contingencies of several consecutive operating points, or monitoring a wider credible set, enlarges the per-window sample and tightens the certificate at the cost of a coarser time resolution. We treat a controller as a black box: it may be a reinforcement learning policy, a model predictive controller, or any other decision rule producing an operating point, that is, a setting of generation, load, and topology. Let =1,…,NC=\1,…,N\ index the credible single-line N-1 contingencies for that operating point. For contingency i, let Vi∈0,1V_i∈\0,1\ be the true post-contingency thermal-violation label, with Vi=1V_i=1 if and only if the full AC power flow after outage i overloads some monitored line beyond its rating. Computing ViV_i requires a full AC solve, which is the cost we wish to avoid. A cheap surrogate assigns each contingency a danger score ri∈ℝr_i , with larger rir_i meaning more dangerous. We use a standard linear screen: the post-contingency line loadings are estimated from the AC base-case flows redistributed by the line outage distribution factors, and rir_i is the worst estimated overload over monitored lines. The score is available for all contingencies without any contingency AC solve. We make no assumption that the surrogate is accurate. I-B The certification problem Fix a safety budget α∈(0,1)α∈(0,1) and a confidence level 1−δ1-δ. The operator wants to act on the controller’s operating point only if the contingencies it does not verify are, with high confidence, safe to skip. Concretely, we seek a decision rule that partitions C into a verified set (on which full AC is run) and a trusted set (skipped, no AC), such that the violation rate among the trusted set is at most α with probability at least 1−δ1-δ, while keeping the verified set, hence the compute cost, small. The challenge is that the joint distribution of (ri,Vi)(r_i,V_i) at deployment is unknown and, because the controller shapes the operating point, generally differs from any historical distribution. We therefore require a guarantee that holds for an arbitrary deployment distribution and does not presume the surrogate is reliable. IV Why Threshold Certificates Fail A natural first design is a threshold certificate: calibrate a single skip threshold on the surrogate score using historical (score, violation) pairs, then trust every deployed contingency whose score falls below it. We summarize why this design, and even its adaptive refinement, does not give an N-1 safety guarantee under deployment shift. This negative result motivates the audited method of Section V. Static threshold certificate Calibrate τ as the largest threshold whose historical skipped-set violation rate is at most α, then skip i:ri≤τ\i:r_i≤τ\ at deployment. This is exactly a split-conformal (risk-controlling) certificate calibrated on a historical holdout [6, 17], the natural distribution-free baseline; it is valid under exchangeability between calibration and deployment. This controls the violation rate at deployment only if the score-conditional violation probability is stable between calibration and deployment. When a learned controller drives the system into a more stressed regime, the same score corresponds to a higher true violation probability, and the skipped set violates above α. We observe exactly this: across our systems the static certificate is valid on IEEE 300-bus but its deployed skipped-set violation reaches 0.50.5 to 0.860.86 on IEEE 118-bus and PEGASE, far above an α=0.15α=0.15 budget. Adaptive threshold certificate Adaptive conformal inference [7] updates the threshold online from realized outcomes to maintain a long-run rate. We implement this for the skip threshold. It improves matters but does not fix them: when the surrogate labels truly violating contingencies as safe, no threshold on the score can exclude them, so the achievable violation rate is floored by the violation rate in the surrogate’s safest score bin. On PEGASE, in the stressed regime, 9797 percent of contingencies that all violate receive a score indicating no predicted overload; the adaptive certificate therefore cannot bring the violation rate below roughly 0.50.5. Figure 1 shows the cumulative skipped-set violation rate of the static and adaptive certificates as the deployment drifts: both exceed the budget on the sharp-transition systems, while IEEE 300-bus, where the surrogate separates violations, stays safe. Figure 1: Threshold certificates fail under shift. Cumulative skipped-set violation rate as the deployment drifts (load multiplier) for the static and adaptive (ACI) threshold certificates. Both exceed the budget α=0.15α=0.15 on IEEE 118-bus and PEGASE, where the cheap surrogate mislabels stressed contingencies as safe; only IEEE 300-bus, where the surrogate separates violations, stays safe. This motivates verification-backed validity. A weighted-conformal certificate [18] could in principle correct for the shift, but only if the deployment-to-history likelihood ratio is known or estimated, which under controller-induced shift is exactly the unknown quantity; the audit in Section V sidesteps this by calibrating on a fresh on-distribution sample instead of modeling the shift. The lesson is that any certificate whose safety rests on the cheap surrogate inherits the surrogate’s blind spots. ASV-N1 instead makes safety rest on real verification, using the surrogate only to decide what to verify. V Audited Selective Verification ASV-N1 wraps a controller and certifies each control window independently. The mechanism has three parts: a cheap triage, an online audit, and a calibrated skip threshold. Algorithm 1 states it. V-A Triage, audit, and calibration Given the window’s contingency set C with scores ri\r_i\, ASV-N1 draws an audit A⊆A that includes each index independently with probability π (a fixed-size uniform sample behaves identically). Crucially, A is drawn from indices only, independently of the scores and labels. Full AC power flow is run on A, revealing Vi:i∈A\V_i:i∈ A\. For a candidate threshold τ, let the skip set be Skip(τ)=i∈:ri≤τSkip(τ)=\i :r_i≤τ\, a set fixed by the scores. Let Uδ(k,n)U_δ(k,n) denote the Clopper-Pearson upper (1−δ)(1-δ) confidence limit for a binomial rate from k violations in n trials. Order the distinct candidate thresholds τ(1)<⋯<τ(M)τ^(1)<…<τ^(M) (a data-independent order, strictest skip set first). Using only the audited contingencies, ASV-N1 tests these thresholds in sequence and accepts τ(j)τ^(j) when Uδ(∑i∈A,ri≤τ(j)Vi,|A∩Skip(τ(j))|)≤αU_δ( _i∈ A,r_i≤τ^(j)V_i,\ |A (τ^(j))|)≤α. It stops at the first threshold that is not accepted and sets τ⋆τ to the largest threshold in the accepted prefix: τ⋆=τ(j⋆),j⋆=maxj:τ(1),…,τ(j) all accepted.τ =τ^(j ), j = \\,j:$τ^(1),…,τ^(j)$ all accepted\,\. (1) This fixed-sequence (prefix) rule, rather than a global maximum over accepted thresholds, is what controls multiplicity and makes the guarantee of Section VI valid even when the audited confidence limit is not monotone in τ; it may be conservative if the accepted set has gaps. ASV-N1 then verifies (full AC) every contingency with ri>τ⋆r_i>τ ; among ri≤τ⋆r_i≤τ , the audited ones are already solved, and the remaining trusted contingencies are skipped. Algorithm 1 ASV-N1 (one control window) Input: contingencies C, scores ri\r_i\, budget α, confidence 1−δ1-δ. 1. Draw audit A⊆A (label-independent random sample); run AC on A to obtain Vi:i∈A\V_i:i∈ A\. 2. Select τ⋆τ by (1) via fixed-sequence testing over the ordered thresholds. 3. Verify (full AC) all i with ri>τ⋆r_i>τ . 4. Trust (skip) all i∉Ai∉ A with ri≤τ⋆r_i≤τ . Output: verified contingencies are AC-checked exactly; the trusted skipped contingencies are risk-controlled at the certified budget (skip-set rate at most α, hence trusted-set rate at most α/(1−f)α/(1-f), with confidence 1−δ1-δ). V-B Adaptive audit sizing The per-window cost is the audit size plus the number of verified contingencies. A larger audit tightens the Clopper-Pearson bound in (1), which can raise τ⋆τ and shrink the verified set, so there is a cost-optimal audit size. Two variants preserve the guarantee. (i) Score-only sizing: choose the audit size as a function of the score distribution alone, for example to place a target number of audited contingencies in the candidate skip region. Because this uses only scores, not labels, Assumption 1 holds and Theorem 1 applies directly. (i) Cost-optimal sizing: grow a fixed grid of G nested audit prefixes and, for each, predict the total AC count (audit plus verify), then select the size minimizing the predicted total. Because this selection peeks at audited labels, validity is restored by a union bound over the grid, replacing δ by δ/Gδ/G in (1); since G is small (a handful of sizes), the resulting widening of the confidence interval is mild. All experiments in this paper use variant (i), cost-optimal sizing with the δ/Gδ/G union-bound correction and G=8G=8 candidate sizes (Proposition 1); variant (i) is stated for completeness. V-C Robustness to controller-induced shift Because the audit is drawn from the current window, the calibration in (1) uses deployment-distribution data. Re-running it every window means the certificate never extrapolates from a stale historical distribution. This is precisely what makes the guarantee robust to the shift the controller induces: the audit observes the controller’s own operating points. The surrogate enters only through τ⋆τ and therefore affects cost, not validity; a useless surrogate yields τ⋆τ that verifies everything, which is expensive but still safe. VI Guarantee We state the per-window safety guarantee. The setting is transductive and distribution-free: we treat each control window as fixed and randomize only the audit, so the guarantee is per window and exact, not asymptotic over time. Here the risk is simply the fraction of skipped contingencies that violate their thermal limits. Assumption 1 (Label-independent audit). The audit is a uniform random sample of the contingency set drawn independently of the violation labels Vi\V_i\. Its size may be fixed in advance or chosen from label-independent information (including the scores ri\r_i\, as in Proposition 1); conditional on the selected size, the audited contingencies form a uniform, label-independent sample. Assumption 2 (Trusted oracle and thermal scope). The AC N-1 power flow is the trusted verifier defining ViV_i, and ViV_i encodes thermal-limit violations under the AC model. Voltage and reactive limits are outside the present statement and are reported separately as a diagnostic. Order the distinct candidate thresholds τ(1)<⋯<τ(M)τ^(1)<…<τ^(M) (this order depends on the scores, which are fixed, not on the labels). Let R(τ)=|Skip(τ)|−1∑i∈Skip(τ)ViR(τ)=|Skip(τ)|^-1 _i (τ)V_i be the true violation rate of the whole skip set. Lemma 1 (Valid audited p-value). Fix τ and condition on the audited skip count n=|A∩Skip(τ)|n=|A (τ)|. Under the null H:R(τ)>αH:R(τ)>α, the audited violation count K is the number of violations in a label-independent size-n sample drawn without replacement from Skip(τ)Skip(τ), hence hypergeometric with mean nR(τ)>nαn\,R(τ)>nα. Consider the left-tail statistic P=Pr[Bin(n,α)≤K]P= [Bin(n,α)≤ K] (equivalently, Uδ(K,n)≤αU_δ(K,n)≤α). The hypergeometric lower tail is dominated by the binomial lower tail at the same mean [27, 28], and since R(τ)>αR(τ)>α places more mass on large K than Bin(n,α)Bin(n,α) does, small values of P occur with at most their nominal probability: Pr[P≤δ∣n]≤δ [P≤δ n]≤δ. Marginalizing over n preserves this, so P is a valid (super-uniform) p-value, and the Clopper-Pearson rule is, if anything, conservative for the without-replacement audit. Theorem 1 (Per-window distribution-free safety). Under Assumptions 1 and 2, with probability at least 1−δ1-δ over the audit, the threshold τ⋆τ selected by (1) with fixed-sequence testing satisfies R(τ⋆)≤α.R(τ )\ ≤\ α. (2) This holds for any fixed (ri,Vi)\(r_i,V_i)\ (arbitrary deployment distribution, including controller-induced shift) and any surrogate, with no assumption on surrogate accuracy or cross-window stationarity. Proof sketch. The candidate-threshold order is data-independent. Fixed-sequence testing that stops at the first non-rejection, using the per-hypothesis level-δ valid p-values of Lemma 1, controls the family-wise error rate at δ under arbitrary dependence [29, 5]. Hence with probability at least 1−δ1-δ every rejected hypothesis is truly false, so R(τ(j))≤αR(τ^(j))≤α for each selected j, and in particular R(τ⋆)≤αR(τ )≤α. This is the transductive Learn-Then-Test instance with the audit as calibration and the fixed skip population as the risk object [5, 6]. A full proof is in Appendix A. ∎ Corollary 1 (Trusted-set bound, deterministic transfer). On the event R(τ⋆)≤αR(τ )≤α of Theorem 1, the trusted set T=i∉A:ri≤τ⋆T=\i∉ A:r_i≤τ \ satisfies, deterministically, 1|T|∑i∈TVi≤R(τ⋆)|Skip(τ⋆)||T|≤α1−f,f=|A∩Skip(τ⋆)||Skip(τ⋆)|, split 1|T| _i∈ TV_i\ &≤\ R(τ )\,|Skip(τ )||T|\ ≤\ α1-f,\\ &f= |A (τ )||Skip(τ )|, split (3) where f is the audited fraction of the skip set (assuming f<1f<1, that is, a nonempty trusted set). This holds with probability at least 1−δ1-δ (the same event). It requires no assumption that T is a clean random subsample after the data-dependent selection of τ⋆τ : the violation count in T is at most the violation count in the whole set Skip(τ⋆)Skip(τ ), which R(τ⋆)≤αR(τ )≤α bounds. We therefore treat the skip-set rate R(τ⋆)R(τ ) as the primary certified quantity; the trusted (unverified) rate is at most α/(1−f)α/(1-f), and with the audit fractions used here (f near 0.150.15 to 0.20.2) this inflation is small. Audited contingencies, being verified, carry no residual risk. Proposition 1 (Cost-optimal audit sizing). Fix a data-independent grid of G candidate audit sizes (nested prefixes of a label-independent sampling order). If the audit size is chosen from this grid by any rule that may depend on the audited labels, then running the selection of (1) at level δ/Gδ/G for each candidate and reporting the chosen one preserves Theorem 1 at level 1−δ1-δ, by a union bound over the G candidates. Our experiments use G=8G=8. Corollary 2 (Long-run budget). Over H windows, applying Theorem 1 at confidence 1−δ/H1-δ/H guarantees every window’s trusted set is safe at level α simultaneously with probability at least 1−δ1-δ. A time-uniform confidence sequence [30] avoids the 1/H1/H factor and handles adaptively chosen windows. No cross-window stationarity is assumed. Skip set versus trusted set Theorem 1 controls the violation rate of the entire skip set Skip(τ⋆)Skip(τ ), which includes the audited contingencies that ASV-N1 has already verified by full AC. The quantity an operator acts on, and the quantity we report in Section VII, is the violation rate of the trusted set T of skipped contingencies that are not verified. Corollary 1 transfers the skip-set guarantee to exactly this trusted set, so the reported trusted-set rate is the certified quantity; the audited contingencies carry no residual risk because their true labels are known. Operational reading In practice, Theorem 1 gives the operator a tunable safety budget with a confidence level. Setting α=0.05α=0.05 and 1−δ=0.951-δ=0.95 means: in each control window, among the full set of contingencies the tool deems skippable, at most five percent are at risk of a post-contingency thermal violation; among the unverified subset it actually trusts without a check, the bound is the slightly looser α/(1−f)α/(1-f), about six percent at the audit fractions used here; and these statements are correct at least ninety-five percent of the time. The operator never trusts the cheap surrogate blindly: every contingency the surrogate cannot vouch for, at the certified threshold, is verified by a full power-flow study, and a small random sample of the rest is verified as an ongoing check. If the surrogate degrades (for example because a new controller drives the system into an unfamiliar regime), the audit detects it and the tool automatically verifies more contingencies, trading compute for the same safety budget rather than silently becoming unsafe. Proposition 2 (Cost and the role of the surrogate). The per-window AC-solve count is |A∪i:ri>τ⋆||A∪\i:r_i>τ \|, that is |A|+|i∉A:ri>τ⋆||A|+|\i∉ A:r_i>τ \| (the audited contingencies above the threshold are counted once, not twice); the AC-solve fraction is this divided by N. Validity is independent of the surrogate; the surrogate enters only through τ⋆τ . A surrogate that ranks violations well admits a larger safe τ⋆τ and lower cost; a useless surrogate yields verify-all, which is maximal cost but still valid. VII Experiments VII-A Setup We evaluate on three public transmission systems using the open-source pandapower solver [8]: IEEE 118-bus, IEEE 300-bus, and the 1354-bus PEGASE system. For each system we equip lines with thermal ratings so that the network is N-1 secure at nominal load with a headroom margin, then create operating points by scaling load with heterogeneous noise; generation follows load only partially, producing congestion. The aim is not to reproduce a specific utility operating condition but to test certificate validity and cost under controlled, reproducible stress regimes on public networks. For every operating point and every credible single-line contingency we compute the cheap surrogate score r (line outage distribution factors applied to the AC base-case flows) and the true label V (a full AC contingency solve). The label encodes thermal violations only: a contingency is labeled violating if its converged AC solution overloads a monitored line. The credible N-1 set excludes single-line outages that island the network, which are non-credible for thermal screening and appear as non-convergent solutions (islanding is detected as failure of the post-contingency AC power flow to converge, that is, the outage disconnects part of the network and is treated as a load-loss event rather than a thermal-screening case); on IEEE 300-bus the per-contingency islanding (non-convergence) rate is about 55 percent. We treat such cases as outside the thermal scope of Assumption 2; a sensitivity test that instead counts every non-convergent case as a violation leaves the certificate within budget (trusted-set violation 0.0000.000 on IEEE 300-bus). For IEEE 118-bus we monitor all 173173 single-line contingencies, a complete N-1 sweep; for IEEE 300-bus we monitor the full credible set of about 268268 of 283283 single-line outages per operating point (about 1515 island) across 100100 operating points whose base case converges; for the 1354-bus PEGASE system we monitor all 17511751 single-line contingencies; all but about two per operating point are credible (almost none island), so this too is a near-complete N-1 sweep, at scale. In operations the credible contingency set is defined by the utility; ASV-N1 certifies the violation rate restricted to whatever set is monitored. The certified result on IEEE 300-bus is the same under the full set and a smaller random sample, confirming the evaluation is not a sampling artifact. Unless stated otherwise the budget is α=0.15α=0.15 and confidence 1−δ=0.91-δ=0.9. The formal certificate (Theorem 1) controls the skip-set violation rate at α; the unverified trusted subset is bounded by α/(1−f)α/(1-f) (Corollary 1). We report the realized (empirical) violation rate among trusted (skipped, unverified) contingencies, which is well within budget, and the AC-solve fraction relative to a full N-1 sweep, where lower is cheaper. Implementation details: the contingency sets are as above; the audit is a label-independent uniform sample whose size is chosen by the adaptive rule of Section V; the budget is swept over α∈0.05,0.10,0.15,0.20α∈\0.05,0.10,0.15,0.20\ and the confidence is 1−δ=0.91-δ=0.9; all random seeds are fixed (data generation and audit). All systems, the pandapower solver [8], and the contingency definitions are public; an anonymized code package that reproduces every table and figure accompanies this submission for review, and a permanent public repository will host it on publication. Experiments ran on a workstation with a 20-core CPU, 121 GB of RAM, and an NVIDIA GB10 GPU; the pandapower AC power flows are CPU Newton-Raphson solves (the GPU was not used for the power-flow oracle). VII-B Validity and cost across systems Table I reports ASV-N1 with adaptive audit sizing at two window granularities. On all three systems the realized trusted-set violation rate is well within the budget, consistent with Theorem 1, in both settings. At the single operating point, which is the real-time control window, ASV-N1 saves 2929 to 7575 percent of the full N-1 solves (AC-solve fraction 0.250.25 to 0.710.71). The savings come from the surrogate: where it reliably separates safe from unsafe contingencies, the audit certifies a large skip set; where it does not, the certificate verifies more, which is why IEEE 300-bus (a well-separated case) saves most. Batching the contingencies of several operating points into one window tightens the Clopper-Pearson bound and raises the savings to 3636 to 8080 percent, at a coarser time resolution; this is a statistical-efficiency option, not a property of a single control decision. The overall violation rates span 44 percent (IEEE 300-bus) to 6363 percent (IEEE 118-bus), so the evaluated operating points range from lightly loaded (where the surrogate separates well and savings are large) to heavily congested, and the savings are not an artifact of a single regime. In wall-clock terms, a single AC N-1 contingency solve takes about 1212 ms on IEEE 300-bus while the surrogate score is a matrix operation costing well under a millisecond, so runtime tracks the AC-solve fraction; the single-operating-point sweep of about 268268 credible contingencies takes roughly 33 s, which ASV-N1 reduces to about 0.80.8 s, and the audit contingencies are independent and parallelizable. Minimum audit size The audit must be large enough for the Clopper-Pearson bound to license any skips. With zero audited violations, certifying a skipped-set rate at α with confidence 1−δ1-δ requires at least nmin=⌈lnδ/ln(1−α)⌉n_ = δ/ (1-α) audited skipped contingencies: nmin=15n_ =15 at α=0.15α=0.15, 2222 at α=0.10α=0.10, and 4545 at α=0.05α=0.05 (all at δ=0.1δ=0.1). For the adaptive rule, which selects the audit size from G=8G=8 candidates and therefore tests each at δ/G=0.0125δ/G=0.0125 (Proposition 1), nminn_ rises to 2727 at α=0.15α=0.15. We floor the single-operating-point audit at 2020 percent of the window, which exceeds nminn_ on every system in Table I. If a window is too small to reach nminn_ , no threshold is certifiable and ASV-N1 verifies the whole window, safe by construction but with no savings. Table I sweeps the budget at the single operating point. Cost rises monotonically as α tightens: the operator buys stricter safety with more AC solves. An AC-solve fraction of 1.001.00 means the window is smaller than the audit needed to certify that budget, so the certificate conservatively verifies everything (safe by construction, no savings). Strict budgets such as α=0.05α=0.05 remain feasible on IEEE 300-bus and PEGASE, and α=0.01α=0.01 on PEGASE; on the smaller systems they are recovered by batching. Throughout, the realized trusted-set violation stays within budget. TABLE I: Budget sweep at the single operating point: AC-solve fraction as the risk budget α tightens (δ=0.1δ=0.1). Cost rises monotonically as α falls; 1.001.00 means the window is too small to certify that α, so ASV-N1 verifies all (safe, no savings). The realized trusted-set violation is within budget in every cell. System α=0.20α=0.20 0.150.15 0.100.10 0.050.05 0.010.01 IEEE 118-bus 0.71 0.71 0.75 1.00 1.00 IEEE 300-bus 0.24 0.25 0.26 0.37 1.00 PEGASE 1354 0.47 0.47 0.47 0.53 0.76 TABLE I: ASV-N1 with adaptive audit sizing (α=0.15α=0.15) at two window granularities. “Single operating point” is the real-time setting (one control decision; about 173173, 268268, and 17491749 credible contingencies for IEEE 118-bus, IEEE 300-bus, and PEGASE, respectively). “Batched” windows contain about 10001000 to 15001500 contingencies depending on system and partitioning: on IEEE 118-bus and 300-bus this pools several consecutive operating points, whereas on PEGASE a single operating point already holds a large credible set, so batching there mainly serves the many-window statistical test rather than increasing the per-window count. Batching tightens the Clopper-Pearson bound and lowers cost at a coarser time resolution. The trusted-set violation is the realized (empirical) rate among unverified skipped contingencies; it is well within budget in both settings, while the formal certificate controls the skip-set rate at α (Theorem 1; trusted subset at most α/(1−f)α/(1-f)). AC-solve fraction is relative to a full sweep (lower is cheaper). Single op. point Batched System Overall viol. trusted AC-frac trusted AC-frac IEEE 118-bus 0.63 0.012 0.71 0.020 0.64 IEEE 300-bus 0.04 0.005 0.25 0.006 0.20 PEGASE 1354 0.34 0.011 0.47 0.010 0.37 Figure 2: Left: safety. A deterministic surrogate screen (no audit) is unsafe under shift on PEGASE (skipped-set N-1 violation 0.340.34 against the 0.150.15 budget), while ASV-N1 is safe on all four testbeds including a genuine controller-induced shift. Right: cost. Adaptive audit sizing lowers the AC-solve fraction on every system (IEEE 300-bus reaches 0.200.20, an 8080 percent reduction), all below the full N-1 sweep, with validity preserved. VII-C Testing the probabilistic guarantee across many windows A single trusted-set violation number does not test an (α,δ)(α,δ) guarantee, which is a statement about the distribution of outcomes across windows: Theorem 1 requires that the per-window violation rate exceed α in at most a δ fraction of windows. We test this directly. For each system we form many control windows by randomly partitioning the labeled pool into windows of about one thousand contingencies, repeat over forty independent seeds (which also re-randomize the audit), and record the realized trusted-set violation rate of each window. Figure 3 plots the resulting cumulative distributions. With α=0.15α=0.15 and 1−δ=0.91-δ=0.9, the fraction of windows that breach α is 0.0580.058 on IEEE 118-bus and 0.0000.000 on IEEE 300-bus and PEGASE, all at or below δ=0.1δ=0.1, across 13601360, 10001000, and 41604160 windows respectively. The empirical breach frequencies are therefore consistent with the stated guarantee, not only in expectation but at the required confidence level, and the breach fraction approaches but does not exceed δ on IEEE 118-bus, indicating the certificate is tight rather than loose at this window size. Window size trades tightness for conservativeness. At the operational window size of a single operating point, with about 173173, 268268, and 17491749 credible contingencies for IEEE 118-bus, IEEE 300-bus, and PEGASE respectively, the Clopper-Pearson interval is wider, so the certified threshold is stricter, the certificate verifies more and skips less, and the trusted-set violation is essentially zero: over 8304083040, 6108061080, and 251760251760 windows for the three systems, the breach fraction is 0.0000.000 everywhere. The certificate is thus even safer, but cheaper savings require aggregating a larger contingency set per window (for example batching several operating points or a wider credible set), which tightens the bound as in Figure 3. Figure 3: The (α,δ)(α,δ) guarantee holds across many windows. Cumulative distribution of the per-window trusted-set violation rate over hundreds of windows and forty seeds per system. The horizontal line at 1−δ=0.91-δ=0.9 meets each curve at or before α=0.15α=0.15, so at most a δ fraction of windows breach the budget (breach fractions: IEEE 118-bus 5.85.8 percent, IEEE 300-bus and PEGASE 0 percent). VII-D Comparison to deterministic screening We compare ASV-N1 against the operational screens practitioners use, on PEGASE under controller-induced shift (the hardest case), at budget α=0.15α=0.15 (Table I). A deterministic screen that skips every contingency the surrogate predicts safe is cheap but unsafe (violation 0.3350.335), because the surrogate’s confidently-safe predictions are wrong under stress; a safety margin does not rescue it (0.3360.336 at a 1010 percent buffer), and a threshold calibrated on historical data breaks catastrophically when the controller shifts the operating point (0.6710.671). Verifying the top-k most severe contingencies at a budget matched to ASV-N1 happens to be safe here (0.0680.068), but it carries no certificate and cannot know the required budget without an audit; under a different shift it degrades to the calibrated-threshold failure. ASV-N1 certifies safety (0.0110.011) at the same cost as top-k. This is the central distinction from fast screening: ASV-N1 provides a guarantee that survives an unreliable surrogate. TABLE I: Operational screens on PEGASE under controller-induced shift (real-time single operating point, α=0.15α=0.15). Deterministic, margin, and historically-calibrated screens are unsafe; top-k at a matched budget is safe only by chance and gives no certificate. ASV-N1 certifies safety at the same cost. Screen Trusted viol. AC-frac Safe? Deterministic LODF screen 0.335 0.01 no Margin screen (10%10\% buffer) 0.336 0.01 no Static calibrated threshold 0.671 0.00 no Top-k severe (matched budget) 0.068 0.47 no cert. ASV-N1 (audited) 0.011 0.47 certified VII-E Adaptive audit sizing Figure 2 (right) compares a fixed audit fraction to the adaptive rule. The adaptive rule lowers the AC-solve fraction on every system while preserving validity; on IEEE 300-bus it falls from 0.300.30 to 0.200.20. All bars remain below the full N-1 sweep baseline, so ASV-N1 always saves work relative to exhaustive verification. Parameter selection The budget α is the operator’s chosen safety level and sets the safety-versus-cost tradeoff: a smaller α certifies a stricter threshold, which verifies more contingencies and raises the AC-solve fraction, while validity holds at every α. Sweeping α∈0.05,0.10,0.15,0.20α∈\0.05,0.10,0.15,0.20\ produces a monotone cost curve that stays below the full-sweep baseline throughout, so an operator can read off the verification cost of a desired safety level directly. The confidence 1−δ1-δ enters only through the Clopper-Pearson width and has a mild logarithmic effect on cost; we use 1−δ=0.91-δ=0.9. The audit size is not a free parameter: the adaptive rule selects it to minimize total cost, so the operator sets only α and δ. VII-F Genuine controller-induced shift The previous experiments induce shift by load scaling. To test a genuine controller-induced shift, we contrast two control policies on IEEE 300-bus. Because the economic controller only becomes security-naive under stress, where the DC optimal dispatch and the AC security limits diverge, this experiment uses a more heavily loaded regime and a representative contingency sample than the scalability study; the certificate is valid over whatever contingency set is monitored. The historical operator uses proportional dispatch, while the deployed controller uses an economic DC optimal power flow redispatch that minimizes generation cost subject to the DC power balance and base-case line limits only, with no N-1 contingency or thermal-security constraints; this is what makes it security-naive under stress and a realistic source of controller-induced shift. The two policies, on the same loads, produce different operating points; the economic, security-naive controller raises the N-1 violation rate from 0.300.30 to 0.610.61 and the mean surrogate score from 3.753.75 to 8.718.71, a strong shift. Table IV reports deploying an operator-calibrated certificate on the controller. The static (no-audit) certificate degrades to a trusted-set violation of 0.1660.166, above the budget, because it was calibrated on the operator distribution. ASV-N1, whose audit observes the controller’s own operating points, holds the trusted-set violation at 0.0430.043, at a higher verification cost (AC-solve fraction 0.840.84). This is the honest and expected behavior: under a strong, genuine controller-induced shift the savings shrink to about 1616 percent, because the audit correctly responds to the less reliable surrogate by verifying more. Safety is maintained throughout; the surrogate’s degradation is paid in compute, not in violations. TABLE IV: Genuine controller-induced shift on IEEE 300-bus (operator proportional dispatch vs. economic DC-OPF controller; α=0.15α=0.15). The operator-calibrated static certificate degrades past the budget; ASV-N1 stays safe. Certificate Trusted-set viol. AC-solve frac. Static (operator-calibrated, no audit) 0.166 0.54 ASV-N1 (online audit on controller) 0.043 0.84 VIII Discussion and Limitations What the guarantee is and is not Theorem 1 is a per-window, high-probability statement, not a per-action worst-case one: it controls the violation rate of the skip set at level α with confidence 1−δ1-δ (Theorem 1), so the unverified trusted subset has violation rate at most α/(1−f)α/(1-f) (Corollary 1), aggregated over windows by Corollary 2. This matches how operators reason about risk budgets and is what makes a distribution-free statement possible without restricting the deployment distribution. It is therefore a careful claim: the certificate controls a violation-rate budget under arbitrary deployment shift, but it does not guarantee that every individual skipped contingency is safe, nor does it provide worst-case N-1 security for each skipped outage. An operator who needs a hard per-contingency guarantee can drive the budget toward zero, at which point ASV-N1 verifies essentially everything. ASV-N1 is thus intended for risk-budgeted screening and triage, not for settings where policy mandates deterministic verification of every credible contingency; a small set of high-impact outages can be excluded from the skip set and always verified, so the certificate complements rather than replaces a deterministic check where one is required. Scope The certified object is N-1 thermal security under the AC model (Assumption 2). Voltage and reactive-power limits are outside the present theorem; in our experiments we track voltage and non-convergence events separately. Extending the audit to a joint thermal-and-voltage label is direct in principle, since the audit already runs full AC, and is a natural next step. The trusted verifier is the AC N-1 solver, which is standard in operations. Cost floor and surrogate quality The audit is an unavoidable cost: even with a perfect surrogate, ASV-N1 pays the audit fraction. Proposition 2 makes the tradeoff explicit, and Figure 2 shows the audit fraction is a small part of the total. The audit contingencies are independent and embarrassingly parallel, so the per-window wall-clock cost scales with available compute rather than with system size; our largest case, the 1354-bus PEGASE system, runs with the same procedure as the smaller ones, indicating the approach scales to realistic networks. A better surrogate raises τ⋆τ and lowers cost; a learned surrogate that resolves the stressed regime where linear screening fails is a promising way to push the savings on the harder systems toward those on IEEE 300-bus, without affecting validity. Integration in an energy management system ASV-N1 is designed to sit between a controller and the dispatch it commits. Concretely, in an energy management system the flow is: (i) a controller (reinforcement learning, model-predictive control, or an optimization such as security-constrained optimal power flow) proposes an operating point; (i) ASV-N1 computes the linear surrogate scores for the operator-defined credible contingency set, draws the audit and runs full AC power flow on the audit plus the above-threshold contingencies, and calibrates τ⋆τ ; (i) it returns the trusted set, the verified violations, and the certified risk. If the realized or certified violation rate exceeds the operator’s budget, the action is rejected or sent back for re-dispatch. The certificate is an overlay, or watchdog layer, that requires no change to the controller and reuses the AC contingency solver an operator already trusts, so it complements rather than replaces security-constrained optimal power flow. Controllers We demonstrated a genuine controller-induced shift with an economic redispatch controller. Wrapping a trained learning-based controller end to end is future work; the method is agnostic to the controller because it certifies the resulting operating point, not the policy. IX Conclusion We presented Audited Selective Verification, a distribution-free certificate that risk-controls the N-1 thermal violation rate of a black-box controller’s action, rather than asserting hard per-contingency N-1 security. By using a cheap surrogate only to decide what to verify, and resting validity on an online audit plus full AC verification, ASV-N1 holds the violation rate of the trusted contingencies at a budget with high confidence, for an arbitrary deployment distribution and any surrogate, including under the shift the controller itself induces. We showed that simpler threshold certificates, static or adaptive, fail when the surrogate is unreliable under stress, and that deterministic screening is unsafe under shift, whereas ASV-N1 keeps the realized violation rate within budget across the tested windows on IEEE 118-bus, IEEE 300-bus, and the 1354-bus PEGASE system. At the single operating point, the real-time control window, it cuts AC N-1 solves by 2929 to 7575 percent, rising to 3636 to 8080 percent when contingencies are batched across operating points, and it degrades gracefully toward verify-all exactly where the surrogate is least trustworthy. The method gives operators a tunable safety budget with a quantified, surrogate-adaptive compute cost, and a verification-backed safety layer behind which a controller can be deployed. As utilities begin to adopt learning-based controllers in operations, ASV-N1 offers a concrete watchdog layer that admits such controllers under an explicit, auditable safety budget rather than on trust, which aligns with operational practice around documented contingency assessment: the certified risk budget and the audit log give operators an auditable trail. Future work includes a joint thermal-and-voltage audit, a learned surrogate to lower cost on stressed systems, and end-to-end evaluation with trained learning-based controllers. References [1] A. J. Wood, B. F. Wollenberg, and G. B. Sheblé, Power Generation, Operation, and Control, 3rd ed. John Wiley and Sons, 2013. [2] A. Marot, B. Donnot, C. Romero, B. Donon, M. Lerousseau, L. Veyrin-Forrer, and I. Guyon, “Learning to run a power network challenge: A retrospective analysis,” arXiv preprint arXiv:2103.03104, 2021. [3] M. Lehna, C. Holzhüter, S. Tomforde, and C. Scholz, “HUGO: Highlighting unseen grid options, combining deep reinforcement learning with a heuristic target topology approach,” arXiv preprint arXiv:2405.00629, 2024. [4] T. Su, T. Wu, J. Zhao, A. Scaglione, and L. Xie, “A review of safe reinforcement learning methods for modern power systems,” arXiv preprint arXiv:2407.00304, 2024. [5] A. N. Angelopoulos, S. Bates, E. J. Candès, M. I. Jordan, and L. Lei, “Learn then test: Calibrating predictive algorithms to achieve risk control,” arXiv preprint arXiv:2110.01052, 2021. [6] S. Bates, A. Angelopoulos, L. Lei, J. Malik, and M. I. Jordan, “Distribution-free, risk-controlling prediction sets,” Journal of the ACM, vol. 68, no. 6, p. 1-34, 2021. [7] I. Gibbs and E. J. Candès, “Adaptive conformal inference under distribution shift,” in Advances in Neural Information Processing Systems (NeurIPS), 2021. [8] L. Thurner, A. Scheidler, F. Schäfer, J.-H. Menke, J. Dollichon, F. Meier, S. Meinecke, and M. Braun, “pandapower: An open-source python tool for convenient modeling, analysis, and optimization of electric power systems,” IEEE Transactions on Power Systems, vol. 33, no. 6, p. 6510-6521, 2018. [9] B. Stott, O. Alsac, and A. J. Monticelli, “Security analysis and optimization,” Proceedings of the IEEE, vol. 75, no. 12, p. 1623-1644, 1987. [10] G. C. Ejebe and B. F. Wollenberg, “Automatic contingency selection,” IEEE Transactions on Power Apparatus and Systems, vol. PAS-98, no. 1, p. 97-109, 1979. [11] A. Monticelli, M. V. F. Pereira, and S. Granville, “Security-constrained optimal power flow with post-contingency corrective rescheduling,” IEEE Transactions on Power Systems, vol. 2, no. 1, p. 175-180, 1987. [12] F. Capitanescu, J. L. M. Ramos, P. Panciatici, D. Kirschen, A. M. Marcolini, L. Platbrood, and L. Wehenkel, “State-of-the-art, challenges, and future trends in security-constrained optimal power flow,” Electric Power Systems Research, vol. 81, no. 8, p. 1731-1741, 2011. [13] P. L. Donti, A. Agarwal, N. V. Bedmutha, L. Pileggi, and J. Z. Kolter, “Adversarially robust learning for security-constrained optimal power flow,” arXiv preprint arXiv:2111.06961, 2021. [14] M. Ni, J. D. McCalley, V. Vittal, and T. Tayyib, “Online risk-based security assessment,” IEEE Transactions on Power Systems, vol. 18, no. 1, p. 258-265, 2003. [15] D. S. Kirschen and D. Jayaweera, “Comparison of risk-based and deterministic security assessments,” IET Generation, Transmission and Distribution, vol. 1, no. 4, p. 527-533, 2007. [16] V. Vovk, A. Gammerman, and G. Shafer, Algorithmic Learning in a Random World. Springer, 2005. [17] A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, and T. Schuster, “Conformal risk control,” in International Conference on Learning Representations (ICLR), 2024. [18] R. J. Tibshirani, R. Foygel Barber, E. J. Candès, and A. Ramdas, “Conformal prediction under covariate shift,” in Advances in Neural Information Processing Systems (NeurIPS), 2019. [19] R. F. Barber, E. J. Candès, A. Ramdas, and R. J. Tibshirani, “Conformal prediction beyond exchangeability,” The Annals of Statistics, vol. 51, no. 2, p. 816-845, 2023. [20] D. Prinster, C. Fannjiang, J. W. Park, K. Cho, A. Liu, S. Saria, and S. Stanton, “Conformal policy control,” arXiv preprint arXiv:2603.02196, 2026. [21] A. Stratigakos, H. Wen, E. Spyrou, and P. Pinson, “Decision-calibrated prediction sets for robust power system operations,” arXiv preprint arXiv:2606.02081, 2026. [22] H. Zhou, Y. Zhang, and W. Luo, “Computationally and sample efficient safe reinforcement learning using adaptive conformal prediction,” arXiv preprint arXiv:2503.17678, 2025. [23] M. Alshiekh, R. Bloem, R. Ehlers, B. Könighofer, S. Niekum, and U. Topcu, “Safe reinforcement learning via shielding,” in AAAI Conference on Artificial Intelligence, 2018. [24] E. Marchesini, B. Donnot, C. Crozier, I. Dytham, C. Merz, L. Schewe, N. Westerbeck, C. Wu, A. Marot, and P. L. Donti, “RL2Grid: Benchmarking reinforcement learning in power grid operations,” arXiv preprint arXiv:2503.23101, 2025. [25] S. Ghamizi, A. Bojchevski, A. Ma, and J. Cao, “SafePowerGraph: Safety-aware evaluation of graph neural networks for transmission power grids,” arXiv preprint arXiv:2407.12421, 2024. [26] M. Eichelbeck, H. Markgraf, and M. Althoff, “CommonPower: A framework for safe data-driven smart grid control,” arXiv preprint arXiv:2406.03231, 2024. [27] W. Hoeffding, “Probability inequalities for sums of bounded random variables,” Journal of the American Statistical Association, vol. 58, no. 301, p. 13-30, 1963. [28] R. J. Serfling, “Probability inequalities for the sum in sampling without replacement,” The Annals of Statistics, vol. 2, no. 1, p. 39-48, 1974. [29] S. Holm, “A simple sequentially rejective multiple test procedure,” Scandinavian Journal of Statistics, vol. 6, no. 2, p. 65-70, 1979. [30] S. R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon, “Time-uniform, nonparametric, nonasymptotic confidence sequences,” The Annals of Statistics, vol. 49, no. 2, p. 1055-1080, 2021. Appendix A Full Proof of Theorem 1 We restate the setting. Condition on the fixed window: the pairs (ri,Vi)i=1N\(r_i,V_i)\_i=1^N are arbitrary fixed quantities. The audit A includes each index independently with probability π (or is a fixed-size uniform sample), drawn independently of (ri,Vi)\(r_i,V_i)\ by Assumption 1. For a threshold τ, Skip(τ)=i:ri≤τSkip(τ)=\i:r_i≤τ\ is fixed, and R(τ)=|Skip(τ)|−1∑i∈Skip(τ)ViR(τ)=|Skip(τ)|^-1 _i (τ)V_i. The candidate thresholds τ(1)<⋯<τ(M)τ^(1)<…<τ^(M) are the distinct score values, ordered independently of the labels. A-A Proof of Lemma 1 Fix τ and condition on n=|A∩Skip(τ)|n=|A (τ)|, a function of A and the scores only. By Assumption 1, given n, the audited subset of Skip(τ)Skip(τ) is a uniform size-n sample without replacement from the finite population Skip(τ)Skip(τ), whose violation indicators are fixed with mean R(τ)R(τ). Hence the audited violation count K is hypergeometric with mean nR(τ)n\,R(τ). Under H:R(τ)>αH:R(τ)>α, [K∣n]>nαE[K n]>nα. Define P=Pr[Bin(n,α)≤K]P= [Bin(n,α)≤ K], the binomial left-tail evaluated at the observed K; this equals the Clopper-Pearson one-sided test that rejects when Uδ(K,n)≤αU_δ(K,n)≤α. By the classical result that the hypergeometric distribution is stochastically dominated, in the lower tail, by the binomial distribution with the same mean [27, 28], the left tail probability under the true hypergeometric law is at most that under Bin(n,R(τ))Bin(n,R(τ)), and since R(τ)>αR(τ)>α shifts mass to larger K, the test statistic P is super-uniform: Pr[P≤δ∣n]≤δ [P≤δ n]≤δ. Taking expectation over n preserves the bound. Thus P is a valid p-value for H, and the Clopper-Pearson rule is, if anything, conservative for the without-replacement audit. □ A-B Proof of Theorem 1 Test the hypotheses Hj:R(τ(j))>αH_j:R(τ^(j))>α in increasing order of j (strictest skip set first). Reject HjH_j if its audited p-value Pj≤δP_j≤δ, and stop at the first non-rejection; let τ⋆τ be the largest threshold in the rejected prefix. This is fixed-sequence (ordered) testing along a data-independent order. Fixed-sequence testing controls the family-wise error rate at δ under arbitrary dependence among the PjP_j: an error requires the first true null in the sequence to be rejected, an event of probability at most δ by Lemma 1 applied to that single hypothesis [29]. Therefore, with probability at least 1−δ1-δ, no true null is rejected, so every rejected HjH_j is false, i.e., R(τ(j))≤αR(τ^(j))≤α for all j in the rejected prefix; in particular R(τ⋆)≤αR(τ )≤α. No assumption on the surrogate or on the deployment distribution was used: the argument conditions on the arbitrary fixed window and uses only the label-independent randomness of the audit. This is the transductive Learn-Then-Test guarantee with the audit as calibration and the fixed skip population as the controlled risk object [5, 6]. □ A-C Proof of Corollary 1 On the event R(τ⋆)≤αR(τ )≤α of Theorem 1, the total violation count in the fixed set Skip(τ⋆)Skip(τ ) is ∑i∈Skip(τ⋆)Vi=R(τ⋆)|Skip(τ⋆)|≤α|Skip(τ⋆)| _i (τ )V_i=R(τ )\,|Skip(τ )|≤α\,|Skip(τ )|. The trusted set T=i∉A:ri≤τ⋆⊆Skip(τ⋆)T=\i∉ A:r_i≤τ \ (τ ), so its violation count is no larger: ∑i∈TVi≤∑i∈Skip(τ⋆)Vi≤α|Skip(τ⋆)| _i∈ TV_i≤ _i (τ )V_i≤α\,|Skip(τ )|. Dividing by |T|=(1−f)|Skip(τ⋆)||T|=(1-f)\,|Skip(τ )| gives 1|T|∑i∈TVi≤α/(1−f) 1|T| _i∈ TV_i≤α/(1-f). This is a deterministic consequence of the event in Theorem 1, so it holds with probability at least 1−δ1-δ; it makes no independence assumption about T after the label-dependent choice of τ⋆τ , which is exactly why it survives the selection. □ A-D Proof of Proposition 1 Let the G candidate audit sizes be n1<⋯<nGn_1<…<n_G, fixed in advance, realized as nested prefixes of a single label-independent random ordering of C. For each g, Theorem 1 applied to audit AgA_g at level δ/Gδ/G gives Pr[R(τ^g)>α]≤δ/G [R( τ_g)>α]≤δ/G. A union bound over g=1,…,Gg=1,…,G yields Pr[∃g:R(τ^g)>α]≤δ [∃ g:R( τ_g)>α]≤δ, so simultaneously all G certified thresholds are valid with probability at least 1−δ1-δ. Any selection rule, including one that inspects audited labels to minimize predicted cost, returns one of these simultaneously-valid thresholds; hence the reported threshold satisfies R≤αR≤α with probability at least 1−δ1-δ. □