Paper deep dive
Budget-Adaptive Routing: Skipping the Weak When the Strong Answers Anyway
Wei Geng, Nitinder Mohan, Jörg Ott
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 7/5/2026, 5:05:54 AM
Summary
The paper introduces 'Budget-Adaptive Routing' for edge-cloud inference collaborations in object detection. It identifies an 'implicit compute tax' in existing weak-conditioned systems where a weak detector must run on every frame before a routing decision is made. The authors propose a 'weak-skipping estimator' that extracts routing signals from raw pixels, allowing the weak model to be bypassed entirely for frames destined for the cloud. They demonstrate that a budget-adaptive approach, which switches between weak-skipping and weak-conditioned placement based on the offload budget, optimizes the accuracy-compute trade-off across different operating regimes. Experimental results on PASCAL VOC show that their method reduces latency and compute while outperforming state-of-the-art selective offloading methods.
Entities (8)
Relation Signals (4)
Budget-Adaptive Routing → evaluatedon → PASCAL VOC
confidence 100% · On PASCAL VOC, our budget-adaptive router traces the upper accuracy envelope
Budget-Adaptive Routing → selectsbetween → Weak-skipping Estimator
confidence 100% · we propose budget-adaptive routing, which selects between them by offload budget via two offline-tuned thresholds.
Budget-Adaptive Routing → selectsbetween → Weak-conditioned Estimator
confidence 100% · we propose budget-adaptive routing, which selects between them by offload budget via two offline-tuned thresholds.
Weak-skipping Estimator → uses → MobileNetV2-Lite
confidence 100% · Our reference instantiation is a highly compressed MobileNetV2-Lite (Sandler et al., 2018) backbone
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Edge-cloud inference collaborations are often designed with a routing estimator that decides whether to offload each frame from weak models at the edge to stronger models in the cloud. Existing systems place the routing estimator after the weak detector, so the weak forward pass still runs even on frames that are later offloaded. In this paper, we argue that this weak-conditioned design can be suboptimal when the offload budget varies. First, we present a competitive weak-skipping estimator (0.153 GFLOPs, about 29x lighter than the weak detector at 4.49 GFLOPs) that extracts routing signal from raw pixels, outperforming the common after-weak placement weak-conditioned baselines. Second, we show that neither weak-skipping nor weak-conditioned placement dominates across the full operating curve, and we propose budget-adaptive routing, which selects between them by offload budget via two offline-tuned thresholds. On PASCAL VOC, our budget-adaptive router traces the upper accuracy envelope of both fixed placements across the operating range. Our method reduces per-frame latency by up to 19.1 ms (about 30% lower at rho = 0.9). Besides outperforming SOTA methods, it is surprisingly stronger than the strong model (+1.7 pp over the strong model's peak mAP) at some operating points with far less compute. Artifacts are available at this https URL
Tags
Links
- Source: https://arxiv.org/abs/2606.30919v1
- Canonical: https://arxiv.org/abs/2606.30919v1
Trouble viewing inline? Open PDF directly →
Full Text
50,115 characters extracted from source content.
Expand or collapse full text
[4.0]by Budget-Adaptive Routing: Skipping the Weak When the Strong Answers Anyway Wei Geng Technical University of MunichMunichGermany wei.geng@tum.de , Nitinder Mohan TU DelftDelftNetherlands n.mohan@tudelft.nl and Jörg Ott Technical University of MunichMunichGermany ott@in.tum.de (2026) Abstract. Edge-cloud inference collaborations are often designed with a routing estimator111In this paper, in most cases we use estimator and router interchangeably. that decides whether to offload each frame from weak models at the edge to stronger models in the cloud. Existing systems place the routing estimator after the weak detector, so the weak forward pass still runs even on frames that are later offloaded. In this paper, we argue that this weak-conditioned design can be suboptimal when the offload budget varies. First, we present a competitive weak-skipping estimator (0.1530.153 GFLOPs, ∼29× 29× lighter than the weak detector at 4.494.49 GFLOPs) that extracts routing signal from raw pixels, outperforming the common after-weak placement weak-conditioned baselines. Second, we show that neither weak-skipping nor weak-conditioned placement dominates across the full operating curve, and we propose budget-adaptive routing, which selects between them by offload budget via two offline-tuned thresholds. On PASCAL VOC, our budget-adaptive router traces the upper accuracy envelope of both fixed placements across the operating range. Our method 222Artifacts are available at https://github.com/ViGeng/bgt-ada reduces per-frame latency by up to 19.119.1 ms (∼30% 30\% lower at ρ=0.9ρ=0.9). Besides outperforming SOTA methods, it is surprisingly stronger than the strong model (+1.7+1.7 p over the strong model’s peak mAP) at some operating points with far less compute. selective offloading, budget-adaptive routing, object detection, cost-accuracy trade-off †journalyear: 2026†copyright: c†conference: Workshop on Networks for AI Computing; August 17–21, 2026; Denver, CO, USA†booktitle: Workshop on Networks for AI Computing (SIGCOMM ’26), August 17–21, 2026, Denver, CO, USA†doi: 10.1145/3789240.3828740†isbn: 979-8-4007-2467-1/26/08†ccs: Networks Cloud computing†ccs: Computer systems organization n-tier architectures†ccs: Computing methodologies Object detection†ccs: Networks Network performance modeling 1. Introduction Edge devices increasingly run visual perception pipelines by offloading computation to the cloud, such as traffic cameras counting vehicles, industrial robots operating on assembly lines, and identity systems checking faces at borders. While some simply stream tasks to the edge or cloud for inference (Geng et al., 2026; Jiang et al., 2018; Li et al., 2020; Canel et al., 2019; Chen et al., 2015), others pair a local/edge weak detector that is fast but limited with a strong detector in the cloud that is more accurate but more costly in latency, compute, bandwidth, and energy. Selective offloading routes each frame so that the cloud complements local inference, reducing overall cost without sacrificing, and often improving, end-to-end accuracy (Cao et al., 2023; Qiu et al., 2024). This routing decision is constrained by an offload budget that caps the fraction of frames sent to the cloud, given network conditions, compute resources, and application requirements. Figure 1. Selective offloading for object detection: a local weak detector and a cloud strong detector with a per-frame router under offload budget ρ. A decade of selective-offloading work (Qiu et al., 2024; Cao et al., 2023; Wang et al., 2024; Teerapittayanon et al., 2016; Huang et al., 2017; Kaya et al., 2019; Geng et al., 2026) runs the weak detector on every frame and feeds its output (proposal scores, top-k box statistics, learned embeddings, or intermediate activations) to a routing estimator that decides whether to escalate to a stronger detector in the cloud. EdgeML (Qiu et al., 2024) regresses on the top-2525 proposal features so that uncertain cases can be improved by strong models. DCSB (Cao et al., 2023) hand-crafts a difficult-case discriminator that directs requests to the strong model at a fixed threshold. Early-exit families such as BranchyNet (Teerapittayanon et al., 2016), MSDNet (Huang et al., 2017), and SDN (Kaya et al., 2019) gate within the weak network at intermediate layers. Despite their architectural variety, all of these methods share one structural commitment: the weak forward pass runs on every frame, because the routing decision depends on its output or intermediate features. This commitment is misaligned with high-budget deployments. Safety-critical settings such as autonomous driving or security run at a high offload budget, say ≥30%≥ 30\% of frames sent to the cloud, and there the weak forward pass is wasted on most frames: they are offloaded and answered by the strong model anyway, so paying for the weak pass only adds unnecessary compute and latency. Structurally, if the estimator is moved before the weak model and predicts from the raw image, the serial dependency on the weak model is removed and the weak pass can be skipped for frames that will most likely be offloaded. We call this a weak-skipping estimator, in contrast to the weak-conditioned estimators of prior work (shown in Fig. 2 left and middle). To our knowledge no previous selective-offloading system has deployed it. The usual assumption is that raw-pixel features are too weak to predict detector failure. Our results challenge this: a 0.150.15 GFLOPs image-only weak-skipping estimator trained on a binary offload-utility target (Eq. 6) is competitive with proposal-feature baselines that require the full 4.494.49 GFLOPs weak forward pass. We attribute this to the assumption that predicting whether a frame is difficult is substantially cheaper than predicting what it contains. Weak-skipping is not, however, strictly better. At low offload budgets the weak forward pass is unavoidable anyway. Since the cloud rarely fires, the weak-conditioned estimator’s richer features come essentially for free. The two placements therefore divide the operating curve into compute regimes: weak-conditioned and weak-skipping win in the low and high-budget bands, respectively, and neither dominates the curve. Taken together, these observations argue for an adaptive router that selects between weak-skipping and weak-conditioned placements as a function of the offload budget, instead of one fixed placement. We confirm it empirically on PASCAL VOC from compute, latency, and accuracy perspectives. This paper contributes: (i) We articulate the weak-first assumption latent in the selective-offloading literature, formalize its hidden cost as an implicit compute tax. (§ 2). (i) We introduce the first competitive weak-skipping estimator for object-detection offloading: a 0.150.15 GFLOPs image-only estimator that matches or exceeds the strongest weak-conditioned baselines (§ 3.1) detection quality-wise. We also build a lightweight weak-conditioned estimator (XGBoost on MORIC) that outperforms current weak-conditioned SOTAs on routing quality at sub-ms inference cost.(Tab. 2). (i) Building on our weak-skipping and weak-conditioned estimators, we propose budget-adaptive routing, which selects between weak-skipping and weak-conditioned placements according to the deployment budget. From our simulated study, our approach outperforms SOTAs. (§ 3.2, § 4). Figure 2. Routing schemes: full partitioned compute (Kang et al., 2017; Kaya et al., 2019; Teerapittayanon et al., 2016; Huang et al., 2017) (left), weak-conditioned after the weak pass (Cao et al., 2023; Qiu et al., 2024) (middle), and our weak-skipping (right), which enables budget-adaptive routing. Estimator placement is the axis: it fires after the weak model (middle) or before it (right). Green marks the keep-local path, red the escalation to the cloud. 2. Problem Statement 2.1. Problem Formulation Let ℐ=(i1,…,iN)I=(i_1,…,i_N) be a stream of image frames, MwM_w a weak local detector with per-frame compute cost CwC_w, and MsM_s a strong cloud detector with per-frame cost CsC_s (compute plus network round-trip). For each frame iti_t the router produces a binary routing decision dt∈0,1d_t∈\0,1\ where dt=0d_t=0 means “return Mw(it)M_w(i_t)” and dt=1d_t=1 means “return Ms(it)M_s(i_t)”. Given a downstream detection utility U(⋅)U(·) (typically mean Average Precision, mAP), an offload budget ρ∈(0,1]ρ∈(0,1], and an estimator producing a per-frame score sts_t, the router solves (1) maxd1:N1N∑t=1NU(it,Mdt)s.t.1N∑t=1Ndt≤ρ. _d_1:N\ 1N _t=1^NU (i_t,M_d_t ) .t. 1N _t=1^Nd_t≤ρ. We instantiate U as mAP@0.5, the standard PASCAL VOC metric (Everingham et al., 2010), for comparability with the VOC benchmark and the weak-conditioned baselines we reproduce. The framework is otherwise metric-agnostic: the proxy reward ΔAP \!AP (Eq. 4) can be computed for any U, so a stricter COCO-style AP@[.5:.95] is a drop-in substitution we leave to future work. 2.2. The Implicit Compute Tax In every weak-conditioned design the weak pass executes on every frame, including those answered by the cloud, as an implicit compute tax. The expected per-frame compute is (2) Tcond(ρ) T_cond(ρ) =Cw+Cecond+ρ⋅Cs, =C_w+C_e^cond+ρ· C_s, (3) Tskip(ρ) T_skip(ρ) =Ceskip+(1−ρ)⋅Cw+ρ⋅Cs, =C_e^skip+(1-ρ)· C_w+ρ· C_s, where CecondC_e^cond and CeskipC_e^skip are the estimator costs for the two placements. The tax is ρCw−(Ceskip−Cecond)ρ\,C_w-(C_e^skip-C_e^cond). With our values (MobileNetV3 (Howard et al., 2019) vs ResNet50 (He et al., 2016), Cw=4.49C_w=4.49, Cecond≈0C_e^cond≈0, Ceskip=0.15C_e^skip=0.15, Cs=280.37C_s=280.37 GFLOPs), the tax reaches 3.903.90 GFLOPs at ρ=0.9ρ=0.9, approaching the full weak-detector cost. In addition, a weak-conditioned router cannot start until the weak forward pass finishes (24.7024.70 ms here), so CwC_w is a serial wall-clock dependency on every frame. The tax reflects a fallback mindset: the cloud as a backstop for the weak detector. Beside the compute and latency costs, it may also increase the jitter of the system (we leave further analysis to a follow-up work), which is undesirable for real-time applications. We don’t have to pay the tax if we treat the cloud as a collaborator that complements local inference instead of a fallback, illustrated as a branch topology in Fig. 2-right. Then the practical question is whether an image-only lightweight estimator can catch enough routing signal without paying the tax? 3. Method 3.1. Weak-skipping Estimator Proxy metric. A weak-skipping estimator can only out-cheap a weak-conditioned one if it learns to predict whether offloading helps, not the contents of the frame. A naive per-frame target like ΔAP \!AP (=cloud AP−local AP=cloud AP-local AP), is misleading because detection AP is computed dataset-wide via a single global Precision-Recall (PR) curve, so a single frame’s contribution depends on every other frame. Inspired by EdgeML (Qiu et al., 2024), we use the contextual offloading reward: for frame iti_t, the per-frame reward (4) ΔAP(it)=AP(swapt→s)−AP(all-weak), \!AP(i_t)\;=\;AP (swap_t→ s )\;-\;AP (all-weak ), i.e., the change in dataset-wide AP@0.5 obtained by replacing the local detections on iti_t with the cloud detections while holding all other frames at their local outputs. Computing it offline over the training set is (N)O(N) via a precomputed-IoU merge swap, and the resulting reward depends only on the dataset, not on the system at inference time. Eq. 4 produces a long-tailed signed signal that is hard to regress directly, as shown by the raw ΔAP \!AP ridge (top) of Fig. 3. We extract two estimator-friendly targets: (5) MORIC+(it)=FΔAP+(ΔAP(it))ΔAP(it)>0,0ΔAP(it)=0,FΔAP−(ΔAP(it))−1ΔAP(it)<0,MORIC^+(i_t)= casesF^+_ ( \!AP(i_t) )& \!AP(i_t)>0,\\ 0& \!AP(i_t)=0,\\ F^-_ ( \!AP(i_t) )-1& \!AP(i_t)<0, cases (6) OffloadBin(it)=[ΔAP(it)>0],OffloadBin(i_t)=1\! [ \!AP(i_t)>0 ], where F+F^+ and F−F^- are the empirical CDFs of ΔAP \!AP over the frames where offloading strictly helps (ΔAP>0 \!AP>0) and strictly hurts (ΔAP<0 \!AP<0), respectively; MORIC+(it)∈[−1,1]MORIC^+(i_t)∈[-1,1] and OffloadBin(it)∈0,1OffloadBin(i_t)∈\0,1\. MORIC+ generalises EdgeML’s MORIC (Qiu et al., 2024) by splitting the CDF at zero, which equalises positive and negative magnitudes and gives a symmetric loss landscape around the routing boundary. OffloadBin instead reduces the problem to a binary class label, which we train with focal loss (Lin et al., 2017b) to handle the positive-class imbalance. On VOC the raw ΔAP \!AP is long-tailed with 62.6%62.6\% of frames at exactly zero, 25.1%25.1\% positive, and 12.3%12.3\% negative (Fig. 3), so direct regression wastes capacity on the zero spike. In other words, the estimator is always lazily and blindly predicting the majority class zero, which already yields good rewards. As a result, it fails to learn the routing boundary. Both transformed targets decouple hardness from content, which is what lets raw pixels suffice. OffloadBin gives our best routing quality and MORIC+ is the softer-signal variant reported in the per-ratio sweep. Figure 3. Per-frame learning targets on VOC test (N=3105N=3105). Top: raw ΔAP \!AP is degenerate, with 62.6%62.6\% of frames exactly at 0 (stem) and short signed tails (25.1%>025.1\%>0, 12.3%<012.3\%<0). Middle: MORIC+ spreads this into a smooth, symmetric regression target on [−1,1][-1,1]. Bottom: OffloadBin ([ΔAP>0]1[ \!AP>0]) is the binary classification target, a ∼1:3 1:3 split (25.1%25.1\% positive). Architecture and calibration. A weak-skipping estimator fskip:I↦s^skipf_skip:I s_skip produces a routing score from the raw image alone, where s^skip∈[0,1] s_skip∈[0,1] estimates P(ΔAP>0)P( \!AP>0) when trained on OffloadBin (or the predicted MORIC+∈[−1,1]MORIC^+∈[-1,1]); a higher score means offloading is more likely to help (not how much to help), instantiating the generic per-frame score sts_t of § 2.1. Our reference instantiation is a highly compressed MobileNetV2-Lite (Sandler et al., 2018) backbone (128×128128×128 input, 0.150.15 GFLOPs, 0.540.54 M parameters) trained on OffloadBin, though the framework is agnostic to the backbone and proxy. A thresholder πρ _ρ (Guo et al., 2017) then maps the score stream to binary decisions dt=πρ(s^skip)d_t= _ρ( s_skip) so that the offloaded fraction matches the requested budget ρ without test-set lookahead. Concretely, πρ _ρ keeps a running estimate of the (1−ρ)(1-ρ) quantile of the incoming scores and offloads any frame scoring above it, nudging the threshold as the stream drifts so the realized offload rate stays near ρ. Ratio control thus stays orthogonal to the score: any estimator emitting a calibrated per-frame score drops into the same thresholder. We leave a full treatment of the thresholder, including drift handling, finite-window error, and the resulting jitter, to follow-up work. 3.2. Budget-adaptive Routing The accuracy side of the placement mirrors the compute side of § 2.2. At low budgets, weak-skipping wastes richer signals from the weak forward pass on most frames because most of them comes for free. At high budgets, weak-conditioned pays the tax on most frames because they will be offloaded anyway. Every fixed-placement router is therefore suboptimal on part of the operating curve. Formulation. A budget-adaptive router holds two estimators, fskip:I↦s^skipf_skip\!:\!I s_skip and fcond:(I,Mw(I))↦s^condf_cond\!:\!(I,M_w(I)) s_cond (s^cond∈[0,1] s_cond∈[0,1], read from the image plus weak output), plus a binary arbiter α(ρ)∈0,1α(ρ)∈\0,1\ and the same thresholder πρ _ρ (§ 3.1), now shared across both estimators. The per-frame decision is (7) dt=πρ(α(ρ)s^skip(it)+(1−α(ρ))s^cond(it,Mw(it))).d_t\;=\; _ρ\! (α(ρ)\, s_skip(i_t)+(1-α(ρ))\, s_cond(i_t,M_w(i_t)) ). (8) α(ρ)=argmaxa∈0,1AP^(ρ,afskip+(1−a)fcond),α(ρ)\;=\; _a∈\0,1\\; AP (ρ,\;a\,f_skip+(1-a)\,f_cond ), where AP AP is offline AP@0.5 on a held-out tuning split. Evaluated per budget, Eq. 8 makes α(ρ)α(ρ) piecewise constant with two crossovers, splitting the operating range into three regimes: (9) α(ρ)=0ρ<ρfrontier(weak-conditioned),1ρfrontier≤ρ<ρceiling(weak-skipping),0ρ≥ρceiling(weak-conditioned).α(ρ)= cases0&ρ< _frontier (weak-conditioned),\\[2.0pt] 1& _frontier≤ρ< _ceiling (weak-skipping),\\[2.0pt] 0&ρ≥ _ceiling (weak-conditioned). cases The two thresholds are the budgets at which the winning placement flips on the tuning split. At the low crossover ρfrontier _frontier the weak pass on offloaded frames turns from a near-free byproduct into pure tax, so weak-skipping starts to win. At the high crossover ρceiling _ceiling the fewer frames still kept local carry enough weight that the richer weak-conditioned signal wins again. We fit both once, offline, by sweeping ρ and reading off the two switch points (on VOC, ρfrontier=0.3 _frontier=0.3 and ρceiling=0.8 _ceiling=0.8). At runtime the arbiter is a constant-time lookup on ρ. 4. Preliminary Evidence We report preliminary results on PASCAL VOC (Everingham et al., 2010) based on the setup in Tab. 1: the weak/strong mAP gap is small which makes routing genuinely difficult and any larger gap can make the benefits more pronounced. Three claims are verified: (i) Weak-skipping estimators are competitive with weak-conditioned baselines on routing quality, and our weak-conditioned estimator outperforms current SOTAs (§ 4.1, Tabs. 2 and 3, Fig. 5). (i) The compute advantage of weak-skipping routing at high offload budget is real GFLOPs- and latency-wise. (§ 4.2, Fig. 4). (i) A budget-adaptive router that selects between the two placements traces the lowest cost across the operating curve and Pareto-dominates either fixed placement on the joint accuracy-compute frontier, outperforming all existing baselines (§ 4.3, Figs. 5 and 3). Table 1. Experiment setup (PASCAL VOC). Detectors (profiled on Nvidia A40) Weak MwM_w fasterrcnn_mobilenet_v3_large_fpn (Ren et al., 2015; Howard et al., 2019; Lin et al., 2017a); Cw=4.49C_w=4.49 GFLOPs / 24.7024.70 ms; mAP =0.760=0.760 Strong MsM_s fasterrcnn_resnet50_fpn_v2 (Ren et al., 2015; He et al., 2016; Lin et al., 2017a); Cs=280.37C_s=280.37 GFLOPs / 44.1244.12 ms; mAP =0.791=0.791 Estimators (ours, § 3) Weak-skipping MobileNetV2-Lite (Sandler et al., 2018) on OffloadBin or MORIC+^\!+ Weak-conditioned XGBoost (Chen and Guestrin, 2016) on MORIC Budget-adaptive offline-tuned ρfrontier _frontier, ρceiling _ceiling Prior weak-conditioned baselines EdgeML (Qiu et al., 2024) proposal-level, after weak detector DCSB (Cao et al., 2023) fixed-rule, after weak detector Reference points Trivial always-weak, always-strong, uniformly random Oracle per-frame ΔAP \!AP ground-truth gain 4.1. Weak-skipping is Empirically Viable Tab. 2 reports an overview across all estimators on VOC Test. Our weak-skipping MobileNetV2-Lite estimator trained on OffloadBin attains a Spearman rank correlation of 0.5570.557 with the per-frame oracle gain, peak end-to-end mAP@0.5 of 0.8080.808, and AUCρ = 0.7950.795 over the offload-budget sweep. It outperforms the strongest weak-conditioned baseline we could construct (XGBoost on MORIC: 0.4720.472 / 0.8040.804 / 0.7940.794) on every metric, despite running before the weak detector and never seeing its features. Both EdgeML (Qiu et al., 2024) and DCSB (Cao et al., 2023) sit well below either family. The MORIC+^\!+ regression target lands between the two families and gives a useful soft signal for the smaller-trunk variant. Table 2. Routing quality on VOC (single-pass, seed 4242, mAP at IoU 0.50.5). Spearman ρs _s vs. the per-frame oracle; Peak mAP and AUCρ (area under the mAP–ρ curve) over ρ∈[0,1]ρ∈[0,1] with step 0.10.1. Green deltas on Peak mAP are relative % over always-weak (0.7600.760). weak-conditioned methods additionally pay Cw=4.49C_w=4.49 GFLOPs upstream. “–” = no continuous score. Estimator ρs _s Peak mAP AUCρ GFLOPs Reference points Always weak 0.0000.000 0.7600.760 0%== 0.7600.760 0 Always strong 0.0000.000 0.7910.791 4.084.08%▲ 0.7910.791 0 Uniform random 0.0000.000 – 0.7770.777 0 Oracle (ΔAP \!AP) – 0.8270.827 8.828.82%▲ 0.8160.816 – Weak-conditioned (after weak) EdgeML (Qiu et al., 2024) 0.1380.138 0.7930.793 4.344.34%▲ 0.7840.784 ≈0≈ 0 DCSB (Cao et al., 2023) 0.3590.359 0.7890.789 3.823.82%▲ – ≈0≈ 0 XGBoost (Chen and Guestrin, 2016) / MORIC (Ours) 0.4720.472 0.8040.804 5.795.79%▲ 0.7940.794 ≈0≈ 0 Weak-skipping (before weak) MV2-Lite + MORIC+^\!+ (Ours) 0.3500.350 0.8010.801 5.395.39%▲ 0.7900.790 0.150.15 MV2-Lite + OffloadBin (Ours) 0.5570.557 0.8080.808 6.326.32%▲ 0.7950.795 0.150.15 Budget-adaptive (ρfrontier=0.3 _frontier=0.3, ρceiling=0.8 _ceiling=0.8) OffloadBin ↔ MORIC (Ours) –‡ 0.8080.808 6.326.32%▲ 0.7960.796 0.15†0.15 † Adaptive uses fcondf_cond for ρ≤0.2ρ≤0.2 and ρ≥0.8ρ≥0.8, fskipf_skip for 0.3≤ρ≤0.70.3≤ρ≤0.7 (offline arbitration, Eq. 8). Reported GFLOPs are the skipping-branch cost; the weak-conditioned branch additionally pays CwC_w. ‡ Spearman is undefined for an arbiter that selects a different score per ρ; both component scores’ ρs _s are reported above. These results indicate that raw images are sufficient for object-detection routing. With 0.150.15 GFLOPs of estimator compute, we obtain a routing signal that exceeds a strong proposal-feature baseline whose signal depends on the weak model’s 4.494.49 GFLOPs. This confirms the suspicion raised in § 1: predicting whether a frame is difficult is substantially cheaper than predicting what it contains. 4.2. The Compute Advantage is Material Figure 4. Expected per-frame cost vs. offload budget ρ on VOC. Each bar stacks estimator CeC_e (blue), weak detector (orange), and cloud ρCsρ C_s (gray). Paired bars are C (weak-conditioned, weak on every frame) and S (weak-skipping, weak on only the (1−ρ)(1-ρ) kept frames); each pair’s label is the cond−-skip saving (pink when weak-skipping costs more). (a) Compute (GFLOPs, broken y-axis; CeC_e too small to see). (b) Latency (ms), which adds the serial weak-pass dependency. Fig. 4 decomposes the expected per-frame cost for weak-conditioned (C) and weak-skipping (S) at each budget ρ into three stacked components: estimator CeC_e (blue), weak detector (orange), and cloud ρCsρ C_s (gray). The cloud band dominates both placements equally (ρCsρ C_s), so the saving comes entirely from the device side: in the S bars the orange weak-detector band shrinks with (1−ρ)Cw(1-ρ)C_w instead of the constant CwC_w paid in C, at the cost of adding a thin blue estimator band (Ce=0.15C_e=0.15 GFLOPs). Compute (panel a). The GFLOP saving is monotone in ρ: from 0.30.3 GFLOPs at ρ=0.1ρ=0.1 through 1.61.6 at ρ=0.4ρ=0.4 and 2.52.5 at ρ=0.6ρ=0.6, to 3.93.9 at ρ=0.9ρ=0.9, approaching one full weak pass (4.494.49 GFLOPs) as ρ→1ρ→1. The breakeven is very low, ρ∗=Ce/Cw≈0.03ρ =C_e/C_w≈0.03: weak-skipping is cheaper for essentially all operating budgets, since the estimator pays for itself once it avoids more than one weak pass in ∼30 30 frames. Latency (panel b). The wall-clock picture differs because it captures a serial dependency. A weak-conditioned router cannot produce a score until the weak forward pass finishes (24.7024.70 ms), whereas a weak-skipping router runs fskipf_skip (3.083.08 ms) in its place on offloaded frames. At ρ=0.1ρ=0.1 weak-skipping is 0.60.6 ms slower (pink label) because the estimator overhead slightly exceeds the small saving from skipping only 10%10\% of weak passes. The crossover sits at ρms∗≈0.12ρ _ms≈0.12. Beyond it the saving grows near-linearly, up to ∼30% 30\% end-to-end latency reduction. For latency-critical deployments with higher ρ , this saving alone justifies the weak-skipping placement. Table 3. End-to-end mAP@0.5 on VOC across offload budgets ρ∈0.1,…,0.9ρ∈\0.1,…,0.9\ (single seed). Each cell is the accuracy at that ρ. ▲ beats and ∙ matches the always-strong model (0.7910.791), an unmarked cell is below it. Bold is the column best. On the budget-adaptive row a superscript marks the selected branch (S == weak-skipping, C == weak-conditioned). ρ 0.10.1 0.20.2 0.30.3 0.40.4 0.50.5 0.60.6 0.70.7 0.80.8 0.90.9 Reference Weak only (constant) 0.7600.760 0.7600.760 0.7600.760 0.7600.760 0.7600.760 0.7600.760 0.7600.760 0.7600.760 0.7600.760 Strong only (constant) 0.7910.791 0.7910.791 0.7910.791 0.7910.791 0.7910.791 0.7910.791 0.7910.791 0.7910.791 0.7910.791 Random offloader 0.7650.765 0.7690.769 0.7720.772 0.7760.776 0.7780.778 0.7810.781 0.7840.784 0.7860.786 0.7880.788 Oracle (ΔAP \!AP) 0.7990.799 ▲ 0.8150.815 ▲ 0.8230.823 ▲ 0.8270.827 ▲ 0.8270.827 ▲ 0.8270.827 ▲ 0.8260.826 ▲ 0.8240.824 ▲ 0.8170.817 ▲ Weak-conditioned EdgeML (Qiu et al., 2024) 0.7600.760 0.7650.765 0.7890.789 0.7930.793 ▲ 0.7910.791 ∙ 0.7910.791 ∙ 0.7910.791 ∙ 0.7910.791 ∙ 0.7910.791 ∙ XGBoost on MORIC (Ours) 0.7760.776 0.7860.786 0.7940.794 ▲ 0.7990.799 ▲ 0.8030.803 ▲ 0.8040.804 ▲ 0.8030.803 ▲ 0.8030.803 ▲ 0.7980.798 ▲ Weak-skipping MV2 + MORIC+^\!+ (Ours) 0.7710.771 0.7780.778 0.7850.785 0.7910.791 ∙ 0.7960.796 ▲ 0.8000.800 ▲ 0.8010.801 ▲ 0.8010.801 ▲ 0.7980.798 ▲ MV2 + OffloadBin (Ours) 0.7730.773 0.7850.785 0.7990.799 ▲ 0.8060.806 ▲ 0.8080.808 ▲ 0.8050.805 ▲ 0.8040.804 ▲ 0.7990.799 ▲ 0.7940.794 ▲ Budget-adaptive (regime in superscript) MV2/OffloadBin ↔ XGB/MORIC (Ours) 0.776C0.776^C 0.786C0.786^C 0.799S0.799^S ▲ 0.806S0.806^S ▲ 0.808S0.808^S ▲ 0.805S0.805^S ▲ 0.804S0.804^S ▲ 0.803C0.803^C ▲ 0.798C0.798^C ▲ Figure 5. End-to-end mAP@0.5 vs. offload budget ρ on VOC. budget-adaptive (teal) is the pointwise upper hull of weak-skipping (blue) and weak-conditioned (orange), with crossovers at ρfrontier≈0.3 _frontier≈0.3, ρceiling≈0.8 _ceiling≈0.8. 4.3. Budget-adaptive Achieves Upper Envelope (1) Neither fixed placement dominates the operating curve. Fig. 5 shows weak-conditioned (XGBoost on MORIC) mAP leads at low and very high budgets (ρ≤0.2ρ≤0.2 and ρ≥0.8ρ≥0.8, by 0.10.1–0.40.4 p), while weak-skipping (MV2-Lite + OffloadBin) dominates in the mid-budget regime (ρ∈0.3,…,0.7ρ∈\0.3,…,0.7\). Tab. 3 shows that two crossovers bracket the skipping band: at ρfrontier≈0.3 _frontier≈0.3 the weak forward pass on offloaded frames becomes a pure tax and skipping it yields both compute savings and better frame selection; at ρceiling≈0.8 _ceiling≈0.8 the few remaining local frames carry enough weight that the richer weak-conditioned signal again dominates the per-frame compute saved by skipping, also shown in Fig. 5 frame selection quality-wise. (2) Budget-adaptive traces the upper envelope of both placements and strictly dominates all baselines. A two-threshold offline arbiter (ρfrontier=0.3 _frontier=0.3, ρceiling=0.8 _ceiling=0.8) produces the per-budget maximum (last row of Tab. 3, teal curve in Fig. 5): it picks weak-conditioned at ρ≤0.2ρ≤0.2 and ρ≥0.8ρ≥0.8, weak-skipping at 0.3≤ρ≤0.70.3≤ρ≤0.7. This envelope achieves the highest peak mAP@0.5 of any router, 0.8080.808, which is +1.7+1.7 p over the strong model itself, and the highest AUC=ρ0.796_ρ=0.796 (Tab. 2). Compared with prior weak-conditioned SOTAs, the budget-adaptive system leads EdgeML (Qiu et al., 2024) by 0.70.7–2.12.1 p across the full budget range (the strong model itself leads the weak model by only 3.13.1 p). The widest gap is at ρ=0.2ρ=0.2 (0.7860.786 vs. 0.7650.765) and the narrowest is at ρ=0.9ρ=0.9 (0.7980.798 vs. 0.7910.791) where EdgeML’s native threshold saturates to always-offload. DCSB (Cao et al., 2023) is a fixed binary rule locked to a single operating point (ρ≈0.79ρ≈0.79, mAP=0.789=0.789). At the same budget the budget-adaptive router reaches 0.8030.803, and its peak (0.8080.808) exceeds DCSB by 1.91.9 p while offering continuous budget tunability. The envelope sits 1.9∼2.91.9 2.9 p below the offline oracle, bounding what routing quality alone can recover. (3) The dominance extends to the joint accuracy-compute frontier inside the skipping band. For 0.3≤ρ≤0.70.3≤ρ≤0.7, budget-adaptive inherits weak-skipping’s compute profile: 2.12.1 GFLOPs/frame less than any weak-conditioned method at ρ=0.5ρ=0.5 and 3.03.0 GFLOPs/frame less at ρ=0.7ρ=0.7 (the top of the band; Fig. 4a), with wall-clock savings reaching 14.214.2 ms (∼26% 26\% reduction) at ρ=0.7ρ=0.7 (Fig. 4b). At ρ<ρfrontierρ< _frontier the weak pass runs on nearly all frames regardless, so switching to weak-conditioned costs no incremental compute. At ρ≥ρceilingρ≥ _ceiling budget-adaptive has the option to trade the skipping compute saving for the 0.40.4 p accuracy lift of weak-conditioned. Inside the skipping band the budget-adaptive router is therefore never worse in accuracy and never worse in compute than either fixed placement alone. 5. Related Work Selective-offloading admits a clean taxonomy along one architectural axis: when does the estimator fire, relative to the weak model? Weak-conditioned estimators. Shown in Fig. 2-middle, the estimator fires after the weak model and consumes its outputs. EdgeML (Qiu et al., 2024) regresses on the top-2525 proposal features and DCSB (Cao et al., 2023) hand-crafts a difficult-case discriminator. Confidence thresholds and learned detection embeddings (Wang et al., 2024) also belong to this class. They exploit rich signals but pay implicit compute tax unconditionally. Mid-network gates and early exit. This is essentially weak-conditioned estimators. BranchyNet (Teerapittayanon et al., 2016), MSDNet (Huang et al., 2017), and SDN (Kaya et al., 2019) fuse the estimator into the weak network and gate at intermediate layers. They reduce average weak-model cost on easy frames but do not skip the ”Early Pass” entirely. Neurosurgeon (Kang et al., 2017) and SPINN (Laskaridis et al., 2020; Li et al., 2018; Matsubara et al., 2022) partition or progressively split a single network across device and cloud, which is related but distinct. Cascaded inference. Classical cascades (Viola and Jones, 2001) and modern cascade detectors (Cai and Vasconcelos, 2018) escalate on uncertainty. As detailed in § 3.2, they use the cheap stage to produce predictions whereas budget-adaptive routing uses it to route. Weak-skipping estimators. The estimator fires before the weak model from the raw image. Early image-complexity heuristics and learned image-level routers in our work belong to this class. They allow the weak pass to be skipped entirely on offloaded frames but have often been considered too weak for detection routing, a premise challenged by our results. 6. Conclusion and Future Work Estimator placement is a design axis the selective-offloading literature has implicitly fixed. We exhibited both a lightweight image-only weak-skipping estimator and a weak-conditioned estimator that outperform the SOTAs that we could reproduce from the literature, on both routing quality and compute. We showed the placement choice is regime-dependent and built a budget-aware budget-adaptive router that traces the upper accuracy envelope of both placements on VOC while saving compute and latency within its skipping band. Several directions remain. A real edge-cloud testbed would replace our offline simulation with measured round-trips (Geng et al., 2025) on real-world hardwares. Testing beyond VOC such as COCO(Lin et al., 2014) would probe how far the raw-pixel routing signal carries. Beyond mAP@0.5, stricter AP@[.5:.95][.5\!:\!.95], class-weighted or downstream-task utility will be experimented. A skip-then-reconsider cascade would run the weak-skipping estimator first and, on frames kept local where the weak pass executes anyway, apply a weak-conditioned estimator to revisit the offload decision on the now-available weak features at no extra detector cost. A head-to-head with early-exit routers across weak-model prefix depths would chart when a shallow prefix already rivals the image-only weak-skipping estimator. Acknowledgements.This work was supported by the Dutch National Growth Fund “Future Network Services”. References (1) Cai and Vasconcelos (2018) Zhaowei Cai and Nuno Vasconcelos. 2018. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. 6154–6162. Canel et al. (2019) Christopher Canel, Thomas Kim, Giulio Zhou, Conglong Li, Hyeontaek Lim, David G Andersen, Michael Kaminsky, and Subramanya Dulloor. 2019. Scaling Video Analytics on Constrained Edge Nodes. In Proceedings of Machine Learning and Systems, A. Talwalkar, V. Smith, and M. Zaharia (Eds.), Vol. 1. 406–417. https://proceedings.mlsys.org/paper_files/paper/2019/file/6bcfac823d40046dca25ef6d6d59c3f-Paper.pdf Cao et al. (2023) Zhiqiang Cao, Zhijun Li, Yongrui Chen, Heng Pan, Youbing Hu, and Jie Liu. 2023. Edge-cloud collaborated object detection via difficult-case discriminator. In 2023 IEEE 43rd International Conference on Distributed Computing Systems (ICDCS). IEEE, 259–270. Chen and Guestrin (2016) Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining. 785–794. Chen et al. (2015) Tiffany Yu-Han Chen, Lenin Ravindranath, Shuo Deng, Paramvir Bahl, and Hari Balakrishnan. 2015. Glimpse: Continuous, real-time object recognition on mobile devices. In Proceedings of the 13th ACM conference on embedded networked sensor systems. 155–168. Everingham et al. (2010) Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. 2010. The pascal visual object classes (voc) challenge. International journal of computer vision 88, 2 (2010), 303–338. Geng et al. (2025) Wei Geng, Oguz Kagan Altas, David Guzman, Giovanni Bartolomeo, Nitinder Mohan, and Joerg Ott. 2025. Poster: KUT: Towards Lightweight On-path Network Assessment for Edge Orchestration. In Proceedings of the 21st International Conference on emerging Networking EXperiments and Technologies. 9–11. Geng et al. (2026) Wei Geng, Xiang Su, Nitinder Mohan, Jörg Ott, and Pan Hui. 2026. SMOOTH: Scalable Multitask Offloading with Backbone Sharing. In 2026 IFIP Networking Conference (IFIP Networking). IFIP, 1–10. Guo et al. (2017) Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In International conference on machine learning. PMLR, 1321–1330. He et al. (2016) Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 770–778. Howard et al. (2019) Andrew Howard, Mark Sandler, Bo Chen, et al. 2019. Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Huang et al. (2017) Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017. Multi-scale dense networks for resource efficient image classification. arXiv preprint arXiv:1703.09844 (2017). Jiang et al. (2018) Junchen Jiang, Ganesh Ananthanarayanan, Peter Bodik, Siddhartha Sen, and Ion Stoica. 2018. Chameleon: scalable adaptation of video analytics. In Proceedings of the 2018 conference of the ACM special interest group on data communication. 253–266. Kang et al. (2017) Yiping Kang, Johann Hauswald, Cao Gao, Austin Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang. 2017. Neurosurgeon: Collaborative intelligence between the cloud and mobile edge. ACM SIGARCH Computer Architecture News 45, 1 (2017), 615–629. Kaya et al. (2019) Yigitcan Kaya, Sanghyun Hong, and Tudor Dumitras. 2019. Shallow-deep networks: Understanding and mitigating network overthinking. In International conference on machine learning. PMLR, 3301–3310. Laskaridis et al. (2020) Stefanos Laskaridis, Stylianos I Venieris, Mario Almeida, Ilias Leontiadis, and Nicholas D Lane. 2020. SPINN: Synergistic progressive inference of neural networks over device and cloud. In Proceedings of the 26th annual international conference on mobile computing and networking. 1–15. Li et al. (2018) En Li, Zhi Zhou, and Xu Chen. 2018. Edge intelligence: On-demand deep learning model co-inference with device-edge synergy. In Proceedings of the 2018 workshop on mobile edge communications. 31–36. Li et al. (2020) Yuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang, Guoqing Harry Xu, and Ravi Netravali. 2020. Reducto: On-camera filtering for resource-efficient real-time video analytics. In Proceedings of the Annual conference of the ACM Special Interest Group on Data Communication on the applications, technologies, architectures, and protocols for computer communication. 359–376. Lin et al. (2017a) Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017a. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2117–2125. Lin et al. (2017b) Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017b. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision. 2980–2988. Lin et al. (2014) Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision (ECCV). Matsubara et al. (2022) Yoshitomo Matsubara, Marco Levorato, and Francesco Restuccia. 2022. Split computing and early exiting for deep learning applications: Survey and research challenges. Comput. Surveys 55, 5 (2022), 1–30. Qiu et al. (2024) Jiaming Qiu, Ruiqi Wang, Brooks Hu, Roch Guérin, and Chenyang Lu. 2024. Optimizing edge offloading decisions for object detection. In 2024 IEEE/ACM Symposium on Edge Computing (SEC). IEEE, 164–177. Ren et al. (2015) Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Advances in Neural Information Processing Systems (NeurIPS). Sandler et al. (2018) Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition. 4510–4520. Teerapittayanon et al. (2016) Surat Teerapittayanon, Bradley McDanel, and Hsiang-Tsung Kung. 2016. Branchynet: Fast inference via early exiting from deep neural networks. In 2016 23rd international conference on pattern recognition (ICPR). IEEE, 2464–2469. Viola and Jones (2001) Paul Viola and Michael Jones. 2001. Rapid object detection using a boosted cascade of simple features. In Proceedings of the 2001 IEEE computer society conference on computer vision and pattern recognition. CVPR 2001, Vol. 1. Ieee, I–I. Wang et al. (2024) Qingyuan Wang, Barry Cardiff, Antoine Frappé, Benoit Larras, and Deepu John. 2024. Tiny models are the computational saver for large models. In European Conference on Computer Vision. Springer, 163–182. Appendix A Supplementary Evidence All numbers below are from the same single-pass VOC evaluation (seed 4242) used in § 4, profiled on the setup of Tab. 1. Peak mAP is the maximum end-to-end accuracy over the offload-budget ρ sweep. mAP@0.5 is the VOC metric and AP@[.5:.95] (COCO-style; written APC_C in table headers) is the stricter metric. These appendices backs and justifies three design choices made in § 3 (the learning target, the estimator backbone, and the claim that routing signal is recoverable from the raw image) and probe robustness to a stricter accuracy metric. A.1. Backbone Ablation Table 4. Backbone ablation for the weak-skipping estimator, controlling for the learning target: both trunks use the identical OffloadBin/focal target and differ only in the backbone (VOC test, seed 4242). GFLOPs and parameters are for the estimator alone. Trunks we trained on other targets are not directly comparable and are omitted. Backbone (OffloadBin/focal) GFLOPs Par. (M) ρs _s Peak mAP MobileNetV2-Lite 0.1530.153 0.540.54 0.5570.557 0.8080.808 EfficientNet-B0-Lite 0.1760.176 0.850.85 0.2350.235 0.7950.795 Our backbone sweep is small, so we report it only briefly. Tab. 4 holds the learning target fixed at OffloadBin/focal and varies the trunk alone: the larger EfficientNet-B0-Lite does not improve routing over the compact MobileNetV2-Lite, with a markedly lower rank correlation and no gain in peak mAP. This is consistent with the premise of § 1 that detecting difficulty needs little capacity, though a broader sweep is left to future work. A.2. Learning-Target Ablation Table 5. Learning-target ablation on the fixed MobileNetV2-Lite weak-skipping backbone (VOC test, seed 4242). ρs _s is the Spearman correlation between the estimator score and the per-frame oracle gain (not the offload budget ρ), and ratio-err is the mean absolute gap between realized and requested offload fractions. Exact target definitions are in our released artifacts. Rows are sorted by ρs _s within each group, best per column in bold. Target / loss Description ρs _s Peak mAP Peak APC_C ratio-err Classification target (focal) OffloadBin Binary [ΔAP>0]1[ \!AP>0]: does the cloud strictly help this frame? Focal loss handles the positive-class imbalance. 0.5570.557 0.8080.808 0.6090.609 0.0090.009 TopQuartile Binary: is the frame’s gain in the top quartile of ΔAP \!AP? A rarer, harder positive class than OffloadBin. 0.3890.389 0.8000.800 0.6060.606 0.0070.007 Regression target MORIC+^\!+ Signed empirical CDF of ΔAP \!AP, split at zero onto [−1,1][-1,1]; this is our reward (Eq. 5). 0.3500.350 0.8010.801 0.6070.607 0.0180.018 HighIoUGain Continuous per-frame gain proxy that up-weights tightly-localised, high-IoU matches. 0.3450.345 0.7940.794 0.6070.607 0.0080.008 MORIC+^\!+ (quantile) The MORIC+^\!+ target, fit with a quantile (pinball) regression loss. 0.3130.313 0.8010.801 0.6080.608 0.0100.010 F1Gain Continuous per-frame proxy: the change in detection F1 when the frame is offloaded. 0.3060.306 0.7910.791 0.6050.605 0.0060.006 MORIC+^\!+ (wing) The MORIC+^\!+ target, fit with a wing regression loss. 0.3030.303 0.7990.799 0.6060.606 0.0240.024 RescueRatio Continuous proxy for the share of a frame’s missed objects that the cloud recovers. 0.2910.291 0.7920.792 0.6050.605 0.0040.004 WorstCaseGain Continuous proxy that emphasises the frame’s worst-case detection loss. 0.2810.281 0.7960.796 0.6070.607 0.0100.010 SigMORIC The MORIC+^\!+ reward passed through a sigmoid squashing. 0.2730.273 0.7970.797 0.6050.605 0.0150.015 MORIC⋆ The MORIC+^\!+ reward under an alternative CDF reshaping. 0.2710.271 0.7970.797 0.6050.605 0.0150.015 Φ -MORIC The MORIC+^\!+ reward under a Φ -based CDF reshaping. 0.2630.263 0.7970.797 0.6050.605 0.0150.015 RescueRatio (wing) The RescueRatio target, fit with a wing regression loss. 0.1690.169 0.7910.791 0.6050.605 0.0110.011 § 3.1 reduces the long-tailed per-frame ΔAP \!AP to a binary OffloadBin label trained with focal loss, rather than regressing a continuous reward. Because the backbone and the target must be chosen jointly, a chicken-and-egg dependency, Tab. 5 fixes the backbone first and sweeps the target. OffloadBin wins by a wide margin (ρs=0.557 _s=0.557, against 0.3890.389 for the next-best target and 0.169∼0.3500.169 0.350 for the continuous-regression variants) and tops both accuracy metrics. Budget tracking is uniformly tight, so this gap reflects the target itself rather than calibration. The result confirms the § 3.1 argument: collapsing whether offloading helps into a binary label decouples hardness from content and lets raw pixels suffice, whereas regressing the signed magnitude wastes capacity on the 62.6%62.6\% zero spike (Fig. 3). A.3. Where Offloading Helps Figure 6. Where offloading helps, by weak-detector stratum (VOC test, seed 4242, N=3105N=3105). Each bar splits a quartile’s frames into Benefit (offloading raises ΔAP \!AP), Harm (lowers it), and Neutral (unchanged). The dominant neutral mass shows the benefit is a sparse partition. (a) Across weak-confidence quartiles benefit falls 6.4×6.4× (Q1→ 4). (b) Across scene crowding (weak detection count) it rises 5.1×5.1×. The right strip is the per-frame weak→ mAP gain in points (p), largest where benefit concentrates. A weak-skipping estimator presumes that which frames benefit from the cloud is both structured and recoverable from the raw image. Fig. 6 speaks to the first half: the benefit is sharply concentrated, with the lowest weak-confidence quartile offload-beneficial 6.4×6.4× more often than the highest (38.9%38.9\% vs. 6.1%6.1\%) and the most crowded scenes 5.1×5.1× more often than the sparsest (46.9%46.9\% vs. 9.2%9.2\%), tabulated in full in Tab. 6. The same strata carry the largest weak→ mAP headroom. These strata are defined by weak-detector outputs, however, so these views localise where offloading helps rather than showing the signal is recoverable before the weak pass. We assume that low confidence and crowding are correlates of scene complexity that is plausibly legible from raw pixels and is established directly by the weak-skipping estimator’s measured routing quality (Tab. 5). Table 6. Offload benefit broken down by weak-detector stratum on the VOC test set (seed 4242, N=3105N=3105 frames). We split the frames into quartiles Q1–Q4 of each weak-detector summary, so that Q1 contains the least-confident or sparsest scenes and Q4 the most-confident or most-crowded. The Benefit, Harm, and Neutral columns give the fraction of frames in each stratum for which offloading respectively raises, lowers, or leaves ΔAP \!AP unchanged, and they sum to one. The final two columns, mAPw and mAPs, report the mean per-frame mAP of the weak and strong detectors within the stratum. Stratum Q Benefit Harm Neutral mAPw mAPs Weak conf. (mean) Q1 0.3890.389 0.1690.169 0.4430.443 0.7500.750 0.8210.821 Q2 0.3120.312 0.1730.173 0.5150.515 0.8470.847 0.8840.884 Q3 0.2460.246 0.1210.121 0.6330.633 0.8950.895 0.9200.920 Q4 0.0610.061 0.0300.030 0.9110.911 0.9610.961 0.9730.973 Weak det. count Q1 0.0920.092 0.0370.037 0.8720.872 0.9280.928 0.9550.955 Q2 0.2170.217 0.0930.093 0.6900.690 0.8640.864 0.9050.905 Q3 0.3430.343 0.1510.151 0.5070.507 0.8310.831 0.8770.877 Q4 0.4690.469 0.2690.269 0.2620.262 0.7840.784 0.8240.824 The benefit structure dictates the target. Tabs. 5 and 6 are two views of one phenomenon. The offload benefit is a sparse partition, not a smooth magnitude field: a dominant neutral mass (62.6%62.6\% of frames at ΔAP=0 \!AP=0; median gain 0 in every stratum) around a minority of beneficial frames (≤47%≤47\% even at best) in low-confidence, crowded scenes. The learnable question is thus whether a frame lies in that region, not by how much it gains, which is why Tab. 5 ranks both classification targets (OffloadBin ρs=0.557 _s=0.557, TopQuartile 0.3890.389) above every regression variant (ρs≤0.350 _s≤0.350), and why routing reduces to recognising a region of image space that a cheap raw-image trunk can read.