Paper deep dive
Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment
Haoran Liu, Mingzhe Liu, Peng Li, Guibin Zan
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Infrared-visible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked by scalar objective metrics that formalize different proxies for information transfer, structure, or source similarity. These proxies often disagree with the judgment that ultimately matters: given the same sources, which of two fused results does a human prefer? Direct pairwise comparison is an established reference protocol for relative subjective assessment, but its cost grows quadratically with the number of algorithms, which prevents routine use. We present the Learned Perceptual Image Fusion Measure (LPIFM), a source-conditioned model that operationalizes the human A/B/Tie comparison protocol as a repeatable, scalable surrogate. LPIFM jointly observes the infrared source, the visible source, and two fused candidates, and predicts whether candidate A is better, candidate B is better, or the two are perceptually equivalent. Supervision comes from a new dense preference corpus that covers every unordered comparison among a broad pool of fusion methods on the scenes of a public benchmark, labeled under a blinded, randomized, two-stage protocol with expert adjudication. Across scene- and method-generalization settings, LPIFM tracks human pairwise decisions closely and reproduces the tie-aware Bradley-Terry rankings derived from human labels; on full method pools it surpasses the strongest conventional metric by a wide margin in both pairwise accuracy and ranking correlation. We release the annotated preference dataset, together with the LPIFM model weights, source code, and evaluation code, to support preference-aligned IVIF assessment. LPIFM offers a practical instrument for human-aligned method comparison and ranking at scale.
Tags
Links
- Source: https://arxiv.org/abs/2608.01301v1
- Canonical: https://arxiv.org/abs/2608.01301v1
Trouble viewing inline? Open PDF directly â
Full Text
71,134 characters extracted from source content.
Expand or collapse full text
1 Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for InfraredâVisible Fusion Assessment Haoran Liu 1,2 , Mingzhe Liu 1,2,* , Peng Li 2 , Guibin Zan 3,* 1 School of Artificial Intelligence and Electronic Engineering, Sichuan Technology and Business University, Chengdu 611745, China 2 Chengdu University of Technology, College of Nuclear Technology and Automation Engineering, Chengdu 610059, China 3 Sigray, Inc., Concord, CA 94520, USA *Corresponding author, liumz@cdut.edu.cn (Mingzhe Liu), gbzan@sigray.com (Guibin Zan) ABSTRACT Infraredâvisible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked by scalar objective metrics that formalize different proxies for information transfer, structure, or source similarity. These proxies often disagree with the judgment that ultimately matters: given the same sources, which of two fused results does a human prefer? Direct pairwise comparison is an established reference protocol for relative subjective assessment, but its cost grows quadratically with the number of algorithms, which prevents routine use. We present the Learned Perceptual Image Fusion Measure (LPIFM), a source-conditioned model that operationalizes the human A/B/Tie comparison protocol as a repeatable, scalable surrogate. LPIFM jointly observes the infrared source, the visible source, and two fused candidates, and predicts whether candidate A is better, candidate B is better, or the two are perceptually equivalent. Supervision comes from a new dense preference corpus that covers every unordered comparison among a broad pool of fusion methods on the scenes of a public benchmark, labeled under a blinded, randomized, two-stage protocol with expert adjudication. Across scene- and method-generalization settings, LPIFM tracks human pairwise decisions closely and reproduces the tie-aware BradleyâTerry rankings derived from human labels; on full method pools it surpasses the strongest conventional metric by a wide margin in both pairwise accuracy and ranking correlation. We release the annotated preference dataset, together with the LPIFM model weights, source code, and evaluation code, to support preference-aligned IVIF assessment. LPIFM offers a practical instrument for human-aligned method comparison and ranking at scale. Keywords: infraredâvisible image fusion; image quality assessment; pairwise preference learning; perceptual metric; subjective evaluation; BradleyâTerry ranking 2 1. INTRODUCTION Fig. 1. Which of two fused results is preferred? Four representative comparisons from the study corpus (Case 1: kettle, SwinFusion vs. RP_SR; Case 2: elecbike, U2Fusion vs. HMSD_GF; Case 3: labMan, ADF vs. RP_SR; Case 4: tricycle, TIF vs. NSCT_SR). For each case, the visible (VI) and infrared (IR) sources are shown above the two fused candidates, and the check marks below indicate which candidate each evaluator prefers. Human observers and LPIFM agree in all four cases, whereas widely used objective metrics (Entropy, Qabf, VIFF, SCD) frequently prefer the candidate that humans reject. Infraredâvisible image fusion (IVIF) combines complementary thermal and reflectance information into a single image and now supports surveillance, driving assistance, and other downstream perception tasks. Because no ideal fused reference exists, the field ranks fusion algorithms with scalar objective metrics, and benchmark studies organize dozens of such metrics into information-theoretic, feature-based, structural-similarity, and human-perception-inspired families [1]. Each metric formalizes a different proxy for fusion quality. The VIFB benchmark explicitly notes that no single implemented metric dominates the others [1], and recent studies continue to report systematic disagreement among objective metrics on IVIF outputs [5â6]. Fig. 1 makes the practical consequence concrete: on typical benchmark scenes, established metrics frequently prefer the fused result that human observers reject. The four cases expose two recurring failure patterns. In Cases 1 and 2, RP_SR and HMSD_GF lose salient infrared content: the pedestrian that stands out in the infrared source of the kettle scene is no longer discernible in RP_SRâs output, and the rider visible in the infrared source of the elecbike scene is swallowed by amplified glare in HMSD_GFâs output. Humans and LPIFM penalize this loss of infrared information, one of the decisive scoring criteria in IVIF, whereas Entropy, Qabf, VIFF, and SCD still prefer these results. In Cases 3 and 4, RP_SR and NSCT_SR introduce conspicuous artifacts (color speckle noise across the labMan scene; a dark blotch on the road and halos around the lights in the tricycle scene), and most of the same metrics reward them, because spurious structure inflates the statistics, gradients, and contrast that these proxies score as transferred information. An evaluation practice built on such proxies can therefore steer method development away from the perceptual quality it is meant to capture. Human comparison answers the relevant question directly. Comparative subjective testing is an established methodology for relative fusion assessment and for validating objective metrics [2], and pairwise human judgments have been aggregated into method rankings with Thurstone and BradleyâTerry models [3]. Among common subjective procedures, forced-choice pairwise comparison has shown the smallest measurement variance under controlled conditions [4]. The obstacle is operational, not conceptual: evaluating n algorithms requires O(nÂČ) comparisons per scene, and every new algorithm, dataset, or ablation re-incurs the full cost of recruiting, instructing, and adjudicating observers. As algorithm pools grow, the community faces a widening gap between the evaluation protocol it trusts and the evaluation protocol it can afford to run. 3 Learning offers a way to remove this recurring cost, and learned IVIF assessors have begun to appear. Existing models, however, answer tasks that differ from the trusted protocol itself. Reference-based pairwise learners such as PieAPP assume a pristine reference image, which IVIF lacks [10]. Recent IVIF-specific assessors predict absolute quality scores, distill conventional metric panels, or elicit scores from multimodal language models [12â15]. None of these directly reproduces the decision that human observers actually make in the trusted protocol: a ternary A- better/B-better/Tie judgment about two candidates conditioned on the same infrared and visible sources. An absolute scorer must invent a scale that observers were never asked to use, and forcing a winner on perceptually indistinguishable pairs discards information that observers deliberately express. In this paper we operationalize the human comparison protocol itself. We present the Learned Perceptual Image Fusion Measure (LPIFM), a source-conditioned model that takes the infrared source, the visible source, and two fused candidates, and predicts the same ternary decision that human observers produce. LPIFM is trained on a new dense preference corpus covering all 6,300 unordered comparisons generated by 25 fusion methods on the 21 scenes of the VIFB benchmark, labeled under a blinded, randomized, two-stage protocol with expert adjudication. A tie-aware objective preserves human indifference as a first-class outcome instead of forcing arbitrary winners. Because LPIFM outputs protocol-level decisions, its predictions plug directly into tie- aware BradleyâTerry aggregation, so a complete, human-aligned ranking of an arbitrary method pool can be computed automatically at negligible marginal cost. Our experiments evaluate LPIFM at both the pair level and the ranking level under explicit scene- and method-generalization settings. Across four evaluation slices, LPIFM attains pairwise accuracy of 0.792â0.840 and Spearman correlation of 0.941â0.977 with human-derived rankings. Its decisive advantage appears exactly where scalable evaluation is needed most: on full 25- method pools that contain many near-ranked competitors, LPIFM maintains accuracy of 0.792â 0.840 while the strongest conventional metric, SCD, falls to 0.629. This advantage has a boundary. On one six-method held-out subset whose members span nearly the full range of human-rated quality, SCD is locally competitive, and we diagnose why this strength does not transfer to realistic pools. These results indicate that trustworthy IVIF evaluation becomes affordable through direct alignment with the trusted protocol, not through better handcrafted proxies. The contributions of this work are as follows. 1. Protocol-aligned problem formulation. We define IVIF quality assessment as source- conditioned ternary preference prediction over two fused candidates, matching the decision space of the trusted human comparison protocol rather than an invented absolute scale. 2. A dense human preference corpus, publicly released. We construct and publicly release an annotated pairwise preference dataset covering all 6,300 unordered comparisons among 25 methods on 21 scenes (13,125 ordered training records), collected under a blinded, randomized, two-stage labeling protocol with expert adjudication. 3. A tie-aware, source-conditioned comparator. LPIFM couples a shared hierarchical encoder with a triadic interaction module and a tie-aware objective, so that indifference is modeled explicitly rather than discarded. 4. Pair-to-ranking evaluation under explicit generalization settings. We evaluate pair decisions and induced method rankings under four scene/method slices with frozen baseline calibration, and we diagnose the one setting in which a conventional metric is locally strong. 5. A practical evaluation instrument. LPIFM, together with the released dataset, model weights, source code, and evaluation code, provides the community with a repeatable surrogate for human preference that can rank arbitrary IVIF method pools at scale. 2. RELATED WORKS 2.1 Objective quality metrics for image fusion Objective fusion metrics score a fused image against its sources through handcrafted proxies. 4 Information-theoretic metrics such as entropy [26] and mutual information [20] reward retained source statistics; feature-based metrics such as Qabf [21] reward transferred gradients; structural metrics such as SSIM [22] and MS-SSIM [29] reward preserved local structure; spatial-frequency and correlation measures follow classical image-quality practice [25]; and human-perception- inspired metrics such as Qcb [23] and Qcv [24] approximate contrast sensitivity. SCD measures the sum of correlations between source-to-fused difference images [7], FMI measures feature mutual information [8], VIFF adapts visual information fidelity to fusion [9], and Nabf quantifies fusion artifacts in the gradient-transfer family [27â28]. These metrics are cheap and deterministic, which explains their ubiquity. However, VIFB reports that no metric dominates across scenes and methods [1], and subsequent studies document persistent disagreement among metrics and with subjective evidence [5â6,49]. Broader reviews of fusion methods and metrics [44,46,50], pseudo- reference IVIF quality evaluation [47], and registrationâfusion coupling [45] further map the surrounding literature. The roots of this misalignment lie in how the scores are computed. These metrics quantify how much source signal survives into the fused image (statistics retained, gradients transferred, structure preserved) and treat every increase as a gain, without regard to whether the added content helps or harms a viewer: noise, halos, and spurious structures raise entropy, gradient energy, and local contrast, so artifacts are routinely scored as rich information. They also aggregate uniformly over pixels, so the loss of a small but decisive region, such as a salient thermal target, barely moves a global score even though it dominates human judgment. And each fused image is scored in isolation against the sources, on a scale that is not calibrated across scenes, whereas the judgment that drives method selection is comparative. Each metric answers a question about signal preservation. None of them answers the comparative question that drives method selection, which is the question this paper targets directly. 2.2 Subjective comparative evaluation and preference aggregation Subjective testing remains the reference standard for perceptual fusion quality. PetroviÄ established comparative subjective tests for fusion assessment and used them to validate objective metrics [2]. Loew et al. collected pairwise human judgments over fusion methods and aggregated them into rankings with Thurstone and BradleyâTerry models [3]. In the broader image-quality literature, Mantiuk et al. compared four subjective methodologies and found forced-choice pairwise comparison to yield the smallest measurement variance [4]. Indifference in paired- comparison aggregation has been treated by tie-aware BradleyâTerry models [16] and by classical analyses of ties in ranking [17]; related work on tie calibration for metric meta-evaluation provides complementary methodological guidance [18]. The strength of this line of work is validity: the protocol elicits exactly the judgment of interest. Its limitation is cost, which scales quadratically with the method pool and linearly with every new benchmark. Our work retains the protocol's decision space and evidence structure while removing the recurring human cost. 2.3 Learned perceptual and fusion-specific quality assessment Learning from human perceptual judgments is established in general image-quality assessment. PieAPP trains on pairwise preferences but assumes a pristine reference image and outputs a single-image error score [10], an assumption IVIF cannot satisfy. Neural Side-by-Side similarly learns from side-by-side human preferences, but for no-reference super-resolution evaluation rather than fusion [48]. Predictive quality models for fused long-wave infrared and visible imagery estimate subjective scores from natural-scene statistics [11]. For IVIF specifically, learned assessment has emerged only recently, and the few existing models answer questions that differ fundamentally from the one posed here. SRT applies a semantic-relation transformer that regresses an absolute quality score for a single fused image [12]; it neither conditions its judgment on a competing candidate nor admits ties, and no public implementation or trained scorer is available for re-evaluation under our protocol. EvaNet learns to approximate and consolidate panels of conventional evaluation signals for efficiency and consistency [13]; because its supervision target is the conventional metric panel itself, the 19 objective metrics compared in Section 4 already represent the judgment it is trained to reproduce, and entering it as an additional baseline would duplicate that control rather than extend it. EVAFusion trains an absolute reward model that grades single fused images, seeded from a small expert-annotated set and expanded 5 with labels produced by a multimodal language model, and its primary product is preference- guided fusion generation rather than method evaluation [14]. FuScore elicits continuous quality scores from multimodal large language models with uncertainty-aware supervision built on the same label lineage [15]; both are absolute per-image scorers whose supervision passes through a language-model intermediary, and neither provides a publicly usable scorer for re-evaluation. These assessors are therefore discussed here rather than entered into the main experiments: each is either mismatched to the source-conditioned pairwise task, redundant with the conventional metric panel it approximates, or not reproducible for a faithful comparison (Section 4.1). LPIFM differs from all of them along three axes at once: it is supervised by dense, directly collected human ternary comparisons rather than by metric panels, expanded labels, or language-model outputs; it performs joint two-candidate inference conditioned on both sources rather than scoring images in isolation; and it preserves ties as an explicit outcome. To our knowledge, LPIFM is among the earliest learned perceptual, fusion-specific quality assessors trained directly on human A/B/Tie pairwise comparisons, and the released dataset, model weights, and code let the community test and extend this class of assessors. 3. METHODS 3.1 Problem formulation Let and denote a registered infraredâvisible source pair, and let ( , ) denote two fused candidates produced from the same sources. We define IVIF quality assessment as source- conditioned ternary preference prediction: a model must map the quadruple ( , , , ) to a decision in A better, B better, Tie. This formulation matches the decision space of the human comparison protocol exactly. It requires no ideal reference, no absolute quality scale, and no assumption that all quality differences are resolvable, because the Tie outcome represents perceptual equivalence explicitly. Pair decisions over a method pool are aggregated into a method- level ranking with a tie-aware BradleyâTerry model (T-BT) [16], so the formulation connects local judgments to the algorithm-selection decisions that evaluation ultimately serves. 3.2 The LPIFM preference dataset Comparison corpus. The corpus is built on the 21 registered infraredâvisible source pairs of the VIFB benchmark [1] and fused outputs from 25 fusion algorithms [30â43]. For each scene, observers compared every unordered pair of distinct methods, yielding 21 Ă C(25,2) = 6,300 human comparison units, each answered as A better, B better, or Tie. For training, every non-self comparison is represented in both candidate orders with the label mirrored, producing 12,600 ordered non-self records; 525 self-pairs labeled Tie are added, for 13,125 records in total. The ordered non-self label distribution is 5,887 A, 5,887 B, and 826 Tie. Throughout the paper we distinguish the 6,300 human comparison units from the 13,125 augmented training records. Presentation and instructions. Each trial displayed four images: the visible source, the infrared source, and the two fused candidates. Candidate order was randomized and method identities were hidden. Observers judged one scene block at a time. Instructions prioritized complementary integration: retention of visible structure, texture, and readability; preservation of infrared salience and contrast; and avoidance of blur, ghosting, halo, blocking artifacts, amplified noise, unsupported structure, and false-color distortion. Two-stage label aggregation. In the first stage, 20 bachelor-level participants provided three independent labels per comparison; majority vote yielded a preliminary label, and an IVIF researcher resolved three-way splits. This stage-1 label was retained as a reference prior, not as a vote in later aggregation. In the second stage, three IVIF researchers independently re-labeled all 6,300 unordered comparisons. The researcher labels formed the primary annotation: unanimous and two-to-one majorities were accepted by default. An additional IVIF expert then performed quality control. Three-way splits (one vote each for A better, B better, and Tie) were mandatory for adjudication and received an expert final label (129/6,300; 2.0%). Two-to-one majorities were eligible for conditional review: the expert could consult domain knowledge and the stage-1 reference and overturn the majority only when it was judged clearly unreasonable. Among the 6 three stage-2 researchers, mean pairwise percent agreement was 76.5% and Fleiss' was 0.58 (pairwise Cohen's 0.54â0.60), indicating moderate inter-rater agreement with a strong majority structure. Viewing-condition records (display calibration, viewing distance, ambient illumination) were not retained and are disclosed as a reproducibility limitation in Section 5. Dataset and code release. We release the annotated preference dataset, including the 6,300 adjudicated unordered comparisons, the mapping to the 13,125 ordered training records, and the scene and method split manifests, together with the LPIFM model weights, source code, and evaluation code, so that the community can train and evaluate preference-aligned IVIF assessors (https://github.com/HaoranLiu507/LPIFM). The source code is released under the GNU Affero General Public License v3.0 (AGPL-3.0), and the model weights and datasets are released under the Creative Commons Attribution Non Commercial Share Alike 4.0 International (C BY-NC- SA 4.0) license. To our knowledge, dense exhaustive pairwise human annotation at this coverage (every method pair on every scene of a public IVIF benchmark) is not otherwise publicly available. 3.3 Architecture Fig. 2. LPIFM architecture. The infrared (IR) and visible (VI) sources are concatenated and passed through a context projection; the two fused candidates A and B are processed by the same shared image encoder (one encoder, three passes). The triadic interaction module applies bidirectional cross-attention (Ă N) among the context stream C and the candidate streams A and B. A shared preference scorer produces scores and , whose difference = â is thresholded into the ternary decision A better / Tie / B better. Fig. 2 shows the architecture. LPIFM has three components, each motivated by a property of the human protocol it must reproduce. Shared image encoder with context projection. The infrared and visible sources are concatenated and projected into a source-context representation C. Both fused candidates are encoded by the same hierarchical backbone (one encoder, three passes), so that candidate representations are extracted identically and no candidate-specific parameters can bias the comparison. We use a pretrained ConvNeXt V2 Base backbone at 384 Ă 384 input resolution [19]. The choice of a hierarchical convolutional backbone is deliberate: the backbone ablation in Section 4.5 shows a spread of more than 11 percentage points across eight architectures, indicating that preference learning depends on hierarchical local structure rather than on generic attention capacity. Triadic interaction. Human observers do not judge candidates in isolation; they look back 7 and forth between each candidate and the sources, and between the two candidates. The triadic interaction module models this behavior with bidirectional cross-attention among the context stream C and the two candidate streams A and B, stacked N times. The module lets each candidate representation encode how it preserves or distorts source information relative to both sources and relative to its competitor, which is precisely the evidence a comparative judgment needs. Removing the module reduces validation accuracy by 1.27 percentage points (Section 4.5). Shared preference scorer. A shared scoring head maps each interacted candidate representation to a scalar preference score, and . The decision statistic is the difference = â . Sharing the scorer across candidates, like sharing the encoder, removes any architectural asymmetry between the A and B roles. 3.4 Tie-aware training objective The objective is defined on the score difference and has three terms. A preference term =softplus(â·/) acts on decisive samples (=+1 if A is preferred, â1 if B is preferred), with temperature controlling penalty sharpness. A margin term , weighted by , requires || to reach a margin on decisive samples, so that clear human preferences map to clearly separated scores. A tie-band term =relu(||â), weighted by , acts only on tie samples and compresses into the band [â,]. The full loss is = + · + · , with defaults =1.0,=0.3,=1.0, =0.5, =1.0. This decomposition separates two questions that scalar metrics conflate: which candidate is better, and whether the pair is decidable at all. The tie band is the operative mechanism for the second question, and the ablations in Section 4.5 show it is the least replaceable design choice in the model: narrowing to 0.1 collapses validation accuracy by 14.02 percentage points, because the model is then forced to pick winners for pairs that humans judged indistinguishable. 3.5 Inference and ranking aggregation At inference, LPIFM computes for a candidate pair and outputs A better if >, B better if <â, and Tie otherwise, with inference threshold =0.3. The model threshold and training band are internal to LPIFM and are distinct from the calibration parameter â used to convert baseline metric scores into ternary decisions (Section 4.1). To rank a pool of methods on a scene set, LPIFM evaluates the O( ) candidate pairs automatically and the resulting ternary decisions are aggregated with tie-aware BradleyâTerry (T-BT) into a method-level strength vector and ranking. The entire pipeline is deterministic given fixed weights, so repeated evaluations of the same pool return identical rankings. 3.6 Implementation details Models were trained for 60 epochs with batch size 4, learning rate 8Ă10â»â” under cosine annealing with 10 warmup epochs and minimum learning rate 1Ă10â»â”, weight decay 5Ă10â»âŽ, EMA and SWA weight averaging, and global seed 200. The full model has 105.24 M parameters and 137.50 GFLOPs at 384 Ă 384 input. Data splits use split seed 42 with validation ratio 0.2. Authoritative run provenance (configuration and split manifests) is archived and included in the release package. The default configuration was independently trained three times (validation accuracies 77.95, 78.01, 80.04; mean ± SD 78.67 ± 1.19); the checkpoint with the highest validation accuracy (80.04%) is used for the main results, and all ablations are reported against the three-run mean. 4. EXPERIMENTS 4.1 Experimental setup Evaluation slices. Training labels exclude two kinds of data. First, five scenes (carLight, carShadow, manCall, running, tricycle) were held out from training updates and used for validation; we refer to them as unseen images (U-I). Because checkpoint selection used validation accuracy on these scenes, we describe U-I results as train-held-out validation generalization, not as an untouched test set. Second, six fusion methods (U2Fusion, Hybrid_MSD, MGFF, FPDE, GTF, NSCT_SR) were excluded from training labels entirely and form the unseen-method (U-M) 8 set. The remaining 16 scenes contribute 5,776 training records; the five U-I scenes contribute 1,805 validation records. Crossing the two factors yields the four evaluation settings of Table 1. Table 1. The four evaluation settings. Setting Image set Method set Evaluation focus U-I/U-M Unseen images Unseen methods Strongest shift: scenes and algorithms both unseen U-I/All-M Unseen images All 25 methods Preference generalization to new scenes over the full pool All-I/U-M All images Unseen methods Isolates the unseen-algorithm factor All-I/All-M All images All 25 methods Overall fit on the complete corpus Note: U = unseen, All = all; I = images, M = methods. All-I settings include training scenes and therefore measure overall fit rather than scene-level generalization. Baselines and frozen calibration. The conventional controls are 19 objective metrics: the 13 metrics published with VIFB [1] plus six recalculated under this protocol (VIFF [9], Nabf [27â 28], MS-SSIM [29], SCD [7], FMI [8], C [25]). Because objective metrics output continuous scores, their pairwise score differences must be converted to ternary decisions. For each scene we compute the median absolute difference (MAD) of method-pair scores and use =· as the tie band. The parameter â = 0.09 was selected once on the training seen-scene Ă seen- method split by matching the objective tie rate to the ground-truth tie rate (relative error 0.0163) and then frozen across all four settings. Re-selecting per test setting is a test-informed oracle and is prohibited in the main results. LPIFM does not use ; it uses its own fixed threshold = 0.3. The full calibration grid is provided as supplementary Table S5. The baseline panel is confined to conventional objective metrics for the reasons detailed in Section 2.3: existing learned IVIF assessors score single images on absolute scales rather than deciding source-conditioned ternary preference, are redundant with the conventional metric panel they approximate, or lack publicly usable implementations for faithful re-evaluation under this protocol. Outcome measures. At the pair level we report three-class accuracy, per-class F1 (1 , 1 , 1 ), and macro-1. 1 and macro-1 are reported only where ground-truth tie support is adequate (at least 30): the U-M slices contain only 2/75 and 13/315 ties, so these columns are omitted there and should be interpreted on the All-M slices (93/1,500 and 413/6,300 ties). At the method level we report Spearman and Kendall between predicted and human T-BT rankings, along with rank tables, mean absolute rank difference (MARD), and top-k overlap. Robustness of the ranking conclusions to the aggregation rule (top-share Ti, normalized win rate T-NR, normalized win share T-NW) is reported in supplementary Table S4. 4.2 Pairwise decision fidelity Tables 2 and 3 report the two full-pool settings, which are the operating conditions closest to practical use: a researcher comparing many methods, most of them competitive, on scenes the assessor has not been tuned to. Table 2. Pairwise classification and T-BT rank correlation on unseen images Ă all methods (U-I/All-M). Method Accâ â â â macro-1â (T-BT)â (T-BT)â Avg_gradient 0.454 0.460 0.486 0.189 0.378 -0.014 0.010 C 0.568 0.538 0.615 0.394 0.516 0.408 0.311 Cross_entropy 0.478 0.503 0.482 0.221 0.402 0.128 0.100 Edge_intensity 0.458 0.463 0.489 0.203 0.385 0.021 0.040 Entropy 0.544 0.553 0.556 0.385 0.498 0.193 0.120 FMI 0.536 0.582 0.544 0.089 0.405 0.343 0.224 MS-SSIM 0.548 0.526 0.600 0.276 0.467 0.420 0.301 Mutinf 0.521 0.532 0.538 0.305 0.458 0.061 0.074 Nabf 0.489 0.511 0.500 0.233 0.415 0.100 0.080 Psnr 0.608 0.623 0.612 0.481 0.572 0.548 0.333 Qabf 0.587 0.611 0.618 0.099 0.443 0.430 0.277 Qcb 0.570 0.601 0.590 0.133 0.441 0.508 0.357 Qcv 0.604 0.593 0.660 0.245 0.499 0.472 0.341 9 Rmse 0.607 0.621 0.612 0.476 0.570 0.531 0.324 SCD 0.629 0.586 0.685 0.439 0.570 0.554 0.411 Spatial_frequency 0.449 0.469 0.474 0.101 0.348 -0.040 -0.033 Ssim 0.593 0.596 0.639 0.196 0.477 0.446 0.290 Variance 0.511 0.492 0.540 0.431 0.488 0.093 0.070 VIFF 0.577 0.579 0.617 0.257 0.484 0.378 0.250 LPIFM 0.792 0.820 0.824 0.215 0.620 0.941 0.830 Note: frozen â = 0.09, calibrated on the training seenĂseen split ( ( â ) = 0.0624 vs. ground-truth 0.0614, relative error 0.0163). Acc is three-class accuracy including ties; / are Spearman/Kendall against the human T-BT ranking. Ground-truth ties in this setting: 93/1,500. Table 3. Pairwise classification and T-BT rank correlation on all images Ă all methods (All-I/All-M). Method Accâ â â â macro-1â (T-BT)â (T-BT)â Avg_gradient 0.478 0.478 0.517 0.180 0.392 0.072 0.107 C 0.543 0.515 0.598 0.294 0.469 0.385 0.273 Cross_entropy 0.453 0.466 0.473 0.196 0.378 0.244 0.197 Edge_intensity 0.479 0.478 0.519 0.185 0.394 0.081 0.127 Entropy 0.533 0.528 0.566 0.328 0.474 0.237 0.193 FMI 0.540 0.566 0.567 0.118 0.417 0.362 0.230 MS-SSIM 0.570 0.540 0.626 0.306 0.491 0.375 0.207 Mutinf 0.535 0.537 0.573 0.212 0.441 0.168 0.140 Nabf 0.455 0.473 0.468 0.216 0.386 -0.056 -0.060 Psnr 0.566 0.568 0.584 0.432 0.528 0.422 0.260 Qabf 0.576 0.587 0.618 0.153 0.453 0.479 0.313 Qcb 0.552 0.563 0.589 0.175 0.442 0.509 0.360 Qcv 0.605 0.589 0.666 0.253 0.502 0.642 0.460 Rmse 0.566 0.568 0.584 0.435 0.529 0.417 0.250 SCD 0.629 0.591 0.693 0.325 0.536 0.717 0.560 Spatial_frequency 0.473 0.479 0.512 0.142 0.377 0.081 0.124 Ssim 0.570 0.558 0.626 0.220 0.468 0.312 0.187 Variance 0.517 0.492 0.557 0.386 0.478 0.173 0.147 VIFF 0.587 0.575 0.636 0.265 0.492 0.502 0.387 LPIFM 0.840 0.859 0.876 0.268 0.667 0.977 0.900 Note: same protocol as Table 2. Ground-truth ties in this setting: 413/6,300. On unseen scenes with the full 25-method pool (Table 2), LPIFM reached an accuracy of 0.792, exceeding the strongest conventional metric, SCD (0.629), by 0.163. On the complete corpus (Table 3), LPIFM reached 0.840 against SCD's 0.629, a gap of 0.211. The gap is not attributable to a single weak baseline: the 19-metric panel spans the major metric families, and its best member in each setting is the one LPIFM is compared against. LPIFM's macro-1 (0.620 and 0.667) also matched or exceeded the best objective values (0.572 and 0.536), indicating that the advantage is not bought by sacrificing any single class. Tie prediction remains the hardest class for all evaluators (LPIFM 1 0.215 and 0.268), which is expected given tie prevalence below 7%, and we return to tie calibration in Section 5. Table 4 summarizes all four settings, including the two U-M slices, and shows both halves of the evidence. LPIFM is stable everywhere: accuracy stays within 0.792â0.840 and (T-BT) within 0.941â0.977. The conventional side is not stable. SCD is genuinely strong on the six-method held- out slices (0.840 on U-I/U-M, where it exceeds LPIFM's 0.800; 0.803 on All-I/U-M, where LPIFM reaches 0.813), yet the same metric falls to 0.629 on both full-pool settings. Section 4.4 diagnoses this contrast; here we note only that a metric whose reliability depends on which methods happen to be compared cannot serve as a general evaluation instrument, and that full method pools, where SCD falls to 0.629, are the settings practitioners face most often. Table 4. Cross-setting comparison of LPIFM against the best conventional baseline. Setting LPIFM Acc LPIFM mF1 LPIFM LPIFM Best Acc Best mF1 Best U-I/U-M 0.800 â 0.943 0.867 0.840 â 0.943 U-I/All-M 0.792 0.620 0.941 0.830 0.629 0.572 0.554 All-I/U-M 0.813 â 0.943 0.867 0.803 â 0.943 All-I/All-M 0.840 0.667 0.977 0.900 0.629 0.536 0.717 10 Note: Best Acc / Best mF1 / Best Ï are the per-column maxima over the 19 objective metrics (excluding LPIFM); the best metric is SCD in all Acc and Ï cells. mF1 = macro-1. Ground-truth tie support on the U-M slices (2/75 and 13/315) is below the reporting threshold of 30, so 1 and macro-1 are omitted there; per-setting full tables are provided as supplementary Tables S1 and S2. 4.3 Ranking fidelity For an evaluation instrument, pair-level accuracy is a means; the end is whether the instrument reproduces the method ranking that humans would produce. Table 5 expands the T-BT strength vectors into full rank lists for the U-I/All-M setting, placing LPIFM beside representative objective metrics against the human ranking. Table 5. T-BT rank table on unseen images Ă all methods (U-I/All-M): fusion-method rankings induced by each evaluator (rank 1 is best). Rank Humans LPIFM SCD FMI Qcv Psnr Qcb Qabf 1 U2Fusion U2Fusion U2Fusion LP_SR CNN ADF LP_SR LP_SR 2 MGFF Hybrid_MSD HMSD_GF CNN HMSD_GF DLF Hybrid_MSD CNN 3 TIF TIF IFCNN Hybrid_MSD IFCNN FPDE GFF NSCT_SR 4 Hybrid_MSD IFCNN LatLRR ADF SwinFusion MSVD CNN Hybrid_MSD 5 GFF GFF VSMWLS NSCT_SR SeAFusion ResNet RP_SR HMSD_GF 6 IFCNN MGFF CNN GFF Hybrid_MSD IFCNN HMSD_GF GFF 7 CNN CNN SwinFusion HMSD_GF TIF TIF NSCT_SR RP_SR 8 HMSD_GF HMSD_GF Hybrid_MSD FPDE VSMWLS Hybrid_MSD MGFF TIF 9 LP_SR LP_SR MGFF IFCNN IFEVIP MGFF U2Fusion IFCNN 10 VSMWLS VSMWLS YDTR MGFF LP_SR VSMWLS TIF MGFF 11 DLF ResNet SeAFusion RP_SR YDTR GFF GFCE SwinFusion 12 ResNet ADF TIF SwinFusion ResNet HMSD_GF IFCNN VSMWLS 13 FPDE DLF GFCE GFCE U2Fusion CNN ADF ADF 14 ADF FPDE DLF U2Fusion LatLRR GTF CBF SeAFusion 15 MSVD RP_SR LP_SR VSMWLS MGFF LP_SR FPDE CBF 16 SeAFusion MSVD FPDE IFEVIP FPDE RP_SR VSMWLS U2Fusion 17 SwinFusion GTF ADF GTF ADF U2Fusion LatLRR FPDE 18 YDTR SwinFusion ResNet LatLRR GFCE YDTR SwinFusion GFCE 19 RP_SR SeAFusion MSVD SeAFusion RP_SR CBF IFEVIP IFEVIP 20 GFCE YDTR IFEVIP TIF DLF GFCE DLF YDTR 21 LatLRR NSCT_SR RP_SR DLF GFF IFEVIP ResNet DLF 22 IFEVIP GFCE GFF YDTR MSVD LatLRR SeAFusion LatLRR 23 NSCT_SR LatLRR CBF CBF NSCT_SR NSCT_SR MSVD GTF 24 CBF CBF NSCT_SR ResNet CBF SeAFusion YDTR ResNet 25 GTF IFEVIP GTF MSVD GTF SwinFusion GTF MSVD Note: ranks follow descending T-BT strength; ties are broken alphabetically. Humans denotes the ranking derived from ground-truth labels. LPIFMâHumans MARD = 1.68; top-1 agreement: yes; top-3 overlap: 2/3. The best objective MARD in this table is SCD at 5.20; the worst is FMI at 6.88. The corresponding All-I/All-M rank table is provided as supplementary Table S3. The rank lists make the correlations auditable method by method, and they expose a failure mode that scalar summaries hide. On scenes never used for training updates, LPIFM reproduced the human top choice (U2Fusion), matched the human ranking to within a mean absolute rank difference of 1.68 positions, and kept every top-10 human method inside its own top 10. The objective metrics did not merely correlate less; they promoted methods that humans placed far down the ranking to rank 1 (FMI and Qcb chose LP_SR, human rank 9; Qcv chose CNN, human rank 7; Psnr chose ADF, human rank 14). An evaluation instrument that selects the wrong best method misdirects method development regardless of its average correlation. Table 6 summarizes rank proximity across all four settings. LPIFM's MARD is lowest or tied- lowest in every setting where the pool is realistic, its top-1 agrees with humans in three of four settings, and its top-3 overlap never falls below 2/3. In the one setting without top-1 agreement (All-I/All-M), the human top two (Hybrid_MSD, U2Fusion) and LPIFM's top two (U2Fusion, Hybrid_MSD) are the same pair transposed, and the top-3 sets are identical. Replacing T-BT with win-rate aggregations (T-NR, T-NW) leaves these conclusions unchanged (LPIFM (T-NR) = 0.940â0.976 across settings; supplementary Table S4), which indicates the ranking fidelity is a property of the predicted decisions, not of one aggregation rule. Table 6. Rank proximity to the human ranking across the four settings. 11 Setting #Metho ds MARD (LPIFM) Top-1 match Top-3 overlap Best objective MARD Humans top-3 LPIFM top-3 U-I/U-M 6 0.33 Yes 3/3 SCD (0.00) U2Fusion > Hybrid_MSD > MGFF U2Fusion > Hybrid_MSD > MGFF U-I/All-M 25 1.68 Yes 2/3 SCD (5.20) U2Fusion > MGFF > TIF U2Fusion > Hybrid_MSD > TIF All-I/U-M 6 0.33 Yes 3/3 SCD (0.33) U2Fusion > Hybrid_MSD > MGFF U2Fusion > Hybrid_MSD > MGFF All-I/All-M 25 1.20 No 3/3 SCD (3.68) Hybrid_MSD > U2Fusion > IFCNN U2Fusion > Hybrid_MSD > IFCNN Note: MARD = mean absolute rank difference against the human ranking (lower is better); best objective MARD is taken over the objective columns of the corresponding rank tables. 4.4 Why SCD is locally strong on the held-out methods Table 4 contains an apparent anomaly: SCD reaches 0.840 accuracy on U-I/U-M and 0.803 on All-I/U-M, competitive with LPIFM, yet only 0.629 on the full pools. Table 7 locates the cause in the composition of the held-out subset. Table 7. Position of the six held-out methods in the human preference ranking (All-I/All-M). Held-out method Human rank (win rate) Within-U-M win rate Mean-SCD rank Hybrid_MSD 1 / 25 (0.822) 0.819 9 / 25 U2Fusion 2 / 25 (0.815) 0.824 1 / 25 MGFF 4 / 25 (0.740) 0.681 7 / 25 FPDE 16 / 25 (0.492) 0.438 17 / 25 NSCT_SR 23 / 25 (0.145) 0.200 24 / 25 GTF 25 / 25 (0.060) 0.038 25 / 25 Note: human ranks derive from method-level win rates over All-I/All-M pairwise labels (ties counted as 0.5); within-U-M win rates count only pairs among the six held-out methods. Two properties, one of the held-out subset and one of SCD itself, jointly explain this contrast, and both are needed. The subset property is its span. The six held-out methods sit at human ranks 1, 2, 4, 16, 23, and 25, with win rates from 0.822 down to 0.060 (Table 7), so most of their pairwise contests pit a clearly strong method against a clearly weak one. Coarse contests of this kind are easier for many evaluators, not only for SCD: several conventional metrics also score visibly higher on the held-out slices than on the full pools (for example, between U-I/All-M and U-I/U- M, Qcv rises from 0.604 to 0.720 and VIFF from 0.577 to 0.653; supplementary Tables S1 and S2). The metric property is where SCDâs share of the gain is concentrated. SCD sums the correlations between source-to-fused difference images [7], so it responds strongly to large, global differences in how much source content survives fusion, and on the held-out pairs its score gaps widen accordingly (median pairwise score difference roughly 0.36, versus roughly 0.20 over all pairs). A resampling analysis quantifies both factors. Over 3,000 random six-method subsets, SCDâs mean accuracy is approximately 0.629, essentially its full-pool value, while its accuracy on the actual held-out subset (approximately 0.803) lies near the 99th percentile; subsets spanning at least 20 human ranks average approximately 0.68, dense mid-tier subsets only about 0.53, and a subsetâs human win-rate span correlates with SCD accuracy at approximately 0.69. The same pattern appears at the ranking level: the Spearman correlation between human win rates and mean SCD is 0.83 on the held-out six, 0.48 on the nineteen training methods, and 0.67 overall. SCD therefore tracks the human order faithfully at the coarse, top-versus-bottom scale that this subset happens to emphasize, and loses resolution exactly where contests move to near-ranked neighbors. These results support a bounded interpretation that credits both factors. SCD genuinely resolves contests between methods far apart in quality, a real strength that follows from its design as a source-difference correlation measure, and the held-out subset is composed almost entirely of the contests that this strength decides; the same discriminative power dissolves among the many near-ranked, mid-tier methods that populate real leaderboards. LPIFM sustains 0.792â0.840 across every subset structure. Its robustness to pool composition therefore appears to come from learning the human decision function itself, not from any fixed signal proxy. The U-M slices should therefore be read as a stress test that a strong handcrafted baseline passes for identifiable, non-transferable reasons, not as evidence that a conventional metric matches LPIFM. 12 4.5 Ablation studies All ablations share the training protocol of Section 3.6 and report best validation accuracy; the default configuration is the three-run mean 78.67 ± 1.19 (%), and single-run deltas smaller than about 1 point should not be read as stable effects. Table 8. Effect of different backbones on LPIFM validation accuracy (8 architectures). Backbone Params (M) FLOPs (B) Best Acc. (%) (p) ConvNeXt-V2-Baseâ 105.24 137.50 78.67±1.19 â SwinV2-Base 104.44 104.76 79.39 +0.72 CAFormer-B36 110.33 200.04 79.22 +0.55 ResNetV2-101Ă1 62.13 1.99 76.51 -2.16 ViT-B/16-384 103.11 119.53 73.52 -5.15 DeiT3-B/16-384 103.13 119.53 72.63 -6.04 FastViT-MA36 60.79 54.88 71.63 -7.04 EfficientNetV2-M 70.93 47.85 67.65 -11.02 Note: â the default configuration; accuracy is the mean ± SD of three independent runs, other rows are single runs. is computed against the default mean. The backbone ablation (Table 8) shows a spread exceeding 11 percentage points across eight architectures. Hierarchical and hybrid backbones (ConvNeXt-V2, SwinV2, CAFormer) clearly outperform pure global-attention transformers (ViT â5.15, DeiT3 â6.04) and efficiency-oriented designs (FastViT â7.04, EfficientNetV2-M â11.02). This pattern indicates that fusion-preference discrimination depends on hierarchical local structure together with cross-scale context: LPIFM is the combination of a preference objective with an appropriate visual prior. SwinV2 and CAFormer exceed the default mean in single runs, but by margins within the run-to-run standard deviation and at higher compute cost, so ConvNeXt-V2-Base is retained as the default. Table 9. Ablation on the capacity of the triadic interaction module. ID Config Depth Embed Heads Params (M) Best Acc. (%) (p) E0 None 0 1024 â 97.15 77.40 -1.27 E1 Nano 1 256 4 88.85 75.96 -2.71 E2 Tiny 1 512 4 91.22 69.20 -9.47 E3 Small 2 512 4 91.94 78.23 -0.44 E4 Medium 2 768 8 96.45 75.29 -3.38 E5â Base 3 1024 8 105.24 78.67±1.19 â E6 Large 4 1024 8 107.93 78.73 +0.06 Note: â the default configuration (three-run mean ± SD); other rows are single runs. E0 removes the triadic interaction module entirely. is computed against the E5 mean. The triadic interaction ablation (Table 9) supports the module's role. Removing it (E0) costs 1.27 percentage points. The more instructive finding is that misconfigured capacity is far more damaging than absence: the Tiny configuration (E2) collapses to 69.20% (â9.47 p), while Small and Large sit within noise of the default. Cross-source interaction thus needs adequate width and depth to help, but the model's performance is primarily driven by the preference objective and the backbone prior, with the interaction module contributing a stable comparison bias at moderate capacity. Table 10. Ablation on pairwise preference loss hyperparameters. ID Setting Changed value Best Acc. (%) (p) P1â Baseline =.,=., =., =., =. 78.67±1.19 â P2 Small margin =0.5 76.45 -2.22 P3 Large margin =2.0 77.17 -1.50 P4 Narrow tie band =0.1 64.65 -14.02 P5 Wide tie band =0.7 75.07 -3.60 P6 Low temperature =0.5 77.73 -0.94 P7 High temperature =2.0 72.74 -5.93 P8 No margin loss =0 78.45 -0.22 P9 Strong margin loss =1.0 79.17 +0.50 13 P10 Weak tie loss =0.5 80.06 +1.39 Note: â the default configuration (three-run mean ± SD); other rows are single runs changing one hyperparameter each. is computed against the P1 mean. The loss ablation (Table 10) isolates the mechanism that most distinguishes LPIFM from conventional evaluators: explicit tie modeling with a correctly sized band. Narrowing the tie band to =0.1 (P4) collapses accuracy to 64.65% (â14.02 p), the most harmful change among the loss ablations, because the model is forced to pick winners on pairs that humans judged indistinguishable. Widening the band to =0.7 (P5) also hurts (â3.60 p) by diluting the preference signal. All remaining perturbations produce sub-noise or marginal changes; the single best run (P10, 80.06%) essentially matches the best default run (80.04%), so the default loss is retained. Table 11. Ablation on learning-rate schedulers and related hyperparameters. ID Setting Key change Best Acc. (%) (p) S1â Cosine + warmup warmup = 10, min. LR = 1Ă10â»â” 78.67±1.19 â S2 OneCycle scheduler = OneCycle 70.08 -8.59 S3 ReduceLROnPlateau scheduler = plateau 79.00 +0.33 S4 No warmup warmup = 0 78.67 0.00 S5 Short warmup warmup = 5 79.06 +0.39 S6 Long warmup warmup = 20 77.56 -1.11 S7 Low start factor start factor = 0.01 78.28 -0.39 S8 High start factor start factor = 0.3 77.23 -1.44 S9 Low minimum LR min. LR = 1Ă10â»â¶ 63.55 -15.12 S10 High minimum LR min. LR = 3Ă10â»â” 78.78 +0.11 Note: â the default configuration (three-run mean ± SD); other rows are single runs changing the stated setting. LR = learning rate. is computed against the S1 mean. The scheduler ablation (Table 11) characterizes the training dynamics that preference learning requires. The cosine family is robust: removing warmup (S4) reproduces the default mean exactly, and short warmup (S5, +0.39 p) or ReduceLROnPlateau (S3, +0.33 p) yield only sub-noise gains. What damages training is suppressing late-stage updates: OneCycle (S2) costs 8.59 percentage points and lowering the minimum learning rate to 1Ă10â»â¶ (S9) costs 15.12, the largest degradation in the entire ablation study. Preference learning evidently depends on sustained, moderate parameter updates late in training, and the default cosine schedule with warmup 10 and minimum learning rate 1Ă10â»â” is retained. 14 4.6 Qualitative case analyses Fig. 3. Per-scene ranking case study on the walking scene. For each evaluator (rows: Entropy, FMI, Qabf, SCD, VIFF, LPIFM, Humans), the three highest-ranked and three lowest-ranked fused results are shown with the corresponding scores (classical metrics: native scalar values; LPIFM and Humans: per-scene normalized win rates). LPIFM's top and bottom sets each share two of three members with the human sets, while classical metrics diverge far more strongly. 15 Fig. 4. Per-scene ranking case study on the man scene, in the format of Fig. 3. LPIFM and human observers agree on LP_SR and HMSD_GF as the leading group in this night scene, while several classical metrics rank strongly artifacted results highly. 16 Fig. 5. Per-scene ranking case study on the snow scene, in the format of Fig. 3. LPIFM's top three (HMSD_GF, Hybrid_MSD, GFCE) coincide with the human top three. 17 Fig. 6. Per-scene ranking case study on the running scene (a validation scene never used for training updates), in the format of Fig. 3. LPIFM and human observers agree on the leading group (Hybrid_MSD, CNN, LP_SR/HMSD_GF) while classical metrics promote methods that humans place near the bottom. Figs. 3â6 examine four scenes at the level a practitioner experiences: which methods does each evaluator send to the top, and which to the bottom? On the walking scene (Fig. 3), LPIFM's top three (IFCNN 1.0000, SeAFusion 0.9375, U2Fusion 0.9375) share two members with the human top three (U2Fusion 0.9583, IFCNN 0.9375, SwinFusion 0.9167), and the two bottom sets 18 likewise share two of three members (CBF and MSVD). On the man scene (Fig. 4), LPIFM places Hybrid_MSD, LP_SR, and HMSD_GF on top, matching two of the human top three and the human leaders' composition. On the snow scene (Fig. 5), LPIFM's top three (HMSD_GF 1.0000, Hybrid_MSD 0.9583, GFCE 0.9167) reproduce the human top three exactly. On the running scene (Fig. 6), a validation scene excluded from training updates, LPIFM's leading group (CNN 0.9792, Hybrid_MSD 0.9792, LP_SR 0.9167) again overlaps the human group (HMSD_GF 0.9792, Hybrid_MSD 0.9792, CNN 0.9167) in two of three positions. Across all four scenes, the classical metric rows repeatedly promote results that humans rank near the bottom, consistent with the quantitative rank tables of Section 4.3. The image content in Figs. 3â6 shows what these disagreements consist of. The results that classical metrics promote frequently carry visible defects: NSCT_SR, whose outputs exhibit speckle and smoke-like spurious structures, appears in the top three of Entropy, FMI, or Qabf in all four scenes, and on the snow and running scenes Entropyâs top three include NSCT_SR and CBF, methods that human observers place at or near the bottom, because their artifacts inflate exactly the statistics such metrics read as information. The bottom sets are as telling as the top sets: humans concentrate artifact-heavy or degraded results such as NSCT_SR, CBF, GTF, and MSVD in their bottom three, yet the classical metricsâ bottom sets frequently omit these methods and instead demote mid-ranked results such as DLF and ResNet (Figs. 3 and 5). LPIFM reproduces both ends: its top sets share at least two of three members with the human top sets on every scene, and its bottom sets likewise share two of three members with the human bottom sets, isolating the same defect-laden methods that humans reject. Four further case studies covering the kettle, elecbike, labMan, and tricycle scenes are provided as supplementary Figs. S1âS4 and show the same qualitative pattern (for example, on the kettle scene LPIFM and humans agree on SwinFusion as the best method for the scene, Fig. S1; on elecbike the top-three sets are identical, Fig. S2). Fig. S4 also documents an informative failure case: on the tricycle scene, LPIFM places GTF in its leading group (0.9375) while human observers rank GTF near the bottom (0.0417). The tricycle scene is a validation scene and GTF is a held-out method, so this case combines both generalization factors; we discuss its implication for scene-level reliability in Section 5. 5. DISCUSSION The central finding of this work is that the human pairwise comparison protocol for IVIF, long treated as trustworthy but unaffordable, can be operationalized as a learned, repeatable instrument without surrendering its decision semantics. LPIFM answers the same ternary question that the protocol puts to human observers, and its agreement with human decisions (0.792â0.840) and human rankings (Ï = 0.941â0.977) held across scene and method generalization slices on this benchmark. We interpret this stability, together with the large remaining gap between conventional metrics and LPIFM on full method pools, as evidence that the limiting factor in current IVIF evaluation is not the sophistication of handcrafted proxies but their distance from the human decision function. Once that function is learned directly, ranking fidelity follows, and it follows robustly across aggregation rules (T-BT, T-NR, T-NW). The ablations suggest three reasons for this margin. First, the decision space matters: the tie band incurs the largest ablation loss (â14.02 p when narrowed), which indicates that agreement with human judgment depends on modeling when humans decline to judge. Second, the visual prior matters: an 11-point backbone spread shows that preference discrimination requires hierarchical local structure, consistent with the fact that fusion defects (halos, ghosting, amplified noise, lost thermal salience) are local, structured phenomena. Third, joint comparison contributes a stable margin: triadic interaction adds a consistent gain at moderate capacity, complementing the supervision and formulation that drive most of the performance. These observations indicate that the dataset and the task definition matter more than any single architectural choice. For the community, the practical significance is an instrument, not just a benchmark result. A researcher developing a fusion method can, at negligible marginal cost, obtain the decision a dense 19 human panel would most likely have produced: which of two variants is preferable on each scene, whether the difference is perceptible at all, and where a new method sits in a human-aligned ranking of an arbitrary pool. The same instrument can serve benchmark maintainers as a meta- evaluation layer, flagging metrics or leaderboard configurations that diverge from human preference. To support this use, we release the annotated preference dataset (6,300 adjudicated comparisons with full split manifests), together with the LPIFM model weights, source code, and evaluation code (code under AGPL-3.0; model weights and datasets under C BY-NC-SA 4.0), so that others can retrain and audit preference-aligned assessors. The release converts a single trained model into a reproducible research direction. Rival explanations for the headline results deserve explicit treatment. The first is that a strong handcrafted metric might suffice. Section 4.4 traces SCDâs local parity to the pairing of an extreme-span held-out subset with SCDâs own strength on coarse contests: its accuracy there sits near the 99th percentile of random six-method subsets, while its expected accuracy on realistic pools equals its full-pool value of 0.629. The second is overfitting to the benchmark. The U-I slices bound this concern on the scene axis (accuracy 0.792â0.800 on scenes never used for gradient updates), with one caveat: the five validation scenes informed checkpoint selection, so they are train-held-out validation evidence, not an untouched test set, and all results are confined to VIFB's 21 scenes and 25 methods. The third is that thousands of pairs might overstate the evidence. The pairs share only 21 independent scenes, so we treat the scene as the proper sampling unit and regard scene-level confidence intervals as necessary in future reporting. Several limitations bound the claims. First, external validity: no external dataset or third untouched split is evaluated here, and generalization beyond VIFB remains to be demonstrated. Second, measure-like consistency: we have not yet quantified candidate-swap antisymmetry, cycle rates, or transitivity of LPIFM's decisions, and until those diagnostics are reported we describe LPIFM as a preference evaluator rather than a scalar measure. Third, tie calibration: 1 (0.215â0.268 on All-M slices) shows ties remain the hardest class, and tie support on small held-out slices is too sparse to assess at all. Fourth, scene-level reliability: the Fig. S4 failure case (GTF on tricycle) shows that per-scene decisions can misfire when scene and method shift combine, so per-scene outputs should be treated with more caution than pool-level rankings. Fifth, most ablation variants are single runs, and the viewing-condition records of the subjective study were not retained. None of these limitations undermines the central evidence, but each defines work that a claim-tight instrument should complete: external validation, consistency diagnostics, scene-bootstrap intervals, and repeated-seed ablations. LPIFM is a surrogate: it reproduces the recorded consensus of one carefully collected human panel, on one benchmark, for one fusion task. New domains, new artifact regimes, and high-stakes applications will still require human validation, and the surrogate should be re-anchored as preference data accumulate. The surrogate changes the cost structure of that loop: human effort can be reserved for anchoring and auditing, while routine method comparison, ablation triage, and leaderboard construction run at machine cost with human-aligned semantics. 6. CONCLUSION This paper asked whether the trusted but costly human pairwise comparison protocol for infraredâvisible image fusion can be turned into a practical, scalable evaluation instrument. LPIFM answers this question by construction. It formulates IVIF assessment as source- conditioned ternary preference prediction, is trained on a dense corpus of 6,300 exhaustively adjudicated human comparisons that we publicly release, and preserves human indifference through an explicit tie band, the design choice whose ablation costs the most accuracy. Across four scene- and method-generalization settings, LPIFM sustained pairwise accuracy of 0.792â 0.840 and human-ranking correlation of 0.941â0.977, exceeded the strongest conventional metric by up to 0.211 accuracy and 0.387 correlation on full method pools, and reproduced human method rankings to within 1.68 mean rank positions on unseen scenes. The one setting in which a handcrafted metric was locally competitive was traced to the pairing of an extreme-span held-out 20 subset with that metricâs strength on coarse contests. Within the boundaries stated in Section 5, these results support LPIFM's intended use: a repeatable, human-preference-aligned surrogate for comparing and ranking IVIF methods at scale, with the released dataset, model weights, and code as the community's basis for auditing, retraining, and extending preference-aligned fusion assessment. Two research directions follow naturally from this work. First, the formulation transfers: source-conditioned preference prediction applies wherever fused images must be judged against shared sources, and collecting analogous preference corpora in neighboring fusion fields, such as CTâMRI fusion in medical imaging, would extend LPIFM into a family of protocol-aligned surrogates across artifact regimes and application semantics. Second, LPIFM can supervise as well as evaluate: a repeatable, human-preference-aligned decision function is precisely the reward signal that reinforcement learning and other preference-driven optimization schemes require, so future fusion models could be trained directly toward human preference, closing the loop between evaluation and method development. DATA AND CODE AVAILABILITY The annotated pairwise preference dataset (6,300 adjudicated unordered comparisons, 13,125 ordered training records, and scene/method split manifests), together with the LPIFM model weights, source code, and evaluation code, is made publicly available at https://github.com/HaoranLiu507/LPIFM. The source code is released under the GNU Affero General Public License v3.0 (AGPL-3.0), and the model weights and datasets are released under the Creative Commons Attribution Non Commercial Share Alike 4.0 International (C BY-NC- SA 4.0) license. SUPPLEMENTARY MATERIAL The following supplementary figures and tables accompany this article. - Fig. S1: per-scene ranking case study, kettle scene. - Fig. S2: per-scene ranking case study, elecbike scene. - Fig. S3: per-scene ranking case study, labMan scene. - Fig. S4: per-scene ranking case study, tricycle scene (failure case). - Table S1: full pairwise/ranking table, U-I/U-M setting. - Table S2: full pairwise/ranking table, All-I/U-M setting. - Table S3: T-BT rank table, All-I/All-M setting. - Table S4: Ti / T-NR / T-NW aggregation robustness, all settings. - Table S5: calibration grid and per-setting ground-truth tie rates. R EFERENCES [1] X. Zhang, P. Ye, G. Xiao, VIFB: A visible and infrared image fusion benchmark, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, p. 468â478. https://doi.org/10.1109/CVPRW50498.2020.00060 [2] V. PetroviÄ, Subjective tests for image fusion evaluation and objective metric validation, Information Fusion 8 (2) (2007) 208â216. https://doi.org/10.1016/j.inffus.2005.05.001 [3] M.H. Loew, J. Bonick, C. Walters, Image fusion for human observers: How should we choose the method?, in: 13th International Conference on Information Fusion, 2010. https://doi.org/10.1109/ICIF.2010.5711838 [4] R.K. Mantiuk, A. Tomaszewska, R. Mantiuk, Comparison of four subjective methods for image quality assessment, Computer Graphics Forum 31 (8) (2012) 2478â2491. https://doi.org/10.1111/j.1467- 8659.2012.03188.x [5] L. Zhang, L. Miao, Z. Zhou, D. Wu, Y. Qian, Y. Wang, et al., Towards the effectiveness of current objective quality metrics for infraredâvisible image fusion and the construction of a comprehensive quality metric, Infrared Physics & Technology (2026) 106560. https://doi.org/10.1016/j.infrared.2026.106560 [6] D. Guan, Y. Wu, T. Liu, A.C. Kot, Y. Gu, Rethinking the evaluation of visible and infrared image fusion, arXiv (2024). https://doi.org/10.48550/arXiv.2410.06811 21 [7] V. AslantaĆ, E. Bendes, A new image quality metric for image fusion: The sum of the correlations of differences, AEU - International Journal of Electronics and Communications 69 (12) (2015) 1890â1896. https://doi.org/10.1016/j.aeue.2015.09.004 [8] M.B.A. Haghighat, A. Aghagolzadeh, H. Seyedarabi, A non-reference image fusion metric based on mutual information of image features, Computers & Electrical Engineering 37 (5) (2011) 744â756. [9] Y. Han, Y. Cai, Y. Cao, X. Xu, A new image fusion performance metric based on visual information fidelity, Information Fusion 14 (2) (2013) 127â135. [10] E. Prashnani, H. Cai, Y. Mostofi, P. Sen, PieAPP: Perceptual image-error assessment through pairwise preference, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, p. 1808â 1817. https://doi.org/10.48550/arXiv.1806.02067 [11] D.E. Moreno-VillamarĂn, H.D. BenĂtez-Restrepo, A.C. Bovik, Predicting the quality of fused long wave infrared and visible light images, IEEE Transactions on Image Processing 26 (7) (2017) 3479â3491. https://doi.org/10.1109/TIP.2017.2695898 [12] Z. Chang, S. Yang, Z. Feng, Q. Gao, S. Wang, Y. Cui, Semantic-relation transformer for visible and infrared fused image quality assessment, Information Fusion 95 (2023) 454â470. https://doi.org/10.1016/j.inffus.2023.02.021 [13] C. Cheng, T. Xu, X.-J. Wu, T. Zhou, H. Li, Z. Tang, J. Kittler, EvaNet: Toward more efficient and consistent infrared and visible image fusion assessment, IEEE Transactions on Pattern Analysis and Machine Intelligence 48 (8) (2026) 9914â9929. https://doi.org/10.1109/TPAMI.2026.3681958 [14] J. Liu, X. Li, Q. Mei, H. Xu, Z. Jiang, L. Ma, R. Liu, X. Fan, Bridging human evaluation to infrared and visible image fusion, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026. https://doi.org/10.48550/arXiv.2603.03871 [15] Y. Guo, J. Gong, Y. Lu, X. Xu, Y. Cheung, W. Su, Bringing multimodal large language models to infrared- visible image fusion quality assessment, arXiv (2026). https://doi.org/10.48550/arXiv.2605.06969 [16] R.R. Davidson, On extending the Bradley-Terry model to accommodate ties in paired comparison experiments, Journal of the American Statistical Association 65 (329) (1970) 317â328. https://doi.org/10.1080/01621459.1970.10481082 [17] M.G. Kendall, The treatment of ties in ranking problems, Biometrika 33 (3) (1945) 239â251. [18] D. Deutsch, G. Foster, M. Freitag, Ties matter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration, in: Proceedings of EMNLP, 2023. [19] S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I.S. Kweon, S. Xie, ConvNeXt V2: Co-designing and scaling ConvNets with masked autoencoders, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. [20] G. Qu, D. Zhang, P. Yan, Information measure for performance of image fusion, Electronics Letters 38 (7) (2002) 313â315. [21] C.S. Xydeas, V. Petrovic, Objective image fusion performance measure, Electronics Letters 36 (4) (2000) 308â309. [22] Z. Wang, A.C. Bovik, H.R. Sheikh, E.P. Simoncelli, Image quality assessment: From error visibility to structural similarity, IEEE Transactions on Image Processing 13 (4) (2004) 600â612. [23] Y. Chen, R.S. Blum, A new automated quality assessment algorithm for image fusion, Image and Vision Computing 27 (10) (2009) 1421â1432. [24] H. Chen, P.K. Varshney, A human perception inspired quality metric for image fusion based on regional information, Information Fusion 8 (2) (2007) 193â207. [25] A.M. Eskicioglu, P.S. Fisher, Image quality measures and their performance, IEEE Transactions on Communications 43 (12) (1995) 2959â2965. [26] J.W. Roberts, J.A. van Aardt, F.B. Ahmed, Assessment of image fusion procedures using entropy, image quality, and multispectral classification, Journal of Applied Remote Sensing 2 (1) (2008) 023522. [27] V. Petrovic, C.S. Xydeas, Objective image fusion performance characterisation, in: Proceedings of the Tenth IEEE International Conference on Computer Vision (ICCV), 2005. https://doi.org/10.1109/ICCV.2005.175 [28] B.K.S. Kumar, Multifocus and multispectral image fusion based on pixel significance using discrete cosine harmonic wavelet transform, Signal, Image and Video Processing 7 (6) (2013) 1125â1143. https://doi.org/10.1007/s11760-012-0361-x [29] Z. Wang, E.P. Simoncelli, A.C. Bovik, Multiscale structural similarity for image quality assessment, in: Proceedings of the 37th Asilomar Conference on Signals, Systems & Computers, 2003. [30] B.K.S. Kumar, Image fusion based on pixel significance using cross bilateral filter, Signal, Image and Video Processing 9 (5) (2015) 1193â1204. [31] D.P. Bavirisetti, G. Xiao, G. Liu, Multi-sensor image fusion based on fourth order partial differential equations, in: 2017 20th International Conference on Information Fusion (Fusion), 2017, p. 1â9. [32] Z. Zhou, M. Dong, X. Xie, Z. Gao, Fusion of infrared and visible images for night-vision context enhancement, 22 Applied Optics 55 (23) (2016) 6480â6490. [33] S. Li, X. Kang, J. Hu, Image fusion with guided filtering, IEEE Transactions on Image Processing 22 (7) (2013) 2864â2875. [34] J. Ma, C. Chen, C. Li, J. Huang, Infrared and visible image fusion via gradient transfer and total variation minimization, Information Fusion 31 (2016) 100â109. [35] Z. Zhou, B. Wang, S. Li, M. Dong, Perceptual fusion of infrared and visible images through a hybrid multi- scale decomposition with Gaussian and bilateral filters, Information Fusion 30 (2016) 15â26. [36] D.P. Bavirisetti, G. Xiao, J. Zhao, R. Dhuli, G. Liu, Multi-scale guided image and video fusion: A fast and efficient approach, Circuits, Systems, and Signal Processing 38 (12) (2019) 5576â5605. [37] Y. Liu, S. Liu, Z. Wang, A general framework for image fusion based on multi-scale transform and sparse representation, Information Fusion 24 (2015) 147â164. [38] J. Ma, Z. Zhou, B. Wang, H. Zong, Infrared and visible image fusion based on visual saliency map and weighted least square optimization, Infrared Physics & Technology 82 (2017) 8â17. [39] J. Ma, L. Tang, F. Fan, J. Huang, X. Mei, Y. Ma, SwinFusion: Cross-domain long-range learning for general image fusion via Swin Transformer, IEEE/CAA Journal of Automatica Sinica (2022). https://doi.org/10.1109/JAS.2022.105686 [40] Y. Zhang, Y. Liu, P. Sun, H. Yan, X. Zhao, L. Zhang, IFCNN: A general image fusion framework based on convolutional neural network, Information Fusion 54 (2020) 99â118. https://doi.org/10.1016/j.inffus.2019.07.011 [41] L. Tang, J. Yuan, J. Ma, Image fusion in the loop of high-level vision tasks: A semantic-aware real-time infrared and visible image fusion network, Information Fusion 82 (2022) 28â42. https://doi.org/10.1016/j.inffus.2021.12.004 [42] H. Xu, J. Ma, J. Jiang, X. Guo, H. Ling, U2Fusion: A unified unsupervised image fusion network, IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (1) (2022) 502â518. https://doi.org/10.1109/TPAMI.2020.3012548 [43] W. Tang, F. He, Y. Liu, YDTR: Infrared and visible image fusion via Y-shape dynamic Transformer, IEEE Transactions on Multimedia 25 (2023) 5413â5428. https://doi.org/10.1109/TMM.2022.3192661 [44] S. Singh, H. Singh, G. Bueno, O. Deniz, S. Singh, H. Monga, P.N. Hrisheekesha, A. Pedraza, A review of image fusion: Methods, applications and performance metrics, Digital Signal Processing (2023). https://doi.org/10.1016/j.dsp.2023.104020 [45] L. Tang, Q. Yan, X. Xiang, L. Fang, J. Ma, C2RF: Bridging multi-modal image registration and fusion via commonality mining and contrastive learning, International Journal of Computer Vision (2025). https://doi.org/10.1007/s11263-025-02427-1 [46] S. Singh, N. Mittal, H. Singh, Classification of various image fusion algorithms and their performance evaluation metrics, in: Computational Intelligence for Machine Learning and Healthcare Informatics, De Gruyter, 2020, Ch. 9. https://doi.org/10.1515/9783110648195-009 [47] X. Meng, C. Chen, Q. Liu, F. Shao, Multi-domain pseudo-reference quality evaluation for infrared and visible image fusion, IET Image Processing (2024). https://doi.org/10.1049/ipr2.13236 [48] V. Khrulkov, A. Babenko, Neural Side-by-Side: Predicting human preferences for no-reference super- resolution evaluation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. https://doi.org/10.1109/CVPR46437.2021.00495 [49] Y. Liu, Z. Qi, J. Cheng, X. Chen, Rethinking the effectiveness of objective evaluation metrics in multi-focus image fusion: A statistic-based approach, IEEE Transactions on Pattern Analysis and Machine Intelligence (2024). https://doi.org/10.1109/TPAMI.2024.3367905 [50] S. Singh, N. Mittal, H. Singh, Review of various image fusion algorithms and image fusion performance metric, Archives of Computational Methods in Engineering (2021). https://doi.org/10.1007/s11831-020-09518-x