Paper deep dive
Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction
Sukriti Tiwari, BHVSP Subrahmanyam, Nidhi Goyal, Sai Amrit Patnaik
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/18/2026, 11:00:29 AM
Summary
The paper introduces a BCI-aware evaluation framework for EEG-to-image reconstruction that distinguishes visual fidelity from semantic recoverability. It identifies limitations in existing metrics (SSIM, LPIPS, CLIP) which suffer from 'harshness' (penalizing semantically correct but visually degraded outputs) and 'semantic blindness' (rewarding plausible but incorrect outputs). The authors propose a framework using four Vision-Language Models (VLMs) to generate Tolerant Perceptual Alignment Scores (T-PAS) and Tolerant Semantic Alignment Scores (T-SAS), which are distilled into a BCI-Coherence Score (BCS). This approach achieves high correlation with human judgments, outperforming traditional pixel-based and representation-based metrics.
Entities (15)
Relation Signals (13)
BCI-Coherence Score → derivedfrom → Tolerant Semantic Alignment Score
confidence 95% · Their consensus is distilled into the BCI-Coherence Score (BCS)
BCI-Coherence Score → derivedfrom → Tolerant Perceptual Alignment Score
confidence 95% · Their consensus is distilled into the BCI-Coherence Score (BCS)
SSIM → suffersfrom → harshness
confidence 92% · Pixel metrics show near-zero correlation with semantic consistency... causing SSIM... to penalize semantically recoverable outputs
LPIPS → suffersfrom → harshness
confidence 92% · causing SSIM, LPIPS, and CLIP to penalize semantically recoverable outputs
CLIP → suffersfrom → semantic_blindness
confidence 92% · causing SSIM, LPIPS, and CLIP to penalize semantically recoverable outputs or reward plausible but incorrect ones
Ovis2-8B → usedin → BCI-Coherence Score
confidence 90% · We annotate all 6,855 image pairs using four recent VLMs: ... and Ovis2-8B.
InternVL3 → usedin → BCI-Coherence Score
confidence 90% · We annotate all 6,855 image pairs using four recent VLMs: InternVL3... to produce T-PAS and T-SAS
SAIL-VL → usedin → BCI-Coherence Score
confidence 90% · We annotate all 6,855 image pairs using four recent VLMs: ... SAIL-VL ...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning. Yet EEG-derived reconstructions are blurry, distorted, and low-detail, causing SSIM, LPIPS, and CLIP to penalize semantically recoverable outputs or reward plausible but incorrect ones. We analyze 6,855 ground-truth/reconstruction pairs from ATM, ENIGMA, BrainVis, and DreamDiffusion using semantic probes, caption harshness and blind-spot rates, and controlled degradations. Pixel metrics show near-zero correlation with semantic consistency, while representation metrics conflate perceptual and semantic errors. We therefore introduce a BCI-aware framework in which four VLMs assess image pairs through structured questions, producing Tolerant Perceptual Alignment Scores (T-PAS) and Tolerant Semantic Alignment Scores (T-SAS). Their consensus is distilled into the BCI-Coherence Score (BCS), a compact evaluator achieving a T-PAS MAE of 0.079 (r = 0.700) and a T-SAS MAE of 0.082 (r = 0.850) on our data. Human validation shows highly reliable joint coherence judgments, with Cohen's kappa = 0.882 +/- 0.174 and Krippendorff's alpha = 0.882, supporting perceptual-semantic recoverability over generic visual similarity. Code and resources are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.12364v1
- Canonical: https://arxiv.org/abs/2607.12364v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
67,899 characters extracted from source content.
Expand or collapse full text
11institutetext: Department of Computer Science and Engineering, Mahindra University, Hyderabad, India 11email: sukriti.tiwari,nidhi.goyal@mahindrauniversity.edu.in 22institutetext: MU-VT Interdisciplinary Advanced Research Centre for Transformative Technologies, Mahindra University, Hyderabad, India 22email: subrahmanyam.bh@mahindrauniversity.edu.in 33institutetext: Avyakt Ehsaas 33email: saiamritp@gmail.com Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction Sukriti Tiwari BHVSP Subrahmanyam Nidhi Goyal and Sai Amrit Patnaik Abstract EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning. Yet EEG-derived reconstructions are blurry, distorted, and low-detail, causing SSIM, LPIPS, and CLIP to penalize semantically recoverable outputs or reward plausible but incorrect ones. We analyze 6,855 ground-truth/reconstruction pairs from ATM, ENIGMA, BrainVis, and DreamDiffusion using semantic probes, caption harshness and blind-spot rates, and controlled degradations. Pixel metrics show near-zero correlation with semantic consistency, while representation metrics conflate perceptual and semantic errors. We therefore introduce a BCI-aware framework in which four VLMs assess image pairs through structured questions, producing Tolerant Perceptual Alignment Scores (T-PAS) and Tolerant Semantic Alignment Scores (T-SAS). Their consensus is distilled into the BCI-Coherence Score (BCS), a compact evaluator achieving T-PAS MAE 0.079 (r=0.700r=0.700) and T-SAS MAE 0.082 (r=0.850r=0.850) on our data. Human validation shows highly reliable joint coherence judgments (κ=0.882±0.174κ=0.882 0.174, α=0.882α=0.882), supporting perceptual–semantic recoverability over generic visual similarity. Code and resources are available at https://sukt03.github.io/BCS/. 1 Introduction EEG-to-image reconstruction is an emerging neural decoding problem that aims to recover visual stimuli from non-invasive brain signals. Unlike conventional image restoration or image generation, the input in this setting is indirect, noisy, and spatially coarse. Consequently, even successful EEG-based reconstructions are often blurry, distorted or only partially faithful to the original stimulus. A reconstruction may fail to match the reference image at the level of pixels, texture, or fine visual detail, while still preserving enough structure or meaning for the original stimulus to remain recognizable. Figure 1: Representative failure patterns in EEG-to-image evaluation: harshness (top) and semantic blindness (bottom). Current evaluation practice in EEG-to-image reconstruction, however, largely inherits metrics from generic image reconstruction, perceptual similarity, and image generation. Commonly used measures include SSIM, PSNR, LPIPS, DISTS, DreamSim, CLIP/DINO-based similarity, and distributional or retrieval-style scores [31, 5, 8, 22, 19]. These metrics are useful in their intended settings, but they were not designed for the weak-signal, degradation-heavy regime of EEG decoding. Pixel and feature-level scores can penalize reconstructions for low visual fidelity even when the intended object or scene remains inferable, while semantic or embedding-based scores can reward broad similarity without verifying whether the reconstruction preserves the specific content of the reference stimulus. Prior neural decoding work has also shown that automatic reconstruction scores can diverge from human judgments [26, 28, 29], yet recent EEG-to-image studies continue to rely on similar evaluation protocols despite the lower spatial fidelity and greater ambiguity of EEG compared with modalities such as fMRI [6, 3]. This mismatch produces two complementary failure patterns. The first is harshness: a metric may assign a low score to a reconstruction that is visually degraded but still semantically recoverable. In neural decoding, such an output may still carry meaningful information if the intended object, scene, or visual concept can be inferred despite imperfect rendering. The second is semantic blindness: a reconstruction may appear perceptually plausible while depicting the wrong object category, scene, or visual content. Similar issues have been observed in image-generation evaluation, where semantic faithfulness often requires explicit object-, relation-, or question-based assessment rather than a single holistic similarity score [13, 9, 17]. Figure 1 illustrates these two cases: harshness, where visual metrics penalize a semantically recoverable reconstruction, and semantic blindness, where a visually plausible reconstruction depicts incorrect semantic content. These observations suggest that EEG-to-image evaluation should not ask only whether a reconstruction is visually similar to the reference, but whether the reference stimulus remains perceptually and semantically recoverable under neural decoding noise. Motivated by this gap, we introduce a BCI-aware evaluation framework that separates visual degradation from semantic loss and measures how well a reconstruction preserves recoverable form, meaning, and their agreement. We use multiple VLMs as structured annotators, analyze their agreement, distill their consensus into a lightweight image-pair evaluator, and validate the resulting score with targeted human judgments on cases where perceptual and semantic assessments disagree. Figure 2 summarizes the proposed evaluation pipeline, including reconstruction outputs, conventional metric analysis, multi-VLM annotation, consensus aggregation, and BCS distillation. Our contributions are as follows: • A perceptual–semantic analysis of EEG reconstruction metrics. We analyze 6,855 ground-truth/reconstruction pairs from EEG-to-image reconstruction outputs spanning ATM, ENIGMA, BrainVis, and DreamDiffusion models. We identify two failure regimes for conventional metrics: harshness, where metrics penalize reconstructions with recoverable semantic content, and blind spots, where metrics assign favorable scores despite semantic mismatch. • A BCI-aware VLM annotation protocol. We introduce a structured annotation framework tailored to the artifacts and ambiguity of EEG-to-image reconstructions. Each image pair is evaluated through paired perceptual and semantic questions, yielding T-PAS (Tolerant Perceptual Alignment Score) and T-SAS (Tolerant Semantic Alignment Score). The protocol separates low-level visual recoverability from semantic recoverability under neural decoding noise and uses multi-VLM agreement and rationale similarity to assess annotation reliability. • BCI-Coherence Score. We introduce BCI-Coherence Score (BCS), a compact evaluator that distills multi-VLM perceptual–semantic annotations into an efficient image-pair scoring model. Given a reference image and an EEG-based reconstruction, BCS predicts BCI-aware consensus scores without repeated access to large VLMs, making the evaluation signal easier to reuse and scale. We evaluate BCS on held-out reconstruction pairs and targeted human validation cases, showing that it preserves perceptual–semantic disagreement patterns that are recognizable to human evaluators. Figure 2: Overview of the proposed BCI-aware coherence evaluation framework. 2 Related Work Neural image reconstruction from brain signals. Visual reconstruction has progressed from decoding hierarchical visual features and optimizing images from fMRI to generative approaches based on adversarial networks, diffusion priors, and multimodal representations [12, 26, 23, 4, 28, 20, 24, 25]. EEG offers greater accessibility and temporal resolution but substantially lower spatial resolution, signal-to-noise ratio, and cross-subject consistency. Early EEG reconstruction systems combined learned neural representations with conditional generation [27, 10], while recent methods align EEG with CLIP or multimodal latent spaces and condition diffusion models on the resulting representations [1, 6, 16, 7, 3, 32]. Their evaluation nevertheless mixes classification, retrieval, distributional quality, pixel fidelity, and embedding similarity. These measures answer different questions and may be influenced by powerful generative priors rather than information reliably decoded from neural activity [18]. Metrics for reconstruction and perceptual similarity. Distributional measures such as FID and KID evaluate generated collections but do not determine whether an individual reconstruction preserves its corresponding stimulus [11, 2]. Full-reference measures range from pixel and structural comparisons such as PSNR and SSIM to learned perceptual metrics including LPIPS, DISTS, PieAPP, and DreamSim [31, 33, 5, 21, 8]. Representation similarities based on CLIP and DINO provide greater semantic sensitivity [22, 19], but may overlook reference-specific errors or conflate visual resemblance with conceptual agreement. CLIP-IQA and Q-Align further demonstrate that vision-language representations can support human-aligned visual-quality assessment [30]. EEG reconstruction, however, requires separating tolerable signal-induced degradation from changes that alter the recoverable stimulus identity or meaning. Aspect-based and model-assisted evaluation. Text-to-image evaluation has increasingly replaced single holistic scores with decomposed tests of objects, attributes, relations, and compositional faithfulness. TIFA, GenEval, T2I-CompBench, Davidsonian Scene Graph, and VQAScore formulate evaluation through structured questions or verifiable semantic propositions [17]. VIEScore and VLM-based quality assessors provide explainable multidimensional judgments, while ImageReward, PickScore, and VisionPrefer distill human or model preferences into reusable evaluators [15, 14]. Recent analyses also show that metric rankings depend strongly on prompts, question construction, and human-rating protocols. These approaches primarily assume prompt-conditioned, visually polished outputs. In contrast, EEG reconstruction is reference-grounded and systematically degraded by neural uncertainty. We therefore introduce a degradation-tolerant protocol that separately measures perceptual recoverability, semantic recoverability, and their joint coherence through the BCI-Coherence Score (BCS). 3 Limitations of Current EEG Reconstruction Metrics We examine whether commonly used EEG-to-image reconstruction metrics capture two properties needed for reliable evaluation: semantic consistency with the reference image and sensitivity to controlled perceptual degradation. 3.1 Caption-Level Semantic Probe To obtain an external semantic reference signal, we caption each ground-truth image and reconstruction using BLIP and compute the SBERT cosine similarity between the two captions: ci=SBERT(BLIP(yi),BLIP(y^i)).c_i=SBERT (BLIP(y_i),BLIP( y_i) ). (1) Here, yiy_i denotes the ground-truth image and y^i y_i denotes the corresponding EEG-based reconstruction. The caption similarity score cic_i is used as a semantic probe rather than a complete reconstruction metric, since captions can miss fine-grained visual details. It nevertheless provides a useful test of whether conventional metrics preserve a basic semantic ordering. For each conventional metric m, we orient its score so that larger values consistently indicate better reconstruction quality: qim=m(yi,y^i),if higher values are better,−m(yi,y^i),if lower values are better.q_i^m= casesm(y_i, y_i),&if higher values are better,\\ -m(y_i, y_i),&if lower values are better. cases (2) This produces an oriented metric score qimq_i^m for each image pair. We then compare the ranking induced by qimq_i^m with the caption-level semantic consistency score cic_i. 3.2 Caption-Based Failure Rates We define two quartile-based failure rates that capture complementary forms of metric misalignment. The first is the caption harshness rate, which measures how often a metric assigns a bottom-quartile score to reconstructions whose captions remain semantically close to the reference: CaptionHR(m)=|i:ci≥Q75(c),qim<Q25(qm)||i:ci≥Q75(c)|.CaptionHR(m)= |\i:c_i≥ Q_75(c),\;q_i^m<Q_25(q^m)\ | |\i:c_i≥ Q_75(c)\ |. (3) A high value indicates that the metric penalizes reconstructions that preserve caption-level semantics, typically because it is sensitive to visual degradation, texture mismatch, or other low-level distortions. The second is the caption blind-spot rate, which measures how often a metric assigns a top-quartile score to reconstructions whose captions are semantically dissimilar to the reference: CaptionBSR(m)=|i:qim≥Q75(qm),ci<Q25(c)||i:qim≥Q75(qm)|.CaptionBSR(m)= |\i:q_i^m≥ Q_75(q^m),\;c_i<Q_25(c)\ | |\i:q_i^m≥ Q_75(q^m)\ |. (4) A high value indicates that the metric can reward reconstructions even when the caption-level semantic content differs from the ground truth. For these quartile-based rates, values near 0.250.25 correspond to behavior close to chance with respect to caption-level semantic consistency. Table 1: Caption-level semantic consistency of conventional reconstruction metrics. Caption-HR measures harshness; Caption-BSR measures blind spots. Lower is better for both. Metric ρcap _cap rcapr_cap Cap-HR ↓ Cap-BSR ↓ MSE -0.012 -0.003 0.257 0.257 SSIM -0.044 -0.059 0.277 0.247 LPIPS 0.109 0.097 0.211 0.188 DISTS 0.229 0.238 0.167 0.159 DreamSim 0.483 0.526 0.071 0.092 OpenCLIP 0.501 0.553 0.060 0.082 DINOv2 0.389 0.503 0.103 0.116 ImageReward 0.309 0.461 0.113 0.152 3.3 Controlled Perceptual Degradation Probe We next test whether the same metrics track perceptual degradation when the corruption process is known. We construct variants of the ground-truth images using six degradations: Gaussian blur, additive noise, downsampling, JPEG compression, color attenuation, and spatial shift. Each degradation is applied at four increasing severity levels, allowing metric behavior to be evaluated without VLM scores or human annotations. For each ground-truth image yiy_i, degradation type d, and severity level k, we generate a corrupted image y~i,k(d) y_i,k^(d). For a metric m, we orient its score so that larger values indicate better quality: qi,km,d=m(yi,y~i,k(d)),if higher values are better,−m(yi,y~i,k(d)),if lower values are better.q_i,k^m,d= casesm\! (y_i, y_i,k^(d) ),&if higher values are better,\\ -m\! (y_i, y_i,k^(d) ),&if lower values are better. cases (5) We then measure whether the oriented score decreases monotonically as degradation severity increases: ρdeg(m,d)=ρSpearman(qi,km,d,−k). _deg(m,d)= _Spearman\! (q_i,k^m,d,-k ). (6) A high value indicates that the metric consistently tracks increasing perceptual corruption, while a low value indicates weak sensitivity or instability for that degradation type. Table 2: Controlled perceptual degradation probe. Entries are Spearman correlations between oriented metric quality and negative degradation severity; higher is better. Metric n Blur Noise Down. JPEG Color Shift MSE 1085 0.738 0.968 0.779 0.612 0.625 0.605 PSNR 1085 0.738 0.968 0.779 0.612 0.625 0.605 SSIM 1085 0.720 0.914 0.844 0.717 0.579 0.233 Edge cosine 1085 0.947 0.905 0.960 0.836 -0.607 0.062 Color hist. 1085 0.764 0.943 0.773 0.727 0.631 0.898 CLIP-L/14 64 0.934 0.823 0.948 0.734 0.890 0.524 SigLIP 64 0.915 0.890 0.927 0.701 0.877 0.489 3.4 Results and Implications Table 1 shows that pixel-level metrics are weak semantic indicators in EEG-to-image reconstruction. MSE and SSIM are nearly uncorrelated with caption-level semantic consistency, with ρcap=−0.012 _cap=-0.012 and ρcap=−0.044 _cap=-0.044, respectively, and their Caption-HR and Caption-BSR values remain close to the chance quartile rate. Representation-based metrics such as OpenCLIP and DreamSim align better with caption semantics, but still compress object identity, scene context, shape, color, texture, and artifacts into a single scalar. The controlled degradation probe in Table 2 shows a complementary limitation: metric behavior depends strongly on the type of visual corruption. Pixel-level metrics track additive noise well but are less stable under JPEG compression, color attenuation, and spatial shift. Edge and color probes are useful for specific degradations but fail outside their intended attribute, while CLIP-L/14 and SigLIP are more robust globally but do not reveal which perceptual property has degraded. Together, the semantic and perceptual probes show that no single metric tracks both semantic consistency and perceptual degradation reliably: metrics sensitive to low-level corruption are weak semantic indicators, while representation metrics that better capture global similarity still obscure specific perceptual and semantic failure modes. 4 BCI-Aware Scoring and Evaluation The analysis in Sec. 3 motivates a protocol that separately measures perceptual and semantic recoverability. The goal is not to judge photorealism, but to determine whether the reference stimulus remains recoverable under the distortions typical of neural decoding. 4.1 Annotation Protocol For each reconstruction pair, the VLM is shown the ground-truth image yiy_i and the generated image y^i y_i. The model is instructed to compare the two images under the assumption that y^i y_i is a noisy EEG-based reconstruction. Each question is answered using a three-level scale: ϕ(no)=0,ϕ(somewhat)=0.5,ϕ(yes)=1.φ( no)=0, φ( somewhat)=0.5, φ( yes)=1. (7) In addition to the discrete answer, the VLM provides a short one to two-line rationale. This rationale is not used directly in score computation, but is used to assess whether different VLMs justify their judgments using similar visual or semantic evidence. The annotation questions are designed to separate visual recoverability from semantic recoverability. Perceptual questions assess whether the reconstruction preserves layout, shape, texture, color, artifact severity, and overall visual interpretability. Semantic questions assess whether the reconstruction preserves category identity, fine-grained identity, functional role, quantity, scene context, and overall meaning. Table 3 lists the exact questions used in the protocol; the full prompt specification is provided in Appendix Sect. 0.B. Table 3: BCI-aware VLM annotation questions. Each question is answered using no, somewhat, or yes. Group ID Question Perceptual P1 Spatial layout: Does the generated image preserve the overall layout and position of the main regions? Perceptual P2 Shape: Is the main object’s shape or silhouette similar? Perceptual P3 Texture/material: Are surface texture or material patterns similar? Perceptual P4 Color: Are the main colors and chromatic appearance similar? Perceptual P5 Artifacts: Is the image free from severe artifacts or noise that obscure the content? Perceptual P6 Holistic visual recoverability: Can the original visual content be broadly recovered from the generated image? Semantic S1 Basic category: Is the main object or scene category correct? Semantic S2 Specific identity: Is the specific subtype or identity correct, rather than only the broad category? Semantic S3 Function/purpose: Does the generated content preserve the object’s role or purpose? Semantic S4 Quantity: Is the number of main objects or entities preserved? Semantic S5 Scene/context: Is the surrounding scene or environment similar? Semantic S6 Holistic semantic recoverability: Can the original meaning or content be understood from the generated image? Each analyzed pair is scored with the same six perceptual and six semantic questions. For each image pair, the perceptual aggregate score T-PAS and semantic aggregate score T-SAS are computed as T-PASi(v)=16∑p=16ϕ(ai,p(v)),T -PAS_i^(v)= 16 _p=1^6φ(a_i,p^(v)), (8) T-SASi(v)=16∑s=16ϕ(ai,s(v)),T -SAS_i^(v)= 16 _s=1^6φ(a_i,s^(v)), (9) where ai,p(v)a_i,p^(v) and ai,s(v)a_i,s^(v) are the answers from VLM v for perceptual and semantic questions, respectively. For downstream consensus supervision, we aggregate across VLMs at the question level: a~i,q=medianv∈ϕ(ai,q(v)), a_i,q=median_v φ(a_i,q^(v)), (10) where V is the set of VLM annotators. With four annotators, this median can take intermediate values such as 0.250.25 or 0.750.75, preserving partial disagreement between models rather than forcing a hard majority label. 4.2 VLM Evaluation Setup We annotate all 6,855 image pairs using four recent VLMs: InternVL3, SAIL-VL, OLA-7B, and Ovis2-8B. We select these VLMs because they are open-weight, reproducible, and representative of strong contemporary VLM performance as demonstrated in Open VLM Leaderboard on Hugging Face. Their open availability allows the annotation protocol to be reproduced without relying on paid proprietary APIs, while their competitive performance on public VLM benchmarks makes them suitable choices for large-scale annotation. The models also differ in calibration and architecture, providing a useful basis for consensus scoring rather than treating any single VLM as an oracle. Each model is evaluated with the same question set and output schema. All final runs produced valid annotations for all image pairs; checkpoint and inference settings are listed in Appendix Sect. 0.C.1. Table 4: Aggregate scores from the four VLM annotators. Scores are averaged over the analyzed image pairs. VLM Mean T-PAS Mean T-SAS InternVL3 0.341 0.252 SAIL-VL 0.213 0.157 OLA-7B 0.433 0.335 Ovis2-8B 0.337 0.345 Table 4 shows that the annotators have different operating points. SAIL-VL is the most conservative, assigning lower perceptual and semantic scores. OLA-7B is more permissive, especially perceptually. InternVL3 and Ovis2-8B are closer in perceptual score, while Ovis2-8B assigns the highest semantic aggregate. This variation motivates using a consensus signal rather than treating any single VLM as the sole authority. Figure 3: Pairwise VLM agreement and reasoning similarity. Left: answer agreement measured by weighted κ. Right: rationale similarity measured by SBERT cosine similarity. The matrices show that model pairs differ in answer calibration, while rationale similarity remains moderate to high across most pairs. 4.3 Agreement and Reliability Analysis We evaluate annotation reliability at two levels: answer agreement and reasoning similarity. Answer agreement is computed over all image-pair/question instances. Each image pair is scored with 12 questions: six perceptual questions and six semantic questions. This gives 82,260 scored question instances in total. For each scored instance, we compare the four VLM answers. The unanimous agreement rate is 44.7%, and Fleiss’ κ is 0.363. This indicates moderate agreement, which is expected because the task requires judging heavily degraded reconstructions rather than clean natural images. Figure 3 shows the pairwise agreement structure across the four VLM annotators. InternVL3 and Ovis2-8B show the strongest agreement, with weighted κ=0.594κ=0.594, and also the highest rationale similarity of 0.700. SAIL-VL and OLA-7B differ more strongly in answer calibration, but the rationale-similarity matrix shows that model explanations remain moderately aligned across most pairs. This suggests that even when discrete answers differ, the VLMs often rely on related visual or semantic evidence. We further inspect which dimensions are stable across models. Table 5 reports representative low and high-agreement dimensions. Table 5: Question-level VLM agreement by annotation dimension. Unan. denotes unanimous answer agreement across the four VLMs. Dimension Type Unan. Fleiss’ κ P6 Holistic visual recoverability Low 0.040 -0.133 P1 Global spatial structure Low 0.058 -0.060 S6 Semantic recoverability Low 0.071 0.012 P5 Artifact absence Low 0.098 -0.035 S2 Subordinate identity High 0.769 0.289 S5 Scene context High 0.784 0.383 S3 Functional role/purpose High 0.819 0.499 P2 Object shape/silhouette High 0.826 0.418 This analysis reveals an important property of the protocol. Concrete attributes such as object shape, functional role, and scene context are relatively stable across VLMs. In contrast, holistic questions such as visual or semantic recoverability are more subjective because they require judging whether a degraded reconstruction is still interpretable. Rather than hiding this uncertainty, the multi-VLM setup exposes it and allows downstream consensus scores to reflect partial agreement. 4.4 Discussion Overall, the four-VLM analysis supports the use of consensus supervision. The models produce valid structured annotations at scale, differ meaningfully in calibration, and show higher agreement on concrete attributes than on holistic recoverability judgments. We therefore use the multi-VLM consensus signal for subsequent analysis and for training the BCI-Coherence Score. Additional metric-to-consensus correlations are reported in Appendix Sect. 0.C.2. 5 BCI-Coherence Score The VLM annotation protocol in Sec. 4 provides detailed perceptual and semantic judgments, but running multiple VLMs for every new reconstruction set is computationally expensive. We therefore introduce the BCI-Coherence Score (BCS), a compact VLM-distilled evaluator that predicts BCI-aware perceptual–semantic scores directly from a ground-truth image and its reconstruction. 5.1 Distillation Target For each image pair (yi,y^i)(y_i, y_i) and annotation question q, we aggregate the four VLM judgments into a consensus target. Let ai,q(v)∈no,somewhat,yesa_i,q^(v)∈\ no, somewhat, yes\ denote the answer from VLM v, mapped to numerical values by ϕ(no)=0,ϕ(somewhat)=0.5,ϕ(yes)=1.φ( no)=0, φ( somewhat)=0.5, φ( yes)=1. (11) The consensus target is the median score across VLMs: zi,q=medianv∈ϕ(ai,q(v)).z_i,q=median_v φ(a_i,q^(v)). (12) With four VLMs, this target can take intermediate values such as 0.250.25 or 0.750.75, preserving partial disagreement between annotators rather than forcing a hard majority label. We train BCS to predict the perceptual and semantic question scores: ^i=fψ(yi,y^i), z_i=f_ψ(y_i, y_i), (13) where fψf_ψ is a lightweight image-pair regression model. 5.2 Model Architecture BCS uses frozen visual encoders to extract image features from the ground-truth and reconstructed images. In our strongest variant, we use a fusion of SigLIP, CLIP, and DINOv3 image encoders. For each encoder EkE_k, we compute normalized embeddings i,kgt=Ek(yi),i,krec=Ek(y^i).e_i,k^gt=E_k(y_i), _i,k^rec=E_k( y_i). (14) We then construct pairwise comparison features: i,k=[i,kgt,i,krec,|i,kgt−i,krec|,i,kgt⊙i,krec,⟨i,kgt,i,krec⟩].h_i,k= [e_i,k^gt,e_i,k^rec,|e_i,k^gt-e_i,k^rec|,e_i,k^gt _i,k^rec, _i,k^gt,e_i,k^rec ]. (15) The final feature vector concatenates these features across encoders: i=[i,1;i,2;⋯;i,K].h_i=[h_i,1;h_i,2;·s;h_i,K]. (16) All visual encoders are kept frozen; only the residual MLP prediction head is trained. The MLP maps ih_i to question-level predictions in [0,1][0,1]. Further architecture and training details are given in Appendix Sect. 0.C.3. 5.3 Training Objective We train BCS on the 6,855 annotated image pairs using a concept-level split to reduce leakage across semantically similar samples. The train, validation, and test proportions are 0.70, 0.15, and 0.15. The model is trained with a weighted regression loss over question targets: ℒ=∑i,qwi,q(z^i,q−zi,q)2∑i,qwi,q.L= _i,qw_i,q ( z_i,q-z_i,q )^2 _i,qw_i,q. (17) The agreement weight wi,qw_i,q gives larger weight to high-consensus VLM labels and lower weight to ambiguous labels. We compute it from the dispersion of the four VLM scores: wi,q=1−Varv∈(ϕ(ai,q(v))).w_i,q=1-Var_v (φ(a_i,q^(v)) ). (18) This reduces the influence of highly disputed examples while still allowing the model to learn from partially agreed annotations. At evaluation time, we report question-level MAE, Pearson correlation, and Spearman correlation. We also compute aggregate perceptual and semantic predictions: T-PAS^i=1||∑q∈z^i,q, T -PAS_i= 1|P| _q z_i,q, (19) T-SAS^i=1||∑q∈z^i,q, T -SAS_i= 1|S| _q z_i,q, (20) where P and S are the six perceptual and six semantic question sets. Because BCS is intended to support evaluation at the perceptual and semantic axis level, aggregate T-PAS and T-SAS prediction is the primary use case; question-level metrics are reported to assess whether the model preserves the underlying annotation structure. Table 6: Held-out performance of BCI-Coherence Score variants. Q-MAE, Q-r, and Q-ρ report question-level prediction error, Pearson correlation, and Spearman correlation over individual VLM annotation questions. PAS and SAS report aggregate perceptual and semantic score prediction, respectively. Lower MAE is better; higher correlation is better. Model Q-MAE Q-r Q-ρ PAS MAE PAS-r SAS MAE SAS-r SigLIP baseline 0.160 0.760 0.724 0.098 0.639 0.103 0.782 3-VLM ordinal ensemble 0.118 0.742 0.739 0.092 0.632 0.098 0.786 4-VLM ordinal ensemble 0.122 0.783 0.788 0.086 0.645 0.091 0.828 4-VLM continuous ensemble 0.137 0.815 0.798 0.079 0.700 0.082 0.850 5.4 Results Table 6 compares several BCS variants on the held-out test split. Q-MAE, Q-r, and Q-ρ measure question-level prediction error, Pearson correlation, and Spearman correlation over individual annotation questions, while PAS and SAS measure aggregate perceptual and semantic score prediction. The SigLIP baseline uses a single frozen visual encoder, while the ensemble variants combine multiple frozen encoders. Ordinal variants predict discretized annotation levels corresponding to the original no, somewhat, and yes scale, whereas the continuous variant predicts the fractional median consensus targets directly, including intermediate values such as 0.250.25 and 0.750.75. The 3-VLM and 4-VLM variants differ in the number of VLM teachers used to form the consensus supervision. The four-teacher continuous ensemble gives the strongest aggregate performance, with PAS MAE of 0.079, PAS-r of 0.700, SAS MAE of 0.082, and SAS-r of 0.850. Although the 3-VLM ordinal ensemble obtains the lowest Q-MAE, its aggregate PAS and SAS performance is weaker than the four-teacher continuous ensemble. This shows that lower question-level error does not necessarily imply better axis-level perceptual or semantic prediction. Since BCS is intended as a reusable continuous evaluator at the perceptual–semantic axis level, we use the four-teacher continuous ensemble as the main BCI-Coherence Score. 5.5 Analysis The results in Table 6 show why aggregate evaluation is more informative than question-level error alone. The ordinal model benefits from snapping predictions to discrete annotation levels, which improves Q-MAE, but the continuous four-teacher ensemble better preserves the fractional consensus signal produced by multiple VLMs. This leads to stronger aggregate T-PAS and T-SAS prediction, which is the intended operating point of BCS. BCS provides three practical benefits. First, it compresses expensive multi-VLM annotation into a compact image-pair model, making BCI-aware evaluation easier to reuse across reconstruction methods. Second, it predicts both perceptual and semantic dimensions rather than collapsing reconstruction quality into a single generic similarity score. Third, its strongest performance is on aggregate semantic scoring: the four-teacher continuous ensemble reaches SAS-r of 0.850, indicating that the learned evaluator captures much of the multi-VLM semantic consensus. The difference between question-level and aggregate performance is also informative. Some individual questions, such as fine-grained color or global structure, remain difficult because they are subjective under reconstruction noise. Aggregating across perceptual or semantic dimensions produces more stable targets, which is the intended use case for BCS. 5.6 Human Agreement We use the human-rated subset from the perceptual-semantic coherence study as a targeted validation set for the perceptual-semantic regimes analyzed in this work. The subset contains 18 reconstructed images derived from 9 unique stimulus instances, with two reconstructions per stimulus. It was stratified using automated perceptual, semantic, and coherence scores: six images were selected from high-perceptual-alignment cases, six from high-semantic-alignment cases, three from the highest-coherence cases, and three from the lowest-coherence cases. This design covers perception-dominant, semantic-dominant, and coherence-extreme regimes, matching the types of disagreement that T-PAS, T-SAS, and BCS are intended to characterize. Human evaluation was conducted with 18 annotators. Each annotator judged paired ground-truth and reconstructed images using structured perceptual, semantic, and holistic coherence questions. Perceptual questions focused on visual realism, shape/layout similarity, and texture/color consistency. Semantic questions focused on category match, key attributes such as object type or count, and whether the reconstruction depicted the same concept or scene. Coherence questions asked whether the reconstruction felt like a faithful reconstruction, whether visual appearance and semantic meaning agreed, and whether a human would accept it as representing the same object or scene. Table 7: Human agreement on perceptual, semantic, and coherence judgments. Cohen’s κ is reported as mean ± standard deviation across annotator comparisons; Krippendorff’s α measures overall reliability. Coherence judgments show substantially higher reliability than isolated perceptual or semantic judgments. Human judgment Cohen’s κ Krippendorff’s α Semantic agreement 0.483±0.3230.483± 0.323 0.574 Perceptual agreement 0.110±0.2410.110± 0.241 0.094 Coherence judgment 0.882±0.1740.882± 0.174 0.882 Table 7 shows that human judgments are most reliable when perceptual and semantic evidence are evaluated jointly. Perceptual agreement alone has low reliability, with κ=0.110±0.241κ=0.110± 0.241 and Krippendorff’s α=0.094α=0.094, indicating that low-level visual fidelity is difficult to judge consistently under EEG reconstruction noise. Semantic agreement is more stable, with κ=0.483±0.323κ=0.483± 0.323 and α=0.574α=0.574, suggesting moderate consensus on high-level content. Coherence judgments are substantially more reliable, reaching κ=0.882±0.174κ=0.882± 0.174 and α=0.882α=0.882. The gap between perceptual and coherence reliability directly motivates the joint T-PAS/T-SAS design: neither axis alone produces stable human judgments under EEG reconstruction noise. The validated subset spans perception-dominant, semantic-dominant, and coherence-extreme regimes, confirming that BCS targets a disagreement structure that human evaluators consistently recognize. 6 Conclusion Standard reconstruction metrics fail systematically in EEG-to-image evaluation: they penalize semantically recoverable outputs and reward visually plausible ones that depict the wrong content. We introduced BCI-Coherence Score, a lightweight evaluator distilled from four-VLM consensus annotations over 6,855 reconstruction pairs. It separately predicts tolerant perceptual and semantic alignment without repeated VLM inference. Human validation supports the joint design: coherence judgments reach κ=0.882κ=0.882 against κ=0.110κ=0.110 for perceptual judgments alone. BCS provides a reusable evaluation signal for future EEG-to-image benchmarking; extending it to fMRI and cross-subject settings remains open. Appendix 0.A Dataset and Pair Construction The evaluated corpus contains 6,855 ground-truth/reconstruction pairs. Pair identifiers encode the reconstruction method, subject where present, concept, rank, and candidate index. BCS uses a concept-level split to reduce semantic leakage across train, validation, and held-out test partitions. Table 8 reports the composition of the evaluated reconstruction corpus. The Pairs column gives the number of reference/reconstruction pairs contributed by each source, while Subjects and Concepts summarize the available diversity within that source. Table 8: Dataset composition. Method Pairs Subjects Concepts ATM 1,990 10 200 ENIGMA 3,980 10 200 brainvis 200 1 200 cvpr40_brainvis 39 1 39 cvpr40_dreamdiffusion 46 1 46 dreamdiffusion 200 1 200 thingseeg_brainvis 200 1 200 thingseeg_dreamdiffusion 200 1 200 ATM and ENIGMA provide 5,970 of the 6,855 pairs and cover ten subjects and 200 concepts each. The remaining sources contribute smaller but methodologically distinct reconstruction sets, including method-specific outputs from BrainVis and DreamDiffusion and their THINGS-EEG/CVPR40 variants. Table 9 groups pairs by thresholded T-PAS and T-SAS. The four quadrants separate jointly coherent reconstructions from two disagreement regimes: high-P/low-S cases, where reconstructions appear visually plausible but depict nonmatching semantic content, and low-P/high-S cases, where semantic content remains recoverable despite weak visual fidelity. Table 9: Perceptual-semantic regimes by method using the 0.5 threshold. Group N High P/High S High P/Low S Low P/High S Low P/Low S overall 6,855 494 (0.072) 207 (0.030) 455 (0.066) 5699 (0.831) ATM 1,990 85 (0.043) 75 (0.038) 115 (0.058) 1715 (0.862) ENIGMA 3,980 153 (0.038) 121 (0.030) 234 (0.059) 3472 (0.872) brainvis 200 110 (0.550) 5 (0.025) 47 (0.235) 38 (0.190) cvpr40_brainvis 39 23 (0.590) 0 (0.000) 9 (0.231) 7 (0.179) cvpr40_dreamdiffusion 46 9 (0.196) 0 (0.000) 4 (0.087) 33 (0.717) dreamdiffusion 200 1 (0.005) 0 (0.000) 1 (0.005) 198 (0.990) thingseeg_brainvis 200 112 (0.560) 6 (0.030) 44 (0.220) 38 (0.190) thingseeg_dreamdiffusion 200 1 (0.005) 0 (0.000) 1 (0.005) 198 (0.990) Most pairs fall in the low-perceptual/low-semantic quadrant, reflecting the difficulty of EEG reconstruction. BrainVis outputs show more high-P/high-S cases than the larger ATM and ENIGMA sets, while DreamDiffusion outputs are mostly low on both axes in this corpus. Although less frequent, the high-P/low-S quadrant captures the semantic-blindness failure mode in which a reconstruction is visually plausible but semantically mismatched. Appendix 0.B VLM Prompt Specification The annotation protocol is configured as a reference-grounded VLM evaluation. Each call presents the reference image and EEG reconstruction in a fixed order and applies the same six perceptual and six semantic questions. The same configuration is used for all VLM annotators. Table 10: Prompt configuration used for VLM-based annotation. Item Configuration Input order Image 1 is always the ground-truth reference stimulus; Image 2 is always the EEG-generated reconstruction. The ordering is fixed across annotation, rationale generation, and score parsing. Evaluation context Reconstructions are judged under EEG-decoding limitations: blur, low resolution, missing details, distortion, low contrast, artifacts, and stylistic mismatch are expected. The prompt explicitly separates this task from generic photorealistic image-generation evaluation. Question set Both images are shown. The VLM answers six perceptual questions P1 to P6 and six semantic questions S1 to S6. Answer set Each question is answered with exactly one of yes, somewhat, or no. These labels are the only values admitted into scoring. Rationale field Each answer is accompanied by a short evidence sentence. Rationales are used for agreement analysis, but the scalar T-PAS/T-SAS scores use only the discrete answer labels. Table 10 defines the run-level annotation contract. It fixes image ordering, defines the expected reconstruction-noise context, and specifies the question and answer format. These constraints limit the tendency of VLM judges to apply generic image-generation criteria to EEG reconstructions and thereby over-penalize expected low-level degradation. The returned annotation is parsed into question-level answer and rationale fields, grouped into perceptual and semantic blocks. The evaluator prompt text used for scoring is reproduced below. 0.B.1 System Instruction Every annotation call used the same task instruction. The system message defined the VLM as an expert evaluator for EEG-to-image reconstruction and specified the input order: Image 1 is the ground-truth reference stimulus shown during EEG recording, and Image 2 is the EEG-generated reconstruction produced by a neural decoding model. The instruction emphasized that EEG-to-image outputs are often blurry, low resolution, missing fine detail, distorted, stylistically different from the reference, noisy, artifact-laden, low contrast, or flat in appearance. The VLM was therefore instructed to judge content preservation under EEG-specific limitations rather than photorealistic image-generation quality. For every question, the VLM selected exactly one answer from yes, somewhat, and no. A yes answer means that the criterion is clearly satisfied, somewhat means that it is partially satisfied with noticeable gaps or errors, and no means that it is mostly not satisfied. Each answer was accompanied by a short reasoning sentence explaining the visible evidence for the label. No answers outside this three-level scale were accepted. 0.B.2 Perceptual Prompts Each perceptual question targets one visual dimension. The VLM is instructed to answer each question independently, tolerate EEG-specific degradation such as blur, noise, and low detail, and penalize only the property specified by the question. P1: Global spatial structure. Does the generated image preserve the coarse spatial organization of the reference, including the approximate position, size, and arrangement of the dominant regions? This prompt asks only where things are in the image and how the space is divided; shape, color, texture, and meaning are not considered. Tolerated deviations include blurry or softened region boundaries, approximate rather than exact positioning, simplified spatial layout, and missing background elements. Penalized errors include placing the dominant object in a completely wrong location, inverting foreground and background relationships, producing a spatial layout unrelated to the reference, or omitting the primary spatial region entirely. P2: Object shape and silhouette. Does the dominant object or figure in the generated image have a shape or silhouette similar to the reference? This prompt asks only about the form of the main object: its outline, contour, and overall geometric structure. Color, texture, position, and meaning are not considered. Tolerated deviations include simplified or smoothed contours, missing fine shape detail, mild deformation of secondary parts, and blurred object boundaries. Penalized errors include a completely wrong dominant-object shape, a silhouette from a different object class, an unrecognizable or collapsed object form, or a shape that implies a different category. P3: Surface texture and material. Does the surface texture and material appearance of the dominant object in the generated image resemble the reference? This prompt asks only about surface quality: smooth versus rough, matte versus shiny, organic versus manufactured, and fine-grained versus coarse. Color, shape, position, and meaning are not considered. Tolerated deviations include reduced texture resolution or fidelity, blurry or flat surface appearance, approximate rather than exact material match, and missing fine surface detail. Penalized errors include a completely wrong surface material, such as fur rendered as metal, a texture pattern belonging to a different object class, or no recoverable surface information. P4: Color and chromatic consistency. Does the dominant color palette of the generated image reasonably match the reference? This prompt asks only about color: dominant hues, approximate saturation, and broad chromatic character of the main object and scene. Texture, shape, position, and meaning are not considered. Tolerated deviations include approximate rather than exact color match, reduced saturation or muted palette, slight hue shift, and missing color variation in secondary regions. Penalized errors include a completely wrong dominant color, a palette associated with a different object type, or chromatic information that is absent or inverted. P5: Absence of dominant artifacts. Is the generated image free from severe visual artifacts that dominate or overwhelm the content? This prompt asks only about artifacts: structured noise, grid patterns, color bleeding, repetitive hallucinated patterns, or generation failures. Blur and low detail are not counted as artifacts. Tolerated deviations include blur, low sharpness, missing detail, flat regions, mild color noise, and soft or undefined edges. Penalized errors include structured artifacts dominating large image regions, repetitive or tiled hallucinated patterns, severe color bleeding across object boundaries, or fragmented incoherent structure from generation failure. P6: Holistic visual recoverability. Taking the generated image as a whole, can a human observer still recover the primary visual content of the reference despite EEG reconstruction degradation? This is a gestalt-level prompt asking whether the overall image, across structure, shape, texture, and color together, carries enough visual evidence to identify what is being depicted. It is not answered from any single dimension alone. Tolerated deviations include any individual visual dimension being weak or missing, low overall visual quality, and degraded or noisy appearance. Penalized errors include no recoverable content across any dimension, an image that is entirely noise or artifact, or an image from which no visual content can be inferred. 0.B.3 Semantic Prompts Semantic prompts ignore photorealism and focus on meaning. The VLM is instructed that a blurry or distorted image can still carry correct semantic content, and that each semantic question should be answered independently. S1: Basic category identity. Does the generated image depict an object or scene belonging to the same basic-level category as the reference? Basic-level categories include animal, vehicle, food, furniture, tool, clothing, building, plant, person, household object, electronic device, natural scene, and indoor scene. This prompt asks only about the broadest categorical identity; specific type, attributes, count, and scene context are not considered. Example judgments include elephant to elephant as yes, elephant to rhinoceros as somewhat, elephant to car as no, apple to orange as yes, and apple to baseball as no. Tolerated deviations include wrong specific type within the correct category, visual degradation making details unclear, and missing secondary objects. Penalized errors include a clearly different basic-level category or categorically different scene type. S2: Subordinate identity. Does the generated image depict the specific type of object shown in the reference, beyond just the basic category? This prompt asks about subordinate-level identity: not just animal but dog versus cat versus elephant, not just vehicle but car versus truck versus bicycle, and not just food but apple versus banana versus pizza. Basic category, function, count, and scene context are not considered. Example judgments include golden retriever to labrador as yes, golden retriever to german shepherd as somewhat, golden retriever to cat as no, sports car to sedan as somewhat, and sports car to truck as no. Tolerated deviations include approximate species or type match, genuinely ambiguous specific type, and visual degradation obscuring fine distinctions. Penalized errors include an unambiguously wrong specific type or a generated image that clearly implies a different subordinate identity. S3: Functional role and purpose. Does the dominant object in the generated image serve the same functional role or purpose as in the reference? This prompt asks what the object is for or what it does, not its category or appearance. Functional roles include seating, transportation, food consumption, cutting or handling, display or communication, containment, and locomotion. Category identity, appearance, count, and scene context are not considered. Example judgments include chair to stool as yes because both provide seating, knife to fork as somewhat because both are utensils with different functions, knife to pen as no, car to bicycle as somewhat because both support transportation, and car to house as no. Tolerated deviations include approximate or related functional role and visual degradation obscuring object details. Penalized errors include a completely different functional role or purpose class. S4: Quantity and cardinality. Does the generated image show approximately the same number of primary objects as the reference? This prompt asks only about quantity: how many main objects are present. Object identity, location, and meaning are not considered. The VLM counts only primary objects, ignores background elements, and ignores secondary or decorative objects. Example judgments include one dog to one dog as yes, one dog to two dogs as no, three apples to two apples as somewhat, three apples to one apple as no, a crowd of people to several people as yes, and one person to a crowd as no. Tolerated deviations include being off by one when the total count is large and ambiguous, missing secondary or background objects, and degradation that makes exact count genuinely unclear. Penalized errors include clearly wrong count or an unambiguous singular/plural inversion. S5: Scene context and environment. Does the generated image imply the same scene context or environment as the reference, independent of the main object? This prompt asks about setting or background context, not the main object. It distinguishes scene pairs such as indoor versus outdoor, natural versus urban or manufactured, water versus land, domestic versus wild, and aerial versus ground level. Main-object identity, function, count, and appearance are not considered. Example judgments include dog in a park to dog in a field as yes, dog in a park to dog in a kitchen as no, car on highway to car on road as yes, car on highway to car underwater as no, and bird on branch to bird over trees as yes. Tolerated deviations include missing background detail, blurry or undefined environment, and approximate setting match. Penalized errors include categorically different context, indoor/outdoor inversion, natural/urban inversion, or a setting implying a completely different environment. S6: Semantic recoverability under noise. Taking the generated image as a whole, can the intended semantic content of the reference be inferred from the generated image despite EEG reconstruction noise? This is a gestalt-level semantic prompt asking whether enough semantic evidence remains across category, identity, function, quantity, and scene to identify what the reconstruction is trying to represent. It is not answered from any single semantic dimension alone. Tolerated deviations include any individual semantic dimension being weak, low visual quality or heavy degradation, partial or approximate semantic evidence, and only one or two semantic dimensions surviving. Penalized errors include no recoverable semantic content across any dimension, a generated image that actively suggests a completely different concept, or an image from which the intended object or scene cannot be inferred. 0.B.4 Answer Mapping and Scores All prompts use the same three-level answer scale. Let ai,q(v)a_i,q^(v) denote the parsed answer from VLM v for image pair i and question q: ϕ(no)=0,ϕ(somewhat)=0.5,ϕ(yes)=1.φ( no)=0, φ( somewhat)=0.5, φ( yes)=1. (21) Per-VLM perceptual and semantic scores are T-PASi(v)=16∑p=16ϕ(ai,p(v)),T-SASi(v)=16∑s=16ϕ(ai,s(v)).T -PAS_i^(v)= 16 _p=1^6φ(a_i,p^(v)), -SAS_i^(v)= 16 _s=1^6φ(a_i,s^(v)). (22) For downstream distillation, consensus targets are formed at the question level: zi,q=medianv∈ϕ(ai,q(v)),z_i,q=median_v φ(a_i,q^(v)), (23) where V is the four-VLM annotator set. This preserves partial disagreement; with four annotators, median targets can take values such as 0.25 and 0.75. All twelve question dimensions enter the per-pair averages, BCS loss, and reported aggregate scores. Appendix 0.C Annotation and Model Details 0.C.1 VLM Annotators The annotation run used four open-weight VLMs with a common prompt. The model set provides complementary calibration behavior while remaining reproducible from public checkpoints. Table 11 lists the checkpoints and inference settings used for the annotation pass. Valid is the fraction of image pairs that produced parseable structured annotations, while mean T-PAS and T-SAS summarize each model’s operating point. Table 11: Final VLM annotator configuration and aggregate operating points. Annotator Checkpoint Dtype Batch Max toks Valid Mean T-PAS Mean T-SAS InternVL3 OpenGVLab/InternVL3-8B bf16 1,2,4,8 80 1.000 0.341 0.252 SAIL-VL BytedanceDouyinContent/SAIL-VL-1d6-8B bf16 24 128 1.000 0.213 0.157 OLA-7B THUdyh/Ola-7b bf16 32 128 1.000 0.433 0.335 Ovis2-8B AIDC-AI/Ovis2-8B bf16 8 128 1.000 0.337 0.345 All four annotators produced valid annotations for the full corpus. Their mean scores differ meaningfully: SAIL-VL is the most conservative on both axes, OLA-7B is more permissive perceptually, and Ovis2-8B gives the highest average semantic score. This variation motivates the four-model consensus, which reduces dependence on any single model’s calibration. 0.C.2 Metric-to-Consensus Correlation Conventional metrics are oriented so larger values indicate better reconstruction quality. Table 12 reports their alignment with the four-VLM consensus T-PAS and T-SAS axes. Spearman correlation measures rank agreement, while Pearson correlation measures linear association with the consensus scores. Table 12: Correlation between conventional metrics and VLM consensus T-PAS/T-SAS. Metric ρ(T-PAS) r(T-PAS) ρ(T-SAS) r(T-SAS) MSE 0.137 0.124 0.074 0.056 PSNR 0.137 0.122 0.074 0.054 SSIM 0.208 0.127 0.060 -0.015 LPIPS 0.144 0.147 0.239 0.228 DISTS 0.363 0.367 0.460 0.447 DreamSim 0.421 0.536 0.627 0.683 OpenCLIP 0.415 0.534 0.676 0.727 DINOv2 0.401 0.571 0.577 0.701 TOPIQ-FR 0.240 0.244 0.297 0.296 PieAPP 0.070 0.073 0.182 0.168 ImageReward 0.255 0.488 0.426 0.625 BLIP-SBERT 0.301 0.409 0.468 0.554 Low-level metrics such as MSE, PSNR, SSIM, and PieAPP align weakly with the VLM consensus, especially on T-SAS. Feature and reward metrics correlate more strongly with semantic consensus, with OpenCLIP, DreamSim, DINOv2, and ImageReward showing the clearest associations. However, these metrics still produce a single scalar similarity score and do not separate perceptual recoverability from semantic recoverability, which motivates T-PAS/T-SAS and BCS. 0.C.3 BCS Training Configuration BCS distills four-VLM consensus into a compact image-pair regressor. The model is trained at the question level: for each image pair i and prompt dimension q, the four VLM answers are first mapped to 0,0.5,1\0,0.5,1\ and aggregated into a consensus target zi,qz_i,q. Agreement weights give more influence to prompt instances where the VLMs agree and less influence to ambiguous instances. For each frozen image encoder EkE_k, BCS computes normalized embeddings for the reference image yiy_i and reconstruction y^i y_i: i,kref=Ek(yi),i,krec=Ek(y^i).e^ref_i,k=E_k(y_i), ^rec_i,k=E_k( y_i). (24) The image-pair feature block for encoder k concatenates the reference embedding, reconstruction embedding, absolute difference, elementwise product, and cosine similarity: i,k=[i,kref;i,krec;|i,kref−i,krec|;i,kref⊙i,krec;⟨i,kref,i,krec⟩].h_i,k= [e^ref_i,k;e^rec_i,k; |e^ref_i,k-e^rec_i,k |;e^ref_i,k ^rec_i,k; ^ref_i,k,e^rec_i,k ]. (25) The final BCS input is the concatenation of these blocks across SigLIP, CLIP, and DINOv3, giving a 9,219-dimensional feature vector. Encoder embeddings are precomputed and cached; all encoders remain frozen during BCS training. The prediction head is a residual MLP with a layer-normalized input projection, hidden dimension 512, two residual blocks, dropout 0.2, and sigmoid outputs for the twelve prompt dimensions. It is optimized with AdamW using learning rate 5×10−45× 10^-4, weight decay 5×10−45× 10^-4, batch size 512, and a maximum of 120 epochs. Training uses early stopping on validation loss with patience 16, cosine learning-rate decay, and gradient-norm clipping at 1.0. The concept-level split uses seed 42 with a 0.70/0.15/0.15 train/validation/test ratio. For the continuous BCS model, the loss is the weighted mean-squared error over the twelve prompt dimensions: ℒBCS=∑i,qwi,q(z^i,q−zi,q)2∑i,qwi,q,L_BCS= _i,qw_i,q ( z_i,q-z_i,q )^2 _i,qw_i,q, (26) where wi,qw_i,q is the agreement weight. The continuous ensemble averages raw predictions from five independently seeded MLPs with seeds 7, 13, 21, 37, and 42. Held-out evaluation reports question-level MAE and correlation, then aggregates predicted perceptual and semantic prompt scores into BCS estimates of T-PAS and T-SAS. Table 13 summarizes the implementation settings for the BCS training run. The encoder row identifies the frozen backbones, the feature-dimension row gives the concatenated pair-feature size, and the split/seed rows define the data partition and ensemble members. The optimization and loss rows summarize the training schedule used for each seed. Table 13: BCS training and architecture details. Item Setting Encoders google/siglip-base-patch16-224; openai/clip-vit-large-patch14; timm/vit_base_patch16_dinov3.lvd1689m Feature dimension 9,219 Split concept split, seed 42, train 0.70, val 0.15, test 0.15 Seeds 7, 13, 21, 37, 42 MLP 2 residual blocks, hidden 512, dropout 0.2 Optimization lr 0.0005, weight decay 0.0005, batch 512, epochs 120, patience 16 Loss MSE weight 1.0, CE weight 0.35 Runtime seed-7 total 27.6s; ensemble uses 5 seeds References [1] Y. Bai, X. Wang, Y. Cao, Y. Ge, C. Yuan, and Y. Shan (2024) DreamDiffusion: high-quality EEG-to-image generation with temporal masked signal modeling and CLIP alignment. In Computer Vision – ECCV 2024, p. 472–488. External Links: Document Cited by: §2. [2] M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton (2018) Demystifying MMD GANs. In International Conference on Learning Representations, Cited by: §2. [3] X. Cao, P. Gong, L. Zhang, and D. Zhang (2026) EEG-CLIP: a transformer-based framework for EEG-guided image generation. Neural Networks 194, p. 108167. External Links: Document Cited by: §1, §2. [4] Z. Chen, J. Qing, T. Xiang, W. L. Yue, and J. H. Zhou (2023) Seeing beyond the brain: conditional diffusion model with sparse masked modeling for vision decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 22710–22720. Cited by: §2. [5] K. Ding, K. Ma, S. Wang, and E. P. Simoncelli (2022) Image quality assessment: unifying structure and texture similarity. IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (5), p. 2567–2581. External Links: Document Cited by: §1, §2. [6] T. Fei and V. R. de Sa (2024) Image reconstruction from electroencephalography using latent diffusion. External Links: 2404.01250, Document Cited by: §1, §2. [7] H. Fu, Z. Shen, J. J. Chin, and H. Wang (2025) BrainVis: exploring the bridge between brain and visual signals via image reconstruction. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, p. 1–5. External Links: Document Cited by: §2. [8] S. Fu, N. R. Tamir, S. Sundaram, L. Chai, R. Zhang, T. Dekel, and P. Isola (2023) DreamSim: learning new dimensions of human visual similarity using synthetic data. In Advances in Neural Information Processing Systems, Cited by: §1, §2. [9] D. Ghosh, H. Hajishirzi, and L. Schmidt (2023) GenEval: an object-focused framework for evaluating text-to-image alignment. External Links: 2310.11513 Cited by: §1. [10] S. Guenther, N. Kosmyna, and P. Maes (2024) Image classification and reconstruction from low-density EEG. Scientific Reports 14, p. 16436. External Links: Document Cited by: §2. [11] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter (2017) GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing Systems, Vol. 30, p. 6626–6637. Cited by: §2. [12] T. Horikawa and Y. Kamitani (2017) Generic decoding of seen and imagined objects using hierarchical visual features. Nature Communications 8, p. 15037. External Links: Document Cited by: §2. [13] Y. Hu, B. Liu, J. Kasai, Y. Wang, M. Ostendorf, R. Krishna, and N. A. Smith (2023) TIFA: accurate and interpretable text-to-image faithfulness evaluation with question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 20406–20417. Cited by: §1. [14] Y. Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy (2023) Pick-a-Pic: an open dataset of user preferences for text-to-image generation. External Links: 2305.01569 Cited by: §2. [15] M. Ku, D. Jiang, C. Wei, X. Yue, and W. Chen (2024) VIEScore: towards explainable metrics for conditional image synthesis evaluation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, p. 12268–12290. External Links: Document Cited by: §2. [16] D. Li, C. Wei, S. Li, J. Zou, and Q. Liu (2024) Visual decoding and reconstruction via EEG embeddings with guided diffusion. In Advances in Neural Information Processing Systems, Vol. 37. Cited by: §2. [17] Z. Lin, D. Pathak, B. Li, J. Li, X. Xia, G. Neubig, P. Zhang, and D. Ramanan (2024) Evaluating text-to-visual generation with image-to-text generation. In Computer Vision – ECCV 2024, p. 366–384. External Links: Document Cited by: §1, §2. [18] D. Mayo, C. Wang, A. Harbin, A. Alabdulkareem, A. E. Shaw, B. Katz, and A. Barbu (2024) BrainBits: how much of the brain are generative reconstruction methods using?. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §2. [19] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. (2023) DINOv2: learning robust visual features without supervision. External Links: 2304.07193 Cited by: §1, §2. [20] F. Ozcelik and R. VanRullen (2023) Natural scene reconstruction from fMRI signals using generative latent diffusion. Scientific Reports 13, p. 15666. External Links: Document Cited by: §2. [21] E. Prashnani, H. Cai, Y. Mostofi, and P. Sen (2018) PieAPP: perceptual image-error assessment through pairwise preference. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 1808–1817. Cited by: §2. [22] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, p. 8748–8763. Cited by: §1, §2. [23] Z. Ren, J. Li, X. Xue, X. Li, F. Yang, Z. Jiao, and X. Gao (2021) Reconstructing seen image from brain activity by visually-guided cognitive representation and adversarial learning. NeuroImage 228, p. 117602. External Links: Document Cited by: §2. [24] P. S. Scotti, A. Banerjee, J. Goode, S. Shabalin, A. Nguyen, E. Cohen, A. J. Dempster, N. Verlinde, E. Yundler, D. Weisberg, K. A. Norman, and T. M. Abraham (2023) Reconstructing the mind’s eye: fMRI-to-image with contrastive learning and diffusion priors. In Advances in Neural Information Processing Systems, Vol. 36, p. 24705–24728. Cited by: §2. [25] P. S. Scotti, M. Tripathy, C. Torrico, R. Kneeland, T. Chen, A. Narang, C. Santhirasegaran, J. Xu, T. Naselaris, K. A. Norman, and T. M. Abraham (2024) MindEye2: shared-subject models enable fMRI-to-image with 1 hour of data. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, p. 44038–44059. Cited by: §2. [26] G. Shen, T. Horikawa, K. Majima, and Y. Kamitani (2019) Deep image reconstruction from human brain activity. PLOS Computational Biology 15 (1), p. e1006633. External Links: Document Cited by: §1, §2. [27] P. Singh, P. Pandey, K. Miyapuram, and S. Raman (2023) EEG2IMAGE: image reconstruction from EEG brain signals. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, p. 1–5. External Links: Document Cited by: §2. [28] Y. Takagi and S. Nishimoto (2023) High-resolution image reconstruction with latent diffusion models from human brain activity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14453–14463. Cited by: §1, §2. [29] R. VanRullen, L. Reddy, and G. A. Rousselet (2021) Natural image reconstruction from fMRI using deep learning: a survey. Neural Networks 144, p. 329–339. External Links: Document Cited by: §1. [30] J. Wang, K. C. K. Chan, and C. C. Loy (2023) Exploring CLIP for assessing the look and feel of images. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, p. 2555–2563. External Links: Document Cited by: §2. [31] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli (2004) Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13 (4), p. 600–612. External Links: Document Cited by: §1, §2. [32] C. K. Wong, Y. Liu, and H. Yan (2026) Brain-to-image retrieval and reconstruction via multimodal EEG alignment. Note: arXiv:2605.23996 External Links: Document Cited by: §2. [33] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang (2018) The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 586–595. Cited by: §2.