Paper deep dive
EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits
Kaifan Zhang, Lihuo He, Yuqi Ji, Junjie Ke, Lukun Wu, Tianhao You, Xinbo Gao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/1/2026, 2:27:49 AM
Summary
The paper introduces EEG-EditBench, a diagnostic benchmark for evaluating EEG-to-image retrieval models using controlled image edits. It analyzes how well eight representative models can distinguish original images from variants with altered object identity, attributes, background, or presence, revealing that standard retrieval accuracy does not guarantee robustness to fine-grained visual changes.
Entities (11)
Relation Signals (12)
EEG-EditBench â containsedittype â Object Identity Edit
confidence 95% · EEG-EditBench defines four edit families... Object Identity Edit
EEG-EditBench â containsedittype â Object Removal
confidence 95% · Object Removal... removes the main object
EEG-EditBench â containsedittype â Background Edit
confidence 95% · Background Edit... replaces the surrounding scene
EEG-EditBench â containsedittype â Attribute Edit
confidence 95% · Attribute Edit... preserves the object identity
EEG-EditBench â usesdataset â THINGS-EEG2
confidence 95% · Built from the 200 THINGS-EEG2 test images
EEG-EditBench â evaluatesmodel â Brain-HIVE
confidence 92% · We evaluate eight representative EEG visual decoding models... Brain-HIVE performs best
EEG-EditBench â evaluatesmodel â ATM
confidence 90% · We evaluate... ATM... on EEG-EditBench
EEG-EditBench â evaluatesmodel â NICE
confidence 90% · We evaluate NICE... on EEG-EditBench
Object Identity Edit â iseasierthan â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent EEG-to-image retrieval models have achieved strong performance in identifying viewed images from semantically diverse candidates. Yet such success does not reveal what visual information supports the match. A model may readily identify a cheetah among tools, plants, and vehicles, but can it still distinguish the viewed cheetah from the same scene with the cheetah replaced by a dog? Motivated by this question, we introduce EEG-EditBench, a diagnostic benchmark that examines this question through controlled edits of object identity, attributes, background, and object presence. Built from the 200 THINGS-EEG2 test images, EEG-EditBench contains 2,137 quality-controlled edits and evaluates eight representative EEG visual decoding models. Our results show that strong standard retrieval does not consistently transfer to edit-based evaluation, with fine-grained attribute changes presenting the greatest challenge. EEG-EditBench reveals model behavior hidden by aggregate retrieval accuracy and provides a controlled basis for studying what visual information EEG-image models preserve. The code and complete dataset are publicly available.
Tags
Links
- Source: https://arxiv.org/abs/2607.27857v1
- Canonical: https://arxiv.org/abs/2607.27857v1
Trouble viewing inline? Open PDF directly â
Full Text
74,696 characters extracted from source content.
Expand or collapse full text
EEG-EditBench: Probing Visual Information in EEGâImage Retrieval Models with Controlled Image Edits Kaifan Zhang1, Lihuo He1 , Yuqi Ji1, Junjie Ke2, Lukun Wu1, Tianhao You3, Xinbo Gao1 Abstract Recent EEG-to-image retrieval models have achieved strong performance in identifying viewed images from semantically diverse candidates. Yet such success does not reveal what visual information supports the match. A model may readily identify a cheetah among tools, plants, and vehicles, but can it still distinguish the viewed cheetah from the same scene with the cheetah replaced by a dog? Motivated by this question, we introduce EEG-EditBench, a diagnostic benchmark that examines this question through controlled edits of object identity, attributes, background, and object presence. Built from the 200 THINGS-EEG2 test images, EEG-EditBench contains 2,137 quality-controlled edits and evaluates eight representative EEG visual decoding models. Our results show that strong standard retrieval does not consistently transfer to edit-based evaluation, with fine-grained attribute changes presenting the greatest challenge. EEG-EditBench reveals model behavior hidden by aggregate retrieval accuracy and provides a controlled basis for studying what visual information EEGâimage models preserve. The code and complete dataset are available through the project repository. Code and Dataset â https://github.com/XiaoZhangYES/EEG-EditBench 1 Introduction Decoding visual information from brain signals is a central problem in neuroscience and brain-computer interface research (Kamitani and Tong 2005; Kay et al. 2008; Naselaris et al. 2009; Cichy et al. 2016). Large-scale natural-image datasets such as THINGS-EEG and THINGS-EEG2 (Grootswagers et al. 2022; Gifford et al. 2022) have supported recent progress in EEG-to-image retrieval, where EEG responses and visual stimuli are mapped into a shared embedding space (Song et al. 2024; Wei et al. 2024; Li et al. 2024; Wu et al. 2025). The standard evaluation is a 200-way zero-shot task that asks a model to retrieve the viewed image from candidates representing distinct object concepts. Recent methods have achieved strong performance under this protocol (Jo et al. 2026; Zheng et al. 2026; Liu et al. 2026). However, the standard 200-way protocol leaves an important question open. Consider an EEG response evoked by viewing an image of a cheetah. In the standard setting, the cheetah image is compared with candidates from different object concepts, such as a notebook, a drum, or an aircraft carrier. If the model retrieves the cheetah correctly, it has clearly learned a useful alignment between EEG and visual representations. Yet this success leaves open which visual distinctions the learned EEGâimage alignment can support. The model may distinguish the target through relatively coarse cues such as âanimal,â âspotted animal,â or âanimal in grassland,â while finer distinctions in fur color, texture, background context, or nearby objects remain untested. Standard retrieval therefore shows that the model can find the viewed image among semantically diverse candidates, while leaving unclear which visual information supports the match. Figure 1: From standard 200-way retrieval to edit-based diagnosis. Standard EEG-image retrieval asks whether a model can identify the viewed image from semantically diverse candidates. EEG-EditBench instead compares each viewed image with controlled variants, enabling factor-level analysis of EEGâimage matching. A single aggregate accuracy cannot resolve this ambiguity. Two models may achieve similar retrieval accuracy while relying on different visual evidence, and a model that performs well on cross-concept candidates may still fail when the alternatives differ only in a specific attribute or contextual detail. Diagnostic evaluation addresses this limitation through controlled examples that isolate particular model capabilities (Ribeiro et al. 2020; Zhao et al. 2022; Thrush et al. 2022; Hsieh et al. 2023). Image editing provides a practical way to construct such examples by changing a specific visual factor while preserving the remaining image content as much as possible (Zhang et al. 2023; Ma et al. 2024). This perspective naturally extends to EEGâimage retrieval, where the EEG query can remain fixed while the candidate image is varied in a controlled manner. Building on this idea, we introduce EEG-EditBench, a controlled edit-based benchmark for probing visual information in EEGâimage retrieval models. As illustrated in Figure 1, EEG-EditBench shifts evaluation from cross-concept retrieval to controlled comparisons between each source image and variants that alter object identity, object attributes, background context, or object presence. Starting from the 200 THINGS-EEG2 test images, we construct quality-controlled edited variants that retain substantial source-image content while differing in a specific visual factor. EEG-EditBench evaluates whether a model can distinguish each viewed image from its controlled variants, both individually and when multiple edited variants compete within the same candidate pool. These complementary settings support factor-level analysis as well as a more challenging joint comparison. We evaluate eight representative EEG visual decoding models on EEG-EditBench (Song et al. 2024; Li et al. 2024; Wei et al. 2024; Wu et al. 2025, 2026; Jo et al. 2026; Zhang et al. 2026b; Zheng et al. 2026) and obtain three main findings: âą High accuracy in standard EEG-to-image retrieval does not guarantee that a model can distinguish the viewed image from its controlled edits. âą Across the evaluated models, object identity and removal edits are distinguished more accurately than edits to color, material, texture, shape, and state. âą The same type of edit can differ in difficulty across image categories. Figure 2: Edit taxonomy of EEG-EditBench. Identity edits are organized by semantic proximity and attribute edits by the changed visual property; background and removal edits target scene context and object presence, respectively. Figure 3: Dataset construction pipeline of EEG-EditBench. A structured scene profile guides edit-target and prompt generation through family-specific branches, followed by image editing. Figure 4: Composition of EEG-EditBench after quality control. The figure summarizes the dataset scale and the distribution of edit families and subtypes. Subtype percentages are computed within each edit family. 2 Related Work EEG-Based Visual Decoding Earlier studies used single-trial EEG decoding and representational analysis to characterize the temporal dynamics of object processing, and explored multimodal and deep visual representation learning (Kaneshiro et al. 2015; Du et al. 2023; Singh et al. 2024). Large-scale natural-image datasets such as THINGS-EEG and THINGS-EEG2 have since supported progress in visual decoding from EEG (Grootswagers et al. 2022; Gifford et al. 2022). Existing methods either align EEG and visual representations for image retrieval (Song et al. 2024; Wei et al. 2024) or combine EEG-derived representations with generative models for image reconstruction (Bai et al. 2024; Li et al. 2024; Zhang et al. 2025, 2026a). More recent work improves EEGâimage alignment through visual priors, hierarchical representations, hyperbolic modeling, adaptive teaching, and self-supervised objectives (Wu et al. 2025; Zheng et al. 2026; Jo et al. 2026; Wu et al. 2026; Zhang et al. 2026b). These advances are commonly evaluated through aggregate retrieval or reconstruction metrics, which measure whether a model recovers the target image but reveal less about the visual information it preserves. Recent work has broadened brain-decoding evaluation toward fine-grained semantic content, multigranular image comparison, and mental imagery (Xia and Oztireli 2025, 2026; Kneeland et al. 2025). Diagnostic Evaluation with Controlled Image Edits Diagnostic benchmarks use controlled examples or hard negatives to expose model capabilities hidden by aggregate performance. This approach has been applied to compositional reasoning, language understanding, and vision-language grounding (Johnson et al. 2017; Ribeiro et al. 2020; Shekhar et al. 2017; Parcalabescu et al. 2022; Zhao et al. 2022; Thrush et al. 2022; Yuksekgonul et al. 2023; Ma et al. 2023; Hsieh et al. 2023). EEG-based retrieval differs because the query is a neural response evoked by viewing an image; factor-level diagnosis therefore varies the image candidates while keeping the EEG query fixed. Image editing benchmarks provide a practical basis for this construction. Existing benchmarks evaluate instruction following, visual realism, and preservation of content unrelated to the requested change (Wang et al. 2023; Zhang et al. 2023; Ma et al. 2024). Recent work expands edit diversity and preservation-oriented evaluation through inversion-based benchmarks, curated high-quality instructionâedit pairs, and unified multi-type editing datasets (Ju et al. 2024; Hui et al. 2025; Yu et al. 2025). Counterfactual image generation similarly emphasizes targeted interventions with minimal unintended changes (Melistas et al. 2024). EEG-EditBench follows these principles and uses quality-controlled edits as controlled candidates for measuring changes in EEGâimage matching. 3 Method Problem Formulation For each concept c, let ece_c denote the EEG response evoked by the original source image IcorigI_c^orig. EEG-EditBench evaluates the same EEG query against a set of valid edited images EcE_c, where Ec(f)âEcE_c^(f) E_c contains edits from edit family fâid,attr,bg,rmfâ\id,attr,bg,rm\, corresponding to object identity, object attribute, background, and object removal. For a given retrieval model, the EEG encoder Ïâ(â )Ï(·) and image encoder Ïâ(â )Ï(·) map EEG responses and images into a shared embedding space. Their similarity is written as sâ(e,I)=âšÏâ(e),Ïâ(I)â©,s(e,I)= Ï(e),Ï(I) , where the inner product denotes the model-specific score used for ranking. EEG-EditBench leaves this scoring function unchanged and modifies only the candidate images used for evaluation. It therefore characterizes how the complete EEGâimage retrieval system behaves under controlled image-side changes, reflecting the combined effects of the EEG encoder, image encoder, and alignment objective. Edit Taxonomy As illustrated in Figure 2, EEG-EditBench defines four edit families that target complementary visual factors. Object Identity Edit. This edit replaces the main object while preserving the composition, viewpoint, background, and visual style. Replacement targets are selected from the remaining THINGS-EEG2 test concepts and grouped into three scene-conditioned subtypes based on semantic distance: near targets are semantically related to the source or share a similar function or affordance; medium targets are less closely related but can still play a similar role in the scene; and far targets are semantically distinct yet remain plausible replacements. Attribute Edit. This edit preserves the object identity and background while changing one visible property of the main object. We consider color, texture, shape, material, and state. Background Edit. This edit replaces the surrounding scene while preserving the foreground object, its appearance, and its spatial arrangement. Object Removal. This edit removes the main object and completes the exposed region with plausible background content, including associated shadows, reflections, or contact boundaries when present. Together, the four edit families evaluate model behavior under changes in object identity, appearance, scene context, and object presence. Dataset Construction Pipeline EEG-EditBench is constructed from the 200 THINGS-EEG2 test images, whose object concepts originate from the THINGS database (Hebart et al. 2019; Gifford et al. 2022), using the four-stage pipeline shown in Figure 3. A vision-language model first produces a structured scene profile describing the main object, visible attributes, spatial layout, and background context, and determines the applicable edit families. It then generates edit targets and self-contained prompts that specify the intended change and the content to preserve. An image editing model uses each source image and prompt to generate the edited variant. All generated images undergo full human review by two independent reviewers. Each image is assessed for whether the requested edit is correctly realized, whether content unrelated to the edit is sufficiently preserved, and whether the result remains visually coherent without salient editing artifacts. Images that fail any of these criteria are excluded, and disagreements between the two reviewers are resolved by a third reviewer. This process yields 2,137 retained edits. The review criteria and adjudication procedure are provided in the supplementary material. The final benchmark contains 979 Attribute Edits, 780 Object Identity Edits, 199 Background Edits, and 179 Object Removals. Figure 4 shows the distribution across edit families and subtypes. Figure 5: Cosine distances between edited images and their corresponding source images. The dashed line and shaded region show the median and interquartile range of the per-image median distances to source images from the other test concepts. Diamonds indicate means. Retrieval 2AFC Accuracy Model Venue Visual Degradation 200-way Original 200-way Null EP-Top1 Identity Attribute Background Removal Overall NICE ICLRâ24 â 19.5 ± 3.8 19.8 ± 4.2 14.0 ± 2.6 79.1 ± 1.8 55.8 ± 3.5 66.6 ± 5.4 84.7 ± 3.7 67.7 ± 2.3 ATM NeurIPSâ24 â 29.8 ± 5.5 22.7 ± 4.4 38.1 ± 5.9 86.7 ± 1.8 77.7 ± 4.1 77.8 ± 3.6 86.0 ± 3.7 81.7 ± 2.7 MB2C ACM Mâ24 â 27.5 ± 6.3 25.7 ± 5.3 26.7 ± 3.2 83.5 ± 2.5 66.8 ± 3.1 77.1 ± 3.7 83.6 ± 3.0 75.3 ± 2.5 UBP CVPRâ25 â 52.8 ± 5.3 46.7 ± 5.6 18.2 ± 2.2 88.8 ± 1.8 65.6 ± 2.1 80.4 ± 2.7 87.9 ± 3.9 77.3 ± 1.8 ATS AAAIâ26 â 56.1 ± 7.3 49.5 ± 6.4 26.7 ± 3.2 89.7 ± 1.7 71.2 ± 2.1 84.8 ± 2.7 94.2 ± 2.3 81.1 ± 1.6 NeuroBridge AAAIâ26 â 62.3 ± 6.5 59.8 ± 7.6 27.6 ± 2.5 91.0 ± 1.1 72.4 ± 1.9 87.2 ± 1.9 96.2 ± 1.1 82.6 ± 1.1 HyFI AAAIâ26 â 68.2 ± 6.5 64.2 ± 7.0 33.4 ± 2.6 92.8 ± 1.1 75.7 ± 1.4 88.3 ± 2.3 97.7 ± ± 1.5 84.9 ± 1.1 Brain-HIVE ICLRâ26 â 75.2 ± ± 6.5 65.5 ± ± 5.7 46.3 ± ± 5.7 95.6 ± ± 1.3 80.5 ± ± 3.6 92.4 ± ± 2.7 97.5 ± 1.2 88.6 ± ± 2.2 Table 1: Main results on standard retrieval and EEG-EditBench. Original and Null denote 200-way Top-1 accuracy using the original and null-edited candidate sets, respectively; EP-Top1 and 2AFC report edit-based performance. A checkmark indicates the use of degraded visual representations. Best results are shown in bold. Object Identity Attribute Model Near Medium Far Overall Color Material Texture Shape State Overall NICE 75.0 ± 2.5 76.5 ± 2.3 86.8 ± 2.2 79.1 ± 1.8 57.3 ± 2.7 51.9 ± 2.8 56.9 ± 4.1 56.0 ± 5.2 56.7 ± 5.2 55.8 ± 3.5 ATM 84.7 ± 2.0 85.4 ± 2.3 90.7 ± 1.6 86.7 ± 1.8 79.5 ± 3.5 76.4 ± 4.7 78.6 ± 4.5 74.3 ± ± 5.8 78.8 ± 4.6 77.7 ± 4.1 MB2C 80.5 ± 2.6 82.0 ± 2.7 88.7 ± 2.8 83.5 ± 2.5 68.2 ± 3.1 64.6 ± 4.3 68.5 ± 3.3 63.5 ± 3.8 68.5 ± 3.1 66.8 ± 3.1 UBP 86.0 ± 1.7 86.1 ± 1.8 95.0 ± 2.3 88.8 ± 1.8 69.3 ± 2.4 66.2 ± 2.8 65.3 ± 4.0 56.8 ± 3.0 68.1 ± 3.2 65.6 ± 2.1 ATS 87.6 ± 1.7 87.8 ± 1.5 94.1 ± 2.2 89.7 ± 1.7 76.3 ± 2.8 69.8 ± 2.8 72.0 ± 2.7 63.2 ± 2.2 72.4 ± 3.0 71.2 ± 2.1 NeuroBridge 88.3 ± 1.3 89.8 ± 1.3 95.7 ± 1.4 91.0 ± 1.1 79.2 ± 3.6 70.2 ± 2.3 71.6 ± 1.9 65.1 ± 2.8 73.5 ± 1.9 72.4 ± 1.9 HyFI 90.1 ± 1.3 92.0 ± 1.2 96.8 ± 1.7 92.8 ± 1.1 81.4 ± 2.4 74.5 ± 2.3 74.9 ± 2.4 69.3 ± 3.0 75.8 ± 2.7 75.7 ± 1.4 Brain-HIVE 94.2 ± ± 1.6 95.4 ± ± 1.7 97.5 ± ± 0.8 95.6 ± ± 1.3 85.0 ± ± 3.4 81.3 ± ± 4.1 81.3 ± ± 4.4 71.8 ± 4.5 80.6 ± ± 4.1 80.5 ± ± 3.6 Table 2: Fine-grained and overall 2AFC accuracy for Object Identity and Attribute edits. Attribute subtypes group material with finish, texture with pattern, shape with size, and physical-state changes under State. Best results are shown in bold. Evaluation Protocol All evaluations use the source images and edited images that pass quality control. Standard 200-way retrieval. For each EEG query, the conventional protocol ranks the corresponding source image against the other 199 THINGS-EEG2 test images. We report whether the source image is ranked first. Null-image 200-way retrieval. To provide a reference for changes introduced by image edit model, we repeat standard 200-way retrieval after replacing each source candidate with its null-edited counterpart, generated without an intended semantic change. Edit-pool retrieval. For each concept c, the candidate pool is c=IcorigâȘEc.P_c=\I_c^orig\âȘ E_c. We define edit-pool Top-1 accuracy (EP-Top1) as the proportion of concepts for which the source image is ranked above all valid edited variants derived from the same source image. Per-family 2AFC accuracy. For edit family f, two-alternative forced-choice accuracy measures how often the original image receives a higher similarity score than an edited variant: A2âAâFâC(f)=âcâIâEc(f)âsâ(ec,Icorig)>sâ(ec,I)âc|Ec(f)|.A_2AFC^(f)= _c _Iâ E_c^(f)1\! \s(e_c,I_c^orig)>s(e_c,I) \ _c|E_c^(f)|. where ââ 1\·\ denotes the indicator function. Ties are not counted as correct. Overall 2AFC accuracy is computed as a micro-average over all valid edits. EP-Top1 evaluates joint competition within the full edit pool, while 2AFC reports performance separately across edit families. 4 Experiments Experimental Setup We conduct experiments on THINGS-EEG2 (Gifford et al. 2022), using recordings from ten subjects and the standard subject-dependent split with 200 held-out test concepts. EEG signals are restricted to the first second after stimulus onset, and the 80 repetitions for each test concept are averaged into one query. We evaluate NICE, MB2C, ATM, UBP, ATS, Brain-HIVE, HyFI, and NeuroBridge using their original model configurations. Each subject-specific model is trained with five random seeds. All methods are evaluated on the same 2,137 quality-controlled edits, which are used only for final testing. We average seeds within each subject and report the mean and standard deviation across subjects. Complete preprocessing, model configurations, training details, and evaluation procedures are provided in the supplementary material. Figure 6: Source-category differences in 2AFC accuracy, averaged across the eight evaluated models. Each cell reports the category-specific 2AFC accuracy minus the all-concept average under the same edit condition, in percentage points. Positive values indicate above-average 2AFC performance within that condition, while negative values indicate below-average 2AFC performance. Figure 7: Target-concept alignment under Object Identity Edit. Panel A reports directional similarity changes for each model: source-image similarity minus edited-image similarity for the source-concept EEG on the horizontal axis, and edited-image similarity minus source-image similarity for the target-concept EEG on the vertical axis. Positive values on both axes indicate that the edited image moves away from the source concept and toward the target concept. Panel B compares the similarity scores of the source and edited images to the target-concept EEG. Error bars denote bootstrap 95% confidence intervals. Semantic Proximity of Edited Images Before evaluating EEG retrieval models, we examine how closely the edited images remain anchored to their source images in CLIP space. The colored boxplots in Figure 5 show the distance between each edited image and its corresponding source image, while the dashed line and shaded region provide a reference for the typical distance between an edited image and images from other test concepts. All distances are computed from OpenCLIP ViT-L/14 embeddings using cosine distance (Radford et al. 2021; Cherti et al. 2023). Across all four edit families, edited images generally remain substantially closer to their own source images than to the cross-concept reference. Attribute Edit and Background Edit induce the smallest shifts, reflecting their preservation of the main object and much of the original composition. Object Identity Edit and Object Removal produce larger changes because they replace or remove the primary semantic content, yet their distances still remain below the typical cross-concept level. These results show that EEG-EditBench constructs challenging candidates that preserve substantial source-image content while introducing controlled visual changes. Overall Benchmark Results Table 1 summarizes performance under original- and null-image 200-way Top-1 accuracy, EP-Top1, and overall 2AFC. The model rankings differ markedly between the standard and edit-based settings. Brain-HIVE performs best in both retrieval metrics, whereas ATM achieves only moderate 200-way accuracy but remains highly competitive in EP-Top1. By contrast, several models with strong standard retrieval performance lose much of their advantage when the candidate pool is restricted to controlled variants of the same source image. These results show that models successful against unrelated distractors can still confuse the viewed image with closely matched edited variants. EEG-EditBench makes this distinction visible by replacing cross-concept distractors with controlled variants of the viewed image. Replacing the original candidates with null-edited images lowers 200-way accuracy for most models, suggesting that regeneration changes visual representations relevant to retrieval. We examine this effect further in the supplementary material. A notable pattern appears among methods that incorporate degraded visual representations. UBP, ATS, NeuroBridge, and HyFI all perform strongly in standard 200-way retrieval, yet their advantages shrink substantially in EP-Top1, where ATM surpasses all four despite its lower standard accuracy. Brain-HIVE, which does not rely on degraded visual features, remains strong in both settings. This shared pattern identifies visual degradation as a design factor worth examining under controlled settings. The overall 2AFC results show a similar separation across models, with Brain-HIVE again achieving the highest accuracy. We next decompose 2AFC performance by edit family and subtype to characterize how model performance varies across edit families and subtypes. Edit-Specific 2AFC Analysis The edit-specific results in Table 1 reveal a consistent hierarchy across edit families. Across the eight models, Object Identity Edit achieves 9.0â23.3 percentage points higher accuracy than Attribute Edit, revealing a consistent cross-model advantage in distinguishing changes to object identity. Object Removal also achieves high accuracy, while Background Edit generally lies between Object Identity and Attribute Edit. Within EEG-EditBench, current models therefore capture changes in object identity and presence more reliably than fine-grained attribute changes. The subtype results in Table 2 further refine this picture. Within Object Identity Edit, far replacements are consistently easier to distinguish than near and medium replacements, while the ordering between near and medium varies across models. Semantically distant replacements are therefore more readily distinguished, whereas near and medium replacements remain more challenging. This pattern establishes semantic proximity as a consistent source of difficulty within identity edits. Performance also varies across attribute subtypes: shape or size edits form the most challenging subtype for nearly all models, while color edits are generally the easiest or among the easiest. The variation across attribute subtypes shows that current EEGâimage representations preserve different visual properties with different reliability. Together, these results provide a finer-grained view of model behavior than the family-level scores. Source-Category Analysis We further examine whether the same edit is equally easy to distinguish across different object categories. We group the 200 test concepts into seven broad source categories and compute category-specific 2AFC accuracy for each edit condition. Figure 6 reports the difference between each category-specific accuracy and the corresponding all-concept average. The Other category contains only three source images and is not discussed further. The results show that different categories respond differently to the same attribute edit. Animals perform relatively well on Material and Texture changes, while Clothing shows clearer advantages on Material, Color, and Shape. Attribute sensitivity therefore depends not only on which property is changed, but also on what kind of object is being edited. The same pattern appears beyond attribute edits. Food performs relatively well on Identity and Removal edits but less well on Background changes, while Vehicles remain below average across several edit conditions. Category differences are largest for near and medium identity replacements and become smaller for far replacements. Source-category differences are therefore most visible when the edited image remains close to the source, while large identity changes are easier to distinguish across categories. Target-Concept Alignment under Object Identity Edit The 2AFC results show whether a model distinguishes an object-identity edit from its source image. We further examine whether these edits induce the intended semantic direction in the EEGâimage joint space. For edits whose target concept has a corresponding EEG query, we compare the source and edited images with both the source-concept EEG and the target-concept EEG. Figure 7 shows a consistent target-alignment pattern across all evaluated models. In Panel A, identity editing reduces alignment with the source-concept EEG and increases alignment with the target-concept EEG, indicating a shift from the original concept toward the specified replacement. In Panel B, the edited image also aligns more strongly with the target-concept EEG than the source image. The consistency across models shows that this pattern is shared by different EEGâimage alignment architectures. Beyond distinguishing the source and edited images, the analysis characterizes the semantic direction of this distinction in the joint space. Together, the two panels show that object identity edits induce structured, target-directed changes in EEGâimage alignment. 5 Discussion and Conclusion The results show that standard retrieval and edit-based evaluation capture different aspects of EEGâimage matching. A model that performs well among semantically diverse candidates may still struggle to distinguish the viewed image from closely matched edits. Across the evaluated models, attribute changes are consistently more difficult than changes in object identity or presence, suggesting that current systems preserve coarse semantic information more reliably than fine-grained visual details. EEG-EditBench is limited by its image-side design. The edited images do not have corresponding EEG recordings, so the benchmark evaluates how a complete EEGâimage retrieval system responds to controlled candidate changes rather than how the brain responds to the edited stimuli themselves. In addition, each test concept is represented by a single source image. Future work can address these limitations by introducing greater within-concept diversity, using additional editing models, collecting neural responses to selected edits, and extending the framework to modalities such as fMRI and MEG. Overall, EEG-EditBench complements standard retrieval with a factor-level view of model behavior under controlled visual changes. The edited variants may also serve as structured hard negatives for developing EEGâimage models that retain more precise and interpretable visual information. References Y. Bai, X. Wang, Y. Cao, Y. Ge, C. Yuan, and Y. Shan (2024) DreamDiffusion: high-quality EEG-to-image generation with temporal masked signal modeling and CLIP alignment. In Computer Vision â ECCV 2024, p. 472â488. External Links: Document Cited by: §2. Black Forest Labs (2026) FLUX.2 [klein] 9B. Note: Hugging Face model card External Links: Link Cited by: §S2.2. M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev (2023) Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 2818â2829. Cited by: §S6.1, §4. R. M. Cichy, A. Khosla, D. Pantazis, A. Torralba, and A. Oliva (2016) Comparison of deep neural networks to spatio-temporal cortical dynamics of human visual object recognition reveals hierarchical correspondence. Scientific Reports 6, p. 27755. External Links: Document Cited by: §1. C. Du, K. Fu, J. Li, and H. He (2023) Decoding visual neural representations by multimodal learning of brain-visual-linguistic features. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (9), p. 10760â10777. External Links: Document Cited by: §2. A. T. Gifford, K. Dwivedi, G. Roig, and R. M. Cichy (2022) A large and rich EEG dataset for modeling human visual object recognition. NeuroImage 264, p. 119754. External Links: Document Cited by: §S2.1, §S3.1, §1, §2, §3, §4. T. Grootswagers, I. Zhou, A. K. Robinson, M. N. Hebart, and T. A. Carlson (2022) Human EEG recordings for 1,854 concepts presented in rapid serial visual presentation streams. Scientific Data 9 (1), p. 3. External Links: Document Cited by: §1, §2. M. Guggenmos, P. Sterzer, and R. M. Cichy (2018) Multivariate pattern analysis for MEG: a comparison of dissimilarity measures. NeuroImage 173, p. 434â447. External Links: Document Cited by: §S3.1. M. N. Hebart, A. H. Dickter, A. Kidder, W. Y. Kwok, A. Corriveau, C. Van Wicklin, and C. I. Baker (2019) THINGS: a database of 1,854 object concepts and more than 26,000 naturalistic object images. PLOS ONE 14 (10), p. e0223792. External Links: Document Cited by: §3. C. Hsieh, J. Zhang, Z. Ma, A. Kembhavi, and R. Krishna (2023) SugarCrepe: fixing hackable benchmarks for vision-language compositionality. In Advances in Neural Information Processing Systems, Vol. 36, p. 31096â31116. Cited by: §1, §2. M. Hui, S. Yang, B. Zhao, Y. Shi, H. Wang, P. Wang, Y. Zhou, and C. Xie (2025) HQ-Edit: a high-quality dataset for instruction-based image editing. In The Thirteenth International Conference on Learning Representations, Cited by: §2. S. Jo, W. Jeong, D. Heo, Y. Hwang, and H. Suk (2026) HyFI: hyperbolic feature interpolation for brain-vision alignment. Proceedings of the AAAI Conference on Artificial Intelligence 40 (7), p. 5575â5583. External Links: Document Cited by: §S3.2, §1, §1, §2. J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick (2017) CLEVR: a diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 2901â2910. External Links: Document Cited by: §2. X. Ju, A. Zeng, Y. Bian, S. Liu, and Q. Xu (2024) PnP inversion: boosting diffusion-based editing with 3 lines of code. In The Twelfth International Conference on Learning Representations, Cited by: §2. Y. Kamitani and F. Tong (2005) Decoding the visual and subjective contents of the human brain. Nature Neuroscience 8 (5), p. 679â685. External Links: Document Cited by: §1. B. Kaneshiro, M. Perreau Guimaraes, H. Kim, A. M. Norcia, and P. Suppes (2015) A representational similarity analysis of the dynamics of object processing using single-trial EEG classification. PLOS ONE 10 (8), p. e0135697. External Links: Document Cited by: §2. K. N. Kay, T. Naselaris, R. J. Prenger, and J. L. Gallant (2008) Identifying natural images from human brain activity. Nature 452 (7185), p. 352â355. External Links: Document Cited by: §1. R. Kneeland, P. S. Scotti, G. St-Yves, J. Breedlove, K. Kay, and T. Naselaris (2025) NSD-Imagery: a benchmark dataset for extending fMRI vision decoding methods to mental imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 28852â28862. Cited by: §2. D. Li, C. Wei, S. Li, J. Zou, H. Qin, and Q. Liu (2024) Visual decoding and reconstruction via EEG embeddings with guided diffusion. In Advances in Neural Information Processing Systems, Vol. 37. Cited by: §S3.2, §1, §1, §2. W. Liu, H. Li, Z. Xu, L. Ma, and H. Li (2026) Leveraging visual blur perception characteristics for EEG decoding. Proceedings of the AAAI Conference on Artificial Intelligence 40 (21), p. 17580â17588. External Links: Document Cited by: §1. Y. Ma, J. Ji, K. Ye, W. Lin, Z. Wang, Y. Zheng, Q. Zhou, X. Sun, and R. Ji (2024) I2EBench: a comprehensive benchmark for instruction-based image editing. In Advances in Neural Information Processing Systems, Vol. 37, p. 41494â41516. Cited by: §1, §2. Z. Ma, J. Hong, M. O. Gul, M. Gandhi, I. Gao, and R. Krishna (2023) CREPE: can vision-language foundation models reason compositionally?. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 10910â10921. Cited by: §2. T. Melistas, N. Spyrou, N. Gkouti, P. Sanchez, A. Vlontzos, Y. Panagakis, G. Papanastasiou, and S. A. Tsaftaris (2024) Benchmarking counterfactual image generation. In The Thirty-Eighth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, Cited by: §2. T. Naselaris, R. J. Prenger, K. N. Kay, M. Oliver, and J. L. Gallant (2009) Bayesian reconstruction of natural images from human brain activity. Neuron 63 (6), p. 902â915. External Links: Document Cited by: §1. L. Parcalabescu, M. Cafagna, L. Muradjan, A. Frank, I. Calixto, and A. Gatt (2022) VALSE: a task-independent benchmark for vision and language models centered on linguistic phenomena. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 8253â8280. External Links: Document Cited by: §2. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021) Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 139, p. 8748â8763. Cited by: §S6.1, §4. M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh (2020) Beyond accuracy: behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 4902â4912. External Links: Document Cited by: §1, §2. R. Shekhar, S. Pezzelle, Y. Klimovich, A. Herbelot, M. Nabi, E. Sangineto, and R. Bernardi (2017) FOIL it! find one mismatch between image and language caption. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 255â265. External Links: Document Cited by: §2. P. Singh, D. Dalal, G. Vashishtha, K. Miyapuram, and S. Raman (2024) Learning robust deep visual representations from EEG brain recordings. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 7553â7562. External Links: Document Cited by: §2. Y. Song, B. Liu, X. Li, N. Shi, Y. Wang, and X. Gao (2024) Decoding natural images from EEG for object recognition. In The Twelfth International Conference on Learning Representations, Cited by: §S3.2, §1, §1, §2. T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross (2022) Winoground: probing vision and language models for visio-linguistic compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 5238â5248. External Links: Document Cited by: §1, §2. S. Wang, C. Saharia, C. Montgomery, J. Pont-Tuset, S. Noy, S. Pellegrini, Y. Onoe, S. Laszlo, D. J. Fleet, R. Soricut, J. Baldridge, M. Norouzi, P. Anderson, and W. Chan (2023) Imagen editor and editbench: advancing and evaluating text-guided image inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 18359â18369. Cited by: §2. W. Wang et al. (2025) InternVL3.5: advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265. Cited by: §S2.2. Y. Wei, L. Cao, H. Li, and Y. Dong (2024) MB2C: multimodal bidirectional cycle consistency for learning robust visual neural representations. In Proceedings of the 32nd ACM International Conference on Multimedia, p. 8992â9000. External Links: Document Cited by: §S3.2, §1, §1, §2. H. Wu, Q. Li, C. Zhang, Z. He, and X. Ying (2025) Bridging the vision-brain gap with an uncertainty-aware blur prior. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 2246â2257. Cited by: §S3.2, §1, §1, §2. L. Wu, J. Li, Z. Ren, K. Zhang, and X. Gao (2026) Shrinking the teacher: an adaptive teaching paradigm for asymmetric EEG-vision alignment. Proceedings of the AAAI Conference on Artificial Intelligence 40 (21), p. 17859â17867. External Links: Document Cited by: §S3.2, §1, §2. W. Xia and C. Oztireli (2025) Exploring the visual feature space for multimodal neural decoding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 4370â4379. Cited by: §2. W. Xia and C. Oztireli (2026) Multigranular evaluation for brain visual decoding. In Proceedings of the AAAI Conference on Artificial Intelligence, p. 17868â17876. External Links: Document Cited by: §2. Q. Yu, W. Chow, Z. Yue, K. Pan, Y. Wu, X. Wan, J. Li, S. Tang, H. Zhang, and Y. Zhuang (2025) AnyEdit: mastering unified high-quality image editing for any idea. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 26125â26135. Cited by: §2. M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou (2023) When and why vision-language models behave like bags-of-words, and what to do about it?. In The Eleventh International Conference on Learning Representations, Cited by: §2. K. Zhang, L. Mo, W. Chen, H. Sun, and Y. Su (2023) MagicBrush: a manually annotated dataset for instruction-guided image editing. In Advances in Neural Information Processing Systems, Vol. 36, p. 31428â31449. Cited by: §1, §2. K. Zhang, L. He, X. Jiang, W. Lu, D. Wang, and X. Gao (2025) Cognitioncapturer: decoding visual stimuli from human eeg signal with multimodal information. In Proceedings of the AAAI Conference on Artificial Intelligence, p. 14486â14493. Cited by: §2. K. Zhang, L. He, J. Ke, Y. Ji, L. Wu, L. Wang, and X. Gao (2026a) CognitionCapturerPro: towards high-fidelity visual decoding from eeg/meg via multi-modal information and asymmetric alignment. arXiv preprint arXiv:2603.12722. Cited by: §2. W. Zhang, S. Wang, Y. Su, X. Li, C. Zhang, and S. Zhong (2026b) NeuroBridge: bio-inspired self-supervised EEG-to-image decoding via cognitive priors and bidirectional semantic alignment. Proceedings of the AAAI Conference on Artificial Intelligence 40 (21), p. 18028â18036. External Links: Document Cited by: §S3.2, §1, §2. T. Zhao, T. Zhang, M. Zhu, H. Shen, K. Lee, X. Lu, and J. Yin (2022) VL-CheckList: evaluating pre-trained vision-language models with objects, attributes and relations. External Links: 2207.00221 Cited by: §1, §2. J. Zheng, H. Jia, M. Li, Y. Zheng, Y. Zeng, Y. Gao, and C. Liang (2026) Learning brain representation with hierarchical visual embeddings. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §S3.2, §1, §1, §2. Supplementary Material Appendix S1 Supplementary Overview This supplementary material provides the construction, reproduction, and analysis details supporting the main paper. We first describe the benchmark taxonomy, generation pipeline, and quality-control procedure, followed by the preprocessing, model-reproduction, and evaluation protocols. We then report a sensitivity analysis for 2AFC aggregation, validity controls for EEG pairing and image-generation effects, and additional diagnostic results on visual edit magnitude, source categories, target-concept transfer, and representative failures. Appendix S2 Benchmark Construction and Quality Control This section describes the benchmark scope and edit taxonomy, the shared construction pipeline, and the human quality-control procedure used to obtain the final dataset. S2.1 Benchmark Scope and Edit Taxonomy EEG-EditBench is constructed from the 200 source images in the THINGS-EEG2 test set (Gifford et al. 2022), with one image representing each test concept. For every source image, the benchmark introduces controlled candidates that change the identity of the main object, one visible attribute, the surrounding background, or the presence of the object itself. During evaluation, the EEG query elicited by the source image is compared with the source and its edited variants. The quality-controlled edited images are used exclusively for final evaluation. Before target generation, each source image receives an independent feasibility assessment for the four edit families. Each family is assigned accept, borderline, or reject. An accept decision proceeds to standard target generation, borderline admits conservative targets that remain reliable under the identified image-specific limitation, and reject closes the corresponding generation branch. This family-specific assessment, together with the availability of suitable targets, determines the final source coverage. Table S1 summarizes the controlled change and preserved content for each edit family. Edit family Controlled change Preserved content Object Identity Main object identity Composition, viewpoint, background, visual style, and spatial role Attribute One visible property of the main object Object identity, background, and non-target content Background Surrounding scene Foreground object, appearance, position, and interaction structure Object Removal Complete main subject and reconstruction of the exposed region Remaining objects and scene context Table S1: Operational definition of the four edit families. Attribute Edit contains five visible-property groups. Color covers changes in hue, tone, brightness, or saturation; Material/Finish covers perceived substance, reflectance, and surface finish; Texture/Pattern covers markings, grain, print, and other surface structure; Shape/Size covers visible geometric or scale changes that preserve the recognizable object category and spatial role; and State/Condition covers physical states such as age, ripeness, cleanliness, damage, openness, fullness, or wetness. For Object Identity Edit, replacement targets are drawn from the remaining THINGS-EEG2 test concepts and assigned to three qualitative, scene-conditioned levels. Near targets belong to a similar broad object family or share a closely related function, affordance, or usage context. Medium targets come from a different object family but can occupy a similar role in the scene, interaction, or spatial arrangement. Far targets are semantically distinct concrete objects that remain plausible replacements in the source composition. These assignments jointly consider semantic relatedness and scene compatibility, including approximate scale, spatial role, support relations, and interactions. Table S2 reports the resulting subtype counts and source-concept coverage. Edit family Subtype Edits Source concepts Object Identity Near 266 198 Medium 280 Far 234 Attribute Color 231 200 Material/Finish 199 Texture/Pattern 198 Shape/Size 156 State/Condition 195 Background â 199 199 Object Removal â 179 179 Overall â 2,137 200 Table S2: Final benchmark composition by edit family and subtype. Source-concept counts report coverage within each family; the complete benchmark covers all 200 test concepts. S2.2 Construction Pipeline The construction figure in the main paper summarizes the four-stage workflow used for all source images. Table S3 provides the input, output, and model used at each stage. The overall workflow is shared across edit families, with family-specific branches for target generation and prompt construction. The complete image-editing settings are reported in Table S4. Stage Input Output Model Visual Scene Profiling Source image and concept label Scene profile and editability InternVL3.5-38B, multimodal Edit Target Generation Scene profile and candidate concepts Family-specific edit targets InternVL3.5-38B, text-only Edit Prompt Generation Scene profile and one target Self-contained editing prompt InternVL3.5-38B, text-only Image Editing Source image and prompt One edited image FLUX.2-klein-9B Table S3: Four-stage benchmark construction pipeline, showing the input, output, and model used at each stage. Edit Target Generation uses four parallel branches. The Attribute branch proposes visible changes to color, material or finish, texture or pattern, shape or size, and state or condition. The Background branch selects a new scene compatible with the foreground subject. The Object Identity branch selects scene-compatible near, medium, and far targets from the remaining test concepts. The Object Removal branch specifies the complete subject to remove and how the exposed region should be reconstructed. Each structured target is then converted into a self-contained editing instruction that states the requested change and the relevant content to preserve, including composition, viewpoint, object placement, visible non-target attributes, and scene relationships. The first three stages use the local OpenGVLab/InternVL3_5-38B checkpoint (Wang and others 2025) with deterministic decoding (do_sample=False). Each response is validated against a stage-specific schema before being written to an incremental JSONL record. Formatting or constraint violations trigger regeneration with corrective feedback for up to two retries, and completed records are retained when a run is resumed. The final stage uses black-forest-labs/FLUX.2-klein-9B (Black Forest Labs 2026) through diffusers.Flux2KleinPipeline. All four edit families use the same source-image-plus-prompt interface and inference settings. Table S4 reports the complete configuration. Setting Value Model FLUX.2-klein-9B Resolution 224 Ă 224 Inference steps 4 Guidance scale 1.0 Generation seed 0, reset for each edit Maximum sequence length 512 Text-encoder output layers 9, 18, and 27 Precision bfloat16 or float16 CPU offload Enabled Table S4: Image-editing inference settings. The shared pipeline produced 2,210 candidate edits, which then entered the quality-control procedure. S2.3 Quality Control and Final Benchmark Composition All 2,210 generated edits underwent full human review. Two reviewers independently evaluated every candidate using side-by-side comparisons of the source and edited images together with the corresponding edit instruction. Reviewers assessed whether the intended edit was visibly fulfilled, whether relevant non-target content was preserved, and whether the edited image contained obvious artifacts or implausible visual relations. Each reviewer assigned a binary pass or fail decision without access to the other reviewerâs judgment. Cases receiving different decisions from the two reviewers were independently assessed by a third reviewer, whose judgment determined the final outcome. Only edits receiving a final pass decision were retained. This process retained 2,137 of the 2,210 generated edits. The final benchmark contains 979 Attribute Edits, 780 Object Identity Edits, 199 Background Edits, and 179 Object Removals. All 200 source concepts remain represented. Appendix S3 Reproduction and Evaluation Details This section documents the dataset and EEG preprocessing, the reproduction settings of the eight evaluated models, and the common evaluation and implementation protocol. S3.1 Dataset and EEG Preprocessing We use THINGS-EEG2 (Gifford et al. 2022), which contains EEG recordings from ten subjects acquired with a BrainVision actiCHamp system during a 5 Hz rapid serial visual presentation paradigm. EEG was originally recorded at 1,000 Hz from 63 channels, and each subject completed four recording sessions. All evaluated models use a subject-dependent setting, with a separate model trained and evaluated for each subject. Table S5 summarizes the training and test partitions. Partition Concepts Images per concept Images Repetitions per image Train 1,654 10 16,540 4 Test 200 1 200 80 Table S5: THINGS-EEG2 partitions used in this work. Each preprocessed EEG trial contains 63 channels and 250 time samples. The training and test concepts are disjoint, and the single image associated with each test concept serves as an EEG-EditBench source image. For preprocessing, trials were first arranged in a fixed 63-channel order, and target and catch trials were removed. EEG was epoched from 0.2 s before to 1.0 s after stimulus onset and baseline-corrected using the prestimulus interval. The signals were then downsampled from 1,000 to 250 Hz, and the 0â1 s post-stimulus interval was retained, yielding 250 samples per channel. Trials were organized by image condition and recording session. For each subject, multivariate noise normalization (MVNN) (Guggenmos et al. 2018) was estimated using only the training partition, and the resulting whitening transform was applied to both training and test EEG. The preprocessed representation of one trial therefore has shape 63Ă25063Ă 250. NICE, MB2C, and ATM use all 63 channels. UBP, ATS, Brain-HIVE, HyFI, and NeuroBridge use the following fixed set of 17 occipital and parietal channels: P7, P5, P3, P1, Pz, P2, P4, P6, P8, PO7, PO3, POz, PO4, PO8, O1, Oz, and O2. All methods use the complete 0â1 s post-stimulus window. At test time, every method averages the 80 repetitions for each image before producing one EEG query embedding. During training, ATM treats the four repetitions of each training image as separate samples, whereas the other seven methods average the four repetitions before model input. S3.2 Evaluated Models and Training Details We evaluate NICE (Song et al. 2024), MB2C (Wei et al. 2024), ATM (Li et al. 2024), UBP (Wu et al. 2025), ATS (Wu et al. 2026), Brain-HIVE (Zheng et al. 2026), HyFI (Jo et al. 2026), and NeuroBridge (Zhang et al. 2026b) in a subject-dependent setting. Each method is trained independently for ten subjects and five random seeds (0â4), yielding 400 subjectâmethodâseed runs. Each reproduction retains its specified EEG encoder, frozen visual representation, and training objective. The resulting EEG and image embeddings are exported to a common evaluator. EEG-EditBench edited images are used only in the final evaluation and do not contribute to training or checkpoint selection. Table S6 summarizes the model-specific representations and objectives. Method EEG encoder Visual target Objective Dim. NICE Temporalâspatial CNN CLIP ViT-L/14 Symmetric contrastive 768 MB2C NICE-style CNN CLIP ViT-L/14 Contrastive, cycle, and adversarial 768 ATM Transformer + ShallowNet OpenCLIP ViT-H/14 image and text Imageâtext alignment 1,024 UBP EEGProjectLayer RN50 multi-level blur Uncertainty-aware contrastive 1,024 ATS EEGProjectLayer RN50 adaptive blur Adaptive-teaching contrastive 768 Brain-HIVE Brain MLP SynCLR, CLIP, and SDXL-VAE Hierarchical visual alignment 1,024 HyFI EEGProjectLayer RN50 semantic and perceptual blur Hyperbolic contrastive 1,024 NeuroBridge EEGProject Augmented-fused RN50 Symmetric contrastive 512 Table S6: Overview of the eight reproduced EEGâimage models. Dim. denotes the shared EEG and image embedding dimension. Table S7 reports the principal optimization settings. Our reproductions are based on the protocols described in the original works and official implementations. Method Epochs Batch Optimization NICE 200 1,000 Adam, 2Ă10â42Ă 10^-4 MB2C †1,000 2,048 Adam / RMSprop ATM 80 1,024 AdamW, 3Ă10â43Ă 10^-4 UBP 50 1,024 AdamW, 1Ă10â41Ă 10^-4 ATS 150 1,024 AdamW, 1Ă10â41Ă 10^-4; StepLR Brain-HIVE 25 1,024 AdamW, 5Ă10â45Ă 10^-4; cosine HyFI †50 1,024 AdamW, 3Ă10â43Ă 10^-4 NeuroBridge 50 1,024 AdamW, 1Ă10â41Ă 10^-4 Table S7: Principal training settings for the reproduced models. Batch denotes the effective batch size of one subject-specific run. CLIP-based image and imageâtext alignment. NICE randomly divides its training data into a 740-example validation subset and the remaining training examples. MB2C augments symmetric InfoNCE with bidirectional cycle consistency, bidirectional adversarial learning, auxiliary classification, and mixup. Its principal settings are a cycle weight of 10, a mixup ratio of 0.75, a gradient-penalty weight of 10, and an auxiliary-classification weight of 1. ATM combines image and text alignment losses with weights 0.99 and 0.01, respectively, using text targets generated from the template This picture is concept. Blur-based visual supervision. UBP constructs three FoveaBlur representations with kernel sizes 45, 51, and 57 and dynamically selects among them according to the positive EEGâimage similarity. Both source and edited images use the medium, kernel-51 representation during evaluation. ATS follows the same three-level adaptive-teaching principle and uses a ShrinkAdapter with a bottleneck ratio of 0.25 to map the 1,024-dimensional RN50 features to 768 dimensions. HyFI combines a semantic FoveaBlur branch with kernel size 51 and a perceptual blur branch with kernel size 31, maps EEG and both visual branches to Lorentz space, and learns a geodesic interpolation of the semantic and perceptual targets. Hierarchical and augmented visual supervision. Brain-HIVE uses 768-dimensional SynCLR ViT-B/16 features, 512-dimensional CLIP ViT-B/32 features, and 1,024-dimensional SDXL-VAE latents. These branches are projected into a shared 1,024-dimensional space and fused during Stage-2 training, which uses bfloat16 and a cosine learning-rate schedule without warmup. NeuroBridge constructs its visual target by fusing deterministic OpenCLIP RN50 features from Gaussian blur, Gaussian noise, low-resolution, and mosaic transformations. The formal configuration uses this visual branch without an additional text target. During training, random smooth EEG augmentation with kernel size 5 is applied with probability 0.3. S3.3 Evaluation Protocol and Implementation Evaluation scores. For every subjectâmethodâseed run, the evaluator receives one EEG query representation for each of the 200 test concepts, together with image representations for the 200 source images and all 2,137 retained edits. Following the original evaluation protocol of each method, we use its native EEGâimage score for retrieval and pairwise comparison. All reported metrics are computed within each model run. Standard 200-way retrieval. For each EEG query, the standard protocol ranks all 200 source images by the corresponding method-native score. Top-1 and Top-5 accuracy measure whether the corresponding source image appears among the first one or five candidates. Their random-ranking references are 1/200=0.5%1/200=0.5\% and 5/200=2.5%5/200=2.5\%, respectively. Edit-pool retrieval. For each test concept c, the complete candidate pool contains its source image and all retained edits derived from that source; let cP_c denote this pool. Edit-pool Top-1 accuracy (EP-Top1) requires the source image to rank above every same-source edit: AEP=1200ââc=1200â[sâ(ec,vcsource)>maxiâĄsâ(ec,vc,iedit)].A_EP= 1200 _c=1^2001\! [s(e_c,v_c^source)> _is(e_c,v_c,i^edit) ]. (S1) Each concept contributes equally. The 200 pools contain 5â14 edited candidates and one source image, giving total pool sizes of 6â15 and a mean size of 11.685. Because pool sizes vary, the random-ranking reference is computed concept by concept: AEPchance=1200ââc=12001|c|=0.08722â8.72%.A_EP^chance= 1200 _c=1^200 1|P_c|=0.08722â 8.72\%. (S2) Per-edit 2AFC accuracy. Each retained edit defines a two-alternative comparison between that edit and its corresponding source image. For edit family f, A2âAâFâC(f)=âcâiâEc(f)â[sâ(ec,vcsource)>sâ(ec,vc,iedit)]âc|Ec(f)|.A_2AFC^(f)= _c _iâ E_c^(f)1\! [s(e_c,v_c^source)>s(e_c,v_c,i^edit) ] _c|E_c^(f)|. (S3) The source must receive a strictly higher similarity; a tie is not counted as correct. Overall 2AFC is the micro-average over all 2,137 edits. The same definition is applied separately to the 979 Attribute, 780 Object Identity, 199 Background, and 179 Object Removal edits. The random-ranking reference is 50%. Table S8 summarizes the three protocols. Protocol Candidates Chance Standard Top-1/5 200 sources 0.5% / 2.5% EP-Top1 Source + edits 8.72% 2AFC Source vs. one edit 50% Table S8: Summary of the three evaluation protocols and their random-ranking references. EP-Top1 compares each source with all same-source edits. The protocols contain 200 concept-level queries, 200 concept-level pools, and 2,137 edit-level pairs, respectively. Aggregation across seeds and subjects. Let ms,rm_s,r denote a metric from subject s and training seed r. We first average the five seeds within each subject, mÂŻs=15ââr=04ms,r, m_s= 15 _r=0^4m_s,r, (S4) and then report the mean and population standard deviation across the ten subject-level values: ÎŒ=110ââs=110mÂŻs,Ï=110ââs=110(mÂŻsâÎŒ)2.ÎŒ= 110 _s=1^10 m_s, Ï= 110 _s=1^10( m_s-ÎŒ)^2. (S5) The reported uncertainty thus reflects variation across subjects after seed averaging, with each subject serving as the unit of analysis. Implementation environment. Table S9 lists the principal software and hardware used for model reproduction and evaluation. Formal runners pass seeds explicitly to Python, NumPy, and PyTorch and record the method, subject, and seed in each runâs metadata. Component Version or configuration Python 3.9.23 PyTorch 2.5.0, CUDA 11.8 build NumPy 1.26.4 SciPy 1.13.1 scikit-learn 1.6.1 PyTorch Lightning 2.6.0 Transformers 4.57.6 OpenCLIP 3.2.0 MNE 1.8.0 GPU NVIDIA RTX 3090, 24 GB Table S9: Principal implementation environment. Computing and caching. The server provides eight RTX 3090 GPUs. Each MB2C run uses two GPUs with data parallelism, whereas each run of the other seven models uses one GPU. Only subject-independent features from frozen image encoders are cached; EEG models and EEG query embeddings are never reused across runs. Brain-HIVE visual caches are isolated by seed because its VAE branch includes stochastic sampling. Before reuse, cache metadata verifies the encoder checkpoint, preprocessing configuration, manifest, input order, file identities, and feature shapes. Figure S1: 2AFC accuracy across sourceâedit visual-distance quintiles, from the smallest changes in Q1 to the largest in Q5. The OpenCLIP cosine-distance intervals are 0.013â0.099, 0.099â0.149, 0.149â0.236, 0.236â0.310, and 0.310â0.575; edit counts are shown below the columns. Figure S2: Absolute source-category 2AFC accuracy across the four edit families, five Attribute subtypes, and three Object Identity distance levels. Rows denote the seven source categories. Figure S3: Representative stable failure cases across the four edit families. Columns denote edit families, and the upper and lower panels show two cases with different source images. Each panel reports the number of models satisfying the stable-failure criterion. Figure S4: Mean rank of the specified target-concept EEG among all 200 EEG concepts for source and edited images. Lower rank indicates stronger target-concept alignment. Appendix S4 Evaluation Sensitivity We examine whether the main 2AFC patterns depend on the weighting assigned to concepts and edit families. S4.1 Alternative 2AFC Aggregation The overall 2AFC result in the main paper gives equal weight to every retained edit, so concepts and edit families with more examples contribute more to the final score. We compare this edit-balanced result with a concept-and-family-balanced alternative that gives equal total weight to each edit family and equal weight to source concepts within each family, using the same 2,137 edits and strict source-over-edit decision rule. Table S10 reports the comparison. The balanced aggregation produces modest changes in absolute accuracy while preserving the main model and edit-family patterns. The maximum model-rank change is one position. Appendix S5 Validity Controls We examine two factors that directly affect the interpretation of edit-based evaluation: dependence on the correct EEGâimage correspondence and sensitivity to characteristics introduced by image generation. S5.1 Dependence on EEGâImage Pairing We first examine how strongly sourceâedit discrimination depends on the correct EEGâimage correspondence. We compare each sourceâedit pair under the correctly matched EEG query and under all 199 non-matching EEG queries. The image candidates remain fixed, and the same mismatch schedule is used for every model, subject, and training run. Correctly paired and mismatched-query 2AFC use the same strict source-over-edit decision rule. We define the pairing gain as Îpair=AcorrectâAmismatch. _pair=A_correct-A_mismatch. Metrics are averaged across seeds within subject, and confidence intervals are obtained by bootstrapping the ten subject-level values. Table S11 reports the overall comparison. Correct pairing improves 2AFC for all eight models, with gains of 13.89â35.14 percentage points. The confidence intervals are above zero in every case, showing that sourceâedit discrimination consistently benefits from the correct EEGâimage correspondence. Table S13 provides the corresponding breakdown by edit family. The family-level results show the same general pattern, with positive gains in nearly all modelâfamily combinations. NICE on Attribute edits is the only case whose confidence interval includes zero. Mismatched-query performance also varies across systems, indicating that model-specific image preferences coexist with the correspondence-dependent effect. S5.2 Sensitivity to Image-Generation Effects We use one null-edited image for each test concept. The null-edit instruction asks the editor to reproduce the source image while preserving its depicted objects, visible attributes, composition, and background, without introducing an intended semantic change. Null edits use the same image-editing model, resolution, inference settings, and generation seed as the semantic edits, and are reviewed side by side with their source images to exclude obvious semantic changes or generation failures. We compare each original image with its null-edited counterpart, and then compare the null-edited image with each semantic edit in EEG-EditBench. The first comparison provides a reference for changes associated with regeneration. In the second, both candidates have passed through the editing pipeline, and we evaluate them with correctly paired and mismatched EEG queries. Pairing gain is the difference between these two EEG conditions. All comparisons use the same edit-level weighting over the 2,137 retained semantic edits, and confidence intervals are computed over subjects. Table S12 reports the results. Original-over-null 2AFC is above the 50% reference for seven models, while NICE remains close to balance, showing that most systems are sensitive to the image-side changes introduced by regeneration. When both candidates are generated images, correctly paired EEG outperforms mismatched EEG for every model, with pairing gains of 16.23â29.29 percentage points. Appendix S6 Additional Diagnostic Analyses We further characterize benchmark behavior through visual edit magnitude, source-category patterns, target-concept alignment, and representative stable failures. S6.1 Evaluation Behavior across Visual Edit Magnitudes We next examine how visual edit magnitude shapes 2AFC difficulty and edit-family differences. 2AFC across visual distances. We measure edit magnitude using cosine distance between the source and edited images in a common OpenCLIP ViT-L/14 feature space (Radford et al. 2021; Cherti et al. 2023) that is independent of the eight evaluated EEGâimage systems. For edit i, di=1âisourceâ ieditâ„isourceâ„2ââ„ieditâ„2.d_i=1- v^source_i·v^edit_i ^source_i _2 ^edit_i _2. (S6) The 2,137 edits are divided into five shared distance quintiles, which are reused for every model, subject, and training seed. Figure S1 reports 2AFC accuracy in each quintile. All eight models show non-decreasing 2AFC accuracy from Q1 to Q5, with differences of 13.89â32.15 percentage points between the two extremes. We also correlate visual distance with the source-minus-edit similarity margin within each run. The model-mean Spearman correlations are positive for all eight systems and range from 0.417 to 0.608. Larger visual changes are therefore consistently associated with easier sourceâedit discrimination. Distance-balanced Attribute comparisons. We next test whether the lower 2AFC accuracy of Attribute edits is primarily associated with their visual-distance distribution. Each comparison is restricted to source concepts represented in both families. Distances are divided into intervals of width 0.05, retaining intervals with at least 20 edits from each family, and the two families receive equal total weight within every retained interval. The unbalanced and distance-balanced estimates use the same eligible edits and differ only in weighting. Table S14 reports comparisons of Attribute with Background, Object Removal, and Object Identity. Distance balancing reduces several absolute differences while preserving the overall ordering across edit families. Across all eight systems, Attribute Edit remains the most challenging, followed by Object Removal and Object Identity. S6.2 Absolute Source-Category Results The main paper reports source-category performance relative to the all-concept average under each edit condition. Figure S2 provides the corresponding absolute 2AFC accuracies. Results are averaged across training seeds within subject, then across subjects within model, and finally across the eight evaluated models. VehiclesâTexture is the most difficult displayed combination at 56.9%, whereas far Object Identity edits reach 91.5%â97.9% across source categories. Category differences remain visible for near and medium replacements but narrow once the identity change is far from the source concept. The Other category contains only three source images and is included for completeness rather than interpreted individually. S6.3 Target-Concept Rank Shift under Object Identity Edits The target-concept analysis in the main paper measures whether Object Identity edits move image representations away from the source concept and toward the specified target concept. We complement this analysis by ranking the specified target-concept EEG among all 200 EEG concepts for both the source and edited images. Lower rank indicates stronger target-concept alignment. Figure S4 reports the results across the complete Object Identity Edit set. The specified target concept moves closer to the front of the complete ranking after editing for all eight models. This candidate-set view complements the similarity analysis in the main paper by showing that Object Identity edits consistently improve the relative position of their intended target concepts. S6.4 Representative Stable Failure Cases We finally examine edits that repeatedly cause one or more systems to prefer the edited image over the viewed source image. For each modelâedit pair, a stable failure requires at least 26 failures among the 50 subjectâseed comparisons, with at least three seeds each producing failures for at least six of the ten subjects. Within each edit family, qualifying edits are ranked by cross-model coverage and total failures, and two cases with different source images are selected. Figure S3 shows the resulting examples. Stable failures occur in all four edit families. The two selected Attribute cases affect all eight models and show the broadest cross-model coverage, while the remaining examples demonstrate persistent failures under changes to scene context, object presence, and object identity. These cases provide concrete illustrations of the model behavior summarized by the benchmark-level results. Appendix S7 Generative AI Use Disclosure Generative AI tools were used to assist with language editing and drafting during manuscript preparation. The authors reviewed and revised all AI-assisted content and take full responsibility for the manuscript, including its claims, analyses, figures, and references. Model Edit-balanced Concept-and-family-balanced NICE 67.7 ± 2.3 71.4 ± 2.7 ATM 81.7 ± 2.7 82.0 ± 2.3 MB2C 75.3 ± 2.5 77.5 ± 2.7 UBP 77.3 ± 1.8 80.6 ± 1.9 ATS 81.1 ± 1.6 84.9 ± 1.6 NeuroBridge 82.6 ± 1.1 86.8 ± 1.0 HyFI 84.9 ± 1.1 88.6 ± 1.1 Brain-HIVE 88.6 ± 2.2 91.6 ± 1.8 Table S10: Overall 2AFC accuracy under the formal edit-balanced aggregation and a concept-and-family-balanced alternative. Values are percentages reported as mean ± population standard deviation across subjects after within-subject aggregation. The best result in each column is shown in bold. Model Correctly paired 2AFC Mismatched-query 2AFC Pairing gain NICE 67.73 [66.34, 69.21] 53.84 [52.07, 55.70] 13.89 [13.19, 14.50] ATM 81.71 [80.05, 83.38] 59.82 [57.57, 62.40] 21.89 [20.56, 23.24] MB2C 75.28 [73.69, 76.79] 48.52 [47.20, 49.86] 26.75 [25.80, 27.74] UBP 77.31 [76.15, 78.42] 49.92 [49.37, 50.45] 27.39 [26.48, 28.31] ATS 81.15 [80.16, 82.20] 46.51 [45.99, 47.08] 34.64 [33.63, 35.70] NeuroBridge 82.59 [81.90, 83.26] 49.95 [49.65, 50.25] 32.64 [32.10, 33.20] HyFI 84.94 [84.27, 85.59] 51.60 [51.27, 51.94] 33.34 [32.60, 34.02] Brain-HIVE 88.57 [87.19, 89.92] 53.43 [52.55, 54.21] 35.14 [34.21, 36.10] Table S11: Overall 2AFC under correctly paired and mismatched EEG queries. Pairing gain is their difference in percentage points. Brackets report subject-level 95% bootstrap confidence intervals. Model Original over null Correct EEG Null over semantic edit Correct EEG Null over semantic edit Mismatched EEG Pairing gain NICE 49.83 [47.25, 52.50] 67.57 [66.80, 68.37] 50.12 [49.30, 50.97] 17.45 [16.87, 18.04] ATM 67.99 [64.24, 71.75] 73.13 [71.52, 74.94] 56.90 [55.93, 57.99] 16.23 [15.44, 17.00] MB2C 67.75 [64.67, 70.76] 67.34 [66.60, 68.01] 49.52 [49.05, 49.94] 17.82 [17.22, 18.30] UBP 56.43 [55.04, 57.83] 73.56 [72.58, 74.47] 47.51 [47.00, 48.03] 26.05 [25.07, 27.05] ATS 61.54 [59.06, 63.80] 75.87 [74.74, 77.01] 53.53 [53.13, 53.92] 22.33 [21.32, 23.34] NeuroBridge 58.66 [56.84, 60.56] 80.23 [79.47, 81.07] 51.30 [50.97, 51.62] 28.93 [28.23, 29.66] HyFI 66.32 [64.12, 68.51] 78.83 [77.85, 79.76] 49.54 [48.87, 50.26] 29.29 [28.41, 30.17] Brain-HIVE 78.10 [74.29, 81.90] 74.38 [73.28, 75.45] 48.08 [47.31, 48.83] 26.30 [25.53, 27.11] Table S12: 2AFC comparisons involving null-edited and semantic-edit candidates. Pairing gain is the difference between correctly paired and mismatched EEG conditions, in percentage points. Brackets report subject-level 95% bootstrap confidence intervals. Model Attribute Background Object Removal Object Identity NICE 0.59 [-0.50, 1.69] 13.10 [11.77, 14.56] 29.60 [27.98, 31.16] 27.17 [26.27, 27.95] ATM 16.08 [15.00, 17.20] 21.26 [19.61, 22.92] 18.28 [16.22, 20.47] 30.17 [28.41, 31.87] MB2C 18.41 [16.84, 19.93] 28.33 [27.05, 29.37] 29.47 [27.69, 31.40] 36.20 [35.21, 37.20] UBP 17.86 [16.95, 18.79] 23.84 [22.87, 24.74] 37.87 [35.92, 39.76] 37.86 [36.64, 39.11] ATS 24.35 [23.21, 25.63] 43.35 [42.44, 44.23] 29.50 [27.99, 30.85] 46.51 [45.06, 47.86] NeuroBridge 21.99 [21.17, 22.87] 42.12 [41.41, 42.93] 24.86 [23.71, 26.13] 45.38 [44.44, 46.22] HyFI 25.36 [24.56, 26.03] 33.57 [32.39, 34.98] 26.98 [25.17, 28.97] 44.75 [43.86, 45.65] Brain-HIVE 27.77 [26.09, 29.44] 36.36 [35.19, 37.39] 31.50 [29.72, 33.08] 44.91 [44.15, 45.67] Table S13: Pairing gain by edit family, in percentage points. Brackets report subject-level 95% bootstrap confidence intervals. Background â - Attribute Object Removal â - Attribute Object Identity â - Attribute Model Unbalanced Distance-balanced Unbalanced Distance-balanced Unbalanced Distance-balanced NICE +10.22 ± 5.74 +8.14 ± 5.39 +16.72 ± 4.09 +14.13 ± 5.13 +14.66 ± 4.32 +8.78 ± 4.37 ATM â-0.84 ± 5.23 â-1.62 ± 4.97 +7.65 ± 4.51 +6.78 ± 3.91 +3.94 ± 2.37 +1.73 ± 3.11 MB2C +9.86 ± 3.28 +9.07 ± 3.62 +7.15 ± 4.39 +5.01 ± 4.42 +9.80 ± 4.47 +5.00 ± 4.44 UBP +13.89 ± 2.42 +12.63 ± 2.67 +17.32 ± 3.97 +13.51 ± 4.14 +15.52 ± 2.07 +11.66 ± 3.32 ATS +12.77 ± 3.59 +11.05 ± 3.75 +15.81 ± 3.39 +12.71 ± 3.66 +10.67 ± 2.09 +7.36 ± 2.52 NeuroBridge +13.25 ± 3.08 +10.80 ± 3.03 +16.48 ± 2.91 +15.47 ± 2.89 +10.13 ± 2.60 +7.04 ± 2.51 HyFI +12.06 ± 2.22 +10.09 ± 2.20 +17.04 ± 3.56 +15.00 ± 3.52 +10.71 ± 2.33 +8.67 ± 3.21 Brain-HIVE +10.36 ± 3.34 +8.41 ± 3.57 +8.37 ± 5.76 +7.15 ± 5.42 +9.02 ± 3.14 +5.78 ± 3.62 Table S14: Attribute-related 2AFC gaps before and after balancing by OpenCLIP visual distance. Values are mean ± population standard deviation across subjects, in percentage points. Positive values indicate higher 2AFC for the non-Attribute family.