Paper deep dive
Banana100: Breaking NR-IQA Metrics by 100 Iterative Image Replications with Nano Banana Pro
Kenan Tang, Praveen Arunshankar, Andong Hua, Anthony Yang, Yao Qin
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/10/2026, 2:09:21 AM
Summary
The paper introduces Banana100, a dataset of 28,000 images generated through 100 iterative editing steps using the Nano Banana Pro model. It identifies a critical failure mode in multi-turn image editing: iterative quality degradation and instruction-following failure. The study demonstrates that popular No-Reference Image Quality Assessment (NR-IQA) metrics, such as BRISQUE, fail to detect this degradation, often assigning better scores to noisy images, which poses risks to the stability and safety of agentic AI systems.
Entities (4)
Relation Signals (3)
Banana100 → contains → degraded images
confidence 100% · we introduce Banana100, a comprehensive dataset of 28,000 degraded images
Nano Banana Pro → generates → Banana100
confidence 100% · We constructed Banana100 by iteratively editing images using Nano Banana Pro.
BRISQUE → failstoevaluate → Nano Banana Pro
confidence 95% · Among 21 popular no-reference image quality assessment (NR-IQA) metrics, none of them consistently assign lower scores to heavily degraded images
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The multi-step, iterative image editing capabilities of multi-modal agentic systems have transformed digital content creation. Although latest image editing models faithfully follow instructions and generate high-quality images in single-turn edits, we identify a critical weakness in multi-turn editing, which is the iterative degradation of image quality. As images are repeatedly edited, minor artifacts accumulate, rapidly leading to a severe accumulation of visible noise and a failure to follow simple editing instructions. To systematically study these failures, we introduce Banana100, a comprehensive dataset of 28,000 degraded images generated through 100 iterative editing steps, including diverse textures and image content. Alarmingly, image quality evaluators fail to detect the degradation. Among 21 popular no-reference image quality assessment (NR-IQA) metrics, none of them consistently assign lower scores to heavily degraded images than to clean ones. The dual failures of generators and evaluators may threaten the stability of future model training and the safety of deployed agentic systems, if the low-quality synthetic data generated by multi-turn edits escape quality filters. We release the full code and data to facilitate the development of more robust models, helping to mitigate the fragility of multi-modal agentic systems.
Tags
Links
- Source: https://arxiv.org/abs/2604.03400v1
- Canonical: https://arxiv.org/abs/2604.03400v1
Trouble viewing inline? Open PDF directly →
Full Text
59,199 characters extracted from source content.
Expand or collapse full text
Banana100: Breaking NR-IQA Metrics by 100 Iterative Image Replications with Nano Banana Pro Kenan Tang, Praveen Arunshankar, Andong Hua, Anthony Yang, Yao Qin University of California, Santa Barbara kenantang@ucsb.edu, yaoqin@ucsb.edu Initial ImageAfter 10 ReplicationsAfter 20 Replications A clean image. BRISQUE = 34.1 Visible noise... BRISQUE = -5.6 Full of static noise!!! BRISQUE = -9.8 Figure 1. Iteratively replicating an image using Nano Banana Pro severely degrades an image, but BRISQUE assigns better (lower) scores to images of worse quality. BRISQUE is a No-Reference Image Quality Assessment (NR-IQA) metric that has been widely used to assess the quality of AI-generated images. This counter-intuitive failure is pervasive across diverse NR-IQA metrics and image textures. Abstract The multi-step, iterative image editing capabilities of multi- modal agentic systems have transformed digital content cre- ation. Although latest image editing models faithfully fol- low instructions and generate high-quality images in single- turn edits, we identify a critical weakness in multi-turn edit- ing, which is the iterative degradation of image quality. As images are repeatedly edited, minor artifacts accumulate, rapidly leading to a severe accumulation of visible noise and a failure to follow simple editing instructions. To sys- tematically study these failures, we introduce Banana100, a comprehensive dataset of 28,000 degraded images gen- erated through 100 iterative editing steps, including diverse textures and image content. Alarmingly, image quality eval- uators fail to detect the degradation. Among 21 popular no- reference image quality assessment (NR-IQA) metrics, none of them consistently assign lower scores to heavily degraded images than to clean ones. The dual failures of generators and evaluators may threaten the stability of future model training and the safety of deployed agentic systems, if the low-quality synthetic data generated by multi-turn edits es- cape quality filters. We release the full code and data to facilitate the development of more robust models, helping to mitigate the fragility of multi-modal agentic systems. 1 1. Introduction AI-based image-text-to-image (IT2T) models has trans- formed digital content creation [8, 11, 38, 39, 48, 58]. These tools allow users to both create new images and iteratively refine them, promising a high degree of creative freedom. This multi-step editing paradigm is further facilitated by the rise of multi-modal agentic systems [36, 62, 63, 72], where autonomous systems composed of a generator (an image editing model) and an evaluator (an image quality assessor) can orchestrate complex image refinement processes. While modern models such as Nano Banana Pro [20] demonstrate impressive image quality in single-turn edits, we identify a critical and underexplored failure mode in the multi-turn scenario, which is iterative degradation. Dur- 1 https : / / huggingface . co / datasets / kenantang / Banana100 1 arXiv:2604.03400v1 [cs.CV] 3 Apr 2026 ing each editing pass, image generators always introduce minor, often imperceptible artifacts [4, 33]. When an output image is fed back into the model again for subsequent edits, these artifacts accumulate into visible quality degradation, such as static noise (Figure 1), greenish tint (Figure 3), or scatter points (Figure 9). Our experiments reveal that af- ter around 5 to 10 steps, Nano Banana Pro quickly starts to suffer from the following two failures: 1. Visual Quality Degradation: High-frequency details are distorted, and visual artifacts emerge in regions that were never targeted for editing. 2. Instruction Following Failure: The model’s capacity to faithfully execute editing prompts progressively deterio- rates, failing to follow even very simple prompts, such as adding an apple on a table (Figure 3). Of greater concern, methods that could potentially serve as the the evaluator component in agentic pipelines prove unreliable for detecting these failure patterns. Out of the 23 popular no-reference image quality assessment metrics (NR-IQA) we examined (Section 4.1), only 2 consistently detected the degradation. Other 21 metrics reported higher quality for noisy images than clean images. As an alarm- ing example, simply replicating an initial image can reach a better (lower) BRISQUE score, despite introducing severe noise and corrupting the original image content (Figure 1). While the clean initial image received a BRISQUE score of 34.1, the noisy image after 20 replications received a far lower (better) BRISQUE score of -9.8. The scores are com- pletely flipped compared to human-perceived image quality. The failures of both the generator and the evaluator allow the degradation to silently leak into datasets without being detected. As an example, the multi-step subset of Pico- Banana-400K [46] exhibited obvious distortions of object textures and human faces, especially after five [5] or six [6] editing steps. 2 The potential negative consequences are pro- found. In particular, we highlight two possible downstream effects. First, on the training side, as AI-edited content proliferates, the future training data may become increas- ingly noisy. If evaluators fail to filtering out noisy data, model collapse could be accelerated in subsequent image generation models [50, 65]. Second, on the inference side, agentic systems are known to be fragile over a long hori- zon [15, 49]. If the degraded images escape the quality checks, the fragility could be further exacerbated. To address these challenges, our three contributions are: 1. Large-scale dataset of iterative degradation: We in- troduce Banana100, a dataset constructed by iteratively editing 13 diverse initial images using 100 editing steps with various instructions, yielding 28,000 images at a cost of $4,000 (Section 2). Other than Nano Banana Pro, 2 The references point to only two example images, but in this dataset, many other images after 5 steps generally suffer from similar degradation. we also confirmed the generalizability of the dataset con- struction pipeline to more IT2T models (Section 4.4). 2. Systematic failure mode taxonomy: With diverse ini- tial images, Banana100 demonstrates multi-step visual quality degradations and instruction-following failure modes, which we systematically categorize into sub- object, object, and image levels (Section 3). 3. Identification of flawed NR-IQA metrics: Beyond generator failure, Banana100 helps to quantitatively identify existing NR-IQA metrics that assign counter- factually good scores for low-quality images (Section 4). This will help researchers avoid falsely reporting an im- provement in image quality when the metrics are actu- ally confounded by model-induced degradation, facili- tating the development of more robust NR-IQA metrics. 2. The Banana100 Dataset We constructed Banana100 by iteratively editing images us- ing Nano Banana Pro. Each initial image was edited by a prompt, and then the output served as the input for the next editing step. Each run consists of 100 editing steps. 2.1. Initial Images We collected a set of high quality initial images with the following 5 requirements. First, the initial images should be in high resolution, with minimal compression artifacts to start with. Second, the initial images should be free from potential copy-right violations. Third, the initial images themselves should be AI-generated, aligning with the re- alistic scenario that a user first generates an image with text prompts and then edits these images multiple times with ad- ditional instructions. Fourth, the images should cover a di- verse range of topics and textures, stress-testing the model’s capability in exact replication. Finally, we deliberately ex- cluded photorealistic faces of humans, as the distortions on real faces are usually visually unpleasant and disturbing [6]. Following these requirements, we curated 13 initial im- ages, all in at least 2K resolution (Table 1). 11 were gen- erated by Nano Banana Pro, with manually refined prompts to cover diverse topics and textures. 2 were generated us- ing SPICE [53], a method that excels at generating highly- resolution and factually correct anime-style images. Note that the conclusions drawn from a deliberately cu- rated set of AI-generated initial images may not directly generalize if the initial images were real-life image with potential compression artifacts, due to the known gap be- tween the two distributions [1]. We leave to future work the exploration of using real-life images for the initial images. 2.2. Iterative Editing Prompts We designed the iterative editing prompts to test the preser- vation of image quality and the evaluation of it with min- imal confounders. One great confounder for the NR-IQA 2 Table 1. The initial images cover diverse content and challenges. The top part of the table includes 11 photorealistic images generated by Nano Banana Pro, and the bottom part includes 2 animation-style images generated by SPICE. The resolutions are width×height. NameImage ContentChallengesResolution BuildingA skyscraperPreservation of highly-regular grid patterns and aerial perspective3392×5056 DongpoA plate of Chinese potstickersPreservation of multi-scale food structure and texture5504×3072 EkphrasisA still life paintingPreservation of diverse textures of the same type of object5632×3072 FogA misty forestPreservation of texture details under lowered color contrast by haze5504×3072 HoliExploding colorful Holi powderPreservation of high color contrast and particle textures5632×3072 LibraryInterior of a libraryPreservation of deep shadows and shafts of light5632×3072 MossTree bark covered in moss and lichenPreservation of soft and non-periodic texture details4800×3584 PeacockA peacock featherPreservation of iridescent texture details4800×3584 RiceRice terraces during sunsetPreservation of reflections and repeated patterns with variations5504×3072 SandA sand dune at twilightPreservation of smooth color gradients5504×3072 TableAn empty wooden tableAddition of diverse objects while preserving the background5504×3072 KokoroA standing animation characterPreservation of asymmetric design and clean stylistic colors1664×2432 YuimanA grid of 9 diverse headshot posesPreservation of 4-colored gradients in the eyes and the grid layout3000×3000 metrics turns out to be the image content. While the initial images are all free from visible noise, some quality metrics provide dramatically different scores for these images. As an example, among all 13 initial images, the Yuiman image has a lowest BRISQUE score of -3.18, while Kokoro has a highest BRISQUE score of 41.1. However, both images were generated with SPICE, and no visible noise is present. Therefore, to minimize the confounding effects of image content on the quality, we primarily conducted the repli- cation runs, where the model was asked to “Produce an exact replica of the provided image, with no alterations.” This focus on a seemingly simple replication task is justi- fied by our pilot study, which revealed that replication leads to noise patterns qualitatively similar to the ones observed with prompts that actually change the semantic content of an image, such as adding objects. Besides this straightforward prompt with the default hy- perparameter set, we also investigated 5 more variants: First, we changed the phrasing of the replication prompt. While the straightforward prompt quantitatively reproduced the failure patterns aligning with general user experience, we would like to test for the sensitivity of vision-language models to the prompt phrasing [29, 44]. Second, we also included multi-step replication opera- tions that transform an image back to its original content using more than one step. For example, horizontally mir- roring an image twice ends up with the original image. This variant was motivated by the observation that when the model is asked to explicitly change one region on the image, the changed region will suffer less from degradation (Section 3.2). Hence, explicitly asking the model to edit the full image might help mitigating the noise accumulation. Third, we further relaxed the requirements on replica- tion by including multi-step reconstruction prompts. These methods are popular in the user community for their poten- tial in denoising a model-edited noisy image. For example, the model is asked to extract simplified color patches in a first step and to extract edge information in a second step. Then, in the third step, the model is asked to reconstruct a photorealistic image from the color patches and the edges. We observed that this method empirically resulted in noise- free images, but the image content was hardly preserved over multiple iterations. Since these methods do not align with the fundamental user requirements of preserving both the quality and the semantic content, we only included a limited number of such runs in the dataset as a reference, but we did not use these runs for image quality assessment. Fourth, we tested with alternative values of three hyper- parameters in the Nano Banana Pro model, including seed, temperature, and resolution. For the seed, either a fixed seed was used throughout the editing steps, or a different seed was provided for each step. This was motivated by obser- vations in our pilot studies that certain images and methods suffer from artifacts when a fixed seed was used throughout the steps, although these artifacts cannot be reliably repro- duced due to the black-box nature of proprietary models. The temperature was either set at 0 or 0.4. The resolution was set to be one of the three options allowed by the API, in- cluding 1K, 2K, and 4K. The resolution could only be cho- sen from these three strings, instead of specified as numeric values. The majority of the dataset was generated with the default resolution of 2K. We used alternative resolutions or interleaved different resolutions (switching periodically in the order of 1K, 2K, and 4K for each step) for a small num- ber of runs, only to investigate the impact of resolution. Finally, to better align with the real use cases while keep- ing confounders minimal, we also used prompts that change only a small region on an image. The Table image was cho- sen for two tasks of adding the same type of fruit (add- apples run) or adding different fruits (add-100-fruits run). All settings above were run with 100 steps, each time in a separate chat session through the Nano Banana Pro API. To keep the cost from quadratically increasing, we did not include all editing steps in a same dialog session. We quali- tatively discuss single-session results in Section 3.3. To en- sure robust analysis, we perform 5 separate runs per setting. 3 However, achieving a full grid search combination is costly. We primarily focused on the replication runs, which was available for all 12 seed images, excluding the Table image that did not include challenging textures and was thus used only for object addition (add-apples and add-100-fruits). Overall, the development and construction of the dataset cost over $4,000, resulting in a dataset of 28,000 total out- put images. The number of images is comparable in the order of magnitude to popular IQA training and evaluation datasets, such as BID [16], CLIVE [18], KonIQ-10k [27], SPAQ [17], Liu13 (deblurring) [37], Min19 (dehazing) [41], AGIQA-3K (image generation) [32], and UHD-IQA [28]. Our dataset is smaller than some of the existing IQA datasets, such as SRIQA-Bench (super-resolution) [14], KADIS-700K [35], and AVA [45]. However, the high im- age resolution in our dataset allows the extraction of multi- ple patches from each image for training or evaluation [23], further increasing the effective size of our dataset. 2.3. Model Selection We selected Nano Banana Pro for its high popularity and its high rank on the Image Edit Arena [7]. While Nano Ba- nana Pro was our primary focus for dataset development, we also tested its successor, Nano Banana 2 [21], together with other open-source models at a smaller scale to demonstrate their qualitative similarities and differences (Section 4.4). We leave the investigation of other agentic image edit- ing systems [36, 62, 63, 72] as future work. However, our focus on the underlying image-editing model deployed in those systems should shed light on the expected degrada- tion behavior of agentic image editing systems. Notably, the evaluation of some agentic image editing systems [62, 72] heavily relied on the NR-IQA metrics such as BRISQUE and NIQE, which we reveal as deeply flawed (Section 4.1). Our dataset is complementary to the existing large scale datasets derived from Nano Banana [46] and Nano Banana Pro [57, 71]. Instead of curating a dataset for the utility of high quality images, we highlight the controlled quality degradation that is unique to our dataset. 2.4. Reasoning Summary Since Nano Banana Pro is a reasoning model, a reasoning trace is generated together with the output image. As Nano Banana Pro does not reveal its full reasoning trace even in the API output, we only included the reasoning summary returned by the API in Banana100. The reasoning summary is broken down into multiple sections. Figure 2 shows an example, in which the final two sections perform evalua- tion, where the model checks if its output aligns with the prompt. In rare cases, the model mentions that the gener- ated output does not align with the prompt and returns to a second round of generation, resulting in a larger number of reasoning summary sections. However, the more predomi- nant pattern is that Nano Banana Pro tends to generate fully confident evaluations, even when the output image totally fails to align with the input text prompt (Section 3). Reasoning Summary Examining the Details I’m currently focused on the painting’s components. I’ve begun dissecting the source material, identifying the individual elements within the still life. I’ve noted the bird, the specific fruits, and other objects. Defining the Components I’m now working to break down the original painting. The intention is to recreate the specific arrangement of objects. The bird, fruit, and other items need careful placement to match the source. Verifying Pixel Accuracy I am currently cross-referencing the initial instructions with the visual out- put. The analysis has revealed an interesting convergence of features. The model appears to have successfully replicated all the required elements, so far. I will proceed with the assessment and proceed to the next step. Evaluating Fidelity of Elements I’m now carefully examining the image, and I’m comparing it to the source material to check for discrepancies. The bird and the chestnuts seem accurately reproduced, based on my initial review. I am now pro- ceeding with a closer inspection. Figure 2. The reasoning summary from Nano Banana Pro ap- pears as clear-cut generation and evaluation sections. The bold text are section titles, copied verbatim from the reasoning sum- mary from the Nano Banana Pro API. In this example, the first two sections are dedicated to image generation, whereas the last two sections are dedicated to the evaluation of a generated image. 3. Analysis of Instruction Following Failures In this section, we qualitatively analyze the failure modes of instruction following. Other than the accumulation of global low-level noise (Figure 1), Nano Banana Pro also failed to follow instructions at three different levels, dubbed as sub-object level, object level, and image level (Fig- ure 3). While non-exhaustive, we list the most obvious fail- ure modes at each level and demonstrate the reasoning sum- mary hallucinations associated to the failures. At least one example image will be provided for each failure mode, and more example images of each failure mode can be easily accessed in our publicly shared dataset. 3.1. Sub-Object-Level Failure Modes In sub-object-level failures, the model failed to faithfully replicate a part of an object. This most frequently happened when a character has a complex and detailed visual design. Simplification Bias. When asked to replicate the image of a character expression grid (Yuiman), the model failed to replicate the exact eye colors after the second step. The original four eye colors (red, orange, purple, and blue) were quickly simplified to only red and blue. In the reasoning summary, we saw that the model only captured the most 4 Mirroring and Rotation Failures Original Eyes Simplified Simplification Bias Counting Failures Previous Step Add an Apple Previous StepAdd a Watermelon Original MirroredRotated Original Multi-SessionSingle-Session Original MonochromeRecolored Original w/oDenoisew/ Denoise Replacement but not Addition Persistent Noise Failure to Reuse Clean Context Monochrome Failures Figure 3. A summary of the failure modes of instruction follow- ing, categorized into sub-object level (blue), object level (yel- low), and image level (green). The images have been cropped and zoomed for visual clarity. As the failures were consistent across different runs and editing steps, we do not report the exact run in- dex and step index for each image here. See Section 3 for details. prominent colors (red and blue) of the eyes, ignoring the other colors (orange and purple). Interestingly, not all grids suffered from the color simplification at the same step. The color gradients on some eyes were preserved in the early steps, but all gradients eventually vanished within 5 steps. This sub-object level failure mode reveals that main- taining character consistency remains an unresolved task. While the consistency might be improved by specifying the character details in the prompt, this approach quickly tum- bles as the number of characters on an image increases. 3.2. Object-Level Failure Modes In object-level failures, the model simply failed to add an object as instructed. Two patterns are listed below. Counting Failures. In the add-apples run, the model was asked to add an apple to the table in each step. In the early steps where the numbers of apples were as small as 7, the model already failed to add one more apple. Moreover, the evaluation section in the reasoning summary mismatched the generation failure. For 3 consecutive editing steps in one run, while the reasoning summary correctly identified 7 apples and confirmed the new total to be 8, the model did not generate a new apple. In the next editing step, the model alternatively added a full row of apple, disregarding the instruction completely. Replacement but not Addition. In the add-100-fruits run, the model was asked to add 100 different fruits to the table, one in each step. Instead of adding the fruit, the model sometimes replaced one of the existing fruit with the new fruit, regardless of the fruit size or the relative position of the fruit (the example shows the replacement of a papaya by a watermelon in the background). The reasoning sum- mary showed that the model did not exhaustively examine each of the existing fruits on the table. Since the full reason- ing trace is not visible, we cannot confirm whether skipping some fruits during reasoning caused this replacement issue. Consistent Background Degradation. Throughout 100 editing steps, the added new object sometimes had re- freshed visual quality, less affected by the worsening noise in the background. This seemed to suggest that editing an image globally might mitigate the noise accumulation and preserve the quality. This motivated us to test the roundtrip decolorization and colorization editing of an im- age as one of the multi-step reconstruction methods (Sec- tion 2.2). In these edits, the model was asked to turn the image monochrome in one editing step and to color the monochrome image in a subsequent editing step, in two sep- arate chat sessions. Although this pair of roundtrip editing steps could not preserve the original colors, this experiment setting was designed to test whether the noise can be re- moved and the quality can be preserved. However, the next subsection shows that this approach did not work. These object-level failure modes reconfirm that handling spatial relationship of objects remains challenging, espe- cially in the presence of model-induced low-level noise. 3.3. Image-Level Failure Modes In image-level failure modes, the model failed in maintain- ing or changing the properties defined on the whole image, such as aspect ratio or orientation. Aspect-Ratio Mismatch. When asked to replicate the im- age, Nano Banana Pro almost always cropped the image in the first step. This might be due to the model requiring the side length of the output to be from a certain set of num- bers. As an example, the resolution of the Ekphrasis image was changed from 5632×3072 to 1408×752, 2816×1504, and 5632×3008 for output resolutions of 1K, 2K, and 4K, respectively. The aspect ratio was changed from 0.545 to 0.534 in all 3 cases by cropping existing pixels in the input. Persistent Noise. The noise introduced over editing steps is persistent, regardless of the prompt phrasing or hyperpa- rameter changes. Notably, explicitly including a denoising instruction in each prompt did not preserve the image qual- ity or content over editing steps. By comparing the “w/o Denoise” and “w/ Denoise” images (both at 20 steps), we saw that both images suffer similarly from an added green tint and a loss of texture. From the reasoning summary, we 5 saw that the model attempted denoising and removing arti- facts, but it failed to denoise the output images at each step. Failure to Reuse Clean Context. One may argue that the multi-session, single-turn setting we adopted prevented the model from reusing the clean images in an earlier genera- tion to eliminate the noise accumulated over the steps. In- deed, as it supports a large context size, the model should be able to use all past context instead of just the most re- cent image. However, when using a single session in the interface for the same object addition task, we saw that the generated result similarly suffered from degradation. Monochrome Failure. When asked to make an image monochrome, the model did not convert the colors strictly to grayscale. Also, the image quality still degraded over the steps, invalidating this two-step reconstruction method. Mirroring and Rotation Failures. For multi-step repli- cation, we chose horizontally mirroring (recovering the original image in every 2 steps) and clock-wise rotation by 90 degrees (recovering the original image in every 4 steps). The mirroring and rotation operations were performed on one realistic image (Ekphrasis) and one animation image (Kokoro). For mirroring, the model had a much lower suc- cess rate on the animation image than the realistic image. For rotation, the success rates were low for both images. For both operations, the image quality degraded similarly as with the naive replication operation. However, the reason- ing summary in each step showed hallucinated confidence. Again, all these full-image operations were motivated by their potential in preserving the image quality over editing steps. Since the obvious failures disqualified these methods from preserving image quality, we did not further quantify the exact failure rate in depth. 4. Noise Quantification and NR-IQA Failures Next, we focused on only the replicate runs for 12 ini- tial images and attempted to use Image Quality Assess- ment (IQA) metrics to quantify the introduced noise. We used a subset of No-Reference IQA (NR-IQA) methods where a score can calculated based on an individual im- age. NR-IQA metrics requiring a reference dataset, such as FID [25], were excluded. Full-Reference IQA (FR-IQA) metrics that require a pair of semantically identical images, such as PSNR [26], LPIPS [67], and SSIM [56], were also excluded. We note that the FR-IQA metrics could be interfered by the change of semantic content on an image (such as an addition of an object). Although we adopted a simplified setting of image replication, such interference makes FR- IQA metrics less suitable than NR-IQA metrics, when the Table 2. A summary of all the No-Reference Image Quality Assessment (NR-IQA) metrics we used for evaluation. In the first part of the table, we show all the NR-IQA metrics imple- mented in the pyiqa Python library [12], with the only exception of MACLIP [34], which is only a placeholder and raises a non- implemented error. The typical range is obtained from the pyiqa library, which do not necessarily correspond to the actual observed range. In the second part of the table, we include two recent NR- IQA metrics based on latest large vision-language models. MetricTypical RangeHigher is Better? ARNIQA [3][0, 1]Yes BRISQUE [42][0, 150]No CLIPIQA [55][0, 1]Yes CNNIQA [30][0, 1]Yes DBCNN [68][0, 1]Yes HyperIQA [51][0, 1]Yes ILNIQE [66][0, 100]No LIQE [69][1, 5]Yes MANIQA [61][0, 1]Yes MUSIQ [31][0, 100]Yes NIMA [52][0, 10]Yes NIQE [43][0, 100]No NRQM [40][0, 10]Yes PaQ-2-PiQ [64][0, 100]Yes PI [9]≥ 0No PIQE [54][0, 100]No Q-Align [59][1, 5]Yes QualiCLIP [2][0, 1]Yes TOPIQ NR [13][0, 1]Yes TReS [19][0, 100]Yes WaDIQaM [10][-1, 0.1]Yes VisualQuality-R1 [60][1, 5]Yes RALI [70][1, 5]Yes end goal is to investigate the quality degradation regard- less of the semantic content. Also, among NR-IQA metrics, the ones that are less interfered by the semantic content are more suitable for the quantification of model-induced noise (more details in Section 4.2). 4.1. NR-IQA Methods Fail to Quantify Degradation The NR-IQA metrics we used are summarized in Table 2. We directly used the models implemented in the pyiqa Python library [12]. When multiple models trained on dif- ferent datasets are available for one metric, we only used the default version as specified on the Model Card page [12]. Since the small degradation over a single step is hard to be precisely judged by humans, we did not obtain Mean Opinion Scores (MOS) for individual images and thus did not use the Pearson Linear Correlation Coefficient (PLCC) and the Spearman Linear Correlation Coefficient (SRCC), two metrics commonly used to rank the performance of NR- IQA models. Instead, we based our evaluation on the ob- servation that the image quality drop after multiple steps is very obvious to bare eyes (Figure 1). This observation aligns with the general experience widely reported by con- temporary users. Based on this observation, we define the normalized score gap ∆ i to be the normalized score of Step i minus the normalized score gap of Step 1 (Figure 4). Here, 6 10 20 30 BRISQUE Score (Lower is Better) 02468101214161820 Step 80 90 Normalized (Higher is Better) 5 = 5.2 10 = 6.5 20 = 6.4 Figure 4. We normalized NR-IQA scores (BRISQUE as an ex- ample) and calculated the difference across steps to quantify the score trend. Please see Section 4.1 for details. The three ∆ values can also be found at the intersection of the second row (Dongpo) and the second column in each heatmap of Figure 5. i can take values from5, 10, 20 but not smaller numbers, because the image quality is unambiguously decreasing for a human observer after a sufficiently large number of editing steps. The initial step was chosen to be 1 instead of 0, in or- der to avoid confounding effects of cropping (Section 3.3). The normalization maps the score from the typical score range to [0, 100], flipping the direction for BRISQUE, IL- NIQE, NIQE, PI, and PIQE such that a higher score consis- tently indicates higher quality. Notably, the normalization does not change the potency of the metric in distinguishing image quality, but it only provides a consistent score scale and direction for the convenience of comparison. Under this definition, a fully successful metric should have all three normalized score gaps (∆ 5 , ∆ 10 , and ∆ 20 ) to be negative. The negative gaps indicate that a metric correctly identifies the image quality as degraded after 4, 9, and 19 steps. However, none of the 21 metrics (which are not based on large VLMs) fully succeeded (Figures 5 and 6). This suggests that the model-induced noise patterns confound these NR-IQA metrics. This could be explained by that these metrics are trained primarily on datasets con- structed with heuristic distortions, such as KonIQ-10k [27], which qualitatively differ from the model-induced noise. 4.2. Two Recent NR-IQA Methods Succeed However, we highlight that RALI [70] and VisualQuality- R1 [60], two recent large-VLM-based metrics, succeed on this task with 0 failure cases, although not free from other failure patterns. RALI is not robust against the change in the image content, exemplified by multiple spikes in the add-100-fruits run (Figure 7). VisualQuality-R1 had scores falling below 1, violating the lower-bound specified in its prompt. Despite these minor issues, the two recent NR-IQA methods successfully identify the accumulation of noise. The success of VisualQuality-R1 might be attributed to its training data covering a diverse mixture of IQA datasets. Building Dongpo Ekphrasis Fog Holi Kokoro Library Moss Peacock Rice Sand Yuiman 5 -2.60.7-5.98.6-6.1-9.10.1-20.9-3.6-7.3-0.8-0.3-2.5-3.1-2.8-3.4-18.6-17.5-5.2-8.17.3 -3.15.2-0.87.5-5.5-7.3-2.9-34.7-2.9-6.1-4.5-0.7-0.9-4.6-4.10.80.1-18.8-5.2-8.47.5 2.98.7-10.44.90.70.7-0.4-25.7-3.4-10.5-1.80.91.7-3.25.93.7-0.3-21.4-2.5-3.732.0 3.67.8-2.65.71.76.3-2.02.62.7-1.4-3.8-0.8-2.1-0.0-5.4-2.10.9-9.9-0.23.030.1 0.97.87.412.3-3.6-3.0-1.1-11.01.93.5-0.2-1.20.11.6-5.70.30.5-9.7-4.0-3.2-0.4 -0.4-1.86.52.45.75.66.56.03.57.6-3.7-1.13.60.8-2.218.2-1.20.77.26.111.3 2.60.2-8.00.53.04.7-2.5-13.5-1.4-6.7-1.0-0.23.70.61.91.8-4.8-13.7-0.41.024.8 2.9-1.30.61.60.11.6-3.1-27.50.0-11.7-3.1-0.7-3.21.5-5.3-18.5-4.7-21.3-2.22.09.1 3.35.62.611.014.36.77.0-2.92.55.01.3-0.12.7-2.20.4-5.5-5.1-4.16.317.028.0 0.66.5-1.52.54.81.8-0.71.6-0.3-0.7-0.4-0.21.11.0-0.3-4.51.3-6.33.53.213.9 -4.3-0.1-8.313.11.3-5.79.1-14.5-5.5-1.7-5.20.22.2-2.02.3-0.0-11.2-4.7-0.8-3.6-2.9 -1.8-10.7-1.1-1.20.92.70.1-11.6-0.2-8.3-2.5-0.70.70.1-3.82.5-1.1-10.9-2.31.0-3.4 Building Dongpo Ekphrasis Fog Holi Kokoro Library Moss Peacock Rice Sand Yuiman 10 -4.83.6-10.06.0-5.4-13.6-1.9-41.8-5.8-10.30.20.1-2.7-4.5-0.9-1.5-45.8-25.0-7.0-12.77.2 -3.36.5-3.26.4-6.7-14.6-7.2-51.3-9.1-11.4-7.2-1.3-0.7-5.4-6.83.1-9.7-28.7-10.3-16.65.9 2.18.3-20.54.7-1.4-9.1-5.0-50.0-12.2-17.5-10.20.22.3-9.32.96.4-8.5-35.4-13.9-16.328.4 9.25.2-8.74.06.610.1-8.42.33.6-10.4-9.0-2.1-2.8-0.3-13.00.8-9.3-12.2-0.98.040.8 4.812.012.514.72.71.4-3.6-16.84.88.8-0.9-1.60.22.9-8.42.0-11.1-10.2-0.84.513.3 2.1-0.85.22.27.64.910.1-24.3-1.16.2-6.5-0.53.6-0.20.217.0-11.2-5.56.48.910.5 3.6-2.1-7.70.13.93.6-5.3-25.5-5.3-6.1-2.0-0.75.41.00.33.3-28.7-19.2-3.31.228.1 11.1-0.1-4.51.04.52.3-5.1-31.91.2-20.0-4.7-1.4-1.62.3-8.4-9.3-37.2-29.0-2.210.718.0 7.26.21.010.916.48.58.2-14.53.94.51.7-0.37.91.62.06.8-21.7-11.23.823.033.6 -0.42.7-9.61.83.3-1.9-1.9-14.9-4.2-1.6-3.2-0.2-0.4-3.5-1.6-9.3-12.6-23.1-3.1-0.713.0 -6.40.8-10.910.72.7-11.011.5-22.3-6.6-2.2-9.20.01.5-4.65.3-1.3-19.7-4.7-0.8-2.9-1.8 -3.5-18.4-4.9-2.1-1.04.64.0-38.4-5.1-13.5-5.4-1.20.80.5-5.82.5-11.7-20.4-4.12.0-3.9 ARNIQA BRISQUE CLIPIQA CNNIQA DBCNN HyperIQA ILNIQE LIQE MANIQA MUSIQ NIMA NIQE NRQM PaQ-2-PiQ PI PIQE Q-Align QualiCLIP TOPIQ NR TReS WaDIQaM Building Dongpo Ekphrasis Fog Holi Kokoro Library Moss Peacock Rice Sand Yuiman 20 -4.88.3-13.58.02.7-18.0-6.0-53.6-8.0-8.6-2.9-0.2-1.2-10.3-1.8-2.4-75.0-27.4-6.0-22.715.3 -1.76.4-12.64.4-3.5-16.7-14.5-63.8-8.2-17.4-8.3-2.4-0.1-9.3-12.74.7-70.4-37.7-13.9-14.16.9 -2.6-3.0-39.71.8-7.1-21.1-14.4-76.3-12.9-27.5-20.0-0.91.2-15.2-3.53.1-56.7-48.0-26.7-21.311.3 14.9-0.4-18.22.312.69.6-11.54.06.9-4.7-17.1-3.6-2.82.0-21.2-1.1-44.7-13.0-1.312.042.2 12.29.117.314.57.94.6-6.0-14.86.85.8-4.0-2.5-0.53.5-12.52.8-69.0-8.21.211.323.0 2.9-12.3-11.81.86.1-1.211.1-50.5-10.52.3-8.4-0.32.5-4.10.39.0-60.9-11.8-3.45.21.3 7.5-6.1-3.3-0.99.97.0-7.7-25.7-2.7-7.0-7.9-1.45.74.5-3.52.4-69.3-19.2-2.110.636.6 14.8-2.2-10.61.414.85.2-12.8-32.88.9-16.2-2.4-3.1-5.12.5-18.7-13.0-92.5-31.25.122.125.7 14.4-2.2-1.610.419.49.08.5-22.96.6-14.3-0.8-1.411.55.7-1.923.7-58.2-13.74.130.035.0 4.1-7.8-31.40.3-0.1-9.60.7-41.7-7.3-9.3-3.7-1.50.0-7.6-6.8-15.4-35.1-37.6-15.3-5.6-0.3 -14.32.0-19.617.4-2.8-22.619.3-33.0-14.8-14.8-17.30.02.7-23.88.50.7-57.8-7.4-13.8-27.0-9.1 -4.7-30.6-12.8-2.5-1.52.63.8-66.1-10.4-13.0-10.5-1.4-0.20.6-7.7-6.7-63.4-20.3-6.02.7-4.9 30 20 10 0 10 20 30 40 20 0 20 40 80 60 40 20 0 20 40 60 80 Figure 5. All 21 NR-IQA metrics fail on identifying degrada- tion, assigning higher normalized scores to images with worse quality. The heatmap shows the gap between the pair of normal- ized scores calculated for the image of a later editing (5, 10, or 20) and the image of the first step. The normalization converts each NR-IQA metric to the same scale of [0, 100], with higher scores corresponding to better image quality. Positive gaps indicate fail- ures and are marked by blue colors in the heatmap. Due to the diversity in the texture types of the 12 initial images, each NR- IQA metric fails on a different set of initial images. ARNIQA BRISQUE CLIPIQA CNNIQA DBCNN HyperIQA ILNIQE LIQE MANIQA MUSIQ NIMA NIQE NRQM PaQ-2-PiQ PI PIQE Q-Align QualiCLIP TOPIQ NR TReS WaDIQaM Building Dongpo Ekphrasis Fog Holi Kokoro Library Moss Peacock Rice Sand Yuiman 0 1 2 3 Number of Positive Figure 6. Aggregated results show that none of the 21 NR-IQA metrics fully succeed on all images. This heatmap overlays the 3 heatmaps from Figure 5, and the brightness of the blue colors correspond to the total number of failures. None of the metrics show a fully white column, corresponding to consistent success. 020406080100 Step 2.5 3.0 3.5 4.0 RALI Score Standard Deviation Mean Figure 7. Despite a consistent drop in image quality, the RALI score (higher is better) fluctuates over the steps. The fluctuation shows that RALI is not robust against the semantic change caused by iterative object addition (add-100-fruits). 4.3. Self-Evaluation is Delayed In the reasoning summary, Nano Banana Pro comments on the original image in the generation section. The comment sometimes mentions the degradation, which can potentially 7 serve a proxy to identify whether the generator is aware of the quality issue, circumventing the evaluator failures. To check whether the model comments on the noise, we use LLM-as-a-judge with Gemini-3-flash (prompt shown in Figure 8). Out of the 100 steps, we looked for the first step where the answer is “yes”, reporting average and stan- dard deviation calculated over 5 replication runs. For the 12 initial images, the smallest identification step is 20± 4 for Holi, and the largest identification step is 37 ± 8 for Rice. These numbers are large compared to the step num- ber where the introduced noise is very obvious, around 5 to 10. This suggests that the generator is not sensitive to the noise it generates, despite the reasoning summary exhibit- ing a certain extent of (heavily hallucinated) self-evaluation. LLM-as-a-Judge Prompt Template Below is a reasoning summary from an image editing model. Please iden- tify if the reasoning summary mentions that the original image is noisy, pixelated, or contains visible artifacts. Output “yes” or “no” only. reasoningsummary Figure 8. The LLM-as-a-judge prompt template to identify whether Nano Banana Pro acknowledges the noise during gen- eration. The reasoning summary, such as one shown in Figure 2, will be inserted to the end of this prompt template. 4.4. Other Image-Editing Models Fail Similarly To examine if noise accumulation is pervasive across mod- els, we follow the image generation and evaluation proto- cols using three alternative models: Nano Banana 2 Fast OriginalNano Banana 2 Fast FLUX.2 [dev]Qwen Image Edit Figure 9. Different noise accumulates during image replication by 3 more models. Nano Banana 2 Fast (without reasoning) gen- erates wrinkles that align with the contours of the objects. FLUX.2 [dev] generates scatter points on many of the objects. Qwen Im- age Edit simplifies the texture and erroneously duplicates objects on the right side of the image to the left side of the image. Building Dongpo Ekphrasis Fog Holi Kokoro Library Moss Peacock Rice Sand Yuiman Nano Banana 2 Fast Building Dongpo Ekphrasis Fog Holi Kokoro Library Moss Peacock Rice Sand Yuiman FLUX.2 [dev] ARNIQA BRISQUE CLIPIQA CNNIQA DBCNN HyperIQA ILNIQE LIQE MANIQA MUSIQ NIMA NIQE NRQM PaQ-2-PiQ PI PIQE Q-Align QualiCLIP TOPIQ NR TReS WaDIQaM Building Dongpo Ekphrasis Fog Holi Kokoro Library Moss Peacock Rice Sand Yuiman Qwen Image Edit 0 1 2 3 Number of Positive Figure 10. Similar to the evaluation of Nano-Banana-Pro gen- erated results, NR-IQA metrics also fail for results from 3 more models. No metric succeeds on all initial images and all models. Interestingly, PI and PIQE fully succeed on Qwen Image Edit, but fails for almost all initial images for Nano Banana 2 Fast. The diverse failure patterns across metrics further confirm the dif- ference of the noise patterns from each model (Figure 9). (without reasoning) [21], FLUX.2 [dev] [8], and Qwen Im- age Edit [47, 58]. We used these models to replicate each of the 12 seed images for 20 steps, repeated for 5 runs. We also used these models for two object addition runs. Over- all, 1,400 new images were created using each model. From the results, we saw that noise similarly accumu- lated over editing steps for each of the models we exam- ined. Notably, the open-source models FLUX.2 [dev] and Qwen Image Edit also suffered from noise, suggesting that the watermarks in the proprietary Nano Banana model fam- ily [22, 24] are not the single cause for quality degradation. However, the noise accumulation patterns differ between these models (Figure 9). A further test using 21 NR-IQA metrics reveal that the metrics again failed on these models, with different failure patterns confirming the qualitatively different nature of the noise patterns (Figure 10). Due to the significant time investment, we did not run the most promis- ing but very large VisualQuality-R1 model on these images. 5. Conclusion Banana100 highlights the fragility of current image gener- ators and evaluators in long-term image editing. By releas- ing 28,000 images that demonstrate quality degradation, we aim to facilitate the development of robust IQA metrics and degradation-free image editors, preventing the unintentional but unchecked pollution of the digital visual ecosystem. 8 References [1] Krzysztof Adamkiewicz, Brian Moser, Stanislav Frolov, To- bias Christian Nauen, Federico Raue, and Andreas Dengel. When pretty isn’t useful: Investigating why modern text-to- image models fail as reliable training data generators. arXiv preprint arXiv:2602.19946, 2026. 2 [2] Lorenzo Agnolucci, Leonardo Galteri, and Marco Bertini. Quality-aware image-text alignment for opinion-unaware image quality assessment. arXiv preprint arXiv:2403.11176, 2024. 6 [3] Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini, and Alberto Del Bimbo.Arniqa: Learning distortion mani- fold for image quality assessment. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 189–198, 2024. 6 [4] Gal Almog, Ariel Shamir, and Ohad Fried. Reed-vae: Re- encode decode training for iterative image editing with dif- fusion models. In Computer Graphics Forum, page e70020. Wiley Online Library, 2025. 2 [5] Apple.10006 attemptAturn5.png. https : / / ml - site . cdn - apple . com / datasets / pico - banana-300k/nb/images/multi-turn/10006_ attemptA_turn5.png, 2026. Accessed: 2026-03-18. 2 [6] Apple.10006attemptAturn6.png. https : / / ml - site . cdn - apple . com / datasets / pico - banana-300k/nb/images/multi-turn/10006_ attemptA_turn6.png, 2026. Accessed: 2026-03-18. 2 [7] Arena AI.Image Editing AI Leaderboard - Best Mod- els Compared. https://arena.ai/leaderboard/ image-edit, 2026. Accessed: 2026-03-14. 4 [8] Black Forest Labs. FLUX.2: Frontier Visual Intelligence. https://bfl.ai/blog/flux-2, 2025. Accessed: 2026-03-14. 1, 8 [9] Yochai Blau, Roey Mechrez, Radu Timofte, Tomer Michaeli, and Lihi Zelnik-Manor. The 2018 pirm challenge on percep- tual image super-resolution. In Proceedings of the European conference on computer vision (ECCV) workshops, pages 0– 0, 2018. 6 [10] Sebastian Bosse, Dominique Maniry, Klaus-Robert M ̈ uller, Thomas Wiegand, and Wojciech Samek. Deep neural net- works for no-reference and full-reference image quality as- sessment. IEEE Transactions on image processing, 27(1): 206–219, 2017. 6 [11] Siyu Cao, Hangting Chen, Peng Chen, Yiji Cheng, Yutao Cui, Xinchi Deng, Ying Dong, Kipper Gong, Tianpeng Gu, Xiusen Gu, et al. Hunyuanimage 3.0 technical report. arXiv preprint arXiv:2509.23951, 2025. 1 [12] Chaofeng Chen.Model Cards for IQA-PyTorch - py- iqa 0.1.13 documentation. https://iqa-pytorch. readthedocs.io/en/latest/ModelCard.html, 2024. Accessed: 2026-03-15. 6 [13] Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu, Liang Liao, Wenxiu Sun, Qiong Yan, and Weisi Lin. Topiq: A top-down approach from semantics to distortions for image quality assessment. IEEE Transactions on Image Processing, 33:2404–2418, 2024. 6 [14] Du Chen, Tianhe Wu, Kede Ma, and Lei Zhang. Toward gen- eralized image quality assessment: Relaxing the perfect ref- erence quality assumption. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 12742– 12752, 2025. 4 [15] Jialong Chen, Xander Xu, Hu Wei, Chuan Chen, and Bing Zhao. Swe-ci: Evaluating agent capabilities in maintain- ing codebases via continuous integration. arXiv preprint arXiv:2603.03823, 2026. 2 [16] Alexandre Ciancio, Eduardo AB Da Silva, Amir Said, Ramin Samadani, Pere Obrador, et al. No-reference blur assessment of digital pictures based on multifeature classifiers. IEEE Transactions on image processing, 20(1):64–75, 2010. 4 [17] Yuming Fang, Hanwei Zhu, Yan Zeng, Kede Ma, and Zhou Wang. Perceptual quality assessment of smartphone pho- tography. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3677–3686, 2020. 4 [18] Deepti Ghadiyaram and Alan C Bovik.Massive online crowdsourced study of subjective and objective picture qual- ity. IEEE transactions on image processing, 25(1):372–387, 2015. 4 [19] S Alireza Golestaneh, Saba Dadsetan, and Kris M Kitani. No-reference image quality assessment via transformers, rel- ative ranking, and self-consistency. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 1220–1230, 2022. 6 [20] Google. Nano Banana Pro: Gemini 3 Pro Image model from Google DeepMind. https://blog.google/ innovation-and-ai/products/nano-banana- pro/, 2025. Accessed: 2026-03-19. 1 [21] Google.Nano Banana 2: Combining Pro capabilities with lightning-fast speed. https://blog.google/ innovation- and- ai/technology/ai/nano- banana-2/, 2026. Accessed: 2026-03-14. 4, 8 [22] Google DeepMind. SynthID - Google DeepMind. https: //deepmind.google/models/synthid/, 2026. Accessed: 2026-03-14. 8 [23] Steve G ̈ oring, Rakesh Rao Ramachandra Rao, and Alexander Raake. Quality assessment of higher resolution images and videos with remote testing. Quality and user experience, 8 (1):2, 2023. 4 [24] Sven Gowal, Rudy Bunel, Florian Stimberg, David Stutz, Guillermo Ortiz-Jimenez, Christina Kouridi, Mel Vecerik, Jamie Hayes, Sylvestre-Alvise Rebuffi, Paul Bernard, et al. Synthid-image: Image watermarking at internet scale. arXiv preprint arXiv:2510.09263, 2025. 8 [25] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems, 30, 2017. 6 [26] Alain Hore and Djemel Ziou. Image quality metrics: Psnr vs. ssim. In 2010 20th international conference on pattern recognition, pages 2366–2369. IEEE, 2010. 6 [27] Vlad Hosu, Hanhe Lin, Tamas Sziranyi, and Dietmar Saupe. Koniq-10k: An ecologically valid database for deep learning 9 of blind image quality assessment. IEEE Transactions on Image Processing, 29:4041–4056, 2020. 4, 7 [28] Vlad Hosu, Lorenzo Agnolucci, Oliver Wiedemann, Daisuke Iso, and Dietmar Saupe. Uhd-iqa benchmark database: Push- ing the boundaries of blind photo quality assessment. In European Conference on Computer Vision, pages 467–482. Springer, 2024. 4 [29] Andong Hua, Kenan Tang, Chenhe Gu, Jindong Gu, Eric Wong, and Yao Qin. Flaw or artifact? rethinking prompt sensitivity in evaluating LLMs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Pro- cessing, pages 19889–19899, Suzhou, China, 2025. Associ- ation for Computational Linguistics. 3 [30] Le Kang, Peng Ye, Yi Li, and David Doermann. Convolu- tional neural networks for no-reference image quality assess- ment. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1733–1740, 2014. 6 [31] Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5148–5157, 2021. 6 [32] Chunyi Li, Zicheng Zhang, Haoning Wu, Wei Sun, Xiongkuo Min, Xiaohong Liu, Guangtao Zhai, and Weisi Lin. Agiqa-3k: An open database for ai-generated image quality assessment. IEEE Transactions on Circuits and Sys- tems for Video Technology, 34(8):6833–6846, 2023. 4 [33] Yucheng Liao, Jiajun Liang, Kaiqian Cui, Baoquan Zhao, Haoran Xie, Wei Liu, Qing Li, and Xudong Mao. Freqedit: Preserving high-frequency features for robust multi-turn im- age editing. arXiv preprint arXiv:2512.01755, 2025. 2 [34] Zhicheng Liao, Dongxu Wu, Zhenshan Shi, Sijie Mai, Han- wei Zhu, Lingyu Zhu, Yuncheng Jiang, and Baoliang Chen. Beyond cosine similarity: Magnitude-aware clip for no- reference image quality assessment. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 6934– 6942, 2026. 6 [35] Hanhe Lin, Vlad Hosu, and Dietmar Saupe. Kadid-10k: A large-scale artificially distorted iqa database.In 2019 Eleventh International Conference on Quality of Multimedia Experience (QoMEX), pages 1–3. IEEE, 2019. 4 [36] Yunlong Lin, ZiXu Lin, Kunjie Lin, Jinbin Bai, Panwang Pan, Chenxin Li, Haoyu Chen, Zhongdao Wang, Xinghao Ding, Wenbo Li, and Shuicheng YAN. Jarvisart: Liberating human artistic creativity via an intelligent photo retouching agent. In The Thirty-ninth Annual Conference on Neural In- formation Processing Systems, 2025. 1, 4 [37] Yiming Liu, Jue Wang, Sunghyun Cho, Adam Finkelstein, and Szymon Rusinkiewicz. A no-reference metric for eval- uating the quality of motion deblurring. ACM Transactions on Graphics, 2013. 4 [38] Zichen Liu, Yue Yu, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Wen Wang, Zhiheng Liu, Qifeng Chen, and Yujun Shen. Magicquill: An intelligent interactive image editing system. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 13072–13082, 2025. 1 [39] Zichen Liu, Yue Yu, Hao Ouyang, Qiuyu Wang, Shuailei Ma, Ka Leong Cheng, Wen Wang, Qingyan Bai, Yuxuan Zhang, Yanhong Zeng, et al. Magicquillv2: Precise and interac- tive image editing with layered visual cues. arXiv preprint arXiv:2512.03046, 2025. 1 [40] Chao Ma, Chih-Yuan Yang, Xiaokang Yang, and Ming- Hsuan Yang. Learning a no-reference quality metric for single-image super-resolution. Computer Vision and Image Understanding, 158:1–16, 2017. 6 [41] Xiongkuo Min, Guangtao Zhai, Ke Gu, Yucheng Zhu, Jiantao Zhou, Guodong Guo, Xiaokang Yang, Xinping Guan, and Wenjun Zhang. Quality evaluation of image de- hazing methods using synthetic hazy images. IEEE Transac- tions on Multimedia, 21(9):2319–2333, 2019. 4 [42] Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-reference image quality assessment in the spatial domain. IEEE Transactions on image processing, 21(12): 4695–4708, 2012. 6 [43] Anish Mittal, Rajiv Soundararajan, and Alan C Bovik. Mak- ing a “completely blind” image quality analyzer. IEEE Sig- nal processing letters, 20(3):209–212, 2012. 6 [44] Wenyi Mo, Tianyu Zhang, Yalong Bai, Bing Su, Ji-Rong Wen, and Qing Yang. Dynamic prompt optimizing for text- to-image generation. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 26627–26636, 2024. 3 [45] Naila Murray, Luca Marchesotti, and Florent Perronnin. Ava: A large-scale database for aesthetic visual analysis. In 2012 IEEE conference on computer vision and pattern recognition, pages 2408–2415. IEEE, 2012. 4 [46] Yusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong, Yinfei Yang, Jiasen Lu, Wenze Hu, and Zhe Gan. Pico- banana-400k: A large-scale dataset for text-guided image editing, 2025. 2, 4 [47] Qwen.Qwen/Qwen-Image-Edit-2511 - Hugging Face. https://huggingface.co/Qwen/Qwen-Image- Edit-2511, 2025. Accessed: 2026-03-14. 8 [48] Team Seedream, Yunpeng Chen, Yu Gao, Lixue Gong, Meng Guo, Qiushan Guo, Zhiyao Guo, Xiaoxia Hou, Weilin Huang, Yixuan Huang, et al. Seedream 4.0: Toward next- generation multimodal image generation.arXiv preprint arXiv:2509.20427, 2025. 1 [49] Shuai Shao, Qihan Ren, Chen Qian, Boyi Wei, Dadi Guo, Yang JingYi, Xinhao Song, Linfeng Zhang, Weinan Zhang, Dongrui Liu, and Jing Shao. Your agent may misevolve: Emergent risks in self-evolving LLM agents. In Socially Re- sponsible and Trustworthy Foundation Models at NeurIPS 2025, 2025. 2 [50] Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Pa- pernot, Ross Anderson, and Yarin Gal. Ai models collapse when trained on recursively generated data. Nature, 631 (8022):755–759, 2024. 2 [51] Shaolin Su, Qingsen Yan, Yu Zhu, Cheng Zhang, Xin Ge, Jinqiu Sun, and Yanning Zhang. Blindly assess image qual- ity in the wild guided by a self-adaptive hyper network. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 3667–3676, 2020. 6 [52] Hossein Talebi and Peyman Milanfar. Nima: Neural image assessment. IEEE transactions on image processing, 27(8): 3998–4011, 2018. 6 10 [53] Kenan Tang, Yanhong Li, and Yao Qin. SPICE: A synergis- tic, precise, iterative, and customizable image editing work- flow. In The Thirty-ninth Annual Conference on Neural In- formation Processing Systems Creative AI Track: Humanity, 2025. 2 [54] Narasimhan Venkatanath, D Praneeth, S Channappayya Sumohana, S Medasani Swarup, et al. Blind image quality evaluation using perception based features. In 2015 twenty first national conference on communications (NCC), pages 1–6. IEEE, 2015. 6 [55] Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Ex- ploring clip for assessing the look and feel of images. In Pro- ceedings of the AAAI conference on artificial intelligence, pages 2555–2563, 2023. 6 [56] Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 6 [57] Xinyu Wei, Kangrui Cen, Hongyang Wei, Zhen Guo, Bairui Li, Zeqing Wang, Jinrui Zhang, and Lei Zhang. Mico-150k: A comprehensive dataset advancing multi-image composi- tion. arXiv preprint arXiv:2512.07348, 2025. 4 [58] Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng-ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, et al. Qwen-image technical report. arXiv preprint arXiv:2508.02324, 2025. 1, 8 [59] Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guang- tao Zhai, and Weisi Lin. Q-align: Teaching LMMs for visual scoring via discrete text-defined levels. In Forty-first Inter- national Conference on Machine Learning, 2024. 6 [60] Tianhe Wu, Jian Zou, Jie Liang, Lei Zhang, and Kede Ma. Visualquality-r1: Reasoning-induced image quality assess- ment via reinforcement learning to rank. In The Thirty-ninth Annual Conference on Neural Information Processing Sys- tems, 2025. 6, 7 [61] Sidi Yang, Tianhe Wu, Shuwei Shi, Shanshan Lao, Yuan Gong, Mingdeng Cao, Jiahao Wang, and Yujiu Yang. Maniqa: Multi-dimension attention network for no-reference image quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1191–1200, 2022. 6 [62] Mingde Yao, Zhiyuan You, Tam-King Man, Menglu Wang, and Tianfan Xue.Photoagent: Agentic photo editing with exploratory visual aesthetic planning. arXiv preprint arXiv:2602.22809, 2026. 1, 4 [63] Ruijie Ye, Jiayi Zhang, Zhuoxin Liu, Zihao Zhu, Siyuan Yang, Li Li, Tianfu Fu, Franck Dernoncourt, Yue Zhao, Jiacheng Zhu, et al. Agent banana: High-fidelity image editing with agentic thinking and tooling. arXiv preprint arXiv:2602.09084, 2026. 1, 4 [64] Zhenqiang Ying, Haoran Niu, Praful Gupta, Dhruv Maha- jan, Deepti Ghadiyaram, and Alan Bovik. From patches to pictures (paq-2-piq): Mapping the perceptual space of pic- ture quality. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3575–3585, 2020. 6 [65] Youngseok Yoon, Dainong Hu, Iain Weissburg, Yao Qin, and Haewon Jeong. Model collapse in the self-consuming chain of diffusion finetuning: A novel perspective from quantita- tive trait modeling. arXiv preprint arXiv:2407.17493, 2024. 2 [66] Lin Zhang, Lei Zhang, and Alan C Bovik. A feature-enriched completely blind image quality evaluator. IEEE Transactions on Image Processing, 24(8):2579–2591, 2015. 6 [67] Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 6 [68] Weixia Zhang, Kede Ma, Jia Yan, Dexiang Deng, and Zhou Wang. Blind image quality assessment using a deep bilinear convolutional neural network. IEEE Transactions on Cir- cuits and Systems for Video Technology, 30(1):36–47, 2018. 6 [69] Weixia Zhang, Guangtao Zhai, Ying Wei, Xiaokang Yang, and Kede Ma. Blind image quality assessment via vision- language correspondence: A multitask learning perspective. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 14071–14081, 2023. 6 [70] Shijie Zhao, Xuanyu Zhang, Weiqi Li, Junlin Li, Li Zhang, Tianfan Xue, and Jian Zhang. Reasoning as representation: Rethinking visual reinforcement learning in image quality assessment. Proceedings of the International Conference on Learning Representations (ICLR), 2026. 6, 7 [71] Jialong Zuo, Haoyou Deng, Hanyu Zhou, Jiaxin Zhu, Yicheng Zhang, Yiwei Zhang, Yongxin Yan, Kaixing Huang, Weisen Chen, Yongtai Deng, et al. Is nano banana pro a low-level vision all-rounder? a comprehensive evaluation on 14 tasks and 40 datasets. arXiv preprint arXiv:2512.15110, 2025. 4 [72] Yushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang, Ren- jie Li, Jian Wang, Yide Zhang, Gengchen Mai, Lihong Wang, James Zou, Xiaoyu Wang, Ming-Hsuan Yang, and Zhengzhong Tu. 4KAgent: Agentic any image to 4k super- resolution. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. 1, 4 11