Paper deep dive
ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation
Junhu Fu, Shuyu Liang, Wutong Li, Chen Ma, Peng Huang, Kehao Wang, Ke Chen, Shengli Lin, Pinghong Zhou, Zeju Li, Yuanyuan Wang, Yi Guo
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 9:07:35 AM
Summary
The paper introduces ColoDiff, a diffusion-based framework for generating dynamic-consistent and content-aware colonoscopy videos to address data scarcity in clinical settings. It features a TimeStream module for inter-frame temporal consistency via cross-frame tokenization, a Content-Aware module for precise control over clinical attributes using noise-injected embeddings and learnable prototypes, and a non-Markovian sampling strategy for efficient real-time generation. Evaluated on public datasets and a hospital database, it improves downstream task performance such as disease diagnosis and lesion segmentation.
Entities (10)
Relation Signals (8)
ColoDiff â contains â TimeStream Module
confidence 95% · At the inter-frame level, our TimeStream module decouples temporal dependency from video sequences
ColoDiff â contains â Content-Aware Module
confidence 95% · At the intra-frame level, our Content-Aware module incorporates noise-injected embeddings
ColoDiff â isbasedon â Diffusion Model
confidence 95% · we propose ColoDiff, a diffusion-based framework
ColoDiff â supports â Colonoscopy Video Generation
confidence 95% · ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation
Content-Aware Module â enables â Content Control
confidence 90% · Content-Aware module... to realize precise control over clinical attributes
TimeStream Module â enables â Temporal Consistency
confidence 90% · TimeStream module decouples temporal dependency... enabling intricate dynamic modeling
ColoDiff â improves â Disease Diagnosis Accuracy
confidence 90% · Incorporating synthetic videos into training promotes discriminative representation learning and improves diagnosis accuracy by 7.1%
ColoDiff â uses â Non-Markovian Sampling
confidence 90% · ColoDiff employs a non-Markovian sampling strategy that cuts steps by over 90%
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Colonoscopy video generation delivers dynamic, information-rich data critical for diagnosing intestinal diseases, particularly in data-scarce scenarios. High-quality video generation demands temporal consistency and precise control over clinical attributes, but faces challenges from irregular intestinal structures, diverse disease representations, and various imaging modalities. To this end, we propose ColoDiff, a diffusion-based framework that generates dynamic-consistent and content-aware colonoscopy videos, aiming to alleviate data shortage and assist clinical analysis. At the inter-frame level, our TimeStream module decouples temporal dependency from video sequences through a cross-frame tokenization mechanism, enabling intricate dynamic modeling despite irregular intestinal structures. At the intra-frame level, our Content-Aware module incorporates noise-injected embeddings and learnable prototypes to realize precise control over clinical attributes, breaking through the coarse guidance of diffusion models. Additionally, ColoDiff employs a non-Markovian sampling strategy that cuts steps by over 90% for real-time generation. ColoDiff is evaluated across three public datasets and one hospital database, based on both generation metrics and downstream tasks including disease diagnosis, modality discrimination, bowel preparation scoring, and lesion segmentation. Extensive experiments show ColoDiff generates videos with smooth transitions and rich dynamics. ColoDiff presents an effort in controllable colonoscopy video generation, revealing the potential of synthetic videos in complementing authentic representation and mitigating data scarcity in clinical settings.
Tags
Links
- Source: https://arxiv.org/abs/2602.23203v1
- Canonical: https://arxiv.org/abs/2602.23203v1
Trouble viewing inline? Open PDF directly â
Full Text
65,787 characters extracted from source content.
Expand or collapse full text
IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 20251 ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation Junhu Fu, Shuyu Liang, Wutong Li, Chen Ma, Peng Huang, Kehao Wang, Ke Chen, Shengli Lin, Pinghong Zhou, Zeju Li, Yuanyuan Wang, Senior Member, IEEE, and Yi Guo, Member, IEEE Abstractâ Colonoscopy video generation delivers dy- namic, information-rich data critical for diagnosing intesti- nal diseases, particularly in data-scarce scenarios. High- quality video generation demands temporal consistency and precise control over clinical attributes, but faces chal- lenges from irregular intestinal structures, diverse dis- ease representations, and various imaging modalities. To this end, we propose ColoDiff, a diffusion-based frame- work that generates dynamic-consistent and content-aware colonoscopy videos, aiming to alleviate data shortage and assist clinical analysis. At the inter-frame level, our TimeStream module decouples temporal dependency from video sequences through a cross-frame tokenization mech- anism, enabling intricate dynamic modeling despite ir- regular intestinal structures. At the intra-frame level, our Content-Aware module incorporates noise-injected embed- dings and learnable prototypes to realize precise control over clinical attributes, breaking through the coarse guid- ance of diffusion models. Additionally, ColoDiff employs a non-Markovian sampling strategy that cuts steps by over 90% for real-time generation. ColoDiff is evaluated across three public datasets and one hospital database, based on both generation metrics and downstream tasks in- cluding disease diagnosis, modality discrimination, bowel preparation scoring, and lesion segmentation. Extensive experiments show ColoDiff generates videos with smooth transitions and rich dynamics. ColoDiff also produces customized contents tailored for diverse tasks, e.g., col- itis, polyps, and adenomas for diagnosis. Incorporat- This work was supported in part by the National Natural Science Foundation of China under Grant 62371139; and in part by the Shanghai Municipal Education Commission under Grant 24KNZNA09. (Corre- sponding authors: Zeju Li; Yuanyuan Wang; Yi Guo.) This work involved human subjects or animals in its research. Ap- proval of all ethical and experimental procedures and protocols was granted by the Ethics Committee of Fudan University Shanghai Cancer Center, Shanghai, China, under Application No. 2509-Exp283, in 2025. Junhu Fu, Shuyu Liang, Wutong Li, Chen Ma, Peng Huang, Kehao Wang, Zeju Li, Yuanyuan Wang, and Yi Guo are with the College of Biomedical Engineering, Fudan University, Shanghai 200433, China (e-mail: jhfu21@m.fudan.edu.cn; syliang22@m.fudan.edu.cn; wtli22@m.fudan.edu.cn;cma24@m.fudan.edu.cn;phuang22@m. fudan.edu.cn;wang kehao@fudan.edu.cn;zejuli@fudan.edu.cn; yywang@fudan.edu.cn; guoyi@fudan.edu.cn). Ke Chen is with the Department of Endoscopy, Fudan Uni- versity Shanghai Cancer Center, Shanghai 200032, China (e-mail: kechen23@m.fudan.edu.cn). Shengli Lin and Pinghong Zhou are with the Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai 200032, China, and also with the Shanghai Collaborative Innovation Center of Endoscopy, Shanghai 200032, China (e-mail: lin.shengli@zs-hospital.sh.cn; zhou.pinghong@zs-hospital.sh.cn). ing synthetic videos into training promotes discrimina- tive representation learning and improves diagnosis accu- racy by 7.1%. ColoDiff presents an effort in controllable colonoscopy video generation, revealing the potential of synthetic videos in complementing authentic representa- tion and mitigating data scarcity in clinical settings. Index Termsâ Colonoscopy video generation, diffusion model, temporal consistency, content controllability. I. INTRODUCTION C OLORECTAL diseases often present with subtle early symptoms yet remain a leading cause of global mortal- ity [1], thus demanding early intervention and improved prog- nosis. Colonoscopy video analysis plays a crucial role in gas- trointestinal diagnosis, providing dynamic, multi-perspective views into mucosal microvascular structures and offering real- time feedback. These capabilities form a solid foundation for a wide range of clinical tasks, from bowel preparation scoring [2], [3], lesion screening or tracking [4], to disease diagnosis [5], as well as auxiliary tasks like modality discrim- ination [6], [7]. Example colonoscopy video sequences with different modalities or diseases are illustrated in Fig. 1(a)-(b), further demonstrating the strength of video analysis in cap- turing temporal-coherent and multi-view features for diverse intestinal conditions. Deep learning methods trained on large- scale data have achieved remarkable success in visual analysis, but collecting adequate colonoscopy videos is impractical in clinical scenarios due to privacy regulation, laborious annotation, and heterogeneous protocols. The lack of high- quality data severely constrains computer-aided diagnosis and treatment, underscoring the urgent need for effective solutions. Colonoscopy data generation has emerged as a promising solution to bridge this gap. In the broader field of image generation, diffusion-based approaches dominate with supe- rior synthesis quality and training stability [8], [9]. As for colonoscopy image generation, recent algorithms [10]â[13] build on latent diffusion models (LDMs) [14], leveraging variational autoencoder (VAE) latent spaces [15], [16] for per- ceptual compression and high-fidelity synthesis. Despite these advances, the models focus on 2D static generation, failing to capture multi-perspective views and dynamic information. In the realm of video generation, transition from static to dynamic arXiv:2602.23203v1 [cs.CV] 26 Feb 2026 2IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2025 (c) Challenge 1: Complex Temporal Modeling (d) Challenge 2: Customized Content Control (e) Challenge 3: Restricted Inference Speed Current: Step-by-step Sampling Expected: Skip-step Sampling Current: Inter-frame Incoherence Current: Limited Content Variation (a) Modality: White-light; Narrow-band (b) Disease: Colitis; Hyperplastic; Adenoma . . .. . .. . . . . .. . .. . . âNarrow-bandâ â â âBBPS 2-3â C u r r e n t M o d e l s E x p e c t e d M o d e l s C u r r e n t M o d e l s E x p e c t e d M o d e l s â â Fig. 1.(a)-(b) Colonoscopy video analysis serves as an important diagnostic approach, integrating diverse imaging modalities and disease manifestations. (c)-(e) Existing challenges for colonoscopy video generation. requires modeling temporal dependencies for inter-frame co- herence. Some approaches employ 3D U-Net to jointly model spatio-temporal features [17], [18], and others propose contin- uous temporal encoding with motion disentanglement [19]â [21]. For colonoscopy video generation specially, it faces more intricate dynamics from irregular intestinal structures, coupled with variable mucosal and microvascular manifesta- tions across imaging modalities and diseases. Although the latest diffusion-based methods show promise in endoscopic video generation by providing synthetic unlabeled data for downstream tasks, they still lack inter-frame consistency and content controllability [22], [23], underscoring the need for customized representation and targeted generation. While recent efforts demonstrate the feasibility of colonoscopy video generation, there remain some critical chal- lenges: (1) Complex temporal modeling. Existing methods employ 3D operations [17], [18] or concatenate frames into pseudo-large volumes [24] for feature extraction, but both in- adequately capture temporal dynamics, leading to inter-frame inconsistency for intricate scenarios like irregular morphol- ogy and scale-varying anatomy (Fig. 1(c)). (2) Customized content control. Conditional diffusion models rely on the time-step index for noise perception and fixed encodings for diverse representations [8], [9]. Such information is insuffi- cient for colonoscopy videos [2], [4], [25], which involve various disease manifestations and imaging modalities, thereby hindering controllable generation (Fig. 1(d)). (3) Restricted inference speed. Diffusion-based video generation requires extensive temporal modeling and hundreds of sampling steps, thus preventing real-time inference (Fig. 1(e)). In this paper, we propose ColoDiff, a diffusion-based frame- work for dynamic-consistent and content-aware colonoscopy video generation. ColoDiff integrates tailored temporal mod- eling, precise content control, and an efficient non-Markovian sampling strategy, enabling real-time synthesis while address- ing the limitations of existing methods. In summary, the main contributions of this paper are as follows: (i) We propose ColoDiff, a diffusion-driven architecture that integrates TimeStream and Content-Aware mod- ules for real-time colonoscopy video generation. By addressing the challenges of temporal modeling, con- tent controllability, and inference efficiency, ColoDiff establishes a new framework for supplementing specific representation and mitigating data scarcity. (i) At the inter-frame level, our TimeStream module ex- plicitly decouples temporal dependencies from video se- quences through a cross-frame tokenization mechanism, capturing motion patterns during endoscope movement. At the intra-frame level, our Content-Aware module in- corporates noise-injected embedding and category-aware prototypes for fine-grained modulation, enabling precise control over clinical attributes. (i) Extensive experiments of generated videos demonstrate enhanced temporal consistency and controllability, with synthetic-augmented data improving disease diagnosis accuracy by 7.1% and segmentation Dice by 6.2%. The results prove ColoDiffâs groundbreaking contribution in complementing customized videos and improving downstream task performance. The above is the Introduction section of this paper. The Related Works, Methodology, Experiments and Results, Dis- cussion, and Conclusion sections are followed in turn. FU et al.: ColoDiff: INTEGRATING DYNAMIC CONSISTENCY WITH CONTENT AWARENESS3 I. RELATED WORKS A. From Static to Dynamic Medical image generation, including colonoscopy image synthesis, has become a privacy-preserving approach to mit- igate data scarcity. Recent algorithms [10]â[13] combine the generative power of diffusion process with the efficiency of perceptual compression [15], [16], enabling the synthesis of high-fidelity images for improved diagnosis. Despite the ad- vances, these methods remain confined to 2D static generation, lacking the ability to capture multi-perspective views and temporal dynamics. This limitation has motivated a growing shift toward dynamic generation, where temporal information is explicitly modeled to better reflect clinical scenarios. Following the trend of dynamic modeling, some algorithms attempt to extend static image models to video synthesis by in- corporating temporal constraints [26]â[28]. These approaches typically integrate fine-tuned temporal layers into existing image models, reducing the need for large-scale video datasets. At the same time, this pattern is inherently limited by the static prior of pre-trained image models, struggling to capture dynamic characteristics, particularly complex motion patterns and long-term temporal dependencies. The key-frame-first approach with intermediate frame interpolation [29], [30], has also been proposed to enhance temporal coherence, while this strategy may lead to inconsistent motion trajectories and fail to introduce effective information. Other algorithms suggest training video generation models from scratch, which allows direct learning of temporal dynamics without static bias. In this setting, U-Net and Transformer architectures remain the predominant choices. Ho et al. and He et al. utilize the standard diffusion setup with a 3D U-Net architecture for video generation, while facing inherent inductive bias like locality and translation equivariance [17], [18]. Additionally, the 3D convolution substantially inflates parameter scales. To achieve better scalability, Sora and Latte leverage a diffusion Transformer (DiT) architecture that operates on spacetime patches of video latent encoding [31], [32]. Zhang et al. propose fully cross-frame interaction, which concatenates all video frames into a large image for joint encoding via Trans- former, but this increases the number of input tokens, leading to high computational complexity [24]. As can be seen, while the transition from static to dynamic has become an important direction, medical video generation still struggles with effectively modeling temporal dynamics. How to achieve inter-frame consistency while maintaining manageable model scale and computational efficiency remains a problem to be solved. B. From Uncontrollable to Customized The latest diffusion-based methods have made promising attempts in endoscopic video generation [22], [23]. Due to the lack of textual descriptions and category labels, the synthetic videos remain uncontrollable. They cannot reliably reflect specific disease types, imaging modalities, or other clinically relevant attributes, and thus are used for unlabeled data augmentation in semi-supervised downstream tasks [22], [23]. This limitation highlights the demand for methods that support customized representation and targeted generation. In the broader field of controllable natural video gener- ation, current approaches primarily rely on textual and vi- sual alignment. For diffusion models based on U-Net, Stable Diffusion injects condition information by inserting cross- attention modules into intermediate layers [14], while Con- trolNet adds trainable condition branches to pre-trained U-Net structure [33]. For diffusion models based on Transformer, U- ViT encodes conditioning as extra tokens that are concatenated with visual tokens and processed through Transformer blocks with long skip connections [34]. Similarly, DiT employs a pure Transformer architecture that incorporates conditioning via adaptive layer normalization, progressively increasing con- ditional influence during training through zero-initialization strategies [32]. To sum up, conditioning in existing diffusion models typically includes time-step indices, textual embed- ding, and class embedding. Unlike natural data, medical videos are inherently con- strained by limited data and lack of paired texts [35], making it difficult to learn robust representations of diverse patterns and achieve high controllability for customized contents. In the absence of textual guidance, time-step and class embedding become the only remaining conditioning, but their coarse gran- ularity precludes fine-grained alignment for precise control. I. METHODOLOGY Fig. 2 illustrates the workflow of ColoDiff, our diffusion- based model with Transformer architecture, supporting colonoscopy video generation with dynamic consistency and content controllability. During training, ColoDiff learns to predict added noise in the diffusion process under conditional guidance, which enables Transformer to learn the denoising pathway. Specifically, the TimeStream module decouples tem- poral dynamics via inter-frame interactions to guarantee video coherence, while the Content-Aware module leverages noise- injected embedding with trainable prototypes for intra-frame clinical attribute control. During inference, ColoDiff achieves customized generation through a non-Markovian chain, guid- ing the denoising process from Gaussian noise. A. TimeStream Module Enhances Dynamic Consistency Colonoscopy videos present complex spatio-temporal pat- terns due to irregular intestinal structures, which challenges temporal consistency modeling. Although Transformers [32], [36], [37] excel at capturing long-range dependency within a frame, efficient cross-frame interaction for temporal model- ing remains computationally expensive [24]. To address this limitation, we introduce the TimeStream module, specifically designed to enhance temporal coherence by modeling inter- frame relationships in a computationally efficient manner. As illustrated in Fig. 2, the Transformer encoder with D latent dimensions would produce feature maps z i t â R FĂPĂD , where i represents the layer index, F represents the number of frames and P denotes the number of patches per frame. Then we use a transpose-like operation to rearrange all to- kens, considering patches with identical spatial location across 4IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2025 T r a n s f o r m e r E n c o d e r L a y e r T i m e S t r e a m M o d u l e T r a n s f o r m e r E n c o d e r L a y e r T i m e S t r e a m M o d u l e . . . N Predicted Noise Training Time Step t L a y e r N o r m L i n e a r a n d R e s h a p e D i f f u s i o n ( N o i s e A d d i n g ) Pachify Input Video MLP Conditioning: Îł ÎČ Î± λ Number of Patches . . . . . . . . . . . . Category Prototypes Class 1 Class 2 Class C . . . . . . . . . . . .. . . α 1 Îł 2 , ÎČ 2 α 2 TimeStream Module Number of Frames . . . S c a l e P o i n t w i s e F e e d f o r w a r d S c a l e , S h i f t L a y e r N o r m S c a l e M u l t i - h e a d A t t e n t i o n S c a l e , S h i f t L a y e r N o r m Transpose Îł 1 , ÎČ 1 A 11 A pq A 1q A p1 A ij A 11 C 11 B 11 A pq B pq C pq A 11 C 11 B 11 A pq B pq C pq Fine-grained Guidance . . . Embedding λ Content-Aware Module Inference Gaussian Noise Synthesized Video Time Step t Embedding Non-Markovian Interative Denoising Category Prototypes Fig. 2. The overall workflow of ColoDiff. E and D denote pre-trained encoder and decoder in VAE, respectively.z t represents the latent features that have been added noise. The light red area shows how the TimeStream module decouples temporal information, and the light blue area shows how the Content-Aware module controls generation process. different frames as sequential inputs. These patches tend to reflect the movement of the same anatomical structures, e.g., specific lesions or capillaries. Therefore, P sequences with length F are formed, and the shape of feature is transformed to R PĂFĂD . In this way, we can feed these P sequences into the subsequent blocks, including multi-head attention (MHA) and MLP-based feedforward layers, in parallel: h i t = MHA n LayerNorm h z i t T io + z i t T ,(1) z i+1 t = MLP LayerNorm h i t + h i t ,(2) where h i t denotes the hidden state at the i-th layer, and z i t T â R PĂFĂD denotes the transpose of z i t . Specifi- cally, in each of the P sequences, the F tokens are from different frames. These tokens initially undergo an attention mechanism, aggregating information from all other tokens via learned attention weights to compute temporal context-aware representations. Following this, they pass through residual connection and layer normalization before entering an MLP that applies a position-wise transformation. The MLP uses a non-linear GELU activation, expanding feature dimensionality before projecting back to the original dimension, and enables each token to integrate both long-range and short-range inter- frame connections. The output z i+1 t of the i-th layer is finally produced after another round of residual connection and normalization, ensuring stable gradients throughout the deep network while enabling complex interactions. To avoid spatial information loss, we employ Transformer encoder layers and TimeStream modules in sequence, with one followed by another, totally repeated N (N =28) times (see top Fig. 2). The TimeStream module leverages prior knowledge that the same anatomical structure generally occupies consistent or adjacent spatial locations across consecutive frames, and its variation over time is continuous and modelable. This enables accurate modeling of irregular intestinal structures and dynamic motion patterns. In addition, this module empowers ColoDiff to efficiently model temporal dynamics by using 2D models for 3D contextual reasoning, without increasing model scale or computational cost. B. Content-Aware Module Provides Precise Guidance To ensure clinical utility, generated colonoscopy videos should align with specific clinical attributes, which requires the content-aware ability of ColoDiff. As a general principle, diffusion models should adapt to inputs with varying noise levels across diverse time steps, but existing frameworks only rely on the time-step index t to perceive noise levels, yielding coarse conditioning without content-aware information. To address this limitation, our Content-Aware module introduces the input video embedding as additional guid- FU et al.: ColoDiff: INTEGRATING DYNAMIC CONSISTENCY WITH CONTENT AWARENESS5 ance. Specifically, the noise-injected data z t is encoded into Embed z i t by a learnable encoder. Compared with time-step index t only providing a global representation for noise level, Embed z i t merges information from noise level, intra-frame visual concepts, and their interaction after noise injection. Hence, it serves as a fine-grained condition with intra-frame spatial information for ColoDiffâs Transformer blocks. This enriched conditioning revises the attention mechanism as Softmax Q i (K i ) T â d k V i + λ· Embed z i t ,(3) where Q i , K i , and V i denote the query, key, and value at the i-th layer. d k denotes the dimension of query and key, serving as a scaling factor. λ is a learnable weighting parameter with zero-initialization, ensuring that Embed z i t produces a progressive influence. According to (3), ColoDiff can better perceive the injected noise level by integrating Embed z i t into the attention mechanism, thus achieving improved noise prediction and content-aware modulation. In addition, to make ColoDiff obtain more specific rep- resentation from various video patterns, our Content-Aware module introduces prototype learning. Different from current fixed encoding schemes [32], [36], we assign a learnable representation vector, i.e., prototype, for each category, as illustrated in Fig. 2. Based on these prototypes, ColoDiff employs scaling parameters Îł, α and bias parameter ÎČ to regulate affine transformations of multi-layer features, thus controlling the content of generated videos. These parameters are also zero-initialized, which ensures training stability while exerting incremental influence of condition information. (4) and (5) present the detailed forward data processing flow, where f n denotes the n-th feature of an input sample; the normalized feature value Ì f n is scaled and shifted under the modulation of Îł and ÎČ to produce the final feature f n out : Ì f n = f n â ÎŒ â Ï 2 + Δ ,(4) f n out = γ· Ì f n + ÎČ.(5) The scaling mechanism of α is identical to that of Îł. During training, as ColoDiff progressively fits the noise patterns specific to different categories, the prototypes become infused with class-discriminative information. To summarize, the noise-injected video embedding en- hances fine-grained visual perception beyond noise level, while the learnable prototypes offer more flexible and class- discriminative representations. By integrating them together, our Content-Aware module enables synthetic videos to accu- rately reflect controllable clinical attributes. C. ColoDiff With Non-Markovian Sampling We build ColoDiff on a state-of-the-art (SOTA) diffusion- based generative model [9], [38], which learns data distribution through a gradual diffusion and denoising process. The frame- work involves a forward process q that incrementally corrupts data by adding Gaussian noise over multiple time steps and a reverse process p Ξ that iteratively denoises samples to recover data similar to the original distribution, where Ξ represents learnable model parameters. Specifically, at a time step t in the forward process, a noisier image x t is calculated from a cleaner image x tâ1 by adding noise: q(x t |x tâ1 ) =N (x t | p 1â ÎČ t x tâ1 ,ÎČ t I),(6) where N denotes the Gaussian distribution, I denotes the identity matrix, and ÎČ t â (0, 1) is a step-varying hyper- parameter. The reverse process uses a noise estimator, Δ Ξ (x k ,k), which predicts the noise within a noisy image to enable its reconstruction into a clean image. Diffusion models typically rely on a Markovian reverse process that denoises data over hundreds of steps for high- quality generation. While acceptable for images, this is com- putationally prohibitive for videos. ColoDiff addresses this by implementing a non-Markovian reverse process [38]. The model first uses its current state x k and the time step k to predict an estimate of the clean image Ë x 0 , which is a linear combination of the noisy input x k and the predicted noise Δ Ξ (x k ,k). Crucially, with this estimate Ë x 0 in hand, the reverse process can then reconstruct any previous state x s (where s < k) in a single step. The integration of non-Markovian strategy allows the sampler to jump between non-adjacent time steps, dramatically accelerating inference [38]. As a result, ColoDiff can generate high-quality videos in real-time by using a drastically reduced number of sampling steps. Additionally, we choose to perform computation in the latent space [14]. It means that during the earlier process, the input video x 0 is first encoded into a low-dimensional space as z 0 by a pre-trained VAE [16] with encoder E . Similarly, in the later process, the final inference will be decoded back to the image space with a deterministic pass through decoder D. Thus, the final loss function of ColoDiff with conditioning regulation can be formulated as [9], [14], [39]: L ColoDiff = E z 0 ,ΔâŒN(0,1),t h â„Δâ Δ Ξ (z t , c,t)â„ 2 2 i ,(7) where t is uniformly sampled from 1,...,T, c denotes the embedded control information from Embed(z i t ) and learnable prototypes, Δ represents the target noise, and Δ Ξ is the model that learns noise. The skip-step sampling enables ColoDiff to drastically reduce inference steps while maintaining generation quality, and the latent representation further improves computational efficiency. These designs make real-time colonoscopy video generation feasible in clinical settings. IV. EXPERIMENTS AND RESULTS A. Experimental Settings 1) Datasets:We collect a total of 4,597 labeled colonoscopy video clips from three public datasets and one hospital database, each corresponding to different tasks. The Colonoscopic dataset contains 152 original videos, annotated with both disease types and imaging modalities, making it well-suited for disease classification and modality discrimina- tion [25]. The HyperKvasir dataset, with 373 original videos, provides annotations of disease types together with Boston Bowel Preparation Scale (BBPS) scores, enabling both disease classification and bowel preparation scoring [2]. Specifically, 6IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2025 TABLE I QUANTITATIVE COMPARISON WITH SOTA ALGORITHMS. BEST RESULTS ARE IN BOLD. TypeMethod Colonoscopic [25]HyperKvasir [2]SUN-SEG [4]Hospital Data FVDâFIDâISâFVDâFIDâISâFVDâFIDâISâFVDâFIDâISâ GAN-based StyleGAN-V [20]2111226.12.1272988.51.9868246.53.9170351.73.42 MoStGAN-V [21]46953.23.3760729.42.7441228.64.0452246.63.59 Diffusion-based LVDM [18]103796.91.9352849.72.1538735.93.7741640.33.12 Endora [22]46113.43.9049619.63.2937517.84.5938923.23.65 FEAT-L [23]35112.3 4.0151121.23.2735613.64.6140218.93.88 ColoDiff (Ours)33912.73.9547316.3 3.4629411.9 4.7833615.7 4.14 the disease labels cover colitis, polyp, and adenoma, while the modality labels distinguish between white-light and narrow- band imaging (WLI vs. NBI). WLI provides a brighter view for clinicians during screening, while NBI displays clearer mu- cosal and microvascular structures during diagnosis [40]. The bowel preparation scores are categorized into BBPS 0-1 and BBPS 2-3, reflecting varying levels of intestinal cleanliness. The SUN-SEG dataset focuses on dense prediction, offering 285 original videos with frame-by-frame segmentation masks, resulting in 49,136 frames for pixel-level annotation [4]. In addition, we collect 578 original videos from Fudan University Shanghai Cancer Center. This hospital database was approved by the Ethics Committee of Fudan University Shanghai Cancer Center, Shanghai, China (No. 2509-Exp283) in 2025. The example videos have been displayed in Fig. 1(a)-(b). For comparison with SOTA methods, experiments are con- ducted on each dataset to make a fair evaluation. For ablation and downstream tasks, experiments are conducted on specific tasks. Taking modality discrimination task as an example, the used WLI and NBI videos are all from Colonoscopic dataset. Regarding data allocation, Colonoscopic, HyperKvasir, and hospital datasets follow an 8:2 split, with 80% videos ran- domly selected for training and the rest for evaluation. The SUN-SEG dataset is separated according to official guidelines. 2) Implementation Details: We conduct all experiments us- ing the PyTorch framework and employ two NVIDIA A100 GPUs for training. The input videos are uniformly adjusted to 128Ă128 with 16 frames. We use the pre-trained VAE based on LDM [14], with a latent space of 16Ă16Ă4. During the training process, the upper limit of iterations is set to be 500,000 with an early stop mechanism to avoid overfitting. We choose a batch size of 32, a learning rate of 1Ă10 â4 , and the AdamW optimizer. Additionally, we implement an expo- nential moving average (EMA) strategy [32], [34] for model parameters, ensuring training stability and output consistency. Prototype vectors match the input category count in number, with each vector dimension configured to 1,024. With these settings, ColoDiff converges in an average training time of 6.5 hours. Hyper-parameters for other compared algorithms follow the specifications in [18], [20]â[23]. 3) Evaluation Metrics: For comparison and ablation experi- ments, we adopt metrics including Fr Ì echet Inception Distance (FID) [41], Fr Ì echet Video Distance (FVD) [42], and Inception Score (IS) [43]. To be specific, FID evaluates the fidelity of generated data, computing the Fr Ì echet distance between two multi-variate Gaussian distributions fitted to deep features of real and generated images: FID =â„ÎŒ r â ÎŒ g â„ 2 2 + Tr(ÎŁ r + ÎŁ g â 2(ÎŁ r ÎŁ g ) 1/2 ),(8) where (ÎŒ r , ÎŁ r ) and (ÎŒ g , ÎŁ g ) are the mean and covariance of the real and generated image features, respectively. FVD extends this idea to the video domain. By employing the same scheme but calculating mean and covariance from spatio- temporal features, it evaluates the temporal consistency of gen- erated videos. Additionally, IS goes beyond fidelity, measuring the diversity and controllability of generated data. It calculates the KL divergence between conditional class distributions and marginal class distributions: IS = exp E xâŒp g [D KL (p(y|x)â„ p(y))] ,(9) where p g denotes the distribution of generated samples, p(y|x) is the class distribution of the generated data x, and p(y) is the marginal class distribution. Lower FID and FVD values indicate better fidelity and temporal consistency, while higher IS values reflect greater diversity and controllability. For downstream tasks, we use precision, recall, F1-score, and accuracy for classification; as well as Dice, mIoU, pre- cision, and recall for segmentation. Together, these metrics comprehensively assess both the generative quality of ColoDiff and its performance in supporting clinical downstream tasks. B. Comparison with SOTA Methods We reproduce five SOTA generative algorithms for colonoscopy video synthesis [18], [20]â[23]. Among them, StyleGAN-V [20] and MoStGAN-V [21] are GAN-based approaches that achieve high temporal coherence through con- tinuous temporal encoding and motion disentanglement. The other three methods are diffusion-based. LVDM [18] creatively introduces conditional latent perturbation and unconditional guidance to mitigate error accumulation, thereby producing high-fidelity videos. Endora [22] and FEAT-L [23] are spe- cially tailored for endoscopic video synthesis, demonstrating superior capability in modeling complex spatio-temporal dy- namics and capturing surgical scene representations. Table I presents the quantitative comparison across four datasets. ColoDiff demonstrates comprehensive superiority over GAN-based approaches [20], [21], with better FVD, FID, and IS scores. Specifically, ColoDiff achieves a 339 FVD on Colonoscopic, lower than StyleGAN-V [20] and MoStGAN-V [21], while also reducing FID to 12.7 and FU et al.: ColoDiff: INTEGRATING DYNAMIC CONSISTENCY WITH CONTENT AWARENESS7 TABLE I ABLATION EXPERIMENT RESULTS FOR TIMESTREAM AND CONTENT-AWARE MODULES. BEST RESULTS ARE IN BOLD. CategoryColitisPolypAdenoma MetricFVDâFIDâISâFVDâFIDâISâFVDâFIDâISâ Temporal Coherence Transformer Encoder Layer51829.13.3750122.93.6165532.93.15 TimeStream Module37216.4 3.9231611.8 4.0832817.5 3.79 Content Controllability One-hot Encoding49120.93.7143014.33.8558125.33.46 Random Encoding48621.63.6444615.13.7956725.83.50 Learnable Prototypes39017.73.8134113.63.8539620.83.55 Content-Aware Module37216.4 3.9231611.8 4.0832817.5 3.79 (a) Compared Methods: Uncontrollable(b) ColoDiff: Controllable S t y l e G A N - V L V D M M o S t G A N - V E n d o r a F E A T â B B P S 2 - 3 â â A d e n o m a â â N B I â â W L I â â P o l y p â Visual Distortion Limited Variation Incoherence Incoherence IncoherenceIncoherenceIncoherence Limited VariationLimited VariationLimited VariationLimited Variation IncoherenceIncoherence IncoherenceIncoherence Visual Distortion Fig. 3. Visual comparison of generated videos on Colonoscopic dataset [25]. (a) Videos generated by other compared methods. The blue boxes indicate regions that exhibit corresponding issues. Row 1: anatomical and textural visual distortion. Row 2: limited content variation. Rows 3-5: inter-frame incoherence. (b) Videos generated by ColoDiff. improving IS to 3.95. Similar trends can be observed on HyperKvasir, SUN-SEG, and the hospital database, confirm- ing that ColoDiff produces colonoscopy videos well aligned with real distribution. In addition, ColoDiff also pushes the boundary of diffusion-based methods [18], [22], [23]. On the hospital database, it achieves an FVD of 336, substantially outperforming the recently proposed Endora [22] and FEAT- L [23], and the results demonstrate ColoDiffâs generalizability to handle complex clinical scenarios. The lower FVD values demonstrate a superior temporal coherence of ColoDiff, while improvements in FID and IS further validate its ability to generate colonoscopy videos with high fidelity and diversity. As shown in Fig. 3(a), compared methods suffer from visual distortion, limited content variation, and temporal incoherence. The related regions are marked with blue boxes. To be specific, Row 1 displays both anatomical and textural visual distortion, presenting deformed lesion shape and discontinuous vascular topology. Row 2 shows almost no shift in viewpoint or illu- mination across frames. Rows 3-4 reveal lesions that pop into existence, while Row 5 shows an abrupt lesion disappearance. In contrast, videos in Fig. 3(b) exhibit improved coherence and controllability. This performance stems from the TimeStream module, which decouples temporal dependencies to capture complex intestinal dynamics, and the Content-Aware module, which enhances fine-grained, class-discriminative representa- tion. Additional examples in the Supplementary Material further demonstrate the diversity, coherence, and controllabil- ity of ColoDiff-generated videos. C. Ablation Experiments Ablation experiments validate the function of TimeStream and Content-Aware modules from both temporal coherence and content controllability. For temporal coherence, all the compared implementations are combined with Content-Aware module. The baseline is pure Transformer encoder layers. The âTimeStream Moduleâ row means employing interlaced Trans- former encoder layers and TimeStream modules. For content controllability, all the compared implementations are com- bined with TimeStream module. Different category represen- tation approaches are compared, including one-hot encoding, random encoding, and learnable prototypes. These approaches retain the modelâs customized content generation ability. The âContent-Aware Moduleâ row means introducing fine-grained guidance besides learnable prototypes. The evaluation metrics 8IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2025 TABLE I ABLATION EXPERIMENT RESULTS FOR NON-MARKOVIAN SAMPLING STRATEGY. THE INFERENCE TIME COLUMN INDICATES THE TIME REQUIRED TO GENERATE A 16-FRAME VIDEO. TimeIfFIDâInferenceFrame StepsMarkovianColitisPolypAdenomaTime (s)Rate T =250Yes16.411.817.513.461.19 T =100No16.712.317.95.173.09 T =50No17.112.518.42.616.13 T =10No17.513.118.70.4932.65 T =5No19.314.221.60.2466.67 (a) Resolution: 128 Ă 128 (b) Resolution: 640 Ă 512 Fig. 4.Examples of videos generated by ColoDiff under different resolution settings. For uniform display, slight stretching deformation is applied to the original images. include FVD, FID, and IS scores across disease categories to assess video generation quality. Table I shows the ablation results. In terms of temporal coherence, TimeStream module achieves remarkable improve- ments compared with pure Transformer encoder layers. It reduces FVD score by over 20% across all categories, e.g., from 518 to 372 for colitis. The substantial decrease in FID scores, e.g., from 32.9 to 17.5 for adenoma, also indicates that generated videos achieve better visual fidelity. These results prove that TimeStream module successfully decouples tempo- ral dependencies, capturing motion patterns during endoscope movement and enhancing coherence of colonoscopy videos. In terms of content controllability, one-hot encoding and random encoding showcase comparable performance, indicat- ing the modelâs potential to generate controllable contents based on specific encodings. Prototype learning constructs a learnable feature vector for each category, helping the model reduce FVD for three diseases to below 400. It demonstrates that learnable vectors offer more customized representations compared to fixed encodings. Furthermore, after the noise- injected video embeddings are introduced as fine-grained guidance, the model performs better across all categories and metrics, even achieving a 4.08 IS score for polyp videos. These experimental results substantiate the utility of TimeStream and Content-Aware modules in enhancing inter-frame consistency and content controllability for colonoscopy videos. (a) Turing Test 56 714526 4541 1944403 1694428 2624335 1844413 2194378 1154482 Ground Truth GeneratedReal Generated Real GeneratedReal Generated Real GeneratedReal Generated Real GeneratedReal Generated Real 500 1000 1500 2000 2500 3000 3500 4000 Judgment of Clinician1Judgment of Clinician2Judgment of Clinician3Judgment of Clinician4 (b) Consistency Test Accuracy 0.985±0.008 0.947±0.017 0.967±0.015 0.993±0.004 0.990±0.004 0.960±0.012 0.975±0.006 Disease DiagnosisModality DiscriminationBowel Preparation Scoring ColitisPolypAdenomaWLINBIBBPS 0-1BBPS 2-3 Fig. 5. The experimental results of Turing test and Consistency test. In addition, we conduct ablation experiments on the non- Markovian sampling strategy employed by ColoDiff. The number of diffusion time steps are scanned from 250 to 100, 50, 10, and 5. As reported in Table I, even when the sampling steps are reduced to 50 or 10, the FID scores across all categories do not exhibit an obvious deterioration. For the 10-step setting, ColoDiff averagely generates 32.65 frames per second for 128Ă128 resolution. Moreover, the Transformer-based design enables ColoDiffâs inherent scala- bility to fit higher resolution and longer sequence, leading to 26.23 frames per second for 640Ă512 resolution with half- precision inference, which exceeds the clinical requirements. Fig. 4 presents colonoscopy videos generated by ColoDiff at different resolution settings, where the 640Ă512 resolution yields clearer mucosal and microvascular structures for visual perception. These results indicate that ColoDiff effectively bal- ances inference time and generation quality, pushing diffusion- based video generation toward real-time performance. D. Clinical Evaluation In the clinical domain, professional evaluation and risk assessment are critical to ensuring the reliability of synthetic data. To address this necessity, we invite four clinicians from hospitals to conduct both Turing and Consistency tests. For Turing test, participants are tasked with distinguishing between real and synthetic colonoscopy videos provided at the same resolution. The test set comprises 4,597 real colonoscopy videos and 4,597 synthetic videos generated by ColoDiff, where the ratio is not disclosed to clinicians in advance. For Consistency test, clinicians perform three classification assess- ments based only on synthetic videos, including disease diag- nosis, modality discrimination, and bowel preparation scoring. 1,000 synthetic videos are generated for each category of each task, and the consistency between cliniciansâ judgments and ColoDiffâs predefined control conditions is quantified to evaluate content controllability. Fig. 5(a) and Fig. 5(b) separately display the results of Turing test and Consistency test. In Fig. 5(a), the left two columns show the judgments of two junior clinicians, while the right two columns represent those of two senior clinicians. For the junior clinicians, nearly half of the videos they identify as âgeneratedâ are actually real (71 vs. 56, 194 vs. 169; FU et al.: ColoDiff: INTEGRATING DYNAMIC CONSISTENCY WITH CONTENT AWARENESS9 TABLE IV QUANTITATIVE RESULTS FOR DOWNSTREAM CLASSIFICATION TASKS (UNIT: %). FOR PAIRED METRICS, MODELS TRAINED WITH REAL VIDEOS (LEFT) vs. REAL + SYNTHETIC VIDEOS (RIGHT) ARE COMPARED. BETTER RESULTS ARE IN BOLD. Metric DiseaseModalityBBPS ColitisPolypAdenomaWLINBI0-12-3 Original +Synthetic Original +Synthetic Original +Synthetic Original +Synthetic Original +Synthetic Original +Synthetic Original +Synthetic Precisionâ79.2 85.985.3 92.873.5 81.088.4 93.793.1 96.886.7 90.387.5 93.8 Recallâ82.1 87.487.7 89.176.8 83.491.2 96.390.6 93.789.2 92.685.3 91.8 F1-scoreâ80.6 86.686.5 90.975.1 82.289.8 95.091.8 95.287.9 91.486.4 92.8 Accuracyâ79.8 85.885.4 91.774.3 83.190.2 94.491.8 95.188.5 92.186.6 92.5 bold values indicate real videos). Although the senior clini- cians exhibit stronger discrimination ability by identifying 262 and 219 synthetic videos respectively, the strictest clinician (3 rd column) still misclassifies over 94.30% of the synthetic videos as real (4,335 out of 4,597). These results confirm that videos generated by ColoDiff are sufficiently realistic to âpass for realâ from the perspective of medical professionals. In Fig. 5(b), the consistency between cliniciansâ assessment and ColoDiffâs specified condition further validates the con- trollability of ColoDiff. Specifically, the average accuracy of four clinicians in disease diagnosis, modality discrimination, and bowel preparation scoring based on synthetic videos is comparable to their performance when evaluating real clinical data. Even for the âpolypâ category, where the lowest accuracy was observed in all three tasks, the consistency accuracy still reaches a high value of 0.947 ± 0.017. The experimental results indicate that synthetic videos effectively capture the represen- tative clinical features of different categories, verifying both the fidelity and controllability of ColoDiffâs output. E. Downstream Task Experiments Benefiting from the Content-Aware module, ColoDiff is capable of generating colonoscopy videos with specified cat- egories, which distinguishes it from other comparative algo- rithms. Thus, ColoDiffâs synthetic videos can be integrated into the original training data for fully supervised re-training towards various downstream tasks. We conduct classification and segmentation tasks as downstream evaluations to demon- strate the potential applications of ColoDiff. In the experi- ments, we uniformly adopt 3D ResNet [44], [45] for classi- fication and SALI [46], a network dedicated to colonoscopy video dense prediction, for lesion segmentation. SALI uses a 2D architecture to handle video-based segmentation, explic- itly modeling inter-frame temporal correlations via short-term alignment and long-term interaction [46]. 1) Downstream Classification Tasks: Table IV summarizes the results of three classification tasks, including disease diag- nosis, modality discrimination, and bowel preparation scoring (BBPS). The results are reported in pairs: left for models trained only with original videos, and right for those trained with equal synthetic videos added. We find that ColoDiff- synthesized videos consistently improve classification perfor- mance for all categories of the three tasks. Especially for disease diagnosis, ColoDiff achieves an average accuracy im- provement of 7.1% (from 79.8% to 86.9%). Additionally, we (a) Disease Diagnosis (c) Bowel Preparation Scoring (b) Modality Discrimination Original Training Set Original Training Set + Equal Synthetic Data Fig. 6.UMAP visualization for downstream classification tasks on test sets, with the left column representing models trained with original training set, and the right column representing models trained with an equal number of synthetic videos added. perform the uniform manifold approximation and projection (UMAP) [47], a dimension-reduction technique, to map the classifierâs last-layer features into a 2D coordinate system. The left column of Fig. 6 shows that point clusters from different categories have noticeable overlapping areas. In contrast, after incorporating generated videos into training, these clusters are dispersed much better in the right column of Fig. 6. The visualization results indicate that generated videos enhance feature robustness and inter-class discrimination, owing to their similarity to real data, diffusion-induced randomness, and re-balancing of the training distribution [48]. 10IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2025 TABLE V QUANTITATIVE RESULTS FOR DOWNSTREAM SEGMENTATION TASKS (UNIT: %). FOR PAIRED METRICS, MODELS TRAINED WITH REAL VIDEOS (LEFT) vs. REAL + SYNTHETIC VIDEOS (RIGHT) ARE COMPARED. BETTER RESULTS ARE IN BOLD. Metric Test-Easy (Seen)Test-Easy (Unseen)Test-Hard (Seen)Test-Hard (Unseen) Original+SyntheticOriginal+SyntheticOriginal+SyntheticOriginal+Synthetic Diceâ91.293.681.386.584.590.772.984.1 mIoUâ89.391.580.485.288.192.677.486.8 Precisionâ89.993.380.883.184.996.178.086.3 Recallâ92.593.981.890.284.185.968.482.0 PolypAdenoma Label Decision data for training Re-decision â AdenomaPolyp PolypAdenomaPrototype â PolypAdenoma Colitis Polyp Adding generated â Colitis Colitis â Fig. 7. Visualized examples of downstream classification performance. The upper part of Fig. 7 shows real cases the classifier initially misclassified. From the constraint of Content-Aware module, ColoDiff can generate videos adhering to specific real-data patterns. With more generated data incorporated into training, the classifier may be guided to make correct re-decisions. Notably, combining category prototypes with diffusion-induced randomness enables ColoDiff to generate non-typical representations, such as adenoma-like polyps, which can challenge the classifier to refine boundary decisions. 2) Downstream Segmentation Tasks: We also perform seg- mentation experiments on SUN-SEG dataset [4]. According to the official distribution: the test set is categorized into Easy and Hard levels based on difficulty; Seen indicates that test images and some training images are from the same video sequence but different frames, while Unseen signifies that test images derive entirely from distinct video sequences. From Table V, the synthetic-augmented training set im- proves the average Dice from 82.5% to 88.7%, with a 6.2% increase. Among these results, the performance on Unseen subsets is more indicative of the modelâs generalization ca- pability. There also exist two notable phenomena. First, for Easy data, models exhibit better Dice than mIoU; but for Hard data, the situation is reversed. This suggests that edge segmen- tation for Hard images is unsatisfactory; although the model locates lesion areas, it fails to accurately perceive the details. In contrast, after incorporating synthetic data into training, the gap between Dice and mIoU for Hard images narrows, indicating an enhancement in the segmentation modelâs ability to capture lesion edges and textures. Second, segmentation Time Fig. 8. Representative examples of downstream segmentation perfor- mance. Translucent blue regions indicate the ground truth, red contours denote results without generated data for training, and green contours show results after integrating generated data for training. performance on Unseen data is much worse than on Seen data, as evidenced by 72.9% Dice for Unseen data vs. 84.5% for Seen data (Hard level). After training with added synthetic data, the Dice score improves to 84.1% for Unseen data vs. 90.7% for Seen data, which to some extent reduces this gap and improves the modelâs robustness to Unseen data. Fig. 8 exhibits sequence-based segmentation examples be- fore and after incorporating synthetic training data. As can be seen, the inclusion of generated data enables the model to segment lesion boundaries more robustly under interference factors like diverse backgrounds, motion artifacts, and specular reflections. The visualization results confirm that generated data facilitates lesion representation learning. Specifically, the randomness of diffusion process, together with the similarity of generated samples to real data, introduces challenging near- boundary cases. By incorporating these cases, the synthetic data re-balances the distribution, regularizes the feature space, and enhances robustness [48], leading to performance gains especially on hard and unseen test conditions. V. DISCUSSION Colonoscopy video generation faces challenges of complex temporal modeling, customized content control, and restricted FU et al.: ColoDiff: INTEGRATING DYNAMIC CONSISTENCY WITH CONTENT AWARENESS11 inference speed. To address these problems, we propose a diffusion-based network with TimeStream and Content-Aware modules, namely, ColoDiff. Experiments show that ColoD- iff is capable of generating temporal-coherent and content- controllable videos in real-time scenarios. A. Towards Temporal-Coherent Colonoscopy Video Generation Temporal-coherent colonoscopy video generation is crucial for supplementing high-quality data and assisting diagnosis. As shown in Table I, StyleGAN-V [20] and MoStGAN-V [21] show poor FVD scores due to inadequate temporal modeling and training instability. Diffusion-based LVDM [18] employs 3D U-Net to capture spatio-temporal motion patterns. Con- strained by CNNâs locality bias, it does not achieve long-range dependency in temporal dimension, leading to inter-frame inconsistency (Row 3, Fig. 3(a)). Although other diffusion- based methods employ Transformer architectures [22], [23], the proposed strategies like fully cross-frame interaction ex- cessively use the modelâs scaling ability [24], leading to the âforgettingâ problem like abrupt lesion appearance or disappearance (Rows 4-5, Fig. 3(a)). In contrast, ColoDiff achieves the best FVD scores across all datasets, e.g., 294 on SUN-SEG dataset, which is 17.4% lower than the best compared method [23]. The results indi- cate that generated videos maintain high temporal coherence even in the presence of irregular intestinal structures. This advantage is largely attributed to our TimeStream module, which treats spatially aligned patches as sequential input tokens. By leveraging the prior knowledge that the same anatomical structure generally occupies consistent or adjacent spatial locations across consecutive frames, the TimeStream module enables accurate modeling of irregular morphology and complex motion. Moreover, by equipping 2D architec- tures with 3D contextual reasoning capabilities, it empowers ColoDiff to work efficiently without increasing model scale or computational cost. B. Towards Content-Controllable Colonoscopy Video Generation Content-controllable generation facilitates customized data synthesis, thus benefiting various downstream tasks. Based on Table I, prototype learning outperforms one-hot and random encoding, indicating the superiority of learnable embeddings for multi-pattern conditioning. This is because prototype learn- ing assigns a unique but adjustable vector to each category. As ColoDiff progressively learns to fit noise patterns specific to different categories, the vectors become class-discriminative. Meanwhile, standard diffusion models only rely on the time- step index t to perceive noise levels, yielding coarse con- ditioning without spatio-temporal information. To overcome this, noise-injected data embedding is further employed as a fine-grained regulation with intra-frame spatial information. Integrating prototypes with noise-injected embeddings, the Content-Aware module refines ColoDiffâs multi-layer features, pushing the performance boundary of polyp-IS to 4.08. According to Tables IV and V, the randomness of diffusion process, combined with the similarity of generated samples to real data, yields feature variations while preserving anatomical fidelity. Incorporating such data into training successfully im- proves downstream task performance, confirming that ColoD- iff produces diverse yet controllable colonoscopy videos. This can be further supported by the UMAP visualization in Fig. 6, where feature clusters from different categories become more separated after incorporating synthetic data. The application of customized prototypes to provide specific feature rep- resentations for different data patterns can be adopted by other generative models. The use of noise-injected embedding as supplementary information to time-step index t, thereby enhancing the model controllability, can also be applied to other diffusion-based models. C. Future Work In our future work, improving the controllability of gen- erative models remains a key focus. Our ambition is to achieve concurrent regulation over multiple variables, includ- ing imaging modality, disease type, etc. Additionally, we aim to compile comprehensive video-text datasets to facilitate more precise control over colonoscopy video generation through multi-modal alignment. VI. CONCLUSION This paper proposes ColoDiff, a diffusion-driven approach with TimeStream and Content-Aware modules for colonoscopy video generation. ColoDiff successfully decouples tempo- ral dependency across frames and precisely controls clini- cal attributes within frames. The non-Markovian sampling strategy further enables real-time synthesis. Comprehensive experiments on multiple benchmarks and downstream tasks confirm ColoDiffâs ability to generate realistic, coherent, and controllable colonoscopy videos. Incorporating synthetic data into training improves disease diagnosis accuracy by 7.1% and lesion segmentation Dice by 6.2%. Our research presents an exploration in leveraging generated videos for colonoscopy analysis, paving the way for future integration of synthetic data into clinical practice. REFERENCES [1] H. Sung, J. Ferlay, R. L. Siegel, M. Laversanne, I. Soerjomataram, A. Je- mal, and F. Bray, âGlobal cancer statistics 2020: Globocan estimates of incidence and mortality worldwide for 36 cancers in 185 countries,â CA: A Cancer Journal for Clinicians, vol. 71, no. 3, p. 209â249, 2021. [2] H. Borgli, V. Thambawita, P. H. Smedsrud, S. Hicks, D. Jha, S. L. Eskeland, K. R. Randel, K. Pogorelov, M. Lux, D. T. D. Nguyen et al., âHyperkvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy,â Scientific Data, vol. 7, no. 1, p. 283, 2020. [3] A. H. Calderwood and B. C. Jacobson, âComprehensive validation of the boston bowel preparation scale,â Gastrointestinal Endoscopy, vol. 72, no. 4, p. 686â692, 2010. [4] G.-P. Ji, G. Xiao, Y.-C. Chou, D.-P. Fan, K. Zhao, G. Chen, and L. Van Gool, âVideo polyp segmentation: A deep learning perspective,â Machine Intelligence Research, vol. 19, no. 6, p. 531â549, 2022. [5] J. Fu, K. Chen, Q. Dou, Y. Gao, Y. He, P. Zhou, S. Lin, Y. Wang, and Y. Guo, âIpnet: An interpretable network with progressive loss for whole-stage colorectal disease diagnosis,â IEEE Transactions on Medical Imaging, vol. 44, no. 2, p. 789â800, 2025. 12IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2025 [6] W. Ma, Y. Zhu, R. Zhang, J. Yang, Y. Hu, Z. Li, and L. Xiang, âToward clinically assisted colorectal polyp recognition via structured cross-modal representation consistency,â in International Conference on Medical Image Computing and Computer-Assisted Intervention, 2022, p. 141â150. [7] K. Yao, Y. Takaki, T. Matsui, A. Iwashita, G. K. Anagnostopoulos, P. Kaye, and K. Ragunath, âClinical application of magnification en- doscopy and narrow-band imaging in the upper gastrointestinal tract: new imaging techniques for detecting and characterizing gastrointesti- nal neoplasia,â Gastrointestinal Endoscopy Clinics of North America, vol. 18, no. 3, p. 415â433, 2008. [8] J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, âDeep unsupervised learning using nonequilibrium thermodynamics,â in International Conference on Machine Learning, 2015, p. 2256â2265. [9] J. Ho, A. Jain, and P. Abbeel, âDenoising diffusion probabilistic models,â Advances in Neural Information Processing Systems, vol. 33, p. 6840â 6851, 2020. [10] V. Sharma, A. Kumar, D. Jha, M. K. Bhuyan, P. K. Das, and U. Bagci, âControlpolypnet: towards controlled colon polyp synthesis for improved polyp segmentation,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, p. 2325â2334. [11] Y. Xie, J. Wang, T. Feng, F. Ma, and Y. Li, âCcis-diff: A generative model with stable diffusion prior for controlled colonoscopy image synthesis,â in 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI), 2025, p. 1â5. [12] V. Sharma, D. Jha, M. Bhuyan, P. K. Das, and U. Bagci, âDiverse image generation with diffusion models and cross class label learning for polyp classification,â arXiv:2502.05444, 2025. [13] M. Chaichuk, S. Gautam, S. Hicks, and E. Tutubalina, âPrompt to polyp: Medical text-conditioned image synthesis with diffusion models,â arXiv:2505.05573, 2025. [14] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, âHigh- resolution image synthesis with latent diffusion models,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2022, p. 10 684â10 695. [15] P. Esser, R. Rombach, and B. Ommer, âTaming transformers for high- resolution image synthesis,â in Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2021, p. 12 873â 12 883. [16] D. P. Kingma and M. Welling, âAuto-encoding variational bayes,â arXiv:1312.6114, 2013. [17] J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, âVideo diffusion models,â Advances in Neural Information Processing Systems, vol. 35, p. 8633â8646, 2022. [18] Y. He, T. Yang, Y. Zhang, Y. Shan, and Q. Chen, âLatent video diffusion models for high-fidelity long video generation,â arXiv:2211.13221, 2022. [19] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, âGenerative adversarial nets,â Advances in Neural Information Processing Systems, vol. 27, 2014. [20] I. Skorokhodov, S. Tulyakov, and M. Elhoseiny, âStylegan-v: A con- tinuous video generator with the price, image quality and perks of stylegan2,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, p. 3626â3636. [21] X. Shen, X. Li, and M. Elhoseiny, âMostgan-v: Video generation with temporal motion styles,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, p. 5652â5661. [22] C. Li, H. Liu, Y. Liu, B. Y. Feng, W. Li, X. Liu, Z. Chen, J. Shao, and Y. Yuan, âEndora: Video generation models as endoscopy simulators,â in International Conference on Medical Image Computing and Computer- Assisted Intervention, 2024, p. 230â240. [23] H. Wang, Z. Yang, H. Zhang, D. Zhao, B. Wei, and Y. Xu, âFeat: Full-dimensional efficient attention transformer for medical video gen- eration,â arXiv:2506.04956, 2025. [24] Y. Zhang, Y. Wei, D. Jiang, X. Zhang, W. Zuo, and Q. Tian, âControlvideo: Training-free controllable text-to-video generation,â arXiv:2305.13077, 2023. [25] P. Mesejo, D. Pizarro, A. Abergel, O. Rouquette, S. Beorchia, L. Poincloux, and A. Bartoli, âComputer-aided classification of gastroin- testinal lesions in regular colonoscopy,â IEEE Transactions on Medical Imaging, vol. 35, no. 9, p. 2051â2063, 2016. [26] J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou, âTune-a-video: One-shot tuning of image diffusion models for text-to-video generation,â in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, p. 7623â7633. [27] U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni et al., âMake-a-video: Text-to-video generation without text-video data,â arXiv:2209.14792, 2022. [28] A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, âAlign your latents: High-resolution video synthesis with latent diffusion models,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, p. 22 563â22 575. [29] T. Zhu, D. Ren, Q. Wang, X. Wu, and W. Zuo, âGenerative inbetweening through frame-wise conditions-driven video generation,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2025, p. 27 968â27 978. [30] D. Danier, F. Zhang, and D. Bull, âLdmvfi: Video frame interpolation with latent diffusion models,â in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 2, 2024, p. 1472â1480. [31] X. Ma, Y. Wang, G. Jia, X. Chen, Z. Liu, Y.-F. Li, C. Chen, and Y. Qiao, âLatte: Latent diffusion transformer for video generation,â arXiv:2401.03048, 2024. [32] W. Peebles and S. Xie, âScalable diffusion models with transformers,â in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, p. 4195â4205. [33] L. Zhang, A. Rao, and M. Agrawala, âAdding conditional control to text-to-image diffusion models,â in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, p. 3836â3847. [34] F. Bao, S. Nie, K. Xue, Y. Cao, C. Li, H. Su, and J. Zhu, âAll are worth words: A vit backbone for diffusion models,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, p. 22 669â22 679. [35] W. Sun, X. You, R. Zheng, Z. Yuan, X. Li, L. He, Q. Li, and L. Sun, âBora: Biomedical generalist video generation model,â arXiv:2407.08944, 2024. [36] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ć. Kaiser, and I. Polosukhin, âAttention is all you need,â Advances in Neural Information Processing Systems, vol. 30, 2017. [37] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., âAn image is worth 16x16 words: Transformers for image recognition at scale,â arXiv:2010.11929, 2020. [38] J. Song, C. Meng, and S. Ermon, âDenoising diffusion implicit models,â arXiv:2010.02502, 2020. [39] P. Dhariwal and A. Nichol, âDiffusion models beat gans on image synthesis,â Advances in Neural Information Processing Systems, vol. 34, p. 8780â8794, 2021. [40] J. Fu, Y. Gao, P. Zhou, Y. Huang, J. Jiao, S. Lin, Y. Wang, and Y. Guo, âD 2 polyp-net: A cross-modal space-guided network for real-time col- orectal polyp detection and diagnosis,â Biomedical Signal Processing and Control, vol. 91, p. 105934, 2024. [41] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, âGans trained by a two time-scale update rule converge to a local nash equilibrium,â Advances in Neural Information Processing Systems, vol. 30, 2017. [42] T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, âTowards accurate generative models of video: A new metric & challenges,â arXiv:1812.01717, 2018. [43] T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, âImproved techniques for training gans,â Advances in Neural Information Processing Systems, vol. 29, 2016. [44] X. Du, Y. Li, Y. Cui, R. Qian, J. Li, and I. Bello, âRevisiting 3d resnets for video recognition,â arXiv:2109.01696, 2021. [45] K. He, X. Zhang, S. Ren, and J. Sun, âDeep residual learning for image recognition,â in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, p. 770â778. [46] Q. Hu, Z. Yi, Y. Zhou, F. Peng, M. Liu, Q. Li, and Z. Wang, âSali: Short- term alignment and long-term interaction network for colonoscopy video polyp segmentation,â in International Conference on Medical Image Computing and Computer-Assisted Intervention, 2024, p. 531â541. [47] L. McInnes, J. Healy, and J. Melville, âUmap: Uniform manifold ap- proximation and projection for dimension reduction,â arXiv:1802.03426, 2018. [48] Y. Sun, W. Tan, Z. Gu, R. He, S. Chen, M. Pang, and B. Yan, âA data-efficient strategy for building high-performing medical foundation models,â Nature Biomedical Engineering, p. 1â13, 2025.