Paper deep dive
A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation
Steven Landgraf, Joceline Hinz, Markus Ulrich
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 8/20/2026, 4:58:54 AM
Summary
This paper presents the first systematic evaluation of Uncertainty Quantification (UQ) methods applied to foundation models for semantic segmentation. The authors fine-tune a lightweight DPT decoder on the pretrained SAM2 encoder to create a baseline model and benchmark four UQ approaches: Monte Carlo Dropout (MCD), Deep Sub-Ensemble (DSE), Test-Time Augmentation (TTA), and Evidential Deep Learning (EDL). Evaluations are conducted on Cityscapes, NYUv2, and out-of-domain settings (Rainy/Foggy Cityscapes), analyzing trade-offs between segmentation accuracy, calibration, uncertainty quality, and inference time.
Entities (12)
Relation Signals (10)
SAM2-DPT → evaluatedon → NYUv2
confidence 98% · benchmark... across... NYUv2
SAM2-DPT → evaluatedon → Cityscapes
confidence 98% · benchmark... across Cityscapes
Deep Sub-Ensemble → evaluatedon → SAM2-DPT
confidence 95% · benchmark four representative UQ approaches... DSE attains the best calibration
Test-Time Augmentation → evaluatedon → SAM2-DPT
confidence 95% · TTA delivers the best calibration
Evidential Deep Learning → evaluatedon → SAM2-DPT
confidence 95% · EDL again provides the second-fastest inference time
SAM2-DPT → evaluatedon → Rainy-Cityscapes
confidence 95% · out-of-domain settings... Rainy-Cityscapes
Monte Carlo Dropout → evaluatedon → SAM2-DPT
confidence 95% · systematic evaluation of UQ methods applied to a foundation model... benchmark four representative UQ approaches
SAM2-DPT → evaluatedon → Foggy-Cityscapes
confidence 95% · out-of-domain settings... Foggy-Cityscapes
DPT → finetunedon →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Foundation models are increasingly breaking what seemed to be impossible not long ago by enabling unprecedented accuracy and cross-domain generalization. Yet their lack of interpretability, tendency to be overconfident, and sensitivity to real-world domain shifts pose critical challenges for safety- and mission-critical applications. Uncertainty quantification (UQ) offers a principled way to address these issues, but its integration into segmentation foundation models has yet to be explored. In this paper we present the first systematic evaluation of UQ methods applied to a foundation model for semantic segmentation. We fine-tune a lightweight DPT decoder on top of the pretrained SAM2 encoder to establish a simple yet competitive baseline and benchmark four representative UQ approaches - Monte Carlo Dropout, Deep Sub-Ensemble, Test-Time Augmentation, and Evidential Deep Learning - across Cityscapes, NYUv2, and two challenging out-of-domain settings. Our analysis compares segmentation accuracy, calibration, uncertainty quality, and inference time, revealing clear trade-offs between predictive performance, reliability, and computational cost. These results highlight both the promise and the current limitations of uncertainty-aware foundation models, pointing to the need for future work that jointly optimizes accuracy, robustness, and efficiency for real-world deployment.
Tags
Links
- Source: https://arxiv.org/abs/2608.18709v1
- Canonical: https://arxiv.org/abs/2608.18709v1
Trouble viewing inline? Open PDF directly →
Full Text
41,574 characters extracted from source content.
Expand or collapse full text
A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation Steven Landgraf Thanks: Corresponding Authors Joceline Hinz* Markus Ulrich Address: Institute of Photogrammetry and Remote Sensing (IPF), Karlsruhe Institute of Technology (KIT), Germany, (steven.landgraf, joceline.hinz, markus.ulrich)@kit.edu Abstract Foundation models are increasingly breaking what seemed to be impossible not long ago by enabling unprecedented accuracy and cross-domain generalization. Yet their lack of interpretability, tendency to be overconfident, and sensitivity to real-world domain shifts pose critical challenges for safety- and mission-critical applications. Uncertainty quantification (UQ) offers a principled way to address these issues, but its integration into segmentation foundation models has yet to be explored. In this paper we present the first systematic evaluation of UQ methods applied to a foundation model for semantic segmentation. We fine-tune a lightweight DPT decoder on top of the pretrained SAM2 encoder to establish a simple yet competitive baseline and benchmark four representative UQ approaches – Monte Carlo Dropout, Deep Sub-Ensemble, Test-Time Augmentation, and Evidential Deep Learning – across Cityscapes, NYUv2, and two challenging out-of-domain settings. Our analysis compares segmentation accuracy, calibration, uncertainty quality, and inference time, revealing clear trade-offs between predictive performance, reliability, and computational cost. These results highlight both the promise and the current limitations of uncertainty-aware foundation models, pointing to the need for future work that jointly optimizes accuracy, robustness, and efficiency for real-world deployment. keywordsUncertainty Quantification, Reliability, Robustness, Foundation Models, Semantic Segmentation 1 Introduction Semantic segmentation is a foundational machine vision task involving assigning class labels to every pixel in an image (34). Recently, segmentation models have evolved rapidly from convolutional neural networks (26) and vision transformers (15) to foundation models trained on internet-scale data (58). These models promise unprecedented accuracy and generalization, enabling zero-shot capabilities across diverse domains and marking a paradigm shift in semantic segmentation. However, even the most potent neural networks remain subject to critical limitations (2). Foundation models in particular can act as black boxes, lacking interpretability (12), failing to recognize out-of-domain inputs (39), producing systematically overconfident outputs (54), or being sensitive to adversarial perturbations (42). It goes without saying that these issues are particularly harmful in safety- and mission-critical applications, where segmentation errors can propagate to downstream decisions with severe consequences. Incorporating uncertainty quantification (UQ) into deep neural networks directly addresses many of the challenges outlined above. Reliable uncertainty estimates not only mitigate overconfidence and sensitivity to domain shifts but also enhance interpretability by indicating the model’s confidence and highlighting areas of potential error, thereby improving the deployability deep learning-based systems. For instance, in safety-critical applications like autonomous driving (33) or medical imaging (38), uncertainty-aware models can flag unreliable predictions to trigger human intervention or additional verification steps, reducing the risk of catastrophic failures. Despite these advantages of UQ, the integration of UQ into foundation models for semantic segmentation remains underexplored. By systematically evaluating how different UQ methods interact with these powerful models, we offer not only improved reliability and explainability but also bridge the gap between state-of-the-art research and real-world deployment. Our contributions can be summarized as follows: 1. We present the first systematic study of UQ methods applied to foundation models for semantic segmentation, evaluating Monte Carlo Dropout (MCD), Deep Sub-Ensemble (DSE), Test-Time Augmentation (TTA), and Evidential Deep Learning (EDL). 2. We establish a simple yet competitive baseline by fine-tuning a DPT head on top of the pretrained SAM2 encoder, and benchmark its in-domain and out-of-domain performance across multiple datasets. 3. We provide a comprehensive comparison between segmentation performance, calibration, uncertainty quality, and inference time for all methods, highlighting trade-offs between accuracy, reliability, and computational cost. 2 Related Work Our work lies at the intersection of semantic segmentation and UQ, focusing on adapting foundation models like the Segment Anything Model (SAM) (20). We first review advancements in segmentation, from CNNs to vision transformers and foundation models. We then discuss key deep learning UQ methods, including recent efforts to tailor them for foundation models. Finally, we highlight the research gap addressed by this paper. 2.1 Semantic Segmentation Convolutional Neural Networks (CNNs). The foundational work in deep learning-based segmentation began with fully convolutional networks, which enable end-to-end pixel-wise classification (30). Subsequent CNN methods introduced richer context and multi-scale features. For instance, DeepLab established dilated convolutions and atrous spatial pyramid pooling to capture image context at multiple scales (3). Similarly, encoder-decoder networks like U-Net used symmetric contracting and expanding paths with skip connections to recover even fine details (43). Together, these CNN-based approaches represented the state of the art for many years and have only recently been challenged by attention-based architectures. Vision Transformers (ViTs). Transformer architecture have been the de-facto standard for natural language processing for a long time (52). More recently, they have also been brought into the vision domain (9). In the context of semantic segmentation, SegFormer stands out with its hierarchical transformer encoder and a lightweight multilayer perceptron decoder, achieving impressive results and high efficiency (56). Mask2Former further unified semantic, instance and panoptic segmentation with a mask-classification transformer, setting new state-of-the-art results (6). Beyond these task-specific advances, the field has shifted toward foundation models trained on internet-scale data. For example, SAM was trained with over 1B masks on 11M images (20). As noted in a recent study (58), adapting these foundation models can yield superior segmentation performance and entirely new capabilities in terms of zero/few-shot segmentation, interactive prompting, or cross-domain generalization that were unseen until now. 2.2 Uncertainty Quantification Overview. A variety of techniques have been proposed to capture predictive uncertainty in deep learning models (32; 11; 21; 50; 51; 28; 36; 1). The most popular methods remain sampling-based due to their ease of use and effectiveness: for instance, MCD (11) treats dropout (49) as a Bayesian approximation. Similarly, Deep Ensembles (21) consist of multiple, individually trained models to obtain state-of-the-art uncertainty results. Unfortunately, these approaches require multiple forward passes and sometimes more training time, which incurs high computational cost. To mitigate this, a number of deterministic, more efficient, methods have been proposed (51; 28; 36; 25; 23). Likewise, EDL enables efficient uncertainty estimation by training the model to infer the parameters of a probability distribution with a single prediction (1). Uncertainty-aware Foundation Models. Being able to estimate the uncertainty of the output of large foundation models is an emerging research field. Recently, 24 studied the synthesis of UQ methods and foundation models in monocular depth estimation and explicitly note that extending this work to other tasks, such as semantic segmentation, is an open opportunity. A few early methods have begun to address UQ for SAM in medical imaging (18; 8; 57). Beyond, USAM proposes a post-hoc training method for additional multilayer perceptrons to estimate the expected uncertainty (19). 29 introduce SUM that quantifies uncertainty in SAM-generated pseudo-labels and uses it to enable uncertainty-aware fine-tuning of the model. Research Gap. While all of these works have made progress toward uncertainty-aware foundation models in semantic segmentation with tailor-made approaches, they also reveal a critical gap: There has yet to be a systematic study of existing UQ methods in combination with SAM, despite its status as a de-facto foundation model for large-scale segmentation. 3 Methodology In this section, we describe how we derive our baseline model from the Segment Anything Model 2 (SAM2) foundation model and how we combine four UQ techniques with it. 3.1 Baseline Model We employ a hybrid architecture that combines the image encoder of the SAM2 (41) with the DPT decoder (40) for semantic segmentation (see Fig. 1). The motivation for this design is twofold: First, by leveraging the SAM2 encoder, which builds on the hierarchical Vision Transformer Hiera (44) pretrained with masked autoencoding (16), we directly harness state-of-the-art foundation model knowledge in the form of robust, multi-scale image representations. Second, coupling this encoder with the DPT decoder allows us to rely on a simple and widely adopted architecture for dense prediction, ensuring interpretability, comparability, and computational efficiency. The SAM2 encoder partitions the image into patches, projects them into tokens with positional embeddings, and processes them through four hierarchical transformer stages, producing multi-scale feature maps. These are fused by the DPT decoder through RefineNet-based (27) fusion modules that progressively upsample and align resolution, before passing them to a lightweight segmentation head. The head simply applies a 3x3-convolution, injects a dropout layer, followed by a linear projection, and lastly uses bilinear upsampling to restore the original image resolution. Training the model is equally simple as its architectural layout: we simply fine-tune the model with the regular cross-entropy loss ℒCE=−1N∑n=1N∑c=1Cyn,clog(p(y^n,c),L_CE=- 1NΣ^N_n=1Σ^C_c=1y_n,c (p( y_n,c) 5.0pt, (1) where ℒCEL_CE represents the loss for a single image, N is the number of pixels, C is the number of classes, yn,cy_n,c is the one-hot encoded ground truth label, and p(y^n,c)p( y_n,c) is the predicted softmax pseudo-probability. Overall, our intent was to maximize the benefits of foundation model pretraining while keeping the rest of the architecture as simple as possible. Figure 1: Schematic illustration of the used model architecture. 3.2 Uncertainty Quantification Our central research question is how to make segmentation foundation models not only accurate but also reliable in real-world deployment where well-calibrated uncertainties are essential. Following (12), existing UQ approaches can be broadly grouped into test-time augmentation, Bayesian methods, ensembles, and deterministic methods. To cover this wide methodological spectrum, we evaluate one representative technique for each group: MCD, DSE, EDL, and TTA. While Fig. 2 provides a schematic overview of how we fused our model architecture with all four UQ approaches, the following will provide a more detailed description of each. Figure 2: A schematic overview of how we combined the four different uncertainty quantification techniques with our model. Monte Carlo Dropout. MCD estimates predictive uncertainty through stochastic sampling with dropout layers. Following 11, dropout can be interpreted as a variational approximation to Bayesian neural networks, where the posterior distribution over weights is intractable and thus approximated by a Bernoulli distribution. By keeping dropout active during inference, each forward pass yields a slightly different softmax output p(y^t)p( y_t). To compute the predictive mean μ μ, we average across all T stochastic samples: μ^=1T∑t=1Tp(y^t). μ= 1TΣ^T_t=1p( y_t) 5.0pt. (2) The corresponding uncertainty can be quantified by the variance as σ^2=1T−1∑t=1T(p(y^t)−μ^)2. σ^2= 1T-1Σ^T_t=1(p( y_t)- μ)^2 5.0pt. (3) Besides the standard deviation σ σ, the predictive entropy can also be computed as a complete measure of the predictive uncertainty, which is composed of a combination of the aleatoric and epistemic uncertainty (55): H(μ^)=−∑c=1Cμ^clog(μ^c).H( μ)=-Σ^C_c=1 μ_c ( μ_c) 5.0pt. (4) Deep Sub-Ensemble. Deep Ensembles (21) are considered a gold standard for UQ (39; 14; 22), as they capture variability across independently trained models. However, their computational cost scales linearly with the number of ensemble members, limiting their practicality in large-scale vision tasks. DSEs approximate the benefits of DEs at reduced cost by sharing most of the architecture and varying only a subset of layers close to the output head (50). In our setup, the SAM2 encoder is shared across all ensemble members, while multiple DPT decoders are initialized independently. During training, only one decoder head is optimized per mini-batch, while the others are frozen. This cycling strategy increases diversity across decoders while keeping training efficient. At inference, each decoder in general produces a different softmax prediction p(y^t)p( y_t) based on the same encoder features. As in MCD, the predictive mean and predictive uncertainty are computed across all T decoder heads (see Eqs. 2 - 4). Evidential Deep Learning. EDL directly models predictive uncertainty by parameterizing a Dirichlet distribution over class probabilities (1). Instead of outputting a single categorical distribution via softmax, the network produces non-negative evidence values ece_c for each class c. These are mapped to the Dirichlet parameters αc=ec+1. _c=e_c+1 5.0pt. (5) The Dirichlet strength is then defined as S=∑c=1Cαc,S=Σ^C_c=1 _c 5.0pt, (6) where C is the number of classes. From here, the belief masses bcb_c and the total uncertainty u can be obtained as bc=ecS,u=CS.b_c= e_cS 5.0pt, u= CS 5.0pt. (7) The expected class probabilities follow as p(yc^)=αcS.p( y_c)= _cS 5.0pt. (8) This representation enables the model to express both confident assignments – with large αc _c values for one class – and uncertain states – with low, evenly distributed αc _c values across classes. Training is performed with an evidential loss function based on the mean squared error, which, for a single sample i, can be calculated as ℒi(Θ)=∑c=1C(yik−p(y^ic))2+p(y^ic)(1−p(y^ic)CLOSESi+1,L_i( )= _c=1^C (y_ik-p( y_ic) )^2+ p( y_ic)(1-p( y_ic)S_i+1 5.0pt, (9) where the first term penalizes prediction errors and the second term incorporates the variance of the predictive distribution, encouraging the model to calibrate its uncertainty appropriately. To prevent overconfidence on misclassified samples, a Kullback-Leibler (KL) divergence regularization is added to obtain the final loss function ℒ(Θ)=∑i=1Nℒi(Θ)+λt∑i=1NKL[D(p(y^i)|a~i)||D(p(y^i)|1)],L( )=Σ^N_i=1L_i( )+ _tΣ^N_i=1KL[D(p( y_i)| a_i)||D(p( y_i)|1)] 5.0pt, (10) where λt _t is an annealing coefficient, t is the current training epoch, D(p(y^i)|1)D(p( y_i)|1) is the uniform Dirichlet distribution, and lastly α~i=yi+(1−yi)⊙αi α_i=y_i+(1-y_i) _i is the Dirichlet parameters after removal of the non-misleading evidence from predicted parameters αi _i. More details on this can be found in 47. To stabilize training with our SAM2–DPT hybrid, we employ a two-head architecture, as shown by Fig. 2: One head predicts segmentation logits, while the second head outputs evidential uncertainty parameters. Empirically, this design significantly improved segmentation performance compared to a single shared head, where segmentation accuracy degraded drastically. Test-Time Augmentation. TTA is applied only during inference to estimate the predictive uncertainty. Consequently, it is applied without modifying the model parameters or the training protocol. During inference, we generate T augmented variants of each input image and feed both the original and its augmented versions into the network for prediction. Akin to MCD and DSE, the final prediction is obtained by the mean and the corresponding uncertainty is quantified by the variance/standard deviation or predictive entropy (cf. Eqs. 2–4). 4 Experimental Setup This section details the unified training configuration, datasets, and augmentations used in all experiments to ensure fair comparisons across UQ methods and robust evaluation under diverse conditions. 4.1 Training Configuration All models are trained for 100 epochs, which was empirically found to be sufficient for convergence. We use the AdamW optimizer (31) with a base learning rate of 3×10−53× 10^-5 for both datasets and a polynomial decay scheduler. As the SAM2 encoder is pre-trained while the decoder is newly initialized, we apply a learning-rate multiplier of 10 to the decoder. The loss function is the standard cross-entropy loss (cf. Eq. 1). Due to GPU memory constraints, we set the batch size to 8 for all experiments. The initial weight decay is set to 1×10−21× 10^-2 for Cityscapes and 1×10−41× 10^-4 for NYUv2. All experiments are conducted on a single NVIDIA A100 GPU with 40 GB VRAM. 4.2 Datasets We evaluate on Cityscapes (7) and NYUv2 (48). Cityscapes contains high-resolution street scenes (2048×1024 px). Following standard practice, we use the 2,875 training and 500 validation images with 19 semantic classes and 1 void label, which is ignored during training. Additionally, we employ the most challenging options of the synthetic Rainy-Cityscapes (17) and Foggy-Cityscapes (45) variants for out-of-domain evaluation. NYUv2 comprises 1,449 RGB-D indoor images captured with a Microsoft Kinect in 464 different rooms. We use the RGB images and the corresponding 40 semantic classes (grouped into 13 super-classes for visualization) split into 795 training and 654 test images at 640×480 px resolution. The four datasets jointly allow testing under both outdoor and indoor conditions as well as in-domain and out-of-domain settings. 4.3 Data Augmentations Regardless of the model, we apply random scaling with a factor between 0.5 and 2.0, random cropping with a crop size of 768×768 px on Cityscapes and 480×640 px on NYUv2, and random horizontal flipping with a flip chance of 50 % during training as data augmentations. 4.4 Uncertainty Quantification Based on rigorous evaluations, we report results based on the following UQ configurations. For TTA, we apply vertical and horizontal flipping as well as scaling during inference to generate samples to compute our predictive mean and uncertainty. We note that vertical flipping can alter spatial priors and may therefore lead to an overestimation of uncertainty in some cases – we nevertheless found the inclusion of vertical flipping to work best. MCD uses the already existing dropout layers in the model architecture with a dropout rate of 20% and we sample ten times during test-time. Similarly, we employ ten DPT heads for DSE. As the predictive entropy delivered slightly better results in our evaluations than the predictive variance, we only report uncertainty metrics based on the former for all models, including the baseline. 4.5 Metrics In terms of metrics, we report the mean Intersection over Union (mIoU), the Expected Calibration Error (ECE) (37; 55), and two uncertainty quality metrics from 35: 1. p(acc.||cer.): The probability that the model is accurate on its output given that the uncertainty is below a specific threshold. 2. p(unc.||ina.): The probability that the uncertainty of the model exceeds a specified threshold given that the prediction is inaccurate. Both metrics are expected to return high values for high-quality uncertainties. Based on our own empiric results and prior work from 22, we opt for the median uncertainty of any given image as the uncertainty threshold. 5 Experiments The following section lays out numerous quantitative as well as qualitative results, encompassing not only in-domain datasets but also out-of-domain settings for a comprehensive set of experiments. 5.1 Quantitative Results Comparison with SOTA. Tables 1 and 2 report segmentation performance on Cityscapes and NYUv2, respectively. Without additional architectural changes or complex training strategies, our simple SAM2-DPT baseline achieves results comparable to recent state-of-the-art methods, demonstrating that fine-tuning a foundation model with a lightweight DPT head can already yield competitive performance across both datasets. For all subsequent experiments, we report results using the SAM2-Tiny backbone due to computational constraints; however, we empirically verified that the observed trends are consistent with larger backbone variants. Cityscapes mIoU ↑ DeepLabV3+ (4) 79.6% U-Net++ (43) 75.5% SegFormer-B0 (56) 76.2% SegFormer-B5 (56) 82.4% Mask2Former (6) 83.3% SAM2-DPT [tiny] (ours) 76.6% SAM2-DPT [large] (ours) 82.4% Table 1: Comparison against previous approaches on Cityscapes. NYUv2 mIoU ↑ TokenFusion (53) 54.2% Omnivore (13) 54.0% EMSANet (46) 53.3% SGNet (5) 51.0% AsymFormer (10) 54.1% SAM2-DPT [tiny] (ours) 43.6% SAM2-DPT [large] (ours) 55.5% Table 2: Comparison against previous approaches on NYUv2. In-Domain Evaluation. Table 3 shows in-domain results on Cityscapes and NYUv2. On Cityscapes, the baseline model achieves competitive mIoU, strong calibration, and offers the fastest inference time. At the same time, it offers the second worst results in terms of uncertainty quality – only undermined by EDL, which performs worst across all metrics. DSE attains the best calibration and uncertainty quality but at the cost of a slight mIoU reduction and substantially slower inference. On NYUv2, the baseline achieves the highest segmentation performance and, unexpectedly, the best uncertainty quality, but shows the poorest calibration – though these results should be interpreted cautiously given the overall low segmentation scores. EDL again provides the second-fastest inference time but performs poorly across all other metrics. TTA delivers the best calibration, while DSE follows the baseline closely in uncertainty quality. in-Domain mIoU [%] ↑ ECE [%] ↓ p(acc.|cer.)p(acc.|cer.) [%] ↑ p(unc.|ina.)p(unc.|ina.) [%] ↑ Inference Time [ms] ↓ Cityscapes Baseline 76.55 0.95 92.61 76.18 58.57 TTA 76.84 1.76 94.29 81.32 202.50 MCD 76.10 1.30 93.52 80.18 656.07 DSE 74.44 0.65 96.91 91.74 315.46 EDL 70.84 4.97 84.40 51.14 73.40 NYUv2 Baseline 43.62 16.80 82.08 81.26 9.49 TTA 43.61 3.01 79.53 77.09 34.52 MCD 38.15 13.15 78.12 77.26 124.75 DSE 39.10 11.21 80.36 80.00 55.40 EDL 37.94 7.13 76.01 74.50 11.15 Table 3: Quantitative in-domain evaluation on Cityscapes and NYUv2 datasets using SAM2-DPT. Out-of-Domain Evaluation. Table 4 shows out-of-domain results on Rainy- and Foggy-Cityscapes. On both datasets, segmentation performance drops markedly compared to the in-domain experiments, highlighting the vulnerability of deep learning models to domain shifts. On Rainy-Cityscapes, TTA achieves the highest mIoU, the best calibration, and also improves the uncertainty quality compared to the baseline. MCD yields slightly higher mIoU than TTA but a substantially worse calibration. DSE performs comparably to the baseline with worse calibration but higher p(unc.||ina.). EDL again shows the weakest results across all metrics. For Foggy-Cityscapes, TTA gives the best mIoU at the cost of notably worse calibration and uncertainty quality, whereas DSE attains the best uncertainty scores, especially p(unc.||ina.). MCD remains close to the baseline with moderate improvements in mIoU but worse calibration. EDL again performs worst on all metrics. Out-of-Domain mIoU [%] ↑ ECE [%] ↓ p(acc.|cer.)p(acc.|cer.) [%] ↑ p(unc.|ina.)p(unc.|ina.) [%] ↑ Rainy-Cityscapes Baseline 46.72 5.09 92.66 80.92 TTA 48.99 0.84 93.40 83.85 MCD 49.37 5.33 92.29 78.28 DSE 45.71 7.07 92.65 83.21 EDL 40.91 12.24 89.68 75.12 Foggy-Cityscapes Baseline 53.96 3.09 89.54 80.11 TTA 56.53 7.46 88.48 77.43 MCD 54.49 5.67 90.39 81.44 DSE 51.14 4.82 93.40 89.49 EDL 50.97 14.26 83.64 67.09 Table 4: Quantitative out-of-domain evaluation on Rainy- and Foggy-Cityscapes using SAM2-DPT. 5.2 Qualitative Results Fig. 3 shows in-domain qualitative results on Cityscapes. The baseline, MCD, and DSE produce the most accurate segmentation predictions in terms of mIoU, whereas TTA performs slightly worse and EDL exhibits the worst results. The corresponding binary accuracy maps reveal largely similar falsely segmented regions across all methods. Regarding the uncertainty, TTA and EDL show a higher and spatially more widespread uncertainty, while the baseline, MCD and DSE display more localized and lower levels of uncertainty. Notably, DSE captures the erroneous regions most effectively, as its uncertainties align well with the binary error maps, indicating a strong correspondence between prediction errors and uncertainty estimates. Figure 3: Qualitative in-domain evaluation on Cityscapes dataset using SAM2-DPT. 5.3 Discussion Our results reveal several important patterns regarding the critical synthesis of UQ and foundation models for segmentation. Although the SAM2–DPT baseline achieves competitive segmentation accuracy with minimal architectural changes, its calibration and uncertainty quality are insufficient and inconsistent across domains. This confirms that high predictive performance does not automatically translate to high reliability. Additionally, we found clear trade-offs w.r.t. the UQ strategies. While DSE offers high uncertainty quality, it comes at a significant cost of inference speed. TTA improves upon the baseline in most cases but also suffers from high computation cost during inference. MCD, despite the highest inference times, often underperforms even the baseline, and EDL yields the least promising results. These observations highlight that there is no single best method and that the choice of UQ strategy highly depends on the deployment context. 6 Conclusion We presented the first systematic evaluation of multiple UQ methods applied to a SAM2-based foundation model for semantic segmentation. Our experiments across Cityscapes, NYUv2 and two out-of-domain settings show that competitive segmentation accuracy is easily attainable with a lightweight decoder, but reliability and efficiency differ markedly between UQ strategies. It is clear that there is no one-size-fits-all solution yet. This underscores the importance of future work not only focusing on predictive performance but also reliability and efficiency. References Amini et al. (2020) A. Amini, W. Schwarting, A. Soleimany, and D. Rus Deep evidential regression. Advances in neural information processing systems 33, p. 14927–14937. Cited by: §2.2, §3.2. Awais et al. (2025) M. Awais, M. Naseer, S. Khan, R. M. Anwer, H. Cholakkal, M. Shah, M. Yang, and F. S. Khan Foundation models defining a new era in vision: a survey and outlook. IEEE Transactions on Pattern Analysis and Machine Intelligence. Cited by: §1. Chen et al. (2017) L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence 40 (4), p. 834–848. Cited by: §2.1. Chen et al. (2018) L. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV), p. 801–818. Cited by: Table 1. Chen et al. (2021) L. Chen, Z. Lin, Z. Wang, Y. Yang, and M. Cheng Spatial information guided convolution for real-time rgbd semantic segmentation. IEEE Transactions on Image Processing 30, p. 2313–2324. Cited by: Table 2. Cheng et al. (2022) B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 1290–1299. Cited by: §2.1, Table 1. Cordts et al. (2016) M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 3213–3223. Cited by: §4.2. Deng et al. (2023) G. Deng, K. Zou, K. Ren, M. Wang, X. Yuan, S. Ying, and H. Fu Sam-u: multi-box prompts triggered uncertainty estimation for reliable sam in medical image. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 368–377. Cited by: §2.2. Dosovitskiy et al. (2020) A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §2.1. Du et al. (2024) S. Du, W. Wang, R. Guo, R. Wang, and S. Tang Asymformer: asymmetrical cross-modal representation learning for mobile platform real-time rgb-d semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 7608–7615. Cited by: Table 2. Gal and Ghahramani (2016) Y. Gal and Z. Ghahramani Dropout as a bayesian approximation: representing model uncertainty in deep learning. In international conference on machine learning, p. 1050–1059. Cited by: §2.2, §3.2. Gawlikowski et al. (2023) J. Gawlikowski, C. R. N. Tassi, M. Ali, J. Lee, M. Humt, J. Feng, A. Kruspe, R. Triebel, P. Jung, R. Roscher, M. Shahzad, W. Yang, R. Bamler, and X. X. Zhu A survey of uncertainty in deep neural networks. Artificial Intelligence Review 56 (Suppl 1), p. 1513–1589. Cited by: §1, §3.2. Girdhar et al. (2022) R. Girdhar, M. Singh, N. Ravi, L. Van Der Maaten, A. Joulin, and I. Misra Omnivore: a single model for many visual modalities. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 16102–16112. Cited by: Table 2. Gustafsson et al. (2020) F. K. Gustafsson, M. Danelljan, and T. B. Schon Evaluating scalable bayesian deep learning methods for robust computer vision. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, p. 318–319. Cited by: §3.2. Han et al. (2022) K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, et al. A survey on vision transformer. IEEE transactions on pattern analysis and machine intelligence 45 (1), p. 87–110. Cited by: §1. He et al. (2022) K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 16000–16009. Cited by: §3.1. Hu et al. (2019) X. Hu, C. Fu, L. Zhu, and P. Heng Depth-attentional features for single-image rain removal. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, p. 8022–8031. Cited by: §4.2. Jiang et al. (2024) M. Jiang, J. Zhou, J. Wu, T. Wang, Y. Jin, and M. Xu Uncertainty-aware adapter: adapting segment anything model (sam) for ambiguous medical image segmentation. arXiv preprint arXiv:2403.10931. Cited by: §2.2. Kaiser et al. (2025) T. Kaiser, T. Norrenbrock, and B. Rosenhahn Uncertainsam: fast and efficient uncertainty quantification of the segment anything model. arXiv preprint arXiv:2505.05049. Cited by: §2.2. Kirillov et al. (2023) A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, et al. Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, p. 4015–4026. Cited by: §2.1, §2. Lakshminarayanan et al. (2017) B. Lakshminarayanan, A. Pritzel, and C. Blundell Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems 30. Cited by: §2.2, §3.2. Landgraf et al. (2025a) S. Landgraf, M. Hillemann, T. Kapler, and M. Ulrich A comparative study on multi-task uncertainty quantification in semantic segmentation and monocular depth estimation. tm-Technisches Messen. Cited by: §3.2, §4.5. Landgraf et al. (2026) S. Landgraf, M. Hillemann, T. Kapler, and M. Ulrich EMUFormer: efficient multi-task uncertainties for reliable joint semantic segmentation and monocular depth estimation. International Journal of Computer Vision 134 (4), p. 142. Cited by: §2.2. Landgraf et al. (2025b) S. Landgraf, R. Qin, and M. Ulrich A critical synthesis of uncertainty quantification and foundation models in monocular depth estimation. arXiv preprint arXiv:2501.08188. Cited by: §2.2. Landgraf et al. (2024) S. Landgraf, K. Wursthorn, M. Hillemann, and M. Ulrich Dudes: deep uncertainty distillation using ensembles for semantic segmentation. PFG–Journal of Photogrammetry, Remote Sensing and Geoinformation Science 92 (2), p. 101–114. Cited by: §2.2. Li et al. (2021) Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou A survey of convolutional neural networks: analysis, applications, and prospects. IEEE transactions on neural networks and learning systems 33 (12), p. 6999–7019. Cited by: §1. Lin et al. (2017) G. Lin, A. Milan, C. Shen, and I. Reid Refinenet: multi-path refinement networks for high-resolution semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 1925–1934. Cited by: §3.1. Liu et al. (2020) J. Liu, Z. Lin, S. Padhy, D. Tran, T. Bedrax Weiss, and B. Lakshminarayanan Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. Advances in neural information processing systems 33, p. 7498–7512. Cited by: §2.2. Liu et al. (2024) K. Liu, B. Price, J. Kuen, Y. Fan, Z. Wei, L. Figueroa, K. Geras, and C. Fernandez-Granda Uncertainty-aware fine-tuning of segmentation foundation models. Advances in Neural Information Processing Systems 37, p. 53317–53389. Cited by: §2.2. Long et al. (2015) J. Long, E. Shelhamer, and T. Darrell Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 3431–3440. Cited by: §2.1. Loshchilov and Hutter (2019) I. Loshchilov and F. Hutter Decoupled weight decay regularization. External Links: 1711.05101, Link Cited by: §4.1. MacKay (1992) D. J. MacKay A practical bayesian framework for backpropagation networks. Neural computation 4 (3), p. 448–472. Cited by: §2.2. McAllister et al. (2017) R. T. McAllister, Y. Gal, A. Kendall, M. Van Der Wilk, A. Shah, R. Cipolla, and A. Weller Concrete problems for autonomous vehicle safety: advantages of bayesian deep learning. In International Joint Conferences on Artificial Intelligence, Inc., Cited by: §1. Mo et al. (2022) Y. Mo, Y. Wu, X. Yang, F. Liu, and Y. Liao Review the state-of-the-art technologies of semantic segmentation based on deep learning. Neurocomputing 493, p. 626–646. Cited by: §1. Mukhoti and Gal (2018) J. Mukhoti and Y. Gal Evaluating bayesian deep learning methods for semantic segmentation. arXiv preprint arXiv:1811.12709. Cited by: §4.5. Mukhoti et al. (2023) J. Mukhoti, A. Kirsch, J. Van Amersfoort, P. H. Torr, and Y. Gal Deep deterministic uncertainty: a new simple baseline. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 24384–24394. Cited by: §2.2. Naeini et al. (2015) M. P. Naeini, G. Cooper, and M. Hauskrecht Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the AAAI conference on artificial intelligence, Vol. 29. Cited by: §4.5. Nair et al. (2020) T. Nair, D. Precup, D. L. Arnold, and T. Arbel Exploring uncertainty measures in deep networks for multiple sclerosis lesion detection and segmentation. Medical image analysis 59, p. 101557. Cited by: §1. Ovadia et al. (2019) Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. Dillon, B. Lakshminarayanan, and J. Snoek Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. Advances in Neural Information Processing Systems 32. Cited by: §1, §3.2. Ranftl et al. (2021) R. Ranftl, A. Bochkovskiy, and V. Koltun Vision transformers for dense prediction. In Proceedings of the IEEE/CVF international conference on computer vision, p. 12179–12188. Cited by: §3.1. Ravi et al. (2024) N. Ravi, V. Gabeur, Y. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al. Sam 2: segment anything in images and videos. arXiv preprint arXiv:2408.00714. Cited by: §3.1. Rawat et al. (2017) M. Rawat, M. Wistuba, and M. Nicolae Harnessing model uncertainty for detecting adversarial examples. In NIPS Workshop on Bayesian Deep Learning, Cited by: §1. Ronneberger et al. (2015) O. Ronneberger, P. Fischer, and T. Brox U-net: convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, p. 234–241. Cited by: §2.1, Table 1. Ryali et al. (2023) C. Ryali, Y. Hu, D. Bolya, C. Wei, H. Fan, P. Huang, V. Aggarwal, A. Chowdhury, O. Poursaeed, J. Hoffman, et al. Hiera: a hierarchical vision transformer without the bells-and-whistles. In International conference on machine learning, p. 29441–29454. Cited by: §3.1. Sakaridis et al. (2018) C. Sakaridis, D. Dai, and L. Van Gool Semantic foggy scene understanding with synthetic data. International Journal of Computer Vision 126 (9), p. 973–992. Cited by: §4.2. Seichter et al. (2022) D. Seichter, S. B. Fischedick, M. Köhler, and H. Groß Efficient multi-task rgb-d scene analysis for indoor environments. In 2022 International joint conference on neural networks (IJCNN), p. 1–10. Cited by: Table 2. Sensoy et al. (2018) M. Sensoy, L. Kaplan, and M. Kandemir Evidential deep learning to quantify classification uncertainty. Advances in neural information processing systems 31. Cited by: §3.2. Silberman et al. (2012) N. Silberman, D. Hoiem, P. Kohli, and R. Fergus Indoor segmentation and support inference from rgbd images. In European conference on computer vision, p. 746–760. Cited by: §4.2. Srivastava et al. (2014) N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research 15 (1), p. 1929–1958. Cited by: §2.2. Valdenegro-Toro (2023) M. Valdenegro-Toro Sub-ensembles for fast uncertainty estimation in neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 4119–4127. Cited by: §2.2, §3.2. Van Amersfoort et al. (2020) J. Van Amersfoort, L. Smith, Y. W. Teh, and Y. Gal Uncertainty estimation using a single deep deterministic neural network. In International conference on machine learning, p. 9690–9700. Cited by: §2.2. Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. Advances in neural information processing systems 30. Cited by: §2.1. Wang et al. (2022) Y. Wang, X. Chen, L. Cao, W. Huang, F. Sun, and Y. Wang Multimodal token fusion for vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 12186–12195. Cited by: Table 2. Wilson and Izmailov (2020) A. G. Wilson and P. Izmailov Bayesian deep learning and a probabilistic perspective of generalization. Advances in Neural Information Processing Systems 33, p. 4697–4708. Cited by: §1. Wolf et al. (2025) D. Wolf, P. Balaji, A. Braun, and M. Ulrich Decoupling of neural network calibration measures. In Pattern Recognition. DAGM GCPR 2024, D. Cremers, Z. Lähner, M. Moeller, M. Nießner, B. Ommer, and R. Triebel (Eds.), Lecture Notes in Computer Science, Vol. 157, p. 117–130. Cited by: §3.2, §4.5. Xie et al. (2021) E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo SegFormer: simple and efficient design for semantic segmentation with transformers. Advances in neural information processing systems 34, p. 12077–12090. Cited by: §2.1, Table 1, Table 1. Zhang et al. (2023) Y. Zhang, S. Hu, C. Jiang, Y. Cheng, and Y. Qi Segment anything model with uncertainty rectification for auto-prompting medical image segmentation. CoRR. Cited by: §2.2. Zhou et al. (2024) T. Zhou, W. Xia, F. Zhang, B. Chang, W. Wang, Y. Yuan, E. Konukoglu, and D. Cremers Image segmentation in foundation model era: a survey. arXiv preprint arXiv:2408.12957. Cited by: §1, §2.1.