Paper deep dive
FSKD: Monocular Forest Structure Inference via LiDAR-to-RGBI Knowledge Distillation
Taimur Khan, Hannes Feilhauer, Muhammad Jazib Zafar
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/2/2026, 11:56:49 PM
Summary
FSKD is a knowledge distillation framework that enables monocular forest structure inference (CHM, PAI, FHD) from RGBI imagery by using a multi-modal teacher (RGBI + LiDAR) to train an RGBI-only SegFormer student. It achieves state-of-the-art zero-shot CHM performance, providing a scalable solution for operational forest monitoring.
Entities (8)
Relation Signals (4)
FSKD â predicts â CHM
confidence 100% · The method jointly predicts CHM, PAI, and FHD
FSKD â trainedon â Saxony
confidence 100% · Trained on 384 km2 of forests in Saxony, Germany
FSKD â uses â SegFormer
confidence 100% · an RGBI-only SegFormer student learns to reproduce these outputs
FSKD â outperforms â HRCHM
confidence 95% · outperforming HRCHM/DAC baselines by 29â46% in MAE
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Very High Resolution (VHR) forest structure data at individual-tree scale is essential for carbon, biodiversity, and ecosystem monitoring. Still, airborne LiDAR remains costly and infrequent despite being the reference for forest structure metrics like Canopy Height Model (CHM), Plant Area Index (PAI), and Foliage Height Diversity (FHD). We propose FSKD: a LiDAR-to-RGB-Infrared (RGBI) knowledge distillation (KD) framework in which a multi-modal teacher fuses RGBI imagery with LiDAR-derived planar metrics and vertical profiles via cross-attention, and an RGBI-only SegFormer student learns to reproduce these outputs. Trained on 384 $km^2$ of forests in Saxony, Germany (20 cm ground sampling distance (GSD)) and evaluated on eight geographically distinct test tiles, the student achieves state-of-the-art (SOTA) zero-shot CHM performance (MedAE 4.17 m, $R^2$=0.51, IoU 0.87), outperforming HRCHM/DAC baselines by 29--46% in MAE (5.81 m vs. 8.14--10.84 m) with stronger correlation coefficients (0.713 vs. 0.166--0.652). Ablations show that multi-modal fusion improves performance by 10--26% over RGBI-only training, and that asymmetric distillation with appropriate model capacity is critical. The method jointly predicts CHM, PAI, and FHD, a multi-metric capability not provided by current monocular CHM estimators, although PAI/FHD transfer remains region-dependent and benefits from local calibration. The framework also remains effective under temporal mismatch (winter LiDAR, summer RGBI), removing strict co-acquisition constraints and enabling scalable 20 cm operational monitoring for workflows such as Digital Twin Germany and national Digital Orthophoto programs.
Tags
Links
- Source: https://arxiv.org/abs/2604.01766v1
- Canonical: https://arxiv.org/abs/2604.01766v1
Trouble viewing inline? Open PDF directly â
Full Text
32,472 characters extracted from source content.
Expand or collapse full text
11institutetext: Helmholtz Centre for Environmental Research â UFZ, Halle (Saale), Germany 11email: taimur.khan@ufz.de 22institutetext: Leipzig University, Leipzig, Germany 33institutetext: Georg-August University of Göttingen, Göttingen, Germany FSKD: Monocular Forest Structure Inference via LiDAR-to-RGBI Knowledge Distillation Taimur Khan Corresponding author. Hannes Feilhauer Muhammad Jazib Zafar Abstract Very High Resolution (VHR) forest structure data at individual-tree scale is essential for carbon, biodiversity, and ecosystem monitoring. Still, airborne LiDAR remains costly and infrequent despite being the reference for forest structure metrics like Canopy Height Model (CHM), Plant Area Index (PAI), and Foliage Height Diversity (FHD). We propose FSKD: a LiDAR-to-RGB-Infrared (RGBI) knowledge distillation (KD) framework in which a multi-modal teacher fuses RGBI imagery with LiDAR-derived planar metrics and vertical profiles via cross-attention, and an RGBI-only SegFormer student learns to reproduce these outputs. Trained on 384 km2 of forests in Saxony, Germany (20 cm ground sampling distance (GSD)) and evaluated on eight geographically distinct test tiles, the student achieves state-of-the-art (SOTA) zero-shot CHM performance (MedAE 4.17 m, R2R^2=0.51, IoU 0.87), outperforming HRCHM/DAC baselines by 29â46% in MAE (5.81 m vs. 8.14â10.84 m) with stronger correlation coefficients (0.713 vs. 0.166â0.652). Ablations show that multi-modal fusion improves performance by 10â26% over RGBI-only training, and that asymmetric distillation with appropriate model capacity is critical. The method jointly predicts CHM, PAI, and FHD, a multi-metric capability not provided by current monocular CHM estimators, although PAI/FHD transfer remains region-dependent and benefits from local calibration. The framework also remains effective under temporal mismatch (winter LiDAR, summer RGBI), removing strict co-acquisition constraints and enabling scalable 20 cm operational monitoring for workflows such as Digital Twin Germany and national Digital Orthophoto programs. Code + Models: TBA upon publication. Data: TBA upon publication. 1 Introduction Forest structure mapping is a 3D perception problem that underpins carbon accounting, biodiversity monitoring, and ecosystem analysis [22, 27]. Airborne LiDAR remains the reference source for fine-scale canopy geometry, but repeated campaigns are costly and typically infrequent [13, 33]. In contrast, very-high-resolution (VHR) aerial RGBI imagery is collected routinely (e.g., 20 cm DOP coverage in Germany [4]) (Figure 1). This creates a clear computer vision question: how much 3D forest structure can be recovered from single-view optical imagery alone [2]? Recent work has shown that monocular canopy height mapping is feasible. Two representative and directly comparable estimators are HRCHM [38] and DAC [34, 43], both designed for CHM prediction from RGB imagery. These are important advances, but they remain height-only methods and do not explicitly target richer vertical structure descriptors. Specifically, HRCHM couples a DINOv2-style representation with a dense prediction decoder for high-resolution CHM mapping [38, 31], while DAC adapts Depth Anything v2 representations to canopy height estimation with substantially lower model complexity [34, 43]. Together, they establish strong monocular CHM baselines, but neither is designed to predict multi-metric vertical structure. That limitation is operationally important. Forest monitoring needs canopy height (CHM), but also metrics such as Plant Area Index (PAI) and Foliage Height Diversity (FHD), which reflect canopy layering and ecosystem function [22]. These quantities are only weakly observable in a single nadir image, so direct RGBI-only supervision is underconstrained for robust transfer. We therefore use learning with privileged information [40] and knowledge distillation [21, 18]. During training, a multi-modal teacher uses RGBI plus LiDAR-derived planar and vertical cues; during inference, a student runs on RGBI only. The design is motivated by cross-modal transfer results in autonomous driving, where LiDAR-aware teachers improve monocular students on geometric tasks [9, 8, 24, 26, 15]. Our framework predicts CHM, PAI, and FHD at 20 cm resolution from RGBI input. We train and evaluate on public data [17] with geographically separated splits to test zero-shot regional transfer. The intended use is practical: frequent structure-map refreshes between LiDAR campaigns for operational workflows such as Digital Twin Germany [3]. This paper makes three contributions: âą A LiDAR-to-RGBI teacherâstudent distillation framework for dense forest structure prediction. âą A multi-target monocular setup (CHM, PAI, FHD), extending beyond height-only inference. âą A CV-oriented evaluation against strong monocular CHM baselines. Figure 1: Multi-modal data acquisition and forest structure derivation framework. (a) RGBI orthophotos and LiDAR are acquired at different times. (b) Inputs are spatially aligned in 2 km Ă 2 km tiles. (c) LiDAR provides privileged structural supervision (CHM, PAI, FHD, PAD) for teacher training, while the deployed student uses RGBI only. 2 Data and Methods 3 Study Area Our experiments are conducted in Saxony (Germany), where forests cover roughly 28% of the state (about 520,000 ha) and include both conifer-dominated and broadleaf systems [12, 20]. This diversity makes the region suitable for testing both dense structural prediction and cross-region transfer behavior. We use ninety-six paired RGBIâLiDAR tiles (2 km Ă 2 km each; 384 km2 total). To reduce spatial leakage, we enforce a geographic split into eighty training, eight validation, and eight test tiles, with test locations held out from training regions (Figure 2). Sampling is stratified by CORINE forest composition [11], yielding a mix dominated by conifer and broadleaf classes with a smaller mixed-forest fraction. Figure 2: Study area and dataset split. Spatial distribution of the 96 RGBIâLiDAR tile pairs in Saxony and neighboring areas with geographically separated train/validation/test regions. 4 Data Sources Aerial Orthophotos (RGBI) We use 20 cm GSD four-band orthophotos (R, G, B, NIR) from GeoSN DOP products [17]. These images represent the operational modality that is available at high refresh rate and large area coverage under national orthophoto programs [4]. The selected tiles are primarily leaf-on (2022â2023). Airborne LiDAR Point Clouds Co-registered LiDAR (LAS/LAZ) is taken from GeoSN products [16]. Acquisition dates are not fully aligned with RGBI (mostly leaf-off LiDAR, typically 2017â2023), which reflects realistic deployment conditions rather than ideal co-acquisition. The LiDAR tiles have a mean point density of 15.82 points/m2, ranging from 0.98 to 24.47 points/m2. 5 Data Preprocessing and Spatial Alignment Training uses 224Ă224 patches (about 45 m Ă 45 m at 20 cm GSD) sampled from forested areas. This patch-based setup keeps training tractable while preserving local canopy structure, as shown in Figure 3. For each patch we prepare: âą RGBI imagery (4 channels), âą planar LiDAR-derived targets (CHM, PAI, FHD) using equations from Table 1, âą vertical Plant Area Density (PAD) profiles. LiDAR point clouds are processed with PyForestScan [32]. Metrics are first computed on a 1 m grid, then aligned to the RGBI reference grid (EPSG:25833), and resampled to 20 cm where required [6, 35]. PAD profiles are retained as vertical supervision and stored at lower effective spatial resolution for memory-efficient training [19]. A validity mask removes invalid pixels before loss computation so supervision is restricted to spatially reliable regions. A sample pre-processing log file is provided in the supplementary material. Table 1: Common canopy structure metrics and their equations[22, 32]. Metric Equation CHM - planar Hcanopy=maxâĄ(HAGpoints)H_canopy= (HAG_points) PAD - vertical PADiâ1,i=lnâĄ(SeSt)â1kâÎâzPAD_i-1,i= ( S_eS_t ) 1k z PAI - planar PAI=âi=1nPADiâ1,iPAI= _i=1^nPAD_i-1,i FHD - planar FHD=ââi=1npiâlnâĄ(pi)FHD=- _i=1^np_i (p_i) P05, P50, P95 PxâxP_x is the height below which xâx%x\% of all points lie. Figure 3: Preprocessing and alignment overview. RGBI and LiDAR products are transformed into pixel-aligned multi-modal training tensors for distillation. 6 Model Architecture As shown in Figure 4, we train a multi-modal teacher and an RGBI-only student. Both output dense CHM, PAI, and FHD maps. Figure 4: Teacherâstudent architecture. A LiDAR-aware teacher supervises an RGBI-only student for CHM/PAI/FHD prediction. 6.1 Teacher Model (Multi-Branch Fusion) The teacher fuses three streams: RGBI features (Swin), planar LiDAR features (ViT), and vertical PAD features (MLP encoder), followed by cross-attention fusion and an FPN decoder [30, 10, 25, 28]. This model is used as a privileged-information supervisor during distillation. Figure 5 illustrates the teacher feature maps and fusion flow. Figure 5: Teacher feature flow. Multi-modal feature extraction, fusion, and decoding in the teacher network. 6.2 Student Model (Monocular RGBI Estimator) The student uses SegFormer (MiT-B2/B5) with four-channel RGBI input and a lightweight regression head [42]. Multi-scale encoder features are fused, then mapped to the three output metrics. A 1Ă1 adapter projects student features into the teacher fusion space for feature distillation. Figure 6 illustrates the feature hierarchy and flow. Studentâ(RGBI)â[CHM~,PAI~,FHD~].Student(RGBI)â[ CHM, PAI, FHD]. Figure 6: Student feature flow. Multi-scale RGBI encoding and dense regression in the RGBI-only student. 7 Knowledge Distillation & Training Training protocol Training is two-stage: (1) train the teacher on multi-modal inputs, then (2) freeze the teacher and train the student with supervised and distillation losses. 7.1 Stage 1: Teacher Teacher optimization uses robust regression over CHM/PAI/FHD plus a CHM gradient term: âteacher=âcâCHM,PAI,FHDârobustâ(Yct,Yc)+λgradââgradâ(YCHMt,YCHM).L_teacher= _câ\CHM,PAI,FHD\L_robust(Y_c^t,Y_c)+ _grad\,L_grad(Y_CHM^t,Y_CHM). (1) Here, ârobustL_robust is a masked Smooth L1 (Huber) regression loss applied per channel with equal weighting across CHM/PAI/FHD. It is quadratic for small residuals and linear for large residuals, improving robustness to outliers while preserving fine-structure sensitivity. âgrad=ââxYCHMtââxYCHMâ1+ââyYCHMtââyYCHMâ1,L_grad= _xY_CHM^t- _xY_CHM _1+ _yY_CHM^t- _yY_CHM _1, (2) where âx _x and ây _y are finite-difference gradients. We include this term because CHM contains the strongest spatial gradient information among the targets; matching CHM gradients improves edge and transition fidelity (e.g., canopy boundaries) and reduces over-smoothing [39]. We use λgrad=0.1 _grad=0.1. The teacher is optimized with AdamW (lr=1Ă10â41Ă 10^-4, weight decay=1Ă10â41Ă 10^-4). 7.2 Stage 2: Student with KD Student training combines supervised output fitting, output distillation from the frozen teacher, feature matching, and vertical-proxy alignment: Lout L_out =âYsâYâ1, = Y^s-Y _1, (3) LKD L_KD =SmoothL1â(Ys,Yt), =SmoothL1(Y^s,Y^t), Lfeat L_feat =âProjâ(Ffuseds)âFfusedtâ22, = (F_fused^s)-F_fused^t _2^2, Lvert L_vert =â„Proj(Ffusedsâ)âFfusedtâ„22. = (F_fused^s\! )-F_fused^t _2^2. âstudent=wsupâLout+wKDâLKD+wfeatâLfeat+wvertâLvert.L_student=w_supL_out+w_KDL_KD+w_featL_feat+w_vertL_vert. (4) We keep wsup=1w_sup=1 and use auxiliary weights in a low range (0â0.5). The selected setting is wKD=0.5w_KD=0.5, wfeat=0.1w_feat=0.1, and wvert=0.1w_vert=0.1, with KD warm-up (wKD=0w_KD=0 in early epochs). The student is optimized with AdamW (lr=2Ă10â42Ă 10^-4, weight decay=1Ă10â41Ă 10^-4). Teacher symbols Y: LiDAR-derived targets; YtY^t: teacher predictions; ârobustL_robust: masked Smooth L1 regression; âgradL_grad: CHM gradient consistency; λgrad _grad: gradient-loss weight. Student symbols YsY^s: student predictions; Ffusedt,FfusedsF_fused^t,F_fused^s: teacher/student fused features; â : downsampling to teacher fusion scale; ProjProj: 1Ă11Ă1 adapter; wâ w_·: student loss weights. We use standard geometric/color augmentation (random patch sampling/cropping only) and spatially separated train/validation/test splits (Tile-pairs:80/8/8). A sample training log file is provided in the supplementary material. 8 Evaluation 8.1 Metrics On validation tiles, we report MAE, RMSE, Bias, and R2R^2 to track fit quality and systematic error during model selection. On held-out test tiles, we additionally report MedAE, Pearson R, IoU, F1, and rMAE to characterize zero-shot transfer in both continuous-value accuracy and spatial agreement. 8.2 Ablations We run comprehensive targeted ablations (Table 2) to measure the contribution of distillation, fusion, backbone scale, and optimization choices [36]. The model weights for each ablation are provided in the supplementary material for reproducibility and further analysis. Table 2: Ablation studies for model architecture and training parameters Study What to remove/test Why No distillation loss Remove KD loss â only RGBI supervision. Does distillation actually help? Asymmetric learning 1. Train teacher for fewer epochs than student. 2. Train teacher for more epochs than student. Does symmetry matter in learning representations? Without multimodal fusion (no cross-attention) Remove multi-modal fusion. Does multi-stage feature fusion via cross-attention matter? Batch size sensitivity Try batch size 8 vs 64. Important for generalization vs. memory. Backbone size Try MiT-B2 vs B5. Smaller vs larger (impact on capacity). 8.3 SOTA CHM Baselines For CHM, we compare against two monocular SOTA families: HRCHM [38] and DAC [34, 43]. We use the model weights provided with the original papers for inference on test tiles. Following the setup in the full paper, HRCHM is evaluated in both native 20 cm and scale-matched 60 cm variants, and DAC outputs are converted to metric CHM before scoring. All baseline outputs are then reprojected to the same 20 cm evaluation grid for pixel-wise comparison. 8.4 Inference Pipeline At deployment, the student performs sliding-window inference over full RGBI tiles and stitches predictions into seamless georeferenced CHM/PAI/FHD rasters. 9 Results 10 Quantitative Performance All results are provided as tables in the supplementary material. 10.1 Core Outcomes We report compact headline metrics here, then interpret what they imply for transfer behavior. Validation summary On validation data, the student retained most of the teacher signal for CHM (teacher/student MAE: 3.88/4.95 m; R2R^2: 0.69/0.61) while reducing bias magnitude. For FHD, teacher and student were nearly identical (R2=0.54R^2=0.54 for both), indicating that this target transfers stably under the training distribution. For PAI, the student outperformed the teacher (MAE: 0.48 vs 0.59; R2R^2: 0.31 vs 0.05), which suggests that distillation helped suppress teacher-specific noise and regularize the final predictor. Zero-shot test summary On the geographically held-out test set, CHM generalized best. It reached MAE 5.81 m, MedAE 4.17 m, R=0.713R=0.713, R2=0.509R^2=0.509, IoU 0.870, and F1 0.930. The main residual error mode is underestimation in taller canopies (bias -2.57 m), but overall spatial structure remains robust. FHD and PAI transferred less reliably (FHD: MAE 0.42, R=0.391R=0.391, rMAE 29.7%; PAI: MAE 0.46, R=0.361R=0.361, rMAE 44.0%). This gap relative to CHM is consistent with domain shift for vertically complex descriptors: height cues transfer more readily across regions than canopy-layer composition cues. Figure 7: Quantitative agreement on held-out test tiles. CHM shows the strongest correlation; PAI and FHD remain harder under zero-shot transfer. 10.2 Ablation Summary Ablations reveal three stable effects. First, multi-modal fusion is a primary driver: removing it degrades all targets, with the largest relative drop for FHD. Second, asymmetric distillation with synchronized teacher/student schedules performs best; extended student training against a weak teacher leads to instability. Third, MiT-B2 is more reliable than MiT-B5 in this data regime, indicating that capacity must match supervision density rather than simply scale upward. The full ablation results table is provided in the supplementary material. 10.3 SOTA CHM Comparison Against HRCHM and DAC baselines, our model achieved the best CHM accuracy on the same test tiles (Table 3): MAE 5.81 m versus 8.14 m (HRCHM-60cm), 9.88 m (HRCHM-full-res), and 10.84 m (DAC-B). Correlation was also strongest (R=0.713R=0.713 vs 0.652/0.452/0.166), with lower systematic underestimation. These results support the central claim: LiDAR-informed distillation transfers stronger structural priors than purely monocular baselines in this setting, even when all methods are evaluated on the same 20 cm target grid. Importantly, this is a stringent comparison: the baselines are strong CHM-focused monocular methods, whereas our model is trained for joint CHM/PAI/FHD prediction and evaluated in a geographically held-out regime. Table 3: Comparison of evaluation metrics with state-of-the-art methods on CHM prediction. Model MAE (m) MedAE (m) RMSE (m) Bias (m) R IoU F1 FSKD (MiT-B2/b64) 5,81±1,445,81± 1,44 4,17±1,624,17± 1,62 7,90 -2,57 0,713 0,870 0,930 HRCHM-aerial (Full Res) 9,88±2,499,88± 2,49 8,83±3,938,83± 3,93 12,89 -8,55 0,452 0,733 0,842 HRCHM-aerial (60cm) 8,14±2,648,14± 2,64 7,46±3,017,46± 3,01 10,34 -6,94 0,652 0,834 0,907 DAC-B 10,84±2,3810,84± 2,38 10,99±3,0910,99± 3,09 12,56 -6,49 0,166 0,801 0,888 11 Qualitative Performance Qualitative maps match the quantitative pattern. CHM predictions preserve stand boundaries and canopy gradients robustly, while PAI/FHD preserve broad structure but compress dynamic range in difficult regions. Figure 8 illustrates tile-scale behavior, whereas Figure 9 and Figure 10 highlights fine-structure transfer at patch scale for different forest types. Figure 8: Representative tile. Normalised (for comparison) student CHM/PAI/FHD maps are spatially coherent with LiDAR-derived references, with strongest agreement for CHM. Figure 9: Broadleaved patch comparison. Ground truth (top), teacher (middle), and student (bottom) show close CHM agreement with clear crown delineation; PAI/FHD preserve major spatial gradients but appear smoother. Figure 10: Coniferous patch comparison. Ground truth (top), teacher (middle), and student (bottom) show strong CHM transfer; PAI/FHD remain directionally correct with compressed amplitude. 12 Discussion What works The teacherâstudent setup reliably transfers LiDAR-informed geometry into an RGBI-only model. The strongest and most stable signal is CHM, where both quantitative and qualitative evidence indicates usable cross-region performance at 20 cm output resolution. Where it breaks PAI and FHD are more sensitive to regional canopy composition and acquisition mismatch. Their weaker transfer indicates that monocular cues alone are insufficient for universally calibrated vertical-structure inference without broader training coverage, stronger priors, or explicit adaptation [14, 37, 5]. Operational implications The student is lightweight at inference and can refresh structural layers whenever new RGBI imagery is available, which is useful for workflows such as Digital Twin Germany [3]. In practice, this is best treated as a LiDAR-complement strategy: frequent optical updates between less frequent LiDAR campaigns, with periodic recalibration to maintain vertical-metric fidelity. Near-term extensions Priority next steps are: (1) expand training coverage across more Saxony tiles/years/seasons (leaf-on vs. leaf-off), (2) add uncertainty estimation for map-level confidence, (3) test additional modalities (e.g., Synthetic Aperture Radar - SAR or multi-temporal RGBI) and feature extractors (e.g. DINO), and (4) integrate crown-level products for downstream workflows [1, 7, 29, 23, 41], as demonstrated in Figure 11. Figure 11: Metric-wise ITC masking. Tree-crown masks derived with DeepTrees [23] can be applied consistently to CHM, PAI, and FHD outputs for per-tree summaries. 13 Conclusion This work shows that LiDAR-to-RGBI knowledge distillation can produce a practical monocular forest-structure estimator with strong CHM performance and useful first-order PAI/FHD signals. The key result is operational: LiDAR-informed structure can be propagated to routine aerial imagery without LiDAR at inference time. The method contributes: (1) a cross-modal teacherâstudent pipeline for CHM/PAI/FHD at 20 cm, (2) stable gains from privileged multi-modal fusion over RGBI-only training, and (3) a deployment path for frequent large-area updates between LiDAR acquisitions. The remaining gap is generalization of vertical structure indices (PAI/FHD). Closing that gap likely requires broader regional training, uncertainty-aware prediction, and targeted domain adaptation rather than architectural scaling alone. Beyond direct mapping, FSKD outputs can also serve as inputs for downstream deep learning tasks, including tree-crown segmentation, tree vitality estimation, species classification, and biomass estimation. Overall, FSKD establishes a practical computer-vision blueprint for converting routine RGBI imagery into actionable 3D forest structure at scale. References [1] F. Alidoost, H. Arefi, and F. Tombari (2019) 2D image-to-3d model: knowledge-based 3d building reconstruction (3dbr) using single aerial images and convolutional neural networks (cnns). Remote sensing 11 (19), p. 2219. Cited by: §12. [2] M. Brandt, J. Chave, S. Li, R. Fensholt, P. Ciais, J. Wigneron, F. Gieseke, S. Saatchi, C. Tucker, and C. Igel (2025) High-resolution sensors and deep learning models for tree resource monitoring. Nature Reviews Electrical Engineering 2 (1), p. 13â26. Cited by: §1. [3] Bundesamt fĂŒr Kartographie und GeodĂ€sie (BKG) () BKG - Digital Twin â bkg.bund.de. Note: https://w.bkg.bund.de/EN/Topics/Digital-Twin/digital-twin.html[Accessed 06-12-2025] Cited by: §1, §12. [4] Bundesamt fĂŒr Kartographie und GeodĂ€sie (BKG) Digitale orthophotos (dop20) der bundesrepublik deutschland, bodenauflösung 20 cm. GeoBasis-DE / BKG. Note: https://gdz.bkg.bund.de/Quellenvermerk: © GeoBasis-DE / BKG (JAHR_DER_AUFNAHME). Abruf ĂŒber das Geodatenzentrum (GDZ). [Accessed 06-12-2025] Cited by: §1, §4. [5] P. Burns, C. R. Hakkenberg, and S. J. Goetz (2024) Multi-resolution gridded maps of vegetation structure from gedi. Scientific Data 11 (1), p. 881. Cited by: §12. [6] A. T. Candan and H. Kalkan (2023) U-net-based rgb and lidar image fusion for road segmentation. Signal, Image and Video Processing 17 (6), p. 2837â2843. Cited by: §5. [7] K. Chen, C. Wang, M. Lu, W. Dai, J. Fan, M. Li, and S. Lei (2023) Integrating topographic skeleton into deep learning for terrain reconstruction from gdem and google earth image. Remote Sensing 15 (18), p. 4490. Cited by: §12. [8] Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, and F. Zhao (2022) Bevdistill: cross-modal bev distillation for multi-view 3d object detection. arXiv preprint arXiv:2211.09386. Cited by: §1. [9] Z. Chong, X. Ma, H. Zhang, Y. Yue, H. Li, Z. Wang, and W. Ouyang (2022) Monodistill: learning spatial features for monocular 3d object detection. arXiv preprint arXiv:2201.10830. Cited by: §1. [10] A. Dosovitskiy (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §6.1. [11] European Environment Agency (EEA) (2019) Corine Land Cover 2018 (vector), version 20. Copernicus Land Monitoring Service. Note: https://sdi.eea.europa.eu/catalogue/copernicus/api/records/960998c1-1870-4e82-8051-6485205ebbac?language=allAccessed on 10-11-2025 External Links: Document, Link Cited by: §3. [12] European State Forest Association () European State Forest Association - Staatsbetrieb Sachsenforst â eustafor.eu. Note: https://eustafor.eu/members/sachsenforst-state-forests-of-saxony/[Accessed 04-12-2025] Cited by: §3. [13] F. E. Fassnacht, C. Mager, L. T. Waser, U. Kanjir, J. SchĂ€fer, A. P. Buhvald, E. Shafeian, F. Schiefer, L. StanÄiÄ, M. Immitzer, et al. (2025) Forest practitionersâ requirements for remote sensing-based canopy height, wood-volume, tree species, and disturbance products. Forestry: An International Journal of Forest Research 98 (2), p. 233â252. Cited by: §1. [14] W. Flynn, S. Grieve, A. Henshaw, H. Owen, R. Buggs, C. Metheringham, W. Plumb, J. Stocks, and E. Lines (2024) UAV-derived greenness and within-crown spatial patterning can detect ash dieback in individual trees. Ecological Solutions and Evidence 5 (2), p. e12343. Cited by: §12. [15] H. Gao, X. Yu, Y. Xu, Q. Ran, and W. Hussain (2024) MonoFG: monocular 3d object detection with knowledge distillation for human-centric autonomous driving systems. ACM Transactions on Autonomous and Adaptive Systems. Cited by: §1. [16] GeoSN Fachliche details (geobasisinformation). Landesvermessung Sachsen. Note: https://w.landesvermessung.sachsen.de/fachliche-details-8645.html[Accessed 2025-12-10] Cited by: §4. [17] GeoSN Luftbild-produkte (offene geodaten). Landesvermessung Sachsen. Note: https://w.geodaten.sachsen.de/luftbild-produkte-3995.html[Accessed 2025-12-10] Cited by: §1, §4. [18] J. Gou, B. Yu, S. J. Maybank, and D. Tao (2021) Knowledge distillation: a survey. International journal of computer vision 129 (6), p. 1789â1819. Cited by: §1. [19] C. R. Harris, K. J. Millman, S. J. van der Walt, R. Gommers, P. Virtanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, R. Kern, M. Picus, S. Hoyer, M. H. van Kerkwijk, M. Brett, A. Haldane, J. F. del RĂo, M. Wiebe, P. Peterson, P. GĂ©rard-Marchant, K. Sheppard, T. Reddy, W. Weckesser, H. Abbasi, C. Gohlke, and T. E. Oliphant (2020-09) Array programming with NumPy. Nature 585 (7825), p. 357â362. External Links: Document, Link Cited by: §5. [20] S. Hering and S. Irrgang (2005) Conversion of substitute tree species stands and pure spruce stands in the ore mountains in saxony. Journal of Forest Science 51 (11), p. 519â525. Cited by: §3. [21] G. Hinton, O. Vinyals, and J. Dean (2015) Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531. Cited by: §1. [22] A. G. Kamoske, K. M. Dahlin, S. C. Stark, and S. P. Serbin (2019) Leaf area density from airborne lidar: comparing sensors and resolutions in a temperate broadleaf forest ecosystem. Forest Ecology and Management 433, p. 364â375. Cited by: §1, §1, Table 1, Table 1. [23] T. Khan, C. Arnold, and H. Grover (2025) DeepTrees: tree crown segmentation and analysis in remote sensing imagery with pytorch. Journal of Open Source Software 10 (114), p. 8056. External Links: Document, Link Cited by: Figure 11, §12. [24] Q. Lan and Q. Tian (2022) Instance, scale, and teacher adaptive knowledge distillation for visual detection in autonomous driving. IEEE Transactions on Intelligent Vehicles 8 (3), p. 2358â2370. Cited by: §1. [25] H. Li and X. Wu (2024) CrossFuse: a novel cross attention mechanism based infrared and visible image fusion approach. Information Fusion 103, p. 102147. Cited by: §6.1. [26] Z. Li, H. Liang, H. Wang, M. Zhao, J. Wang, and X. Zheng (2023) MKD-cooper: cooperative 3d object detection for autonomous driving via multi-teacher knowledge distillation. IEEE Transactions on Intelligent Vehicles 9 (1), p. 1490â1500. Cited by: §1. [27] K. Lim, P. Treitz, M. Wulder, B. St-Onge, and M. Flood (2003) LiDAR remote sensing of forest structure. Progress in physical geography 27 (1), p. 88â106. Cited by: §1. [28] T. Lin, P. DollĂĄr, R. Girshick, K. He, B. Hariharan, and S. Belongie (2017) Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 2117â2125. Cited by: §6.1. [29] Y. Lin, H. Li, L. Jing, H. Ding, and S. Tian (2024) Individual tree crown delineation using airborne lidar data and aerial imagery in the taigaâtundra ecotone. Remote Sensing 16 (21), p. 3920. Cited by: §12. [30] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin transformer: hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, p. 10012â10022. Cited by: §6.1. [31] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. (2023) Dinov2: learning robust visual features without supervision. arXiv preprint arXiv:2304.07193. Cited by: §1. [32] J. E. H. Percival and B. P. Leamon (2025) PyForestScan: a python library for calculating forest structural metrics from lidar point cloud data. Journal of Open Source Software 10 (106), p. 7314. Cited by: Table 1, Table 1, §5. [33] N. Puletti, M. Grotti, C. Ferrara, and F. Chianucci (2020) Lidar-based estimates of aboveground biomass through ground, aerial, and satellite observation: a case study in a mediterranean forest. Journal of Applied Remote Sensing 14 (4), p. 044501â044501. Cited by: §1. [34] D. Rege Cambrin, I. Corley, and P. Garza (2024) Depth any canopy: leveraging depth foundation models for canopy height estimation. In European Conference on Computer Vision, p. 71â86. Cited by: §1, §1, §8.3. [35] C. Ressl, N. Pfeifer, and G. Mandlburger (2012) Applying 3d affine transformation and least squares matching for airborne laser scanning strips adjustment without gnss/imu trajectory data. The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences 38, p. 67â72. Cited by: §5. [36] S. Sheikholeslami (2019) Ablation programming for machine learning. Cited by: §8.2. [37] S. Tan, Y. Zhang, J. Qi, Y. Su, Q. Ma, and J. Qiu (2024) Exploring the potential of gedi in characterizing tree height composition based on advanced radiative transfer model simulations. Journal of Remote Sensing 4, p. 0132. Cited by: §12. [38] J. Tolan, H. Yang, B. Nosarzewski, G. Couairon, H. V. Vo, J. Brandt, J. Spore, S. Majumdar, D. Haziza, J. Vamaraju, et al. (2024) Very high resolution canopy height maps from rgb imagery using self-supervised vision transformer and convolutional decoder trained on aerial lidar. Remote Sensing of Environment 300, p. 113888. Cited by: §1, §1, §8.3. [39] H. Tong (2023) Functional linear regression with huber loss. Journal of Complexity 74, p. 101696. Cited by: §7.1. [40] V. Vapnik, R. Izmailov, et al. (2015) Learning using privileged information: similarity control and knowledge transfer.. J. Mach. Learn. Res. 16 (1), p. 2023â2049. Cited by: §1. [41] B. Xiang, M. Wielgosz, S. Puliti, K. KrĂĄl, M. KrĆŻÄek, A. Missarov, and R. Astrup (2025) ForestFormer3D: a unified framework for end-to-end segmentation of forest lidar 3d point clouds. arXiv preprint arXiv:2506.16991. Cited by: §12. [42] E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo (2021) SegFormer: simple and efficient design for semantic segmentation with transformers. Advances in neural information processing systems 34, p. 12077â12090. Cited by: §6.2. [43] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao (2024) Depth anything v2. Advances in Neural Information Processing Systems 37, p. 21875â21911. Cited by: §1, §1, §8.3.