Abstract
Geotechnical engineering decisions in tunnelling and slope design depend on a representative description of fractured rock masses. While established rock-mass classification systems can be grounded in measurable parameters, field practice also relies on "visual" interpretation, the kind of fast, experience-driven judgement often summarized as the "geologist¿s eye". This thesis explores whether a modern computer-vision foundation model like SAM (Segment anything Model) can be turned into an "electronic pocket geologist" for this setting by learning to (i) classify visually distinct rock-mass structure domains and (ii) segment the corresponding areas directly on outcrop imagery. The study uses photogrammetric bench-slope imagery from the portal area of Zentrum am Berg (Erzberg, Austria). From 4000×2250 pixel Ultra High Definition photographs, ten representative slope images are selected and sliced into 400 × 400 pixel tiles. A total of 503 tiles are manually annotated into three structural quality classes: "good", "medium", and "bad" plus background. This results in a VOC (Visual Object Class) type dataset where the unannotated image serves as SAM input and the annotated image serves as the ground truth mask, giving the true classes and their boundaries. SAM as a masking algorithm returns a segmentation mask for a certain object in the image, designated by a series of prompts. However, since the mask is not assigned any class, two SAM adaptation strategies are implemented and compared. In the "semantic" pathway, SAM predicts a full-image multiclass mask, with every pixel assigned a class, losing its promptable ability. In the "proposal"-based pathway, SAM retains the promptable nature of the original model, returning a mask and its corresponding class based on the user prompt input. The image - mask pairs of the dataset and the semantic/proposal adaptations are used for the purpose of retraining SAM to respond with better masks and additional class features. To keep retraining feasible on limited data and hardware, various types of adapters have been added to the main body of SAM in order to enhance its learning performance. Across experiments, proposal-based SAM variants consistently outperform semantic SAM. The proposal checkpoints achieve routinely higher evaluation metric values in for mean Dice, IoU and recall, as well as more favorable training slopes. Two ideal proposal variants have been identified as having the most potential for further research. A qualitative comparison of the semantic vari- ant against ilastik¿s Random-Forest-based pixel classification shows markedly weaker boundary adherence and shape consistency, reinforcing the advantage of transformer-based segmentation for structurally complex rock textures. The results demonstrate that SAM can be retrained to produce meaningful rock-mass structure segmentation and classification under a constrained hardware and dataset budget. However, because the dataset is geographically and geologically narrow, direct transfer to other rock types is not recommended without further data diversification. Extending the workflow to broader lithologies, image resolutions and eventually to newer SAM variants and real geotechnical targets (e.g. joint/surface-condition cues), is a clear next step.
| Translated title of the contribution | Gesteinsklassensegmentation von 2D Aufschlussaufnahmen mithilfe des sam (segment anything) Modells |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisors/Advisors |
|
| Award date | 26 Jun 2026 |
| Publication status | Published - 2026 |
Bibliographical note
no embargoKeywords
- rock mass classification
- rock mass segmentation
- outcrop imagery
- computer vision
- transformers
- Segment Anything Model (SAM)
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver