Ograniczanie wyników
Czasopisma help
Autorzy help
Lata help
Preferencje help
Widoczny [Schowaj] Abstrakt
Liczba wyników

Znaleziono wyników: 22

Liczba wyników na stronie
first rewind previous Strona / 2 next fast forward last
Wyniki wyszukiwania
Wyszukiwano:
w słowach kluczowych:  semantic segmentation
help Sortuj według:

help Ogranicz wyniki do:
first rewind previous Strona / 2 next fast forward last
EN
A method for semantic segmentation of RGB images captured by UAVs to detect railway infrastructure elements, including tracks, level crossings, and surrounding vegetation is proposed. The study was conducted at the Lukasiewicz Research Network - Institute of Aviation, where a proprietary, manually annotated UAV RGB dataset was created. Five deep neural network architectures were trained and compared: DeepLabV3+, Feature Pyramid Network (FPN), LinkNet, Pyramid Attention Network (PAN) and X-Unet. These models were chosen for their distinct approaches to semantic segmentation and feature processing. Training was performed on a desktop computer with an NVIDIA GeForce RTX 3080 GPU and tests were made also on an NVIDIA Jetson AGX Orin to assess deployment feasibility under real-time conditions. Experimental results confirm the strong performance of the analyzed models in segmenting railway tracks and surrounding vegetation. FPN achieved the highest scores, followed by X-Unet, DeepLabV3+, LinkNet, and PAN. All models operated reliably on the NVIDIA Jetson AGX Orin edge platform. The proposed solution can support remote monitoring of railway infrastructure and vegetation. It can also be adapted to other applications by adjusting the training dataset and object categories. This research demonstrates the potential of deep learning as a powerful tool for analyzing UAV RGB imagery in engineering and environmental contexts.
PL
Artykuł przedstawia koncepcję wykorzystania generatywnego modelu językowego jako inteligentnego planisty tras dla bezzałogowych statków powietrznych (BSP). Zaprojektowany system integruje modele językowe z analizą danych geoprzestrzennych oraz algorytmami trasowania, umożliwiając automatyczne tworzenie tras na podstawie komend tekstowych operatora.
EN
The article presents the concept of using a generative language model as an intelligent planner for unmanned aerial vehicle (UAV) flight paths. The proposed system integrates language models with geospatial data analysis and routing algorithms to automatically generate flight routes based on the operator’s text commands.
EN
Under the long-term action of train loads and complex environmental conditions, the surfaces of railway tracks are prone to defects such as cracks, spalling, and pitting, which seriously threaten the safety of railway operations. Semantic segmentation can achieve pixel-level positioning and morphological characterization of defects. However, existing methods still struggle to model strongly directional structures and multi-scale defects while maintaining a balance between accuracy and efficiency in rail-surface inspection. To address the above issues, this paper proposes a lightweight semantic segmentation network for railway track surface defects (SFAB-Net) based on spatial fusion and adaptive bottleneck feature enhancement. This network effectively characterizes the features of slender cracks along the rail direction using the direction-sensitive Spatial-Fusion module and combines them with the simplified spatial pyramid pooling module to achieve multi-scale context aggregation. In the decoding stage, an adaptive feature reconstruction mechanism and spatial-channel joint attention are introduced to enhance multi-scale feature fusion and suppress background interference. Experimental results on the NEU-DET dataset and a self-built rail surface image dataset show that SFAB-Net outperforms several representative methods in segmentation accuracy and robustness, and has strong potential for engineering applications.
PL
Segmentacja semantyczna ma kluczowe znaczenie w zastosowaniu metodyki Heritage Building Information Modelling (HBIM), umożliwiając nie tylko precyzyjną dokumentację geometrii, ale także zachowanie wartości kulturowych i historycznych. Technologie takie jak naziemne skanowanie laserowe (TLS) czy fotogrametria stanowią podstawę do tworzenia tzw. cyfrowych bliźniaków, jednak same modele 3D, pozbawione semantycznego wzbogacenia, nie odzwierciedlają pełnego kontekstu obiektów zabytkowych. Manualna segmentacja, przeprowadzana we współpracy z konserwatorami i historykami architektury, pozostaje nieodzownym elementem procesu HBIM. Pozwala uchwycić niuanse dotyczące materiałów, technik budowlanych czy detali artystycznych, które zautomatyzowane algorytmy często pomijają. W efekcie cyfrowe kopie, wiernie odzwierciedlające nie tylko układ przestrzenny, ale także integralność kulturową i historyczną obiektów, stają się narzędziem niezbędnym do efektywnej konserwacji i wieloaspektowego zarządzania dziedzictwem. W artykule omówiono metodologie i wyzwania związane z segmentacją semantyczną w HBIM, podkreślając jej podwójną rolę: w tworzeniu cyfrowych bliźniaków i wspieraniu współpracy interdyscyplinarnej. Łącząc precyzję geometryczną z głębią semantyczną, HBIM wspiera wieloaspektowość tworzonych modeli i zapewnia zachowanie integralności kulturowej.
EN
Semantic segmentation is crucial in Heritage Building Information Modelling (HBIM), enabling not only precise geometric documentation, but also the preservation of cultural and historical values. Technologies such as terrestrial laser scanning (TLS) or photogrammetry provide the basis for so-called “digital twinning” of heritage, but 3D models alone—lacking semantic enrichment—do not reflect the full context of heritage buildings. Manual segmentation, carried out in collaboration with conservators and architectural historians, remains an indispensable part of the HBIM process. It captures the nuances of materials, construction techniques or artistic details that automated algorithms often miss. The result is digital twins that faithfully reflect not only the spatial layout but also the cultural and historical integrity of the buildings, becoming an essential tool for effective conservation and multi-faceted heritage management. This paper discusses the methodologies and challenges of semantic segmentation in HBIM, highlighting its dual role in creating meaningful digital twins and fostering interdisciplinary collaboration. By combining geometric precision with semantic depth, HBIM supports interdisciplinary collaboration and ensures the preservation of cultural integrity.
EN
Effective multi-scale feature representation and focused attention on critical objects are essential for accurate perception of waterborne navigation scenes. To address the insufficient exploitation of multi-scale information in existing methods that leads to imprecise segmentation, this study proposes a real-time semantic segmentation method for waterborne navigation scenes through multi-scale information enhancement and importance-weighted optimization. First, DDRNet-23-slim is selected as the backbone network for feature extraction. An edge-guided branch is embedded into its shallow layers, and a Dynamic Feature Fusion Module (DFFM) is constructed by integrating a lightweight hybrid attention mechanism, effectively enhancing multi-scale feature interaction capabilities. Second, the loss function is improved using an importance-weighted strategy to prioritize critical objects during training. Finally, a parameter-free attention mechanism is introduced in the upsampling stage, maintaining real-time performance while ensuring segmentation stability for key objects under complex background interference. Evaluations on the On_Water and Seaships datasets demonstrate that the proposed method achieves mIoU scores of 83.1% and 73.2%, respectively, with ship segmentation accuracy reaching 88.2% on On_Water. The inference speed attains 69.1 FPS, outperforming mainstream real-time segmentation models (e.g., DDRNet, STDC) in balancing accuracy and efficiency. Notably, it exhibits stronger robustness in complex inland river scenarios with dense shore structures and numerous small targets.
EN
The global demand rise in lithium, driven by the expansion of the new energy sector and electric vehicle markets, underscores the urgent need for efficient mineral processing of spodumene, the primary lithium ore. Traditional methods encounter challenges such as low-grade ores, fine particle sizes, complex mineral intergrowth, and high energy consumption with notable environmental impacts. This study addresses these issues by developing an intelligent X-ray Transmission (XRT) pre-sorting system enhanced with the novel Mineral-DeepLabV3+ deep learning algorithm. The approach integrates machine vision and artificial intelligence to deliver lightweight, precise ore segmentation under industrial mining conditions, incorporating innovations such as the MobileNetV4 backbone, optimized Atrous Spatial Pyramid Pooling module, and CSWin attention mechanism. Key findings demonstrate that the Mineral-DeepLabV3+ model achieves a mean intersection over union of 95.19 %, boosts segmentation accuracy by 3.28 %, and reduces parameter count by over 44 % compared to baseline models, while maintaining fast inference speed. Semi-industrial trials confirm its superior performance across varying ore sizes, processing rates, and conveyor speeds, achieving Li₂O concentrate grades up to 2.27% and waste rejection rates up to 66%. This technology significantly enhances resource efficiency, reduces environmental footprint, and advances operational sustainability in spodumene beneficiation. The proposed framework offers a scalable solution for driving low-carbon, intelligent practices in mineral processing.
EN
Accurate segmentation of leaf regions plays a vital role in plant phenotyping and agricultural analysis. This paper presents AKDUNet, a lightweight UNet-based architecture that integrates attention gates and knowledge distillation to improve segmentation performance while minimizing computational complexity. The architecture replaces traditional skip connections with attention gates to focus on salient spatial features and employs a two-stage training pipeline, where a compact student model learns from a deeper teacher model using a tailored distillation loss function. AKDUNet is evaluated on two benchmark datasets (CWFID and Sunflower) and outperforms a range of state-of-the-art models, including UNet++, Inception UNet, VGG-based UNets, SDUNet, INSCA UNet, and SegFormer. Ablation studies confirm the advantages of attention modules, and qualitative analyses using Grad-CAM visualizations reveal the model’s ability to effectively focus on crucial leaf structures. The results demonstrate that AKDUNet is not only computationally efficient but also highly accurate, making it suitable for real-time deployment in resource-constrained agricultural environments.
EN
The building extraction from remote sensing (RS) images has been a significant area of research in the photogrammetric and remote sensing communities, especially with the development of deep learning for over a decade. With the availability of multi-source data from RS images, accurately identifying buildings with different spatial image resolutions has become a challenging task. In this study, we assessed how the unalignment of image resolution between the training and testing datasets affects the ability to extract buildings. Image resolution plays a crucial role in the performance of building extraction. Our experiments found that as the image resolution decreased from 10 cm to 50 cm, the efficiency of building segmentation reduced from 0.759 to 0.585 according to the IoU metric. Besides, the ability and accuracy of building segmentation significantly decreased when the difference in image resolution between the training and testing datasets increased. In the case study, we use the model trained on a 10 cm resolution dataset to predict for 50 cm resolution data, the IoU drops significantly to 0.299. This research offers important insights into building segmentation tasks using multi-source data from satellite, airborne, and UAV images.
EN
Semantic segmentation is important for robots navigating with 3D LiDARs, but the generation of training datasets requires tedious manual effort. In this paper, we introduce a set of strategies to efficiently generate large datasets by combining real and synthetic data samples. More specifically, the method populates recorded empty scenes with navigation-relevant obstacles generated synthetically, thus combining two domains: real life and synthetic. Our approach requires no manual annotation, no detailed knowledge about actual data feature distribution, and no real-life data of objects of interest. We validate the proposed method in the underground parking scenario and compare it with available open-source datasets. The experiments show superiority to the off-the-shelf datasets containing similar data characteristics but also highlight the difficulty of achieving the level of manually annotated datasets. We also show that combining generated and annotated data improves the performance visibly, especially for cases with rare occurrences of objects of interest. Our solution is suitable for direct application in robotic systems.
EN
This research presents an application of the Mask R-CNN algorithm for apple detection and semantic segmentation, aiming to enhance automation in the agricultural sector. Despite the growing use of deep learning techniques in object detection tasks, their application in agricultural contexts, specifically for fruit detection and semantic segmentation, remains relatively unexplored. This study evaluates the performance of the Mask R-CNN algorithm through a series of numerical experiments, with metrics including mean intersection over union (mIoU), F1 score, accuracy, and a confusion matrix analysis. Our results demonstrated that the Mask R-CNN model was effective in detecting and segmenting apples with a high degree of precision, achieving an mIoU of 0.551, an F1 score of 0.704, and an accuracy of 0.957. However, areas for potential improvement were also identified, such as reducing the model's false negative rate. This study provides insights into the application of deep learning algorithms in the agricultural sector, paving the way for more efficient and automated fruit harvesting systems.
PL
Artykuł ten przedstawia zastosowanie algorytmu Mask R-CNN do wykrywania i semantycznej segmentacji jabłek, mając na celu zwiększenie automatyzacji w sektorze rolniczym. Pomimo rosnącego wykorzystania technik uczenia głębokiego w zadaniach detekcji obiektów, ich stosowanie w kontekstach rolniczych, szczególnie w wykrywaniu i semantycznej segmentacji owoców, pozostaje stosunkowo niezbadane. Niniejsze badanie ocenia wydajność algorytmu Mask R-CNN poprzez serię eksperymentów numerycznych, wykorzystując metryki takie mIoU, wynik F1, dokładność oraz analizę macierzy pomyłek. Nasze wyniki wykazały, że model Mask R-CNN był skuteczny w wykrywaniu i segmentacji jabłek z dużą dokładnością, osiągając mIoU wynoszące 0.551, wynik F1 równy 0.704 oraz dokładność 0.957. Jednakże zidentyfikowano również obszary potencjalnych ulepszeń, takie jak zmniejszenie fałszywie negatywnego wskaźnika modelu. To badanie dostarcza wglądów w zastosowanie algorytmów uczenia głębokiego w sektorze rolniczym, torując drogę do bardziej wydajnych i zautomatyzowanych systemów zbierania owoców.
EN
In computer vision, Convolutional Neural Networks (CNNs) have become a foundation for image analysis. They excel in tasks such as object recognition, classification, and more, semantic segmentation. In order to achieve better accuracy, it is crucial to apply normalization techniques to the network for enhancing overall performance. This paper introduces an innovative approach that incorporates Batch Group Normalization (BGN) into the popular U-Net for binary semantic segmentation, with a particular focus on aerial road detection. Our research primarily focuses on evaluating the BGN-UNet’s performance compared to traditional normalization techniques, such as Batch Normalization (BN) and Group Normalization (GN). With a batch size of 2, the U-Net model enhanced with Batch Group Normalization (BGN-UNet) achieves a remarkable Mean IoU of 98.4% in aerial road segmentation, demonstrating its superior accuracy in this task.
EN
Automatic crack detection in construction facilities is a challenging yet crucial task. However, existing deep learning (DL)-based semantic segmentation methods for this field are based on fully supervised learning models and pixel-level manual annotation, which are time-consuming and labor-intensive. To solve this problem, this paper proposes a novel crack semantic segmentation network using weakly supervised approach and mixed-label training strategy. Firstly, an image patch-level classifier of crack is trained to generate a coarse localization map for automatic pseudo-labeling of cracks combined with a thresholding-based method. Then, we integrated the pseudo-annotated with manual-annotated samples with a ratio of 4:1 to train the crack segmentation network with a mixed-label training strategy, in which the manual labels were assigned with a higher weight value. The experimental data on two public datasets demonstrate that our proposed method achieves a comparable accuracy with the fully supervised methods, reducing over 65% of the manual annotation workload.
EN
The purpose of this article is to present a novel approach for recording information contained in an image in a structured form and performing image similarity assessment with use of these data structures. The solution presented in this document relies on an analysis of results produced by pre-trained semantic segmentation algorithms. These outcomes can be transformed to a set of vectors representing some characteristics of each class of objects detected in the provided image. These data structures can contain meaningful information about algorithm detections, such as the object’s position on the image, the object’s size compared to the overall image size or the object’s dominant colors, etc. Vectors prepared as described previously can be further compared with other image embeddings using many mathematical tools like distance measures. Moreover, the approach described in this article allows the user to define a value of weight tied to each characteristic. This provides the ability to make a subset of features more important than others and have a greater impact on the final value of image similarity.
PL
Celem niniejszego artykułu jest zaprezentowanie nowatorskiego sposobu zapisywania informacji zawartych na obrazach w ustrukturyzowanej formie oraz przeprowadzania procesu szacowania podobieństwa obrazów z użyciem wspomnianych struktur danych. Rozwiązanie zaprezentowane w tym dokumencie opiera swoje działanie na analizie wyników otrzymanych od wstępnie wytrenowanych algorytmów segmentacji semantycznej. Rezultaty te mogą zostać przetransformowane do postaci zbioru wektorów, których wartości będą reprezentowały cechy obiektów wykrytych na dostarczonych obrazach. Takie struktury danych mogą zawierać istotne informacje na temat detekcji algorytmu np.: położenie wykrytego obiektu na obrazie, rozmiar wykrytego obiektu w porównaniu do wielkości całej grafiki, kolor dominujący itp. Przygotowane w taki sposób wektorowe reprezentacje obrazów mogą być porównywane między sobą przy użyciu wielu narzędzi matematycznych takich jak miary odległości. Co więcej zaprezentowane w niniejszym artykule podejście pozwala decydentowi zdefiniować wartość wagi każdej z cech dla poszczególnych klas obiektów. Pozwala to modelować preferencje decyzyjne oraz sprawia, że podzbiór cech obiektów może mieć większy wpływ na ostateczną wartość podobieństwa obrazów od pozostałych parametrów.
EN
Satellite imagery plays an important role in detecting algal blooms because of its ability to cover larger geographical regions. Excess growth of Sea surface algae, characterized by the presence of Chlorophyll-a (Chl-a), is considered to be harmful. The detection of algal growth at an earlier stage may prevent hazardous effects on the aquatic environment. Semantic segmentation of algal blooms is helpful in the quantization of algal blooms. A rule-based semantic segmentation approach for the segregation of sea surface algal blooms is proposed. Bloom concentrations are classified into three different concentrations, namely, low, medium, and high. The chl_nn band in the Sentinel-3 satellite images is used for experimentation. The chl_nn band has exclusive details of the presence of chlorophyll concentrations. A dataset is proposed for the semantic segmentation of algal blooms. The devised rule-based semantic segmentation approach has produced an average accuracy of 98%. A set of 100 images is randomly selected for testing. The tests are repeated on 5 different image sets. The results are validated by the pixel comparison method. The proposed work is compared with other relevant works. The Arabian Sea near the coastal districts of Udupi and Mangaluru has been considered as the area of study. The methodology can be adapted to monitor the life cycle of blooms and their hazardous effects on aquatic life.
EN
Composite materials are prone to various kinds of defects in their service life, among which delamination is a very hazardous type of damage. The traditional visual inspection techniques often fail to detect delamination in composite structures. Guided Lamb waves are increasingly being applied for the identification of delamination in these structures. Scanning laser Doppler vibrometry can measure the full wavefield of guided Lamb waves, such full wavefield contains rich information about defects. In this research work, a novel deep learning-based semantic segmentation technique is applied for delamination identification on full wavefield data. A big dataset of full wavefield images resulting from the interaction with delamination of random shape, size, and location was utilised and fed into the proposed deep learning model. The main motive of this research work is to investigate the applicability of deep learning-based approach for delamination identification in composite structures by using only the animations of guided Lamb waves. It is verified that the performance of the proposed deep learning model is good. Moreover, it enables better automation of identification of delamination, which can further produce damage maps without the intervention of the user. Furthermore, the developed deep learning model also indicates the capability of generalising well to the experimental data.
EN
Mushrooms are a rich source of antioxidants and nutritional values. Edible mushrooms, however, are susceptible to various diseases such as dry bubble, wet bubble, cobweb, bacterial blotches, and mites. Farmers face significant production losses due to these diseases affecting mushrooms. The manual detection of these diseases relies on expertise, knowledge of diseases, and human effort. Therefore, there is a need for computer-aided methods, which serve as optimal substitutes for detecting and segmenting diseases. In this paper, we propose a semantic segmentation approach based on the Random Forest machine learning technique for the detection and segmentation of mushroom diseases. Our focus lies in extracting a combination of different features, including Gabor, Bouda, Kayyali, Gaussian, Canny edge, Roberts, Sobel, Scharr, Prewitt, Median, and Variance. We employ constant mean-variance thresholding and the Pearson correlation coefficient to extract significant features, aiming to enhance computational speed and reduce complexity in training the Random Forest classifier. Our results indicate that semantic segmentation based on Random Forest outperforms other methods such as Support Vector Machine (SVM), Naïve Bayes, K-means, and Region of Interest in terms of accuracy. Additionally, it exhibits superior precision, recall, and F1 score compared to SVM. It is worth noting that deep learning-based semantic segmentation methods were not considered due to the limited availability of diseased mushroom images.
17
Content available remote Urban scene semantic segmentation using the U-Net model
EN
Vision-based semantic segmentation of complex urban street scenes is a very important function during autonomous driving (AD), which will become an important technology in industrialized countries in the near future. Today, advanced driver assistance systems (ADAS) improve traffic safety thanks to the application of solutions that enable detecting objects, recognising road signs, segmenting the road, etc. The basis for these functionalities is the adoption of various classifiers. This publication presents solutions utilising convolutional neural networks, such as MobileNet and ResNet50, which were used as encoders in the U-Net model to semantically segment images of complex urban scenes taken from the publicly available Cityscapes dataset. Some modifications of the encoder/decoder architecture of the U-Net model were also proposed and the result was named the MU-Net. During tests carried out on 500 images, the MU-Net model produced slightly better segmentation results than the universal MobileNet and ResNet networks, as measured by the Jaccard index, which amounted to 88.85%. The experiments showed that the MobileNet network had the best ratio of accuracy to the number of parameters used and at the same time was the least sensitive to unusual phenomena occurring in images.
EN
The liver is a vital organ of the human body and hepatic cancer is one of the major causes of cancer deaths. Early and rapid diagnosis can reduce the mortality rate. It can be achieved through computerized cancer diagnosis and surgery planning systems. Segmentation plays a major role in these systems. This work evaluated the efficacy of the SegNet model in liver and particle swarm optimization-based clustering technique in liver lesion segmentation. Over 2400 CT images were used for training the deep learning network and ten CT datasets for validating the algorithm. The segmentation results were satisfactory. The values for Dice Coefficient and volumetric overlap error achieved were 0.940 ± 0.022 and 0.112 ± 0.038, respectively for liver and the results for lesion delineation were 0.4629 ± 0.287 and 0.6986 ± 0.203, respectively. The proposed method is effective for liver segmentation. However, lesion segmentation needs to be further improved for better accuracy.
EN
The paper is focused on automatic segmentation task of bone structures out of CT data series of pelvic region. The authors trained and compared four different models of deep neural networks (FCN, PSPNet, U-net and Segnet) to perform the segmentation task of three following classes: background, patient outline and bones. The mean and class-wise Intersection over Union (IoU), Dice coefficient and pixel accuracy measures were evaluated for each network outcome. In the initial phase all of the networks were trained for 10 epochs. The most exact segmentation results were obtained with the use of U-net model, with mean IoU value equal to 93.2%. The results where further outperformed with the U-net model modification with ResNet50 model used as the encoder, trained by 30 epochs, which obtained following result: mIoU measure – 96.92%, “bone” class IoU – 92.87%, mDice coefficient – 98.41%, mDice coefficient for “bone” – 96.31%, mAccuracy – 99.85% and Accuracy for “bone” class – 99.92%.
EN
For brain tumour treatment plans, the diagnoses and predictions made by medical doctors and radiologists are dependent on medical imaging. Obtaining clinically meaningful information from various imaging modalities such as computerized tomography (CT), positron emission tomography (PET) and magnetic resonance (MR) scans are the core methods in software and advanced screening utilized by radiologists. In this paper, a universal and complex framework for two parts of the dose control process – tumours detection and tumours area segmentation from medical images is introduced. The framework formed the implementation of methods to detect glioma tumour from CT and PET scans. Two deep learning pre-trained models: VGG19 and VGG19-BN were investigated and utilized to fuse CT and PET examinations results. Mask R-CNN (region-based convolutional neural network) was used for tumour detection – output of the model is bounding box coordinates for each object in the image – tumour. U-Net was used to perform semantic segmentation – segment malignant cells and tumour area. Transfer learning technique was used to increase the accuracy of models while having a limited collection of the dataset. Data augmentation methods were applied to generate and increase the number of training samples. The implemented framework can be utilized for other use-cases that combine object detection and area segmentation from grayscale and RGB images, especially to shape computer-aided diagnosis (CADx) and computer-aided detection (CADe) systems in the healthcare industry to facilitate and assist doctors and medical care providers.
first rewind previous Strona / 2 next fast forward last
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.