Preferencje help
Widoczny [Schowaj] Abstrakt
Liczba wyników

Znaleziono wyników: 10

Liczba wyników na stronie
first rewind previous Strona / 1 next fast forward last
Wyniki wyszukiwania
Wyszukiwano:
w słowach kluczowych:  segmentacja semantyczna
help Sortuj według:

help Ogranicz wyniki do:
first rewind previous Strona / 1 next fast forward last
PL
Artykuł przedstawia koncepcję wykorzystania generatywnego modelu językowego jako inteligentnego planisty tras dla bezzałogowych statków powietrznych (BSP). Zaprojektowany system integruje modele językowe z analizą danych geoprzestrzennych oraz algorytmami trasowania, umożliwiając automatyczne tworzenie tras na podstawie komend tekstowych operatora.
EN
The article presents the concept of using a generative language model as an intelligent planner for unmanned aerial vehicle (UAV) flight paths. The proposed system integrates language models with geospatial data analysis and routing algorithms to automatically generate flight routes based on the operator’s text commands.
PL
Segmentacja semantyczna ma kluczowe znaczenie w zastosowaniu metodyki Heritage Building Information Modelling (HBIM), umożliwiając nie tylko precyzyjną dokumentację geometrii, ale także zachowanie wartości kulturowych i historycznych. Technologie takie jak naziemne skanowanie laserowe (TLS) czy fotogrametria stanowią podstawę do tworzenia tzw. cyfrowych bliźniaków, jednak same modele 3D, pozbawione semantycznego wzbogacenia, nie odzwierciedlają pełnego kontekstu obiektów zabytkowych. Manualna segmentacja, przeprowadzana we współpracy z konserwatorami i historykami architektury, pozostaje nieodzownym elementem procesu HBIM. Pozwala uchwycić niuanse dotyczące materiałów, technik budowlanych czy detali artystycznych, które zautomatyzowane algorytmy często pomijają. W efekcie cyfrowe kopie, wiernie odzwierciedlające nie tylko układ przestrzenny, ale także integralność kulturową i historyczną obiektów, stają się narzędziem niezbędnym do efektywnej konserwacji i wieloaspektowego zarządzania dziedzictwem. W artykule omówiono metodologie i wyzwania związane z segmentacją semantyczną w HBIM, podkreślając jej podwójną rolę: w tworzeniu cyfrowych bliźniaków i wspieraniu współpracy interdyscyplinarnej. Łącząc precyzję geometryczną z głębią semantyczną, HBIM wspiera wieloaspektowość tworzonych modeli i zapewnia zachowanie integralności kulturowej.
EN
Semantic segmentation is crucial in Heritage Building Information Modelling (HBIM), enabling not only precise geometric documentation, but also the preservation of cultural and historical values. Technologies such as terrestrial laser scanning (TLS) or photogrammetry provide the basis for so-called “digital twinning” of heritage, but 3D models alone—lacking semantic enrichment—do not reflect the full context of heritage buildings. Manual segmentation, carried out in collaboration with conservators and architectural historians, remains an indispensable part of the HBIM process. It captures the nuances of materials, construction techniques or artistic details that automated algorithms often miss. The result is digital twins that faithfully reflect not only the spatial layout but also the cultural and historical integrity of the buildings, becoming an essential tool for effective conservation and multi-faceted heritage management. This paper discusses the methodologies and challenges of semantic segmentation in HBIM, highlighting its dual role in creating meaningful digital twins and fostering interdisciplinary collaboration. By combining geometric precision with semantic depth, HBIM supports interdisciplinary collaboration and ensures the preservation of cultural integrity.
EN
The building extraction from remote sensing (RS) images has been a significant area of research in the photogrammetric and remote sensing communities, especially with the development of deep learning for over a decade. With the availability of multi-source data from RS images, accurately identifying buildings with different spatial image resolutions has become a challenging task. In this study, we assessed how the unalignment of image resolution between the training and testing datasets affects the ability to extract buildings. Image resolution plays a crucial role in the performance of building extraction. Our experiments found that as the image resolution decreased from 10 cm to 50 cm, the efficiency of building segmentation reduced from 0.759 to 0.585 according to the IoU metric. Besides, the ability and accuracy of building segmentation significantly decreased when the difference in image resolution between the training and testing datasets increased. In the case study, we use the model trained on a 10 cm resolution dataset to predict for 50 cm resolution data, the IoU drops significantly to 0.299. This research offers important insights into building segmentation tasks using multi-source data from satellite, airborne, and UAV images.
EN
This research presents an application of the Mask R-CNN algorithm for apple detection and semantic segmentation, aiming to enhance automation in the agricultural sector. Despite the growing use of deep learning techniques in object detection tasks, their application in agricultural contexts, specifically for fruit detection and semantic segmentation, remains relatively unexplored. This study evaluates the performance of the Mask R-CNN algorithm through a series of numerical experiments, with metrics including mean intersection over union (mIoU), F1 score, accuracy, and a confusion matrix analysis. Our results demonstrated that the Mask R-CNN model was effective in detecting and segmenting apples with a high degree of precision, achieving an mIoU of 0.551, an F1 score of 0.704, and an accuracy of 0.957. However, areas for potential improvement were also identified, such as reducing the model's false negative rate. This study provides insights into the application of deep learning algorithms in the agricultural sector, paving the way for more efficient and automated fruit harvesting systems.
PL
Artykuł ten przedstawia zastosowanie algorytmu Mask R-CNN do wykrywania i semantycznej segmentacji jabłek, mając na celu zwiększenie automatyzacji w sektorze rolniczym. Pomimo rosnącego wykorzystania technik uczenia głębokiego w zadaniach detekcji obiektów, ich stosowanie w kontekstach rolniczych, szczególnie w wykrywaniu i semantycznej segmentacji owoców, pozostaje stosunkowo niezbadane. Niniejsze badanie ocenia wydajność algorytmu Mask R-CNN poprzez serię eksperymentów numerycznych, wykorzystując metryki takie mIoU, wynik F1, dokładność oraz analizę macierzy pomyłek. Nasze wyniki wykazały, że model Mask R-CNN był skuteczny w wykrywaniu i segmentacji jabłek z dużą dokładnością, osiągając mIoU wynoszące 0.551, wynik F1 równy 0.704 oraz dokładność 0.957. Jednakże zidentyfikowano również obszary potencjalnych ulepszeń, takie jak zmniejszenie fałszywie negatywnego wskaźnika modelu. To badanie dostarcza wglądów w zastosowanie algorytmów uczenia głębokiego w sektorze rolniczym, torując drogę do bardziej wydajnych i zautomatyzowanych systemów zbierania owoców.
EN
The purpose of this article is to present a novel approach for recording information contained in an image in a structured form and performing image similarity assessment with use of these data structures. The solution presented in this document relies on an analysis of results produced by pre-trained semantic segmentation algorithms. These outcomes can be transformed to a set of vectors representing some characteristics of each class of objects detected in the provided image. These data structures can contain meaningful information about algorithm detections, such as the object’s position on the image, the object’s size compared to the overall image size or the object’s dominant colors, etc. Vectors prepared as described previously can be further compared with other image embeddings using many mathematical tools like distance measures. Moreover, the approach described in this article allows the user to define a value of weight tied to each characteristic. This provides the ability to make a subset of features more important than others and have a greater impact on the final value of image similarity.
PL
Celem niniejszego artykułu jest zaprezentowanie nowatorskiego sposobu zapisywania informacji zawartych na obrazach w ustrukturyzowanej formie oraz przeprowadzania procesu szacowania podobieństwa obrazów z użyciem wspomnianych struktur danych. Rozwiązanie zaprezentowane w tym dokumencie opiera swoje działanie na analizie wyników otrzymanych od wstępnie wytrenowanych algorytmów segmentacji semantycznej. Rezultaty te mogą zostać przetransformowane do postaci zbioru wektorów, których wartości będą reprezentowały cechy obiektów wykrytych na dostarczonych obrazach. Takie struktury danych mogą zawierać istotne informacje na temat detekcji algorytmu np.: położenie wykrytego obiektu na obrazie, rozmiar wykrytego obiektu w porównaniu do wielkości całej grafiki, kolor dominujący itp. Przygotowane w taki sposób wektorowe reprezentacje obrazów mogą być porównywane między sobą przy użyciu wielu narzędzi matematycznych takich jak miary odległości. Co więcej zaprezentowane w niniejszym artykule podejście pozwala decydentowi zdefiniować wartość wagi każdej z cech dla poszczególnych klas obiektów. Pozwala to modelować preferencje decyzyjne oraz sprawia, że podzbiór cech obiektów może mieć większy wpływ na ostateczną wartość podobieństwa obrazów od pozostałych parametrów.
6
Content available remote Urban scene semantic segmentation using the U-Net model
EN
Vision-based semantic segmentation of complex urban street scenes is a very important function during autonomous driving (AD), which will become an important technology in industrialized countries in the near future. Today, advanced driver assistance systems (ADAS) improve traffic safety thanks to the application of solutions that enable detecting objects, recognising road signs, segmenting the road, etc. The basis for these functionalities is the adoption of various classifiers. This publication presents solutions utilising convolutional neural networks, such as MobileNet and ResNet50, which were used as encoders in the U-Net model to semantically segment images of complex urban scenes taken from the publicly available Cityscapes dataset. Some modifications of the encoder/decoder architecture of the U-Net model were also proposed and the result was named the MU-Net. During tests carried out on 500 images, the MU-Net model produced slightly better segmentation results than the universal MobileNet and ResNet networks, as measured by the Jaccard index, which amounted to 88.85%. The experiments showed that the MobileNet network had the best ratio of accuracy to the number of parameters used and at the same time was the least sensitive to unusual phenomena occurring in images.
EN
The paper is focused on automatic segmentation task of bone structures out of CT data series of pelvic region. The authors trained and compared four different models of deep neural networks (FCN, PSPNet, U-net and Segnet) to perform the segmentation task of three following classes: background, patient outline and bones. The mean and class-wise Intersection over Union (IoU), Dice coefficient and pixel accuracy measures were evaluated for each network outcome. In the initial phase all of the networks were trained for 10 epochs. The most exact segmentation results were obtained with the use of U-net model, with mean IoU value equal to 93.2%. The results where further outperformed with the U-net model modification with ResNet50 model used as the encoder, trained by 30 epochs, which obtained following result: mIoU measure – 96.92%, “bone” class IoU – 92.87%, mDice coefficient – 98.41%, mDice coefficient for “bone” – 96.31%, mAccuracy – 99.85% and Accuracy for “bone” class – 99.92%.
EN
For brain tumour treatment plans, the diagnoses and predictions made by medical doctors and radiologists are dependent on medical imaging. Obtaining clinically meaningful information from various imaging modalities such as computerized tomography (CT), positron emission tomography (PET) and magnetic resonance (MR) scans are the core methods in software and advanced screening utilized by radiologists. In this paper, a universal and complex framework for two parts of the dose control process – tumours detection and tumours area segmentation from medical images is introduced. The framework formed the implementation of methods to detect glioma tumour from CT and PET scans. Two deep learning pre-trained models: VGG19 and VGG19-BN were investigated and utilized to fuse CT and PET examinations results. Mask R-CNN (region-based convolutional neural network) was used for tumour detection – output of the model is bounding box coordinates for each object in the image – tumour. U-Net was used to perform semantic segmentation – segment malignant cells and tumour area. Transfer learning technique was used to increase the accuracy of models while having a limited collection of the dataset. Data augmentation methods were applied to generate and increase the number of training samples. The implemented framework can be utilized for other use-cases that combine object detection and area segmentation from grayscale and RGB images, especially to shape computer-aided diagnosis (CADx) and computer-aided detection (CADe) systems in the healthcare industry to facilitate and assist doctors and medical care providers.
EN
Background: Breast cancer is a deadly disease responsible for statistical yearly global death. Identification of cancer tumors is quite tasking, as a result, concerted efforts are thus devoted. Clinicians have used ultrasounds as a diagnostic tool for breast cancer, though, poor image quality is a major limitation when segmenting breast ultrasound. To address this problem, we present a semantic segmentation method for breast ultrasound (BUS) images. Method: The BUS images were resized and then enhanced with the contrast limited adaptive histogram equalization method. Subsequently, the variant enhanced block was used to encode the preprocessed image. Finally, the concatenated convolutions produced the segmentation mask. Results: The proposed method was evaluated with two datasets. The datasets contain 264 and 830 BUS images respectively. Dice measure (DM), Jaccard measure, and Hausdroff distance were used to evaluate the methods. Results indicate that the proposed method achieves high DM with 89.73% for malignant and 89.62% for benign BUSs. Moreover, the results obtained validate the capacity of the proposed method to achieve higher DM in comparison with reported methods. Conclusion: The proposed algorithm provides a deep learning segmentation procedure that can segment tumors in BUS images effectively and efficiently.
10
EN
Segmentation of lesions from fundus images is an essential prerequisite for accurate severity assessment of diabetic retinopathy. Due to variation in morphologies, number and size of lesions, the manual grading process becomes extremely challenging and time-consuming. This necessitates the need of an automatic segmentation system that can precisely define the region of interest boundaries and assist ophthalmologists in speedy diagnosis along with diabetic retinopathy severity grading. The paper presents a modified U-Net architecture based on residual network and employs periodic shuffling with sub-pixel convolution initialized to convolution nearest neighbour resize. The proposed architecture has been trained and validated for microaneurysm and hard exudate segmentation on two publicly available datasets namely IDRiD and e-ophtha. For IDRiD dataset, the network obtains 99.88% accuracy, 99.85% sensitivity, 99.95% specificity and dice score of 0.9998 for both microaneurysm and exudate segmentation. Further, when trained on e-ophtha and validated on IDRiD dataset, the network shows 99.98% accuracy, 99.88% sensitivity, 99.89% specificity and dice score of 0.9998 for microaneurysm segmentation. For exudates segmen-tation, the model obtained 99.98% accuracy, 99.88% sensitivity, 99.89% specificity and dice score of 0.9999, when trained on e-ophtha and validated on IDRiD dataset. In comparison to existing literature, the proposed model provides state-of-the-art results for retinal lesion segmentation.
first rewind previous Strona / 1 next fast forward last
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.