Accurate indoor camera localization is crucial for applications in augmented reality, robotics, and autonomous navigation. While single-image deep learning models for 6-DOF pose regression have shown competitive results on established benchmarks, their development still requires extensive data annotation and hyperparameter tuning. In this work, we investigate the combination of advanced network architectures, transfer learning, and synthetic data to improve single-image indoor pose regression. Our approach employs a ResNet50 backbone pre-trained on the Places365 dataset and further trained and evaluated on established benchmarks. To enhance the training data, synthetic images are generated from 3D BIM models using Unreal Engine, with alignment procedures ensuring accurate correspondence between synthetic and real environments. Real RGB images are preprocessed to resemble synthetic data, enabling effective cross-domain evaluation. Experiments demonstrate that both architectural design and pretraining significantly influence model performance. On the UniMelb dataset (real-to-real scenario), the model achieves 0.21 m and 0.80° errors, surpassing baseline accuracy. We also present cross-validation and synthetic-to-synthetic experiments, providing insights into factors affecting performance and interactions between architecture, pretraining, and dataset characteristics.
Artykuł przedstawia możliwości wykorzystania technik uczenia maszynowego oraz danych teledetekcyjnych w procesie aktualizacji Bazy Danych Obiektów Topograficznych BDOT10k. Badania przeprowadzono na obszarze dzielnicy Dębniki w Krakowie, wykorzystując wysokorozdzielczą ortofotomapę lotniczą, dane z lotniczego skanowania laserowego (ALS) oraz referencyjne dane BDOT10k. Automatyczna detekcja budynków została wykonana w środowisku ArcGIS Pro, z zastosowaniem modelu Deep Learning, co pozwoliło na identyfikację 1351 obiektów budowlanych w porównaniu do 1250 budynków zarejestrowanych w bazie referencyjnej. Przeprowadzono analizę zgodności geometrycznej wykrytych obiektów z danymi BDOT10k oraz ocenę różnic powierzchniowych. Dodatkowo, na podstawie danych ALS oszacowano liczbę kondygnacji budynków, przyjmując wysokość jednej kondygnacji równą 3 m. Uzyskane wyniki wskazują, że integracja danych teledetekcyjnych i algorytmów uczenia maszynowego może stanowić narzędzie wspomagające proces aktualizacji krajowych baz danych przestrzennych.
EN
The article presents the potential of using machine learning techniques and remote sensing data in the process of updating the Topographic Objects Database (BDOT10k). The study was conducted in the Dębniki district of Kraków using highresolution aerial orthophotos, airborne laser scanning (ALS) data, and reference BDOT10k data. Automatic building detection was performed in the ArcGIS Pro environment using a deep learning model, which enabled the identification of 1,351 building objects compared to 1,250 buildings recorded in the reference database. A geometric consistency analysis between the detected objects and the BDOT10k data was carried out, along with an assessment of area differences. In addition, the number of building storeys was estimated based on ALS data, assuming a storey height of 3 m. The obtained results indicate that the integration of remote sensing data and machine learning algorithms may constitute a supporting tool for the process of updating national spatial databases.
In response to the need for rapid and precise chemical identification, this work presents a novel system that synergizes optical tomography with machine learning. A deep residual neural network (ResNet) was trained using Fourier transform infrared spectroscopy (FTIR) measurements from 120 distinct substances. The system's efficacy is confirmed by its high recognition quality, marked by a top-3 accuracy of 0.88 (defined as the correct class being among the three highest probability predictions) and an AUC of 0.99. Offering significant improvements in spectral feature extraction and classification accuracy over existing approaches, this system is well-suited for demanding applications such as industrial quality control and environmental monitoring.
PL
W odpowiedzi na potrzebę szybkiej i precyzyjnej identyfikacji chemicznej, w niniejszej pracy przedstawiono nowatorski system, który łączy tomografię optyczną z uczeniem maszynowym. Głęboka resztkowa sieć neuronowa (ResNet) została wytrenowana z wykorzystaniem pomiarów w podczerwieni z transformacją Fouriera (FTIR) dla 120 różnych substancji. Skuteczność systemu potwierdza wysoka jakość rozpoznawania, charakteryzująca się dokładnością top-3 na poziomie 0.88 (zdefiniowaną jako obecność właściwej klasy w trzech najwyżej ocenionych prawdopodobieństwach) oraz wartością AUC 0.99. Oferując znaczącą poprawę w ekstrakcji cech widmowych i dokładności klasyfikacji w porównaniu z istniejącymi podejściami, system ten jest doskonale przystosowany do wymagających zastosowań, takich jak przemysłowa kontrola jakości i monitoring środowiska.
This paper proposes a secure deep learning based analytical approach for voltage stability classification in smart grid systems. The approach integrates machine learning (ML) and deep learning (DL) models to estimate the Fast Voltage Stability Index (FVSI) of the IEEE 30-bus power system under varying operating conditions. Several classification techniques are evaluated, including Support Vector Machines (SVM), Naïve Bayes, K-Nearest Neighbors (KNN), and a Convolutional Neural Network (CNN). A comparative analysis demonstrates the effectiveness of the CNN-based model in capturing complex nonlinear relationships for accurate voltage stability classification.
PL
W niniejszej pracy zaproponowano bezpieczny algorytm analityczny oparty na głębokim uczeniu do klasyfikacji stabilności napięciowej w inteligentnych systemach elektroenergetycznych. Algorytm integruje modele uczenia maszynowego (ML) oraz głębokiego uczenia (DL) w celu predykcji wskaźnika Fast Voltage Stability Index (FVSI) dla testowego systemu IEEE 30-bus w zmiennych warunkach pracy. Przeanalizowano różne techniki klasyfikacyjne, w tym Support Vector Machines (SVM), Naïve Bayes, K-Nearest Neighbors (KNN), a model oparty na konwolucyjnej sieci neuronowej (CNN) osiągnął najwyższą skuteczność.
YOLO object detectors recently became a key component of vision systems in many domains. The family of available YOLO models consists of multiple versions, each in various variants. The research reported in this paper aims to validate the applicability of members of this family to detect objects located within the robot workspace. In our experiments, we used our custom dataset and the COCO2017 dataset. To test the robustness of investigated detectors, the images of these datasets were subject to distortions. The results of our experiments, including variations of training/testing configurations and models, may support the choice of the appropriate YOLO version for robotic vision tasks.
PL
Detektory obiektów YOLO stały się ostatnimi czasy kluczowym elementem systemów wizyjnych w wielu dziedzinach. Rodzina dostępnych modeli YOLO składa się z wielu wersji, z których każda występuje w różnych wariantach. Badania opisane w niniejszej pracy mają na celu zweryfikowanie przydatności członków tej rodziny do wykrywania obiektów znajdujących się w przestrzeni roboczej robota. W eksperymentach wykorzystano nasz własny zbiór danych oraz zbiór COCO2017. Aby przetestować odporność badanych detektorów, obrazy z tych zbiorów poddano zniekształceniom. Wyniki eksperymentów, uwzględniające różne konfiguracje treningowe/testowe oraz modele, mogą stanowić wsparcie przy wyborze odpowiedniej wersji YOLO dla zadań związanych z wizją robotyczną.
A method for semantic segmentation of RGB images captured by UAVs to detect railway infrastructure elements, including tracks, level crossings, and surrounding vegetation is proposed. The study was conducted at the Lukasiewicz Research Network - Institute of Aviation, where a proprietary, manually annotated UAV RGB dataset was created. Five deep neural network architectures were trained and compared: DeepLabV3+, Feature Pyramid Network (FPN), LinkNet, Pyramid Attention Network (PAN) and X-Unet. These models were chosen for their distinct approaches to semantic segmentation and feature processing. Training was performed on a desktop computer with an NVIDIA GeForce RTX 3080 GPU and tests were made also on an NVIDIA Jetson AGX Orin to assess deployment feasibility under real-time conditions. Experimental results confirm the strong performance of the analyzed models in segmenting railway tracks and surrounding vegetation. FPN achieved the highest scores, followed by X-Unet, DeepLabV3+, LinkNet, and PAN. All models operated reliably on the NVIDIA Jetson AGX Orin edge platform. The proposed solution can support remote monitoring of railway infrastructure and vegetation. It can also be adapted to other applications by adjusting the training dataset and object categories. This research demonstrates the potential of deep learning as a powerful tool for analyzing UAV RGB imagery in engineering and environmental contexts.
Background: Recent advancements in supply chain management, supported by information technology, have enabled reductions in inventory levels, among other operational improvements. Nevertheless, inventory-related challenges persist. Economic factors, particularly various forms of uncertainty, often necessitate holding inventory to ensure product availability. Demand variability, which is frequently unpredictable, remains a major challenge in numerous industries and requires the creation of safety buffers. Addressing this issue calls for increasingly sophisticated forecasting methods within replenishment models. Forecasts based solely on traditional time series methods offer limited improvements, whereas advanced approaches using machine learning and deep neural networks provide significantly greater potential. These models are capable of identifying factors that influence customer purchasing decisions, leading to more accurate demand forecasts and, consequently, a stronger foundation for improving replenishment processes. Objective: The primary aim of the research was to illustrate the extent to which advanced demand forecasting models can improve replenishment efficiency, particularly by reducing inventory levels. Methods: Simulation techniques were employed to replicate the replenishment process under various forecasting scenarios based on historical data. This dataset consisted of demand patterns for 1,000 Walmart stock-keeping units (SKUs), publicly released by the retailer for research purposes. The forecast methods examined included a benchmark arithmetic mean model, seven traditional time series-based models, and five advanced models employing machine learning and deep learning techniques. All simulations were conducted using the Reorder Cycle replenishment model with a uniform inventory review cycle across products. The second control parameter, the maximum inventory level (S), was fixed for the benchmark model and dynamically adjusted for the remaining twelve models according to their respective forecasts. A total of 13,000 replenishment simulation cycles were performed. Key performance indicators included average inventory and the service level (SLα), defined as the probability of fully satisfying demand within a replenishment cycle. The parameter S was calibrated to ensure a consistent service level across models. Consequently, the inventory level index was the primary measure of replenishment efficiency, enabling comparisons between the twelve forecasting models and the benchmark. Additionally, the relationship between this index and forecast quality improvement was analyzed using forecast error measurements, specifically the root mean square error (RMSE). Results: The findings confirm that inventory levels can be reduced by an average of 10% through the application of machine learning and deep neural network-based forecasting methods, without compromising service quality. The magnitude of the reduction varied depending on specific temporal demand patterns. Conclusions: The observed improvements can be attributed to two main factors: increased forecast accuracy and the dynamic adjustment of the maximum inventory level (S), based on current forecasts and their associated error estimates. Furthermore, replenishment efficiency may be enhanced further by selecting the most appropriate forecasting method for individual products during specific time periods.
Facial recognition technology finds applications in security, surveillance, and social media. Existing research explores the use of machine learning and deep learning for face recognition, emphasizing the need for improved accuracy. This paper proposes a system for suspect identification using facial recognition. The system leverages ensemble learning by integrating seamlessly with OpenAI’s advanced technologies and is supported by a robust cloud infrastructure. The comparison of the proposed ensemble model to individual models like VGG-Face, Facenet, Facenet512, Deepface, DeepID, ArcFace, and SFace uses multiple detectors and the Labelled Faces in the Wild (LFW) dataset. The results show that the ensemble model offers the most efficient processing time across all sample sizes. In contrast, models like VGG-Face and DeepID exhibit a steeper increase in processing time, suggesting lower scalability. For instance, at a sample size of 50, the local test completes in 61.3 seconds, while the cloud API test takes 67.2 seconds. This highlights the faster processing speed of the local test across all sample sizes. FaceNet, VGG-Face, and ArcFace models are chosen in ensemble model where in all of them have accuracy above 95% in every face detector test. Facenet512 model has 98.4% among the selected ensemble model whereas ensemble of these models shows 98.8 accuracy.
Skin lesion segmentation is a critical task in dermatology, essential for the accurate diagnosis and treatment of various skin conditions, including skin cancer. The precise identification of lesion boundaries in medical images significantly helps in early detection and effective management of these conditions. In this study, a U-Net model was employed to perform segmentation of skin lesions, using its advanced encoder-decoder architecture and skip connections to capture fine details and spatial hierarchies within the images. The model yielded an overall accuracy of 95.47%, a precision of 97.21%, a recall of 84.04%, and an F1 score of 90.15%. These results affirm the U-Net model’s proficiency in accurately segmenting skin lesions. The precise boundary delineation provided by the model can help healthcare professionals detect malignant lesions, which often exhibit irregular boundaries, thereby improving diagnostic accuracy. The integration of this model into clinical practice can enhance diagnostic accuracy and efficiency, reduce the workload on health-care professionals, and improve patient outcomes. The promising performance of the U-Net model emphasizes its potential to revolutionize dermatological diagnostics and support healthcare professionals in delivering timely and precise patient care.
Dynamic infrared thermography is emerging as a noninvasive technique for monitoring microvascular health, yet its interpretation remains largely qualitative and labor-intensive. This work systematically benchmarks four deep learning architectures: 2D CNN, 3D CNN, CNN-LSTM, and CNN-Transformer, evaluated for automated DIRT sequence classification in a clinically relevant cohort of post-COVID-19 and post-myocardial infarction patients. The study introduces a rigorous pipeline encompassing thermal image acquisition, standardized preprocessing, tailored data augmentation, and stratified cross-validation to ensure reliable evaluation. Purely spatial models such as the 2D CNN underperform, achieving a macro F1 score of 73.5% and accuracy of 80.1%, while temporally aware models yield substantial gains: CNN-LSTM reaches a macro F1 score of 91.4% and accuracy of 92.7%, and the CNN-Transformer achieves 88.8% and 90.6% prior to hyperparameter optimization. After automated hyperparameter optimization, both models converge to a macro F1 score of 93.8% and accuracy of 94.8%, with the Transformer requiring less than half the parameters. Functional ANOVA analysis highlights that learning rate is the most influential factor for LSTM tuning, while dropout dominates for the Transformer. These findings establish a foundation for robust, sequence-aware DIRT analysis, demonstrating that modern deep learning models, when rigorously validated, can transform DIRT into a quantitative biomarker for longitudinal vascular assessment.
This paper investigates the robustness of traffic sign classification models against real-world visual disturbances. We conduct a comparative evaluation of three distinct architectures: a standard CNN, a hybrid CNN enhanced with Kolmogorov-Arnold dense layers (CNN-KAN), and a fully convolutional Kolmogorov-Arnold Network (CKAN). Unlike traditional CNNs, the KA-based models utilize learnable activation functions, potentially offering improved resilience. The experiments were conducted using the German Traffic Sign Recognition Benchmark (GTSRB) dataset, containing 43 classes of traffic signs. Models were trained and tested on both original images and versions degraded by controlled disturbances, including rotation, blur, brightness variation, and simulated rain. The results demonstrate that the proposed CNN-KAN model provides consistently superior performance under small-to-moderate rotations (up to 20 degrees) and moderate brightness increases, achieving the highest accuracy in all rain-mask scenarios. It remains competitive under blur, where it ranks second only to the standard CNN. Performance decreases were observed only at extreme brightness levels, where both the standard CNN and CKAN maintained higher stability. Overall, the findings highlight the potential of Kolmogorov-Arnold-based architectures for improving robustness in traffic sign recognition systems operating under realistic and dynamically changing environmental conditions.
Gaze estimation plays a central role in computer vision and human-computer interaction, enabling applications in assistive systems, attention modeling, and human-robot collaboration. However, existing datasets often rely on infrared-based hardware, are collected in constrained laboratory environments, or lack precise synchronization between stimuli and gaze data, which limits model generalization to real-world conditions. To address these challenges, we present HybridGaze - an open-source eye tracking dataset collected using a Tobii tracker combined with a standard RGB webcam. The recordings are processed into eye images and facial landmarks, providing synchronized gaze annotations and facial information across a variety of visual tasks. By capturing gaze data in naturalistic settings, the dataset reflects real-world visual behavior and serves as a valuable benchmark for gaze estimation research. Furthermore, we introduce GazeModalNet, a multi-stream neural network that estimates gaze direction from two complementary sources: eye images and facial landmarks. Together, the dataset and model establish a strong foundation for developing robust, multimodal gaze estimation systems beyond laboratory constraints.
Detecting weapons in public spaces remains a significant challenge in computer vision and public safety applications. While deep learning models have achieved great progress in general object detection, there is still a lack of focused studies on class-specific detection tasks, in particular those using new architectures such as transformers. In this work, a comprehensive evaluation of the state-of-the-art deep learning object detection approaches is conducted, including convolution and transformer-based architectures. Therefore, a dedicated large-scale dataset that combines images from multiple public sources is introduced, with a focus on three main weapons categories, enabling a more targeted evaluation. Furthermore, in the paper, the effectiveness of the best-performing architecture is further improved with proposed modifications, including architectural changes and determining a suitable loss function. Finally, the obtained detection approach achieves superior detection results, as evidenced by all performance criteria.
This work presents a system for automatic detection of various stages of diabetic retinopathy (DR) based on fundus images of patients. The system was built based on a relatively new and little-used image database: ”Dataset of fundus images for the study of diabetic retinopathy” version v3 CastilloBenitez21. The primary dataset was expanded using clinical fundus photographs acquired from the Department of Nephrology at Wroclaw Medical University. The diagnostic system was developed based on various variants of convolutional neural networks (CNNs) that were pre-trained on ImageNet data. The CNN classifier, based on VGG16 with transfer learning, proved to be effective and gave a global accuracy of 83.15%. The evaluation of discrimination between the non-DR and the DR state resulted in an accuracy of 89.7%, with a sensitivity of 94.9%, a specificity of 88.3%, and a Matthews Correlation Coefficient of 0.7665.
This article considers the problem of fish monitoring in an underwater environment, where many problems might occur, including occlusion, pose changes, and complexity of the scene. Recognizing fish behavior is very important to develop various types of technologies able to provide more precise estimations and monitoring of fish populations in a long term. In this paper, we propose a novel method for underwater fish monitoring (shape modeling and pose estimation). Two main aspects of underwater image processing will be studied: classification and localisation. Additionally, we extract key point features from fish patterns. The fish position and motion are not sufficient features to avoid scene problems. Skeleton extraction could offer us a large range of additional information. It models an object as a set of points of a certain manifold. The 3-dimensional fish pose, along the track of its 3D motion, could depend on curve segments of the underlying manifold. Faster reccurent conventional neural networks (faster R-CNNs) will be used to extract the fish skeleton in different poses. Also, a 3-dimensional trajectory of multiple fish will be derived using a Kalman filter based on the previous feature matching process. The simulation is made for live fish in a fish tank. Experimental results show that our method outperforms relevant models in terms of precision, achieving a minimal accuracy of 94.2%.
Brain tumours are aggressive malignant diseases, both in children and adults, representing 86 to 92 percent of all primary and almost half of secondary Central Nervous System (CNS) tumours. For individuals with malignant brain or central nervous system (CNS) tumours, the 5-year survival rate is about 34% for males and 36% for women. Brain tumours can be classified into several types, including benign, malignant, pituitary, etc. This study proposes a new architecture named Multilayered Max-Norm Regularization CNN (MMNR-CNN) and investigates the performance of this model for the classification of brain tumours in multi-modal MRI images. The model incorporates Markov Random Field (MRF) for bias field correction, and Monte Carlo Dropout to quantify prediction uncertainty through stochastic forward passes, enhancing the model's reliability in clinical decisionmaking. Furthermore, we integrate Explainable AI (XAI) techniques using Gradient-weighted Class Activation Mapping (Grad-CAM) to visually interpret the regions of MRI scans that contribute most to the classification decisions. We present a complete analysis of the Multilayered Max-Norm Regularization model trained on augmented brain image data and compare the performance on different values of regularization parameters that lead to the automatic selection of spatially important features for the classification task. This increases the generalization and robustness of the training dataset through augmentation. The model is trained using the Br35H database and the Figshare database. Both are used primarily for research in brain tumour detection and classification. The obtained performance metrics are the best in the literature, with a testing accuracy of 99.88% and 100 % precision.
PL
Guzy mózgu to agresywne choroby złośliwe, zarówno u dzieci, jak i u dorosłych, stanowiące 86–92% wszystkich pierwotnych i prawie połowę wtórnych guzów ośrodkowego układu nerwowego (OUN). U osób ze złośliwymi guzami mózgu lub ośrodkowego układu nerwowego (OUN) 5-letni wskaźnik przeżycia wynosi około 34% dla mężczyzn i 36% dla kobiet. Guzy mózgu można podzielić na kilka typów, w tym łagodne, złośliwe, przysadkowe itp. W niniejszym badaniu zaproponowano nową architekturę o nazwie Wielowarstwową Techniką Ograniczenia Normy Maksymalnej (Multilayered MaxNorm Regularization CNN – MMNR-CNN) i zbadano wydajność tego modelu w klasyfikacji guzów mózgu w wielomodalnych obrazach MRI. Model wykorzystuje losowe pole Markowa (MRF) do korekcji pola błędu oraz metodę Monte Carlo Dropout do ilościowego określania niepewności prognozy poprzez stochastyczne przejścia do przodu, zwiększając niezawodność modelu w podejmowaniu decyzji klinicznych. Ponadto integrujemy techniki XAI (Exploreable AI) wykorzystujące Gradient-weighted Class Activation Mapping (Grad-CAM), aby wizualnie zinterpretować obszary skanów MRI, które mają największy wpływ na decyzje klasyfikacyjne. Przedstawiamy pełną analizę modelu MMNR wytrenowanego na rozszerzonych danych obrazowych mózgu i porównujemy wydajność przy różnych wartościach parametrów regularyzacji, które prowadzą do automatycznego wyboru cech istotnych przestrzennie dla zadania klasyfikacji. Zwiększa to generalizację i odporność zbioru danych treningowych poprzez rozbudowę. Model jest trenowany z wykorzystaniem bazy danych Br35H i Figshare. Obie są wykorzystywane głównie w badaniach nad wykrywaniem i klasyfikacją guzów mózgu. Uzyskane wskaźniki wydajności są najlepsze w literaturze, z dokładnością testowania na poziomie 99,88% i 100% precyzją.
Agricultural monitoring plays an important role in ensuring food security and sustainable farming practices. This project focuses on the task of paddy field detection on the satellite images collected from the agricultural lands of Andhra Pradesh. Using satellite images of Sentinel-2 and deep learning techniques, our approach aims to improve the accuracy and efficiency of paddy land identification. The project employs a deep learning model, which is the EfficientDet trained on a carefully annotated dataset, to detect the paddy fields in the region. The utilization of remote sensing technology allows for scalable and timely monitoring across vast agricultural lands. The selected model architecture, combined with fine-tuning strategies, ensures adaptability to the unique spatial and seasonal characteristics of South Indian agriculture. Results from the project showcase the capability of the proposed approach in accurately identifying and detecting paddy crops. The integration of advanced technologies for precision agriculture contributes to informed decision-making, resource optimization, and overall sustainability in the farming sector. To collect the ground truth data, we used the AP GIS portal, which is supervised by the Andhra Pradesh agriculture department, which has the data of the percentage of paddy lands in small villages. The places with more than 95 percent of paddy lands are selected for better data samples. The collected samples are cropped and labelled, and trained on the model architecture and verified for accuracy. This project not only advances the field of agricultural monitoring but also holds significant importance for crop management and supporting the livelihoods of farmers of our state. The proposed model achieved an accuracy of 86%, demonstrating its reliability in detecting paddy fields from Sentinel-2 satellite imagery.
PL
Monitorowanie rolnictwa odgrywa kluczową rolę w zapewnieniu bezpieczeństwa żywnościowego i zrównoważonych praktyk rolniczych. Niniejszy projekt koncentruje się na wykrywaniu pól ryżowych na obrazach satelitarnych zebranych z terenów rolniczych stanu Andhra Pradesh. Wykorzystując obrazy Sentinel-2 oraz techniki głębokiego uczenia, nasze podejście poprawia dokładność i efektywność identyfikacji terenów uprawnych ryżu. Model EfficientDet, przeszkolony na starannie oznakowanym zbiorze danych, umożliwia skuteczne wykrywanie pól ryżowych, a zastosowanie technologii teledetekcji pozwala na skalowalne i terminowe monitorowanie rolnictwa. Wybrana architektura modelu, w połączeniu ze strategiami dostrajania, zapewnia dostosowanie do unikalnych cech przestrzennych i sezonowych rolnictwa południowych Indii. Wyniki projektu potwierdzają skuteczność proponowanego podejścia w precyzyjnym wykrywaniu upraw ryżu. W celu zebrania danych referencyjnych wykorzystano portal AP GIS nadzorowany przez Departament Rolnictwa Andhra Pradesh, który zawiera informacje o procentowym udziale pól ryżowych w małych wsiach. Do analizy wybrano miejsca, gdzie pola ryżowe stanowią ponad 95% powierzchni, co pozwoliło uzyskać wysokiej jakości próbki danych. Zebrane próbki zostały przycięte, oznaczone, przeszkolone na architekturze modelu i zweryfikowane pod kątem dokładności. Projekt ten nie tylko przyczynia się do rozwoju monitorowania rolnictwa, ale także ma istotne znaczenie dla zarządzania uprawami oraz wspierania środków utrzymania rolników w naszym stanie. Proponowany model osiągnął dokładność na poziomie 86%, co potwierdza jego niezawodność w wykrywaniu pól ryżowych na obrazach satelitarnych Sentinel-2.
Modern decision support systems (DSS) increasingly must analyze large-scale, complex, and heterogeneous data streams in real time. Highperformance, AI-driven processing methods – particularly deep neural networks – offer effective solutions. This study examines the integration of contemporary architectures – YOLOv8, ResNet-50, EfficientNet-B3, and the Vision Transformer (ViT) – to enhance DSS capabilities. The models are benchmarked on a representative image-classification task using the COCO dataset for training and evaluation. Empirical results indicate that the transformer-based model (ViT) attains the highest accuracy, whereas the one-stage architecture (YOLOv8) achieves the fastest inference. EfficientNetB3 and ResNet-50 exhibit intermediate trade-offs between accuracy and speed. Deployment considerations across DSS scenarios are outlined: YOLOv8 is appropriate for real-time, resource-constrained environments; ResNet-50 provides balanced performance; EfficientNet-B3 offers strong accuracy with moderate computational demand; and ViT delivers the best accuracy when ample data and computational resources are available. The findings are discussed in the context of DSS workflows, illustrating how the model outputs can directly inform and improve decision-making processes.
PL
Nowoczesne systemy wspomagania decyzji (DSS) muszą coraz częściej analizować w czasie rzeczywistym duże, złożone i niejednorodne strumienie danych. Skutecznym rozwiązaniem są wysokowydajne metody przetwarzania oparte na sztucznej inteligencji, w szczególności głębokie sieci neuronowe. W niniejszym badaniu przeanalizowano integrację współczesnych architektur – YOLOv8, ResNet-50, EfficientNet-B3 i Vision Transformer (ViT) – w celu zwiększenia możliwości systemów DSS. Modele są porównywane w reprezentatywnym zadaniu klasyfikacji obrazów przy użyciu zbioru danych COCO do szkolenia i oceny. Wyniki empiryczne wskazują, że model oparty na transformatorze (ViT) osiąga najwyższą dokładność, podczas gdy architektura jednostopniowa (YOLOv8) zapewnia najszybsze wnioskowanie. EfficientNet-B3 i ResNet-50 wykazują pośredni kompromis między dokładnością a szybkością. Przedstawiono kwestie związane z wdrażaniem w różnych scenariuszach DSS: YOLOv8 jest odpowiedni do zastosowań w czasie rzeczywistym, z ograniczonymi zasobami, EfficientNet-B3 i ResNet-50 wykazują pośredni kompromis między dokładnością a szybkością. Przedstawiono rozważania dotyczące wdrażania w różnych scenariuszach DSS: YOLOv8 jest odpowiedni dla środowisk działających w czasie rzeczywistym, o ograniczonych zasobach; ResNet-50 zapewnia zrównoważoną wydajność; EfficientNet-B3 oferuje wysoką dokładność przy umiarkowanych wymaganiach obliczeniowych; a ViT zapewnia najlepszą dokładność, gdy dostępne są duże zasoby danych i mocy obliczeniowej. Wyniki są omawiane w kontekście przepływów pracy DSS, ilustrując, w jaki sposób wyniki modelu mogą bezpośrednio wpływać na procesy decyzyjne i je usprawniać.
19
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Sudden cardiac arrest (SCA) is a life-threatening arrhythmic event in which the heart abruptly loses its ability to pump blood effectively. If the arrest is not reversed within minutes, it progresses to sudden cardiac death (SCD), a fatal outcome. responsible for approximately 18% of all global deaths. This review focuses on supervised learning methods for predicting SCA episodes at risk of progressing to SCD, based on electrocardiogram (ECG) analysis. It establishes the clinical significance of SCD prediction before examining signal processing techniques for extracting relevant characteristics from ECG signals. Subsequently, the analysis focuses on machine and deep learning approaches, particularly their roles in pattern recognition and predictive modeling. Furthermore, the investigation explores emerging supervised classifiers with potential applications in SCD prediction. Finally, the review concludes by addressing current challenges and future research directions, with emphasis on three critical aspects: (1) development of robust predictive models, (2) integration of multi-source data, and (3) implementation of personalized healthcare strategies. This synthesis of existing knowledge combined with novel methodological insights offers valuable guidance for advancing SCD prediction research and improving clinical outcomes.
20
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Atrial fibrillation (AF) is the most common form of arrhythmia, significantly increasing the risk of stroke, heart failure, and other cardiovascular complications. Although AF detection methods have achieved accuracies exceeding 98%, AF onset prediction remains underexplored. Paroxysmal AF, an early stage of AF progression, often goes undetected even with continuous monitoring beyond 24 h, and its transition to sustained AF is associated with increased mortality and severe complications. Notably, approximately 15% of the 5 million critically ill patients annually hospitalized in United States intensive care units (ICUs) experience new-onset AF, highlighting the urgent need for early AF onset prediction. This study proposes a two-stage deep learning framework for AF prediction using RR intervals (RRIs). The first stage extracts features using a convolutional and bidirectional long short-term memory (BiLSTM) network, while the second stage employs another BiLSTM with a fully connected classifier to predict AF onset one hour in advance. In subject-wise testing, the model achieved a sensitivity of 0.936, specificity of 0.893, F1-score of 0.906, and an area under the receiver operating characteristic curve (AUROC) of 0.980. In external independent dataset validation, it achieved a sensitivity of 0.848, specificity of 0.978, F1-score of 0.938, AUROC of 0.976, and an area under the precision-recall curve (AUPRC) of 0.966. Our approach demonstrates: (1) state-of-the-art predictive performance, (2) lightweight computational complexity despite a large number of parameters, (3) flexible training through the two-stage design, (4) the ability to identify high-risk RRI segments using masking techniques to enhance clinical interpretation, and (5) a robust AF onset prediction framework capable of predicting AF up to one hour in advance using one hour of input data - providing sufficient lead time for preventive interventions.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.