According to the Convention on the International Regulations for Preventing Collisions at Sea (COLREGs), vessels must maintain an effective lookout by sight, hearing, and all available technical means, taking into account prevailing circumstances and environmental conditions. In this context, automated visual surveillance systems represent an important enhancement of conventional navigational tools, enabling early detection and interpretation of surrounding vessel behaviour. This study proposes a video surveillance system based on the lightweight YOLOv8n deep learning model for the detection and classification of eight vessel aspects. The model was trained on an initial dataset of 925 images, which was further expanded to 1,742 annotated images. The dataset was designed to reflect real maritime operating conditions, including different times of day, weather scenarios, vessel categories, and geographical regions. To improve robustness, data augmentation techniques such as colour space transformations, geometric modifications, and classification-specific augmentations were applied. Class imbalance was mitigated through the use of class weighting. The paper also describes the system architecture and camera configuration, providing effective surveillance coverage up to 6 nautical miles. The proposed approach enables not only vessel aspect recognition but also the estimation of relative and true bearings, thereby contributing to improved situational awareness and collision avoidance in maritime navigation.
W artykule przedstawiono zastosowanie metod wizji komputerowej oraz głębokiego uczenia w zadaniu segmentacji odprysków w elementach konstrukcyjnych betonowych. Pomimo dynamicznego rozwoju metod segmentacji obrazu, w literaturze nadal brakuje kompleksowych analiz porównawczych modeli w kontekście segmentacji nieregularnych defektów, takich jak odpryski betonu. W badaniu przeanalizowano skuteczność wybranych modeli segmentacyjnych, w tym DeepLabV3+, YOLO (warianty seg), U-Net oraz HRNet. Analizę przeprowadzono na zbiorze 138 obrazów przedstawiających uszkodzenia betonu, podzielonym na zbiory treningowe i testowe. Modele oceniono przy użyciu standardowych metryk, takich jak Intersection over Union (IoU), Dice coefficient, precision, recall oraz F1-score. Uzyskane wyniki wskazują na wyraźną przewagę modelu DeepLabV3+, który osiągnął najwyższe wartości metryk jakościowych, w tym IoU = 0,709 oraz Dice = 0,825. Wyniki potwierdzają, że zaawansowane architektury segmentacyjne mogą znacząco zwiększyć dokładność i obiektywność identyfikacji uszkodzeń betonowych, stanowiąc podstawę do rozwoju zautomatyzowanych systemów diagnostycznych.
EN
The article presents the application of computer vision and deep learning methods for the segmentation of spalling defects in concrete structural elements. Despite the dynamic development of image segmentation methods, there is still a lack of comprehensive comparative analyses of models in the context of segmenting irregular defects such as concrete spalling. The study evaluates the performance of selected segmentation models, including DeepLabV3+, YOLO (segmentation variants), U-Net, and HRNet. The analysis was conducted on a dataset of 138 images depicting concrete damage, divided into training and testing subsets. The models were assessed using standard evaluation metrics such as Intersection over Union (IoU), Dice coefficient, precision, recall, and F1-score. The results indicate a clear advantage of the DeepLabV3+ model, which achieved the highest performance values, including IoU = 0.709 and Dice = 0.825. The findings confirm that advanced segmentation architectures can significantly improve the accuracy and objectivity of concrete damage identification, providing a foundation for the development of automated diagnostic systems.
Wraz z dynamicznym rozwojem technologii modelowania informacji o budynku (BIM) algorytmy sztucznej inteligencji (AI) znajdują coraz szersze zastosowanie w analizie i wspomaganiu kluczowych procesów realizacji inwestycji budowlanych, obejmujących m.in. planowanie, wytwarzanie dokumentacji, kontrolę jakości oraz zarządzanie ryzykiem. Artykuł przedstawia jakościową analizę porównawczą wybranych rozwiązań AI, koncentrując się na ich mechanizmach technicznych, wymaganiach wdrożeniowych oraz efektach praktycznego zastosowania. W treści pracy omówiono zastosowanie algorytmów uczenia maszynowego w automatyzacji harmonogramowania, technologii widzenia komputerowego i mapowania 3D w monitorowaniu postępu prac, a także skanowania laserowego w kontroli jakości robót. Przedstawiono również rolę AI w analizie dokumentacji oraz prognozowaniu ryzyka realizacyjnego. Na zakończenie wskazano najważniejsze korzyści oraz ograniczenia związane z wdrażaniem przedstawionych rozwiązań w praktyce.
EN
With the rapid development of building information modelling (BIM) technology, artificial intelligence (AI) algorithms are increasingly being used in the analysis and support of key construction investment processes, including planning, documentation production, quality control and risk management. This article presents a qualitative comparative analysis of selected AI solutions, focusing on their technical mechanisms, implementation requirements and practical application effects. The paper discusses the use of machine learning algorithms in scheduling automation, computer vision and 3D mapping technologies in monitoring work progress, as well as laser scanning in quality control. The role of AI in document analysis and implementation risk forecasting is also presented. Finally, the most important benefits and limitations associated with the practical implementation of the presented solutions are indicated.
YOLO object detectors recently became a key component of vision systems in many domains. The family of available YOLO models consists of multiple versions, each in various variants. The research reported in this paper aims to validate the applicability of members of this family to detect objects located within the robot workspace. In our experiments, we used our custom dataset and the COCO2017 dataset. To test the robustness of investigated detectors, the images of these datasets were subject to distortions. The results of our experiments, including variations of training/testing configurations and models, may support the choice of the appropriate YOLO version for robotic vision tasks.
PL
Detektory obiektów YOLO stały się ostatnimi czasy kluczowym elementem systemów wizyjnych w wielu dziedzinach. Rodzina dostępnych modeli YOLO składa się z wielu wersji, z których każda występuje w różnych wariantach. Badania opisane w niniejszej pracy mają na celu zweryfikowanie przydatności członków tej rodziny do wykrywania obiektów znajdujących się w przestrzeni roboczej robota. W eksperymentach wykorzystano nasz własny zbiór danych oraz zbiór COCO2017. Aby przetestować odporność badanych detektorów, obrazy z tych zbiorów poddano zniekształceniom. Wyniki eksperymentów, uwzględniające różne konfiguracje treningowe/testowe oraz modele, mogą stanowić wsparcie przy wyborze odpowiedniej wersji YOLO dla zadań związanych z wizją robotyczną.
Despite many years of effort and research, the current waste management problem remains. So far, no fully effective waste management system has been developed. Many programs and projects improve statistics on the percentage of waste recycled every year. Modern computer vision techniques supported by artificial intelligence are worth using in these efforts. In the article, we present a method of identifying plastic waste based on the asymmetry analysis of the image’s histogram containing the waste. The method is simple but effective ( 94% ), allowing it to be implemented on devices with low computing power, particularly microcomputers. Such devices will be used both at home and in waste-sorting plants.
A process of water purification using electrical coagulation and its optimisation using artificial intelligence tools was presented. Experimental data was analysed and correlated. The experimental studies used were developed to optimise iron coagulation during water treatment. A neural network was developed to optimise iron coagulation and machine learning was performed. Neural network tests were conducted and methods for its practical application were proposed.
PL
Przedstawiono proces oczyszczania wody z wykorzystaniem koagulacji elektrycznej oraz jego optymalizację przy użyciu narzędzi sztucznej inteligencji. Przeanalizowano i skorelowano dane eksperymentalne. Przeprowadzono badania eksperymentalne mające na celu optymalizację koagulacji żelaza podczas uzdatniania wody. Opracowano sieć neuronową w celu optymalizacji koagulacji żelaza oraz przeprowadzono proces uczenia maszynowego. Przeprowadzono testy sieci neuronowej i zaproponowano metody jej praktycznego zastosowania.
This paper presents a comparative analysis of state-of-the-art multi-object tracking algorithms applied in UAV-based video surveillance systems. The performance results of three advanced tracking methods – DeepSORT, ByteTrack, and StrongSORT – integrated with the YOLOv8 object detector are presented. A mathematical description and experimental simulations were conducted to evaluate the accuracy, stability, and computational performance of the algorithms in dynamic and complex scenes. The obtained results indicate that the StrongSORT + YOLOv8 combination provides the best balance between accuracy and robustness, whereas the ByteTrack method demonstrates high track continuity in high-density environments. The proposed approach can be utilized to enhance the efficiency of UAV-based autonomous monitoring systems.
PL
W tym artykule przedstawiono analizę porównawczą najnowocześniejszych algorytmów śledzenia wielu obiektów stosowanych w systemach monitoringu wizyjnego opartych na bezzałogowych statkach powietrznych (UAV). Przedstawiono wyniki działania trzech zaawansowa nych metod śledzenia – DeepSORT, ByteTrack i StrongSORT – zintegrowanych z detektorem obiektów YOLOv8. Przeprowadzono opis matematyczny i symulacje eksperymentalne w celu oceny dokładności, stabilności i wydajności obliczeniowej algorytmów w dynamicznych i złożonych scenach. Uzyskane wyniki wskazują, że połączenie StrongSORT + YOLOv8 zapewnia najlepszą równowagę między dokładnością a odpornością, podczas gdy metoda ByteTrack wykazuje wysoką ciągłość śledzenia w środowiskach o dużej gęstości. Proponowane podejście może być wykorzystane do zwiększenia wydajności autonomicznych systemów monitorowania opartych na bezzałogowych statkach powietrznych.
This study investigates the development, adoption, and implications of artificial intelligence (AI) models by analysing a comprehensive dataset of over 316,000 models hosted on the Hugging Face platform. Focusing on two dominant model architectures - transformers and diffusion models - it examines their distribution across tasks, user engagement patterns, and practical applications in domains such as natural language processing, computer vision, audio processing, and generative media. The research highlights the growing prominence of generative AI, the role of open-source platforms in shaping model accessibility, and the divergence in use trends between foundational and emerging AI tools. Drawing on correlations between downloads, likes, citations, and model size, the paper discusses how each library’s community-driven dynamics shape their respective strengths. Finally, the paper discusses implications for business strategy and adoption, encompassing practical considerations like infrastructure requirements and ethical challenges, and underscores the potential for these evolving model ecosystems to drive innovative, human-centric AI solutions across diverse sectors.
The article explores the application of stereo images and neural networks for tracking designated manufacturing employees, with a focus on optimizing production processes. The primary objective of this study is to present a novel approach for the automatic construction of movement trajectories of employees, an issue of significant importance for improving production efficiency. By analyzing these trajectories, valuable insights can be gained regarding the optimal arrangement of workstations, facilitating adjustments to their positions in alignment with the actual workflows. This approach aligns closely with principles of lean manufacturing, offering a method to enhance operational efficiency. The authors propose a solution based on U-Net type neural networks for classifying objects within stereo images, alongside the integration of stereoscopic imaging with regression techniques for accurate 3D localization of objects. A thorough analysis of the proposed AI models is presented, accompanied by the results of practical tests conducted under varying configuration parameters of the image processing system. The study highlights the novelty of the approach, contributing to the advancement of automated monitoring and optimization in manufacturing environments.
Object segmentation in multidimensional data spaces is a pivotal component of modern computational analysis. Frequently, accurate segmentation hinges on the detection and localisation of object boundaries. The targeted objects often exhibit spherical symmetry. This paper introduces an algorithm for the automatic detection of n-dimensional hyperspheres embedded in (n+1)-dimensional Euclidean space. The algorithm utilises an evolutionary computation strategy to estimate hypersphere parameters from extensive point clouds. This method demonstrates notable advantages over traditional approaches such as the Hough transform and active surface models. Preliminary results suggest strong potential of the proposed technique as well as its adaptability to broader classes of hypersurfaces, offering a promising extension for future exploration.
Effective multi-scale feature representation and focused attention on critical objects are essential for accurate perception of waterborne navigation scenes. To address the insufficient exploitation of multi-scale information in existing methods that leads to imprecise segmentation, this study proposes a real-time semantic segmentation method for waterborne navigation scenes through multi-scale information enhancement and importance-weighted optimization. First, DDRNet-23-slim is selected as the backbone network for feature extraction. An edge-guided branch is embedded into its shallow layers, and a Dynamic Feature Fusion Module (DFFM) is constructed by integrating a lightweight hybrid attention mechanism, effectively enhancing multi-scale feature interaction capabilities. Second, the loss function is improved using an importance-weighted strategy to prioritize critical objects during training. Finally, a parameter-free attention mechanism is introduced in the upsampling stage, maintaining real-time performance while ensuring segmentation stability for key objects under complex background interference. Evaluations on the On_Water and Seaships datasets demonstrate that the proposed method achieves mIoU scores of 83.1% and 73.2%, respectively, with ship segmentation accuracy reaching 88.2% on On_Water. The inference speed attains 69.1 FPS, outperforming mainstream real-time segmentation models (e.g., DDRNet, STDC) in balancing accuracy and efficiency. Notably, it exhibits stronger robustness in complex inland river scenarios with dense shore structures and numerous small targets.
W pracy przedstawiono aplikację do rozpoznawania pionowych znaków drogowych z użyciem modelu sztucznej inteligencji, zaprojektowaną w celu poprawy bezpieczeństwa ruchu drogowego. Model został przetrenowany na przygotowanym zbiorze danych obejmującym zdywersyfikowane obrazy, wzbogacone technikami augmentacji. Aplikacja umożliwia wykrywanie znaków drogowych z kamery internetowej oraz nagrań wideo. Model sztucznej inteligencji wykazuje potencjał do zastosowań w systemach wsparcia kierowców i technologii autonomicznych pojazdów.
EN
This paper presents an application for recognizing vertical traffic signs using an artificial intelligence model, designed to enhance road safety. The model was trained on a prepared dataset comprising diversified images, enriched with augmentation techniques. The application enables the detection of traffic signs from webcam feeds and video recordings. The artificial intelligence model shows potential for use in driver assistance systems and autonomous vehicle technologies.
Over the past few years, AI development has impacted fields like computer vision, image description, and generation. The article explored AI's capability to create descriptions and generate images, comparing these with human perception. Images were examined using eye tracking in a VR art gallery and on a desktop. The study involved expert and AI descriptions of BITSCOPE project images, followed by AI-generated images based on those descriptions, focusing on gaze plot metrics.
PL
W ciągu ostatnich kilku lat rozwój sztucznej inteligencji (SI) przyczynił się do postępów w takich dziedzinach jak widzenie komputerowe, opisywanie i generowanie obrazów. Analizy skupiły się na zdolności SI do tworzenia opisów i generowania obrazów, porównując je z ludzką percepcją. Obrazy były badane za pomocą śledzenia ruchu gałek ocznych w galerii sztuki VR oraz w środowisku stacjonarnym. Badanie obejmowało opisy obrazów projektu BITSCOPE dokonane przez eksperta i SI, a następnie generowane przez SI obrazy na podstawie tych opisów, koncentrując się na metrykach śledzenia wzroku.
The article examines the current issues of beer production related to the yeast foam formation during fermentation, and discusses the importance of controlling the fermentation stages to ensure the proper product quality. Computer vision technologies were applied to identify the stages of the main fermentation. Based on the analysis of the computer vision algorithms, the K-means method was used for image clustering. The systematic description of the algorithm for detecting contaminated foam based on the K-means method is provided.
PL
W artykule omówiono bieżące problemy produkcji piwa związane z powstawaniem piany drożdżowej podczas fermentacji oraz omówiono znaczenie kontrolowania etapów fermentacji w celu zapewnienia odpowiedniej jakości produktu. Zastosowano technologie wizji komputerowej w celu zidentyfikowania etapów głównej fermentacji. Na podstawie analizy algorytmów wizji komputerowej do klasteryzacji obrazów zastosowano metodę K-means. Przedstawiono systematyczny opis algorytmu wykrywania zanieczyszczonej piany w oparciu o metodę K-means.
The paper presents a study of the possibilities of using modern machine learning methods based on the YOLO (You Only Look Once) algorithm in the detection and classification of fire hazards based on camera image recognition. The paper aims to develop an automation system for effectively identifying fire and smoke to develop effective protection of forest complexes. The YOLOv8 model was used in the detection process, which turned out to be a highly effective object detection model in real-time. The paper presents the process of preparing image data sets for the construction of the YOLO model. In the final part of the paper, many tests were carried out to assess the effectiveness and precision of the developed fire detection and fire prediction models. The results of these tests confirmed that the detection model works very precisely and can accurately identify fiery and smoky areas in camera images.
PL
W pracy zaprezentowano badanie możliwości zastosowania nowoczesnych metod uczenia maszynowego w oparciu o algorytm YOLO (You Only Look Once) w detekcji i klasyfikacji zagrożenia pożarowego na podstawie rozpoznawania obrazu pozyskanego z kamery. Praca ma na celu określenie efektywności działania algorytmu w automatyzacji identyfikacji ognia i zadymienia dla potrzeb opracowania skutecznej ochrony kompleksów leśnych. W procesie detekcji zastosowano model YOLOv8, który okazał się modelem wykrywania obiektów o wysokiej skuteczności w czasie rzeczywistym. W pracy zaprezentowano proces przygotowania zbiorów danych obrazowych dla potrzeb budowy modelu YOLO.
Semantic segmentation of plant images is crucial for various agricultural applications and creates the need to develop more demanding models that are capable of handling images in a diverse range of conditions. This paper introduces an extended DeepLabV3+ model with a channel-wise attention mechanism, designed to provide precise semantic segmentation while emphasizing crucial features. It leverages semantic information with global context and is capable of handling object scale variations within the image. The proposed approach aims to provide a well generalized model that may be adapted to various field conditions by training and tests performed on multiple datasets, including Eschikon wheat segmentation (EWS), humans in the loop (HIL), computer vision problems in plant phenotyping (CVPPP), and a custom “botanic mixed set” dataset. Incorporating an ensemble training paradigm, the proposed architecture achieved an intersection over union (IoU) score of 0.846, 0.665 and 0.975 onEWS, HIL plant segmentation, and CVPPP datasets, respectively. The trained model exhibited robustness to variations in lighting, backgrounds, and subject angles, showcasing its adaptability to real-world applications.
The coordinate rotation digital computer (CORDIC) algorithm is a popular method used in many fields of science and technology. Unfortunately, it is a time-consuming process for central processing units (CPUs) and graphics processing units (GPUs), and even for specialized digital signal processing (DSP) solutions. The CORDIC algorithm is an alternative for Newton-Raphson numerical calculation and for the FPGA based resource-expensive look-up-table (LUT) method. Various modifications of the CORDIC algorithm allow to speed up the operation of hardware in edge computing devices.With that context taken into consideration, this article presents a fast and accurate square root floating point (SQRT FP) CORDIC function which can be implemented in field programmable gate arrays (FPGAs). The proposed algorithm offers low-complexity, decent accuracy and speed, and is sufficient for digital signal processing (DSP) applications, such as digital filters, accelerators for neural networks, machine learning and computer vision applications, and intelligent robotic systems.
Bioprinting is the technology that combines the use of living matter and biomaterials to manufacture biological models, tissues, and structures layer by layer for applications in regenerative medicine, drug testing, and tissue engineering. Among bioprinting techniques, extrusion-based methods are the most widely used because of their relative simplicity, affordability, and ability to handle as wide range of biomaterials, including those with high viscosities. However, achieving consistent print quality remains a challenge, as the rheological properties of bioinks are highly variable and sensitive to environmental factors such as temperature. A critical aspect of print quality is maintaining a consistent and predictable line width, as pre-programmed trajectories and design fidelity rely on this parameter being well controlled. This work introduces a closed-loop control system for Extrusion-Based Bioprinting (EBB), utilizing real-time computer vision. The system employs a camera that is placed to monitor the line width immediately after extrusion, enabling real-time feedback to adjust the feedrate of the extrusion mechanism. This approach ensures consistent line widths across a wide range of materials and conditions, addressing the variability that traditionally hampers EBB. The method was validated using a Pluronic hydrogel, achieving closed-loop control over a wide range of target line widths. These findings demonstrate the potential for automated, robust bioprinting with improved reproducibility and precision, advancing the reliability of this technology for biomedical applications.
This study compares two artificial intelligence approaches for parking occupancy detection: computer vision and convolutional neural networks (CNN). A dataset of 1,000 parking images was captured and labeled, using OpenCV in Python for computer vision processing and the YOLO V5 model for CNN. Results showed that the YOLO V5 model achieved 88% precision and 82% sensitivity, outperforming the computer vision method, which achieved 80% precision and 79% sensitivity. The research suggests that while CNNs offer superior performance, computer vision is a more economical option in contexts with limited resources. Future research will evaluate the YOLOv7 version to reduce false positives and combine techniques to balance accuracy and efficiency under variable conditions.
The hydraulic properties of unsaturated soil provide important information on fluid movement in geological systems. The degree of saturation of unsaturated soil, recognized as one of the most critical hydraulic properties, is crucial for geotechnical problems such as slope stability, bearing capacity, and seepage. Computer vision for digital image processing approaches has been widely used recently, with applications in many aspects of construction and geotechnical engineering. The objective of this research is to use the digital image processing (DIP) method to determine the degree of saturation of unsaturated sand. The RGB image of the soil column is processed into a grayscale channel and becomes a matrix for processing calculation. The fluctuation of pixel values extracted from computer vision is correlated with the change in the degree of saturation of the sand. Hence, DIP techniques can be considered as an effective tool to establish the soil-water characteristic curve (SWCC) of unsaturated soil. Furthermore, employing matrix processing techniques can help improve the reliability of the SWCC derived from the correlation between image pixel values and the degree of saturation of unsaturated soil by identifying and mitigating the influence of outliers. The results demonstrate a strong correlation between the degree of saturation obtained through digital image processing techniques and the measured values from resistivity probe techniques.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.