Wraz z dynamicznym rozwojem technologii modelowania informacji o budynku (BIM) algorytmy sztucznej inteligencji (AI) znajdują coraz szersze zastosowanie w analizie i wspomaganiu kluczowych procesów realizacji inwestycji budowlanych, obejmujących m.in. planowanie, wytwarzanie dokumentacji, kontrolę jakości oraz zarządzanie ryzykiem. Artykuł przedstawia jakościową analizę porównawczą wybranych rozwiązań AI, koncentrując się na ich mechanizmach technicznych, wymaganiach wdrożeniowych oraz efektach praktycznego zastosowania. W treści pracy omówiono zastosowanie algorytmów uczenia maszynowego w automatyzacji harmonogramowania, technologii widzenia komputerowego i mapowania 3D w monitorowaniu postępu prac, a także skanowania laserowego w kontroli jakości robót. Przedstawiono również rolę AI w analizie dokumentacji oraz prognozowaniu ryzyka realizacyjnego. Na zakończenie wskazano najważniejsze korzyści oraz ograniczenia związane z wdrażaniem przedstawionych rozwiązań w praktyce.
EN
With the rapid development of building information modelling (BIM) technology, artificial intelligence (AI) algorithms are increasingly being used in the analysis and support of key construction investment processes, including planning, documentation production, quality control and risk management. This article presents a qualitative comparative analysis of selected AI solutions, focusing on their technical mechanisms, implementation requirements and practical application effects. The paper discusses the use of machine learning algorithms in scheduling automation, computer vision and 3D mapping technologies in monitoring work progress, as well as laser scanning in quality control. The role of AI in document analysis and implementation risk forecasting is also presented. Finally, the most important benefits and limitations associated with the practical implementation of the presented solutions are indicated.
YOLO object detectors recently became a key component of vision systems in many domains. The family of available YOLO models consists of multiple versions, each in various variants. The research reported in this paper aims to validate the applicability of members of this family to detect objects located within the robot workspace. In our experiments, we used our custom dataset and the COCO2017 dataset. To test the robustness of investigated detectors, the images of these datasets were subject to distortions. The results of our experiments, including variations of training/testing configurations and models, may support the choice of the appropriate YOLO version for robotic vision tasks.
PL
Detektory obiektów YOLO stały się ostatnimi czasy kluczowym elementem systemów wizyjnych w wielu dziedzinach. Rodzina dostępnych modeli YOLO składa się z wielu wersji, z których każda występuje w różnych wariantach. Badania opisane w niniejszej pracy mają na celu zweryfikowanie przydatności członków tej rodziny do wykrywania obiektów znajdujących się w przestrzeni roboczej robota. W eksperymentach wykorzystano nasz własny zbiór danych oraz zbiór COCO2017. Aby przetestować odporność badanych detektorów, obrazy z tych zbiorów poddano zniekształceniom. Wyniki eksperymentów, uwzględniające różne konfiguracje treningowe/testowe oraz modele, mogą stanowić wsparcie przy wyborze odpowiedniej wersji YOLO dla zadań związanych z wizją robotyczną.
Over the past few years, AI development has impacted fields like computer vision, image description, and generation. The article explored AI's capability to create descriptions and generate images, comparing these with human perception. Images were examined using eye tracking in a VR art gallery and on a desktop. The study involved expert and AI descriptions of BITSCOPE project images, followed by AI-generated images based on those descriptions, focusing on gaze plot metrics.
PL
W ciągu ostatnich kilku lat rozwój sztucznej inteligencji (SI) przyczynił się do postępów w takich dziedzinach jak widzenie komputerowe, opisywanie i generowanie obrazów. Analizy skupiły się na zdolności SI do tworzenia opisów i generowania obrazów, porównując je z ludzką percepcją. Obrazy były badane za pomocą śledzenia ruchu gałek ocznych w galerii sztuki VR oraz w środowisku stacjonarnym. Badanie obejmowało opisy obrazów projektu BITSCOPE dokonane przez eksperta i SI, a następnie generowane przez SI obrazy na podstawie tych opisów, koncentrując się na metrykach śledzenia wzroku.
The article examines the current issues of beer production related to the yeast foam formation during fermentation, and discusses the importance of controlling the fermentation stages to ensure the proper product quality. Computer vision technologies were applied to identify the stages of the main fermentation. Based on the analysis of the computer vision algorithms, the K-means method was used for image clustering. The systematic description of the algorithm for detecting contaminated foam based on the K-means method is provided.
PL
W artykule omówiono bieżące problemy produkcji piwa związane z powstawaniem piany drożdżowej podczas fermentacji oraz omówiono znaczenie kontrolowania etapów fermentacji w celu zapewnienia odpowiedniej jakości produktu. Zastosowano technologie wizji komputerowej w celu zidentyfikowania etapów głównej fermentacji. Na podstawie analizy algorytmów wizji komputerowej do klasteryzacji obrazów zastosowano metodę K-means. Przedstawiono systematyczny opis algorytmu wykrywania zanieczyszczonej piany w oparciu o metodę K-means.
To address the challenges in the CO2 injection process, CO2 microbubble dispersion has been proposed as an alternative to traditional methods, such as miscible injection and water-alternating-gas (WAG) injection. This study presents an AI-assisted model for detecting CO2 microbubbles, powered by the YOLOv8 algorithm, renowned for its high-accuracy predictions. Conventional image processing techniques often struggle with detecting microbubbles, particularly in cases involving overlapping bubbles, variations in size, and low-contrast images, which can lead to inaccuracies in bubble identification and measurement. In contrast, YOLOv8’s advanced detection capabilities offer a more robust solution by precisely localizing and classifying microbubbles, even in challenging scenarios. The model’s performance was rigorously evaluated, demonstrating its effectiveness as a valuable tool for microbubble analysis. The detection images processed using YOLOv8 illustrate its ability to accurately detect and classify bubbles of varying sizes, generating precise bounding boxes around each identified bubble. This combination of data visualization and advanced detection techniques underscores the efficacy of YOLOv8 in microbubble analysis, enabling accurate measurement and detailed characterization of bubble size distributions—an essential factor in optimizing chemical engineering processes.
In the field of concrete structure health monitoring, accurately and swiftly identifying damage characteristics stands as a pivotal task. To enhance the accuracy and efficiency of concrete damage identification, this research proposes an improved Self-Organizing Map algorithm based on visual sensing. By optimizing feature extraction and representation methods, introducing novel learning strategies, and incorporating spatial attention mechanisms, the model becomes adept at capturing and identifying concrete damage features more effectively. Additionally, employing stochastic gradient descent as an optimization algorithm enhances the model training efficiency. Experimental results showcase that the model exhibits a detection time of merely 0.8 seconds, while demonstrating outstanding fitting and clustering performance, achieving an actual accuracy of 98.2%. Compared to methods based on digital image monitoring and deep learning detection, it shows an improvement of 12.7% and 31.8%, respectively. The proposed enhanced model significantly augments the accuracy and efficiency of concrete damage identification, providing an effective solution for the health monitoring of concrete structures, particularly in scenarios requiring large-scale and real-time monitoring. This advancement elevates the practicality and convenience of concrete damage detection, propelling progress in the field of building safety.
This paper explores the application of deep learning and computer vision techniques for automated classification and detection of electronic waste (e-waste). A system based on convolutional neural networks (CNN) and faster R-CNN is developed for analyzing e-waste images and extracting information about equipment type and dimensions. The experiment is conducted on a dataset of 500 real-world images of three key e-waste categories – refrigerators, kitchen stoves and TVs. Results demonstrate high classification accuracy of 92% using CNN and 91% detection accuracy with R-CNN. The obtained data enables more precise waste collection planning. The main conclusion is that deep learning holds great potential for improving e-waste management systems.
PL
Artykuł ten bada zastosowanie technik głębokiego uczenia i widzenia komputerowego do automatycznej klasyfikacji i detekcji elektronicznych odpadów (e-odpadów). Opracowany zostaje system oparty na splotowych sieciach neuronowych (CNN) i szybszym R-CNN do analizy obrazów e-odpadów oraz wydobycia informacji o typie i wymiarach sprzętu. Eksperyment przeprowadzony jest na zbiorze danych 500 realnych obrazów trzech kluczowych kategorii e-odpadów – lodówek, kuchenek kuchennych i telewizorów. Wyniki wykazują wysoką dokładność klasyfikacji na poziomie 92% przy użyciu CNN oraz dokładność detekcji na poziomie 91% przy użyciu R-CNN. Uzyskane dane umożliwiają bardziej precyzyjne planowanie zbierania odpadów. Głównym wnioskiem jest, że głębokie uczenie ma duży potencjał do poprawy systemów zarządzania e odpadami.
W artykule przedstawiony został system obrazowej analizy zachowania dystansu społecznego za pomocą współczesnych algorytmów detekcyjnych opartych na konwolucyjnych sieciach neuronowych. Algorytm wykonywany jest na procesorze graficznym (GPU), dzięki czemu wykonany system może zostać zaimplementowany na komputerze PC średniej klasy. Wynik detekcji obrazowany jest graficznie poprzez objęcie wykrytych w analizowanej scenie osób ramkami w kolorze zależnym od wyznaczonego dystansu.
EN
The article presents a system of visual analysis of social distancing behavior using modern detection algorithms based on convolutional neural networks. The algorithm is executed on a graphics processor (GPU), so that the system made can be implemented on a mid-range PC. The detection result is graphically illustrated by covering the people detected in the analyzed scene with frames in a color depending on the determined distance.
Oil is used for lubrication and cooling in every standard jet engine. Therefore, hydraulic installations are one of main parts of most of component test rigs and in some cases, they could be large and complicated. Removing sources of leakages is significant task for engineers and technicians. Oil leakages generate costs, reduce reliability of tests and are difficult to detect with use of classic sensors. This paper describes implementation of computer vision methods in the aviation component test laboratory. Three algorithms were proposed and successfully tested.
PL
Olej jest wykorzystywany do smarowania i chłodzenia w każdym silniku odrzutowym. Z tego względu instalacje olejowe s ˛a jednymi z głównych części stanowisk badawczych, a usuwanie przyczyn wycieków jest znaczącym zadaniem inżynierów i techników. Wycieki oleju generują koszty, ograniczają wiarygodność testów i są trudne do wykrycia przy pomocy klasycznych czujników pomiarowych. Dokument opisuje implementację metod widzenia maszynowego w lotniczych laboratoriach badawczych. W ramach prac zostały zaproponowane i przetestowane trzy algorytmy.
10
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
This paper presents the effectiveness of different multimodal neural networks in captioning newspaper scan images. These methods were evaluated on a dataset created for the Temporal Image Caption Retrieval Competition, which is a part of the FedCSIS 2023 conference. The task was to predict a relevant caption for a picture taken from a newspaper, chosen from a given list of captions. The results we obtained show the promising potential of image captioning using CLIP architectures and emphasize the importance of developing new multimodal methods for problems that combine multiple disciplines, such as computer vision with natural language processing.
This paper describes an image caption generation system using deep neural networks. The model is trained to maximize the probability of generated sentence, given the image. The model utilizes transfer learning in the form of pretrained convolutional neural networks to preprocess the image data. The datasets are composed of a still photographs and associated with it, five captions in English language. Constructed model is compared to other similarly constructed models using BLEU score system and ways to further improve its performance are proposed.
PL
W tym artykule opisano system generujący podpisy do zdjęć z wykorzystaniem głębokich sieci neuronowych. Model jest trenowany pod kątem maksymalizacji prawdopodobieństwa wygenerowanego zdania, dla zadanego obrazu. Model wykorzystuje uczenie transferowe w postaci wytrenowanych wstępnie neuronowych sieci konwolucyjnych. Zbiory danych wykorzystane do trenowania modelu składają się z fotografii, oraz przypisanych do niej pięciu zdań w języku angielskim. Skonstruowany model jest potem porównany z innymi modelami o podobnej konstrukcji z wykorzystaniem punktacji BLEU.
The paper describes visualization steps of the surface of internal structures of the human body during stereo-endoscopic and laparoscopic operations using modern computer vision techniques. The presented stages make it possible to obtain three-dimensional representation (more useful for representation and analysis), which is especially important for assessing the state of the examined area and for training health care specialists. The direction of further research is the development of training tools using the proposed approaches.
PL
W pracy opisano etapy wizualizacji powierzchni struktur wewnętrznych ciała ludzkiego podczas operacji stereo-endoskopowych i laparoskopowych z wykorzystaniem nowoczesnych technik widzenia komputerowego. Przedstawione etapy pozwalają na uzyskanie trójwymiarowej reprezentacji (bardziej przydatnej do reprezentacji i analizy), co jest szczególnie istotne dla oceny stanu badanego obszaru oraz dla szkolenia specjalistów ochrony zdrowia. Kierunkiem dalszych badań jest opracowanie narzędzi szkoleniowych wykorzystujących proponowane podejście.
Tracking of small objects in any given airspace is an integral part of modern security systems. In these systems, there are embedded methods that employ the techniques based on either radio waves, or acoustic signals, or light radiation. The computer vision operation, springing from the light radiation-based technique, has prompted interest in its research. This operation has the advantage of being less expensive than radars and acoustic systems. In addition, it can solve complex security problems by detecting and tracking humans, vehicles, and flying objects. Therefore, this article evaluates the usefulness of the varying computer vision algorithms for tracking of small flying objects.
PL
W artykule przedstawiono analizę metod śledzenia bezzałogowych statków powietrznych, wykorzystujących techniki widzenia komputerowego.
Detection of small objects in the airspace is a crucial task in the military. In the era of today’s unmanned aerial vehicles (UAVs) technology, many military units are exposed to recognition and observation through flying objects. They are often equipped with optoelectronic warhead making a way to collect essential and secret data of the military unit. Modern technical solutions make it possible to implement some methods facilitating detection of flying objects. A lot of them utilize computer vision techniques based on image processing algorithm. Therefore, in this article, we present an analysis of the most promising algorithm for detection of small flying objects.
PL
W artykule przedstawiono analizę metod wykrywania bezzałogowych statków powietrznych wykorzystujących techniki widzenia komputerowego.
The identity of a language being spoken has been tackled over the years via statistical models on audio samples. A drawback of these approaches is the unavailability of phonetically transcribed data for all languages. This work proposes an approach based on image classification that utilized image representations of audio samples. Our model used Neural Networks and deep learning algorithms to analyse and classify three languages. The input to our network is a Spectrogram that was processed through the networks to extract local visual and temporal features for language prediction. From the model, we achieved 95.56 % accuracy on the test samples from the 3 languages.
This article is devoted to works on using natural user interfaces (NUI) in computer support systems of aircraft service. The concept of such interfaces involves the usage in human-machine communication the same measures as in the communication between people, that is sound or gesture. In the case of gesture communication, it is indispensable to adopt methods related to computer vision algorithms. One of them is a three-dimensional reconstruction of objects based on processing techniques of a pair of two-dimensional images. The above method and the results of its application were presented to obtain a three-dimensional cloud of points describing the hand shape. The obtained software will constitute an element of gesture classifier based on the analysis of the spatial location of the acquired points of the cloud.
PL
Artykuł dotyczy prac nad wykorzystaniem naturalnych interfejsów użytkownika w komputerowych systemach wspomagania obsługi statków powietrznych. Koncepcja tego typu interfejsów zakłada wykorzystanie w komunikacji człowiek-komputer takich samych środków jak w komunikacji między ludźmi, a więc głosu lub gestu. W przypadku komunikacji za pomocą gestów konieczne jest zastosowanie metod związanych z algorytmami komputerowego widzenia. Jedną z nich jest trójwymiarowa rekonstrukcja obiektów oparta na technikach przetwarzania pary dwuwymiarowych obrazów. Przedstawiono tę metodę oraz wyniki jej zastosowania w celu uzyskania trójwymiarowej chmury punktów opisujących kształt dłoni. Uzyskane oprogramowanie będzie stanowić element klasyfikatora gestów opartego na analizie lokalizacji przestrzennej otrzymanych punktów chmury.
Badanie skierowane na wyznaczenie zależności pomiędzy wysokością nadchodzącej fali a prędkością strumienia w otwartych kanałach z użyciem narzędzi widzenia komputerowego. Autorzy korzystają z modelowania komputerowego oraz badań eksperymentalnych do sprawdzenia możliwości wyznaczenia prędkości strumienia poprzez pomiar wysokość fali padającej na częściowo zanurzoną sztuczną przeszkodę znajdującą się na otwartym kanale.
Projection of a complicated geometry of industrial objects is the complex issue, which requires properly planned and prepared measurements. Such objects must be accurately inventoried, but their complicated nature often makes the access and the visibility of their entire surface very difficult. Documentation of measurements is often prepared in the form of sketches, plans or maps, which are amended with photographic documentation. The objective of this paper is to test the possibilities to apply laser scanning and the network of digital images for inventory and monitoring of technical conditions of industrial objects. Processing of a precise documentation acquired basing on terrestrial laser scanning data or dense points clouds generated from digital images still causes many difficulties and problems. Although data processing algorithms have been intensively developed with respect to generation of high resolution orthoimages or precise vector drawings, the existing problems are still connected with limitations related to imperfections of both techniques of measurements.
19
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Artykuł dotyczy procesu budowy taniego systemu rejestracji ruchu. Rozwiązanie to wykorzystuje kamery PlayStation 3 Eye. Artykuł ten pokazuje, w jaki sposób wykonać synchronizacje wielu kamer i jak stworzyć oprogramowanie do rejestracji i przetwarzania danych wideo w czasie rzeczywistym. Niniejszy artykuł prezentuje także algorytm i rezultaty wyszukiwania oraz śledzenia ruchu jednokolorowych obiektów na podstawie obrazów z dwóch zsynchronizowanych kamer.
EN
This article concerns the creation process of a cheap optical motion capture system. The solution uses PlayStation 3 Eye cameras The paper shows how to synchronise multiple cameras and how to develop software for capturing and processing real-time video data. The article presents an algorithm and the results of findings and tracking of mono-colour objects based on images from two synchronised cameras, too.
In this paper we describe Bayesian inference-based approach to the solution of parametric identification problem in the context of updating of a finite element model of a structure. The proposed inverse solution is based on Monte Carlo filter and on the comparison of structure displacements extracted using digital image correlation method during a quasi-static loading and the corresponding displacements predicted by finite element method program. Our approach is applied to the problem of material model parameter identification of an aluminum laboratory-scale frame. The results are also verified by comparing the Monte Carlo filter-based solution with the analytical solution obtained using Kalman filter.
PL
Artykuł przedstawia zastosowanie podejścia opartego na wnioskowaniu bayesowskim do problemu identyfikacji parametrycznej w kontekście strojenia modelu MES konstrukcji. Proponowane rozwiązanie odwrotne opiera się na filtrze Monte Carlo oraz porównaniu przemieszczeń konstrukcji otrzymanych metodą korelacji obrazów cyfrowych podczas quasi statycznej próby obciążeniowej i odpowiadających im przemieszczeń przewidywanych przez program oparty na metodzie elementów skończonych. Nasze podejście zostało zastosowane do identyfikacji parametru modelu materiału aluminiowej ramki laboratoryjnej. Otrzymane wyniki porównano z wynikami otrzymanymi za pomocą filtru Kalmana.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.