In the modern world, the quality of video information services is important. The greatest difficulties arise in the case of providing remote services using wireless information and communication systems (ICS). This concerns the imbalance between the intensity of the video stream and the speed of data transmission in the ICS. Elimination of the imbalance is possible using compression methods. However, compression of the bit volume of video data is achieved with a loss of integrity. Hence, a relevant scientific and applied problem is the further improvement of coding methods based on the elimination of various types of redundancy. At the same time, most methods are characterized by achieving the desired level of compression by introducing distortions. This necessitates the development of a compression direction based on the controlled elimination of the number of different types of redundancy. Accordingly, to create conditions for increasing the possibilities for detecting characteristic dependencies, the technological apparatus of converting video segments to spectral space is used. Therefore, the research goal of the article is to develop compression methods based on a controlled reduction in the number of different types of redundancy in spectral space. The article outlines the main stages of developing a method for structural block coding of transform tuples. To increase the compression level without loss of integrity, it is proposed to identify characteristic structural dependencies for a set of transforms. Such sets are divided into tuples according to the parametric data of the structural and spectral features of the transforms. Further, such a space will be referred to as the spectral and parametric description of transforms (SPDT). Comparative evaluation of compression methods is carried out according to the system of indicators: compression level - integrity level. Under such conditions, the advantage of the created method over basic analogues in terms of compression level is from 12 to 19%.
PL
We współczesnym świecie jakość usług związanych z przekazem obrazu ma duże znaczenie. Największe trudności pojawiają się w przypadku świadczenia usług zdalnych z wykorzystaniem bezprzewodowych systemów informacyjno-komunikacyjnych (ICS). Dotyczy to nierównowagi między intensywnością strumienia wideo a prędkością transmisji danych w ICS. Wyeliminowanie tej nierównowagi jest możliwe dzięki zastosowaniu metod kompresji. Jednak kompresja objętości bitowej danych wideo wiąże się z utratą integralności. Stąd też istotnym problemem naukowym i praktycznym jest dalsze udoskonalanie metod kodowania opartych na eliminacji różnych rodzajów nadmiarowości. Jednocześnie większość metod charakteryzuje się osiąganiem pożądanego poziomu kompresji poprzez wprowadzanie zniekształceń. Wymaga to opracowania kierunku kompresji opartego na kontrolowanej eliminacji różnych rodzajów nadmiarowości. W związku z tym, aby stworzyć warunki do zwiększenia możliwości wykrywania charakterystycznych zależności, wykorzystuje się aparaturę technologiczną przekształcającą segmenty wideo do przestrzeni spektralnej. W związku z tym celem badawczym niniejszego artykułu jest opracowanie metod kompresji opartych na kontrolowanym zmniejszeniu liczby różnych rodzajów nadmiarowości w przestrzeni spektralnej. W artykule nakreślono główne etapy opracowywania metody strukturalnego kodowania blokowego krotek transformacji. Aby zwiększyć stopień kompresji bez utraty integralności, proponuje się zidentyfikowanie charakterystycznych zależności strukturalnych dla zbioru transformacji. Zbiory takie dzieli się na krotki zgodnie z danymi parametrycznymi cech strukturalnych i spektralnych transformacji. Ponadto taka przestrzeń będzie określana jako spektralny i parametryczny opis transformacji (SPOT). Ocena porównawcza metod kompresji przeprowadzana jest zgodnie z systemem wskaźników: poziom kompresji – poziom integralności. W takich warunkach przewaga opracowanej metody nad podstawowymi odpowiednikami pod względem poziomu kompresji wynosi od 12 do 19%.
Foresight can be viewed as an approach to managing uncertainty—an instrument that enables foreseeing while actively shaping the future under conditions of unpredictability. The rapid development of artificial intelligence (AI) has introduced new opportunities for foresight research. Although AI methods have not traditionally been part of the foresight canon, they offer significant potential for future applications. Integrating machine learning (ML) techniques into foresight research appears to be a natural progression. AI provides transformative capabilities by analysing complex datasets, uncovering hidden relationships, and generating data-driven recommendations. This work investigates the integration of AI tools into technology foresight projects by reviewing existing literature on their combined application. The analysis identifies the most frequently used AI and foresight methods, along with their primary objectives, providing a structured overview of current practices. Empirical analysis, based on data from a technology foresight project, demonstrates how AI can be utilised to enhance data analysis, thereby supporting theoretical considerations and complementing the traditional expert panel approach for technology clustering. The AI-assisted process provides a scalable alternative to traditional methods, with code tools, enhancing perspectives on identifying technology clusters, selecting key attributes, and incorporating expert self-assessment. However, the value of the proposed approaches lies more in a posteriori analysis, which can be utilised in future foresight projects regarding the attributes used for evaluation or the selection of expert panels. The diversity of proposed analyses demonstrates various interpretation possibilities but does not fundamentally influence the achievement of the main goal, which is the identification of key technologies.
Automation of map production is an important subject of work of many scientists. Particular attention should be paid to dot distribution maps, the editing of which is very time-consuming and complicated due to the lack of support in GIS programs. The aim of this research was to develop a method supporting automation of dot map creation. In the study, equal-size spectral clustering algorithm was used, which was modified with a function equalizing the number of points in clusters. Spatial data on residential buildings and statistical data on the population were integrated to calculate theoretical population distributions. These data were entered into the spectral clustering algorithm based on a predefined dot value. The output clusters were then visualized in ArcGIS Pro, where manual adjustments, such as the definition of the dot size and the dispersion of overlapping markers, completed the map editing process. The results showed that the algorithm successfully created clusters representing the population distribution with an acceptable margin of error of the dot map of less than 5% for the entire county (study area – County of Pszczyna, Poland). The adaptation of the equal-size spectral clustering algorithm for cartographic purposes shows its potential to support automation of the dot distribution map editing process. The study found that reducing the input dot value slightly below the target value improved clustering precision, resulting in more consistent clusters. Despite these successes, the method has limitations, including partial reliance on manual corrections in densely populated areas where overlapping dots could not be fully automatically resolved. These issues underscore the need for further refinement to achieve full automation.
Selecting an optimal mining method is a complex and critical decision in underground mining, influenced by multiple geological, technical, and economic parameters. This study introduces a novel frame-work that combines Hierarchical Clustering (HC) and Correspondence Analysis (CA) to enhance the selection process by evaluating the consistency and similarity among outcomes from both first-pass methods (UBC and Nicholas) and several multi-criteria decision-making (MCDM) techniques (including AHP, EDAS, PROMETHEE II, AHP-PROMETHEE, TOPSIS, and VIKOR). The proposed HC-CA approach identifies consistent conflicts among the considered mining methods and quantifies the agreement among the initial assumptions of the adopted selection procedures. A case study of a Pb-Zn deposit demonstrates that the framework can effectively detect consistent and co-occurring (i.e., conflicting) solutions, such as Cut-and-Fill Stoping, Shrinkage Stoping, and Sublevel Stoping. The results show that the adopted design criteria align more closely with the UBC selection method, compared to the Nicholas selection procedure for the considered deposit. Additionally, applying the HC-CA approach to the input matrices prior to applying the MCDM methods can yield different results, compared to subjecting the MCDM output scores to the proposed framework. This integrative approach extends traditional selection procedures and links them with commonly used MCDM methodologies and unsupervised machine learning methods by enabling flexible strategy development, with the inclusion of considering mixed-mining-method scenarios tailored to the deposit. Additionally, the approach offers improved decision support in early project stages by visualizing affinities among different assumptions and hence potentially mitigating biases during the following design stage.
When the hydraulic motor fault occurs, it is not easy to be detected, and the leakage degree will gradually increase. In order to avoid bigger accidents caused by the hydraulic motor fault, the accident is excluded in the embryonic stage, and the hydraulic motor fault prediction method based on fuzzy neural network is used to predict the hydraulic motor fault. The feature vector is output in the global mean pooling layer, and the feature vector matrix between the health state feature vector library and the samples to be measured is constructed. The dynamic cluster graph is obtained by fuzzy clustering, so as to realize the fault diagnosis of the hydraulic motor. The results show that the accuracy of training set, verification set and test set is higher than 99.8%. The accuracy of diagnosis classification is 99.00%, which is better than other comparison models. In this study, the number of training samples can be appropriately increased or decreased according to the curve complexity of the detection target, so as to improve the feature extraction capability of the convolutional layer and increase the classification accuracy.
Algorytmy klasteryzacji pełnią kluczową rolę w analizie dużego wolumenu ruchu i wykrywaniu wzorców nieznanych ataków. W pracy omówiono problem grupowania danych zebranych przez rozproszony system pułapek sieciowych. W proponowanym podejściu, zgodność przepływu sieciowego z sygnaturą ataku jest równoważna nadaniu mu odpowiedniej etykiety. Dzięki temu możliwe jest zastosowanie algorytmów uczenia częściowo nadzorowanego oraz poprawa jakości wyników klasteryzacji. W artykule porównano wyniki algorytmów uczenia przeprowadzonego bez oraz z częściowym nadzorem w zadaniu grupowania przepływów sieciowych.
EN
Clustering algorithms play a crucial role in detecting patterns of unknown attacks. The paper discusses the problem of clustering data collected by a distributed system of network traps. In the proposed approach, the flow conformity to the attack signature is equivalent to assigning it an appropriate label. This allows for the application of semi-supervised learning algorithms and the improvement of clustering quality. The article compares the results of learning algorithms conducted with and without partial supervision in clustering network flows.
An experiment is presented of cluster-wise modelling, related to modelling and analysis of internal migration flows in Poland at the municipality level (some 2500 units). The experiment was carried out within a project, in which it was assumed that migration flows linearly depend upon unemployment. This simple dependence was positively verified, and the associated model error maps, corresponding to the consecutive years over two decades (2003-2022) provide a very clear and telling spatial image. Yet, the models obtained were statistically rather feeble. So, it was decided to experiment, in particular, with a set of analogous models, identified for subsets (clusters) of municipalities, employing a simile of classical k-means clustering procedure. Given the known dependence of the outcome from the k-means-like procedure upon the starting point, various initial configurations were considered. The results are exemplified in this paper, the problems appearing indicated, and some broader conclusions are drawn therefrom.
The prediction of physicochemical parameters of novel compounds and their blends is an emerging problem in chemical nanoengineering, especially in the fields of fine chemicals and specialty polymers. Hansen solubility parameters in practice (HSPiP) software enables one to get valuable data from the known chemical structure on the basis of appropriate, standardized experiment, e.g. dissolution or swelling in a particular set of solvents. Herein we present a brief description of the recent applications of HSPiP software, supported by experimental and theoretical methods, including machine learning clusterization, aimed to solve technological and engineering issues. The present contribution aims to show the combined way of solvent clusterization for HSPiP studies for the determination of solubility parameters for novel surfactants, polymer blending toward drug nanocarriers as well as design of silicone-based polymer nanomatrix. The examples comprise the recent contribution of the groups’ research, previously published or studied in current projects of our groups.
The 5G enhanced mobile broadband (eMBB) category offers faster data rates, network capacity, and user experiences than prior generations. This research aims to boost the 5G uplink user equipment (UE) user data transfer rate. We use Python to build frameworks and analyze data. A 250-m-radius centre-excited picocell base station (PBS) is investigated to support 15 clients. Cell-range Poisson distribution determines user position. All UEs send channel state information (CSI) to the PBS, which evaluates signal transmission channel conditions. The study uses Rayleigh, Rician, free space path, and long-distance route loss models. This inquiry produces a channel state dataset and then it is formulated dataset is dynamic. For service-specific requirements, UEs use k-means clustering. Clustering concatenates bandwidth, enhancing system efficiency and UE sum rate. The research includes observations from simulation findings, in which UEs are grouped by channel gain, achievable data rate, and minimum service-required data rate. Users in cluster 3 achieve the highest cumulative rate of 9.09 Mbps after clustering with an average of 7.16 Mbps. Bandwidth concatenation increased system capacity, meeting each UE service needs. After evaluating performance criteria for different clustering models, k-means remains the best algorithm for the framework. The methodology was carefully designed to satisfy study goals. This paper investigates beamforming and dynamic clustering to improve user fairness and performance.
10
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
The accurate interpretation of well-logging data is a crucial stage in the exploration of gas- and oil-bearing reservoirs. Geological formations, such as the Miocene deposits, present many challenges related to thin layers, whose thickness is often less than the measurement resolution. This research emphasizes the potential of utilizing electrofacies in such challenging environments. The application of electrofacies not only allows for the grouping of intervals with similar physical characteristics but can also be useful for estimating porosity and permeability parameters. For this purpose, various clustering methods were tested, including the 2D indexed and probabilized self-organizing map (IPSOM) method with and without supervision. Subsequently, the usefulness of the obtained results to improve the estimation of porosity and permeability parameters with the help of artificial neural networks was verified. As a result of the conducted analyses, significantly better results were obtained compared to classical petrophysical interpretation. The calculated porosity and permeability parameters were characterized by much greater variability and alignment with laboratory measurements on porosity and permeability. The best results were obtained for the IPSOM method, but the other methods did not differ significantly. In conclusion, the studies have shown a positive result of applying clustering methods, including the IPSOM method, to improve the estimation of permeability and porosity parameters in complicated, thinly-layered formations.
Business process management has become increasingly important in the manufacturing sector, playing a vital role in fostering productivity and facilitating organizational adaptability to technological advancements. Although each company’s business process models vary in uniqueness and complexity, certain similarities can be identified from these differences. This study employs a systematic literature review to aggregate and summarize findings on business process models, modeling languages, clustering techniques, and foundational clustering principles. Results demonstrate the dearth of research on the grouping of business process models in the manufacturing industry. Several studies have focused on sectors such as services, trading, and insurance. Research specifically addressing clustering in the manufacturing sector is limited. Existing clustering efforts in manufacturing revolve around groupings related to product defects, industrial locations, business ecosystems, and similar factors. Analysis of the methods, scope, and criteria used in grouping business process models in the manufacturing industry indicates that most approaches rely on structural or graphic similarities. Follow-up research is lacking once these business process clusters are identified. This study proposes a novel approach to grouping, integrating business process modeling with the implementation of a management information system. Business process management relies on integrating departments within the company through an enterprise resource planning (ERP) system. The next step involves proposing a conceptual framework to categorize business processes and assess the comparability of models in the manufacturing sector. Future research directions are also delineated.
Background: Global supply chains are confronted with the challenge of ensuring on-time deliveries while simultaneously enhancing supply chain resilience. Conventional methods aim to address the complexities of modern supply chains, promoting the transition to intelligent and data-driven strategies. Methods: This research represents an innovative methodology for predicting the risk of late deliveries in supply chains. The presented framework combines clustering and multiclassification techniques, where the clustering phase is executed through hyperparameter optimization and a novel metaheuristic called RIME. In the multiclassification phase, five distinct deep learning models are employed, namely, Generative Adversarial Network (GAN), Convolutional Neural Network Long Short-Term Memory (CNN-LSTM), within Ensemble learning via bagging, Ensemble learning stacking, and Ensemble learning within boosting. The three ensemble learning models are based in GAN and CNN-LSTM. Result: This paper presents a systematic evaluation of diverse models in a risk of late delivery prediction framework. This evaluation demonstrates that Ensemble learning stacking provides the higher accuracy by 0.926, showcasing its prowess in precise predictions. Notably, Ensemble learning bagging and Ensemble learning boosting exhibit strong precision. Regression metrics reveal Ensemble learning stacking and Ensemble learning bagging's superior error minimization (MSE 0.11, MAE 0.09). This metric demonstrates that the proposed model can predict the risk level of late delivery in a supply chain with high precision. Conclusion: This paper introduces an innovative clustering and multiclassification-based framework for predicting the risk of late deliveries. The ability of prediction late deliveries risk helps organizations to enhance supply chain resilience by adopting a proactive management risks strategy, optimizing operational processes, and elevating customer satisfaction.
This research is concerned with the fusion of artificial intelligence (AI) and machine learning within multi-hierarchical caching systems, specifically targeting vehicular and edge caching domains. This study introduces an innovative architecture harmonizing Thompson sampling learning-based caching policies with advanced vehicle clustering and content-popularity prediction methods (TS-MMCM). Simulations show substantial performance improvements and a big impact of the proposed approach on system efficiency in dynamic network environments. The proposal demonstrates a notable gain in cache hit rates and decreased latency levels, highlighting the potential of AI to improve caching techniques in dynamic network environments.
This study aims to analyze energy consumption patterns across selected nations from Africa, America, Asia/Middle East, and Europe, with a focus on the types of energy sources used. Covering 46 countries, the research spans the years 2000 to 2018 and examines the distribution and changes in energy consumption by source and type. The regions studied include diverse countries such as Austria, Sweden, Czechia, and Croatia in Europe; Algeria, Egypt, and South Africa in Africa; China, India, and Saudi Arabia in Asia/Middle East; and Brazil, Canada, and the United States in the Americas, with Australia and New Zealand representing Oceania. Utilizing data from the BP Statistical Review of World Energy and the SHIFT Data Portal, along with key indicators maintained by Our World in Data, the study employs methods such as descriptive statistics, cluster analysis using k-means, and time-series clustering with dynamic time warping (DTW). The analysis highlights regional similarities and variances in energy use, providing new insights into the complex relationship between energy consumption patterns and factors such as economic growth, national policies, and geopolitical contexts. This research addresses a significant gap in the existing literature by offering a detailed comparative analysis of how different nations manage and consume energy. It contributes to the broader discourse on sustainable energy policies and economic development in the face of global energy challenges.
W artykule przedstawiono analizę statystyczną wieloletnich danych (wartości godzinowe zapotrzebowania na energię elektryczną) z KSE oraz analizę możliwości zastosowania sztucznej sieci neuronowej samoorganizującej się (Self Organizing Map) do podziału dobowych profili zapotrzebowania na energię elektryczną w KSE. Artykuł kończy podsumowanie oraz wnioski z wykonanych analiz statystycznych oraz badań związanych z zastosowaniem SOM do grupowania profili zapotrzebowania na energię.
EN
The article presents a statistical analysis of long-term data (hourly values of electricity demand) from the NPS and an analysis of the possibility of using a self-organizing artificial neural network (Self Organizing Map) to divide daily profiles of electricity demand in the NPS. The article concludes with a summary and conclusions from the conducted statistical analyses and studies related to the application of SOM for clustering electricity demand profiles.
In this paper, we attempt to generalize the ability to achieve quality inferences of survey data for a larger population through data augmentation and unification. Data augmentation techniques have proven effective in enhancing models’ performance by expanding the dataset’s size. We employ ML data augmentation, unification, and clustering techniques. First, we augment the limited survey data size using data augmentation technique(s). Second, we carry out data unification, followed by clustering for inferencing. We took two benchmark survey datasets to demonstrate the effectiveness of augmentation and unification. The first dataset contains information on aspiring student entrepreneurs’ characteristics, while the second dataset comprises survey datarelated to breast cancer. We compare the inferences drawn from the original survey data with those derived from the transformed data using the proposed scheme. The results of this study indicate that the machine learning approach, data augmentation with the unification of data followed by clustering, can be beneficial for generalizing the inferences drawn from the survey data.
The scope of this paper is that it investigates and proposes a new clustering method thattakes into account the timing characteristics of frequently used feature words and thesemantic similarity of microblog short texts as well as designing and implementing mi-croblog topic detection and detection based on clustering results. The aim of the proposedresearch is to provide a new cluster overlap reduction method based on the divisions ofsemantic memberships to solve limited semantic expression and diversify short microblogcontents. First, by defining the time-series frequent word set of the microblog text, a fea-ture word selection method for hot topics is given; then, for the existence of initial clusters,according to the time-series recurring feature word set, to obtain the initial clustering ofthe microblog.
Machine learning-based classification algorithms allow communication and computing (2C) task allocation to network edge servers. This article considers poisoning of classifiable 2C data features in two scenarios: noise-like jamming and targeted data falsification. These attacks have a fatal effect on classification in the feature areas with unclear decision boundary. We propose training and noise detection using the Silhouette Score to detect and mitigate attacks. We demonstrate effectiveness of our methods.
PL
Algorytmy klasyfikacji oparte na uczeniu maszynowym umożliwiają zoptymalizowaną alokację zadań telekomunikacyjnych i obliczeniowych (2C) do serwerów brzegowych sieci. W artykule omówiono ataki zatruwające, które mają negatywny wpływ na klasyfikację zadań 2C w obszarach, w których granica decyzyjna jest niejasna. Proponujemy metodę trenowania modelu oraz wykorzystanie testu Silhouette do wykrywania i unikania ataków. Wykazujemy skuteczność tych metod wobec rozważanych ataków.
The article presents the application of selected clustering algorithms for detecting anomalies in financial data compared to several dedicated algorithms for this problem. To apply clustering algorithms for anomaly detection, the Determine Abnormal Clusters Algorithm (DACA) was developed and implemented. This parameterized script (DACA) allows clusters containing anomalies to be automatically detected on the basis of defined distance measures. This kind of operation allows clustering algorithms to be quickly and efficiently adapted to anomaly detection. The prepared test environment has allowed for the comparison of selected clustering algorithms. K-Means, Hierarchical Cluster Analysis, K-Medoids, and anomaly detection: Stochastic Outlier Selection, Isolation Forest, Elliptic Envelope. The research has been carried out on real financial data, in particular on the income declared in the asset declarations of the targeted professional group. The experience of financial experts has been used to assess anomalies. Furthermore, the results have been evaluated according to a number of popular classification and clustering measures. The highest result for the investigated financial problem was provided by the K-Medoids algorithm in combination with the DACA script. It is worthwhile to conduct future research on the introduced solutions as an ensemble method.
PL
Artykuł przedstawia zastosowanie wybranych algorytmów klasteryzacji do wykrywania anomalii w danych finansowych w porównaniu do kilku dedykowanych algorytmów dla tego problemu. W celu wykorzystania algorytmów klasteryzacji do wykrywania anomalii opracowano i zaimplementowano Determine Abnormal Clusters Algorithm (DACA). Ten sparametryzowany skrypt umożliwia na automatyczne wykrycie klastrów zawierających anomalie, na podstawie zdefiniowanych miar odległości. Takie działanie pozwala na szybkie i skuteczne dostosowanie algorytmów klasteryzacji do wyszukiwania anomalii. Przygotowane środowisko badawcze pozwoliło na porównanie wybranych algorytmów klasteryzacji: Hierarchical Cluster Analysis, K-Means, K-Medoids oraz wykrywania anomalii: Stochastic Outlier Selection, Isolation Forest, Elliptic Envelope, Badania przeprowadzono na rzeczywistych danych finansowych, w szczególności dotyczących dochodów zadeklarowanych w oświadczeniach majątkowych wybranej grupy zawodowej. Wykorzystano doświadczenie ekspertów finansowych do oceny anomalii. Ponadto, wyniki oceniono na podstawie wielu popularnych miar klasyfikacji i klasteryzacji. Najlepsze wyniki dla badanego problemu finansowego przedstawił algorytm K-Medoids w połączeniu ze skryptem DACA. W przyszłości warto przebadać metody złożone oparte o przedstawione rozwiązanie.
This paper researches the detection method in halftone images using AdaBoost algorithm for training a object detector. Modified Haar-like features are used as features in weak classifiers. The method has been tested on The Yale B Face Database where images are obtained under 65 different illumination conditions. Experimental research on face detection method was carried out using the Matlab environment.
PL
W artykule zbadano metodę detekcji w obrazach półtonowych z wykorzystaniem algorytmu AdaBoost do szkolenia detektora obiektów. Zmodyfikowane cechy podobne do Haara są używane jako cechy w słabych klasyfikatorach. Metoda została przetestowana w bazie danych Yale B Face Database, gdzie obrazy są uzyskiwane w 65 różnych warunkach oświetleniowych. Badania eksperymentalne metody detekcji twarzy przeprowadzono z wykorzystaniem środowiska Matlab.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.