The paper considers the problems of automating the construction of classification trees based on the scheme of branched feature selection.The object of research is classification trees. The subject of research is methods, algorithms, and schemes for constructing classification trees. The aimof this work is to build an effective method (scheme) for synthesizing classification tree models based on a group assessmentof the importance of discrete features within a branched attribute selection. A method forconstructing classification trees is proposed, which for a given training sample determinesthe individual information content (importance) of groups of features (and their combinations) in relation to the initial value of the classification function (data from the training sample). The developed logical tree method, when constructing the next node of the classification tree, tries to identify a groupof the most closely interrelated discrete features, this reduces the overall structural complexity of the model (the number of levels of the classification tree), speeds up calculations when recognizing objects based on the model, and also increases the generalizing properties of the model and its enterprise.The proposed scheme for selecting groups of discrete traits allows using the constructed decision tree to assess the informative value (importance)of traits. The developed classification tree method is implemented programmatically and studied when solving the problem of classifying discrete objects represented by a set of features. The conducted experiments confirmed the operability of the proposed mathematical support and allow us to recommendit for use in practice in solving applied problems of classification of discrete objects based on logical classification trees.Prospects for further research may consist in creating a modified method of the logical classification tree by effectively iterating and evaluating sets of elementary features basedon the proposed method, optimizing its software implementations, and experimentally studying the proposed method on a wider set of applied problems.
PL
W artykule rozważono problemy związane z automatyzacją tworzenia drzew klasyfikacyjnych w oparciu o schemat rozgałęzionego wyboru cech. Przedmiotem badań są drzewa klasyfikacyjne. Tematem badań są metody, algorytmy i schematy tworzenia drzew klasyfikacyjnych. Celem niniejszej pracy jest opracowanie skutecznej metody (schematu) syntezy modeli drzew klasyfikacyjnych w oparciu o grupową ocenę znaczeniacech dyskretnychw ramach rozgałęzionego wyboru atrybutów. Zaproponowano metodę konstruowania drzew klasyfikacyjnych, która dla danej próbki szkoleniowej określa indywidualną zawartość informacyjną (znaczenie) grup cech (i ich kombinacji) w odniesieniu do wartości początkowej funkcji klasyfikacyjnej (danez próbki szkoleniowej). Opracowana metoda drzewa logicznego, podczas konstruowania kolejnego węzła drzewa klasyfikacyjnego, próbuje zidentyfikować grupę najbardziej powiązanych ze sobą cech dyskretnych, co zmniejsza ogólną złożoność strukturalną modelu (liczbę poziomów drzewa klasyfikacyjnego), przyspiesza obliczenia podczas rozpoznawania obiektów na podstawie modelu, a także zwiększa właściwości uogólniające modelu i jego przedsiębiorstwa. Proponowany schemat wyboru grup cech dyskretnych pozwala na wykorzystanie skonstruowanego drzewa decyzyjnego do oceny wartości informacyjnej (znaczenia) cech. Opracowana metoda drzewa klasyfikacyjnego została zaimplementowana programowo i zbadana podczas rozwiązywania problemu klasyfikacji obiektów dyskretnych reprezentowanych przez zbiór cech. Przeprowadzone eksperymenty potwierdziły funkcjonalność proponowanego wsparcia matematycznego i pozwalają nam polecić je do praktycznego zastosowania w rozwiązywaniu problemów stosowanych związanych z klasyfikacją obiektów dyskretnych w oparciu o logiczne drzewa klasyfikacyjne. Perspektywydalszych badań mogą polegać na stworzeniu zmodyfikowanej metody logicznego drzewa klasyfikacyjnego poprzez skuteczne iterowanie i ocenę zestawów cech elementarnych w oparciu o proponowaną metodę, optymalizację jej implementacji programowych oraz eksperymentalne badanie proponowanej metody na szerszym zestawie problemów stosowanych.
The radio stations now use a hybrid method of broadcasting. Transmission between the studio and the transmitter is digital. In the transmitter digital signal is converted to analog and this signal is transmitted by AM or FM modulation. The identified problem are transmission errors in the digital path. Single errors in transmission are not unusual and happen in most cases systems. However, repeated and grouping errors are usually a sign of track damage transmission line and require intervention. In the article presents an example solution to automatically detect some errors that manifest themselves in the identified way. To identify transmission errors, basic statistical methods were used that allowed to create a reference data set on the basis of which it is possible to identify the fact occurrence of an error based on patterns and algorithms from the machine learning area. The environments related to the radio industry are keenly interested in researching this phenomenon because now is not available solutions allowing for automatic detection and identification of this type problems.
W artykule przedstawiono innowacyjny algorytm ukrywania danych w obrazach oparty na klasyfikacji. Metoda pozwala na niezauważalne dla ludzkiego oka ukrycie danych, a jednocześnie zapewnia możliwość ich późniejszego wykrycia i rozpoznania bez konieczności posiadania klucza lub oryginalnego obrazu. Przeprowadzone eksperymenty potwierdzają skuteczność i niezawodność tej metody w porównaniu z innymi algorytmami. Proponowane rozwiązanie może się przyczynić do rozwoju takich dziedzin jak: bezpieczeństwo informacyjne, ochrona prywatności, autoryzacja i znakowanie wodne.
EN
The article presents an innovative algorithm for data hiding in images based on classification. The method allows imperceptible data hiding to the human eye while also providing the ability to later detect and recognise the hidden data without needing a key or the original image. The conducted experiments confirm the effectiveness and reliability of this method compared to other algorithms. The proposed solution can contribute to developing fields such as information security, privacy protection, authentication, and watermarking.
In this study, the effectiveness of six machine learning and eight deep learning algorithms in analyzing electroencephalogram (EEG) signals for detecting epileptic seizures has been investigated. The study utilizes 14 channels in the EMOTIV EPOC+ device which is based on international 10-20 system. To find out the most informative and sensitive channel, one of the 14 channels has been dropped one at a time. The accuracy values were determined for all the methods using two different publicly available datasets: the Guinea-Bissau epilepsy dataset and the Nigeria epilepsy dataset. In case of machine learning models, the performance of SVM classifier performs best with maximum accuracy of 83.2% (Guinea-Bissau) and 77% (Nigeria) without excluding any channels. No significant performance degradation has been observed for single channel exclusion of this classifier. Among the deep learning models, the four best performing models in terms of accuracy are CNN-LSTM (92.5%), IC-RNN (91.8%), ChronoNet (91.1%) and C-DRNN (88.6%). After excluding one channel at a time and investigating their effect on the performance of the four DL models, it has been observed that the most significant and most sensitive channels lie within the frontal and parietal zone. This finding will be very useful in practice as it indicates that the electrodes in the frontal and parietal zone should be placed with absolute precision for accurate diagnosis of the diseases. In addition, this study also explore the effectiveness of the selected classifiers in detecting seizure in case of failure of any particular EEG signal channel.
PL
W tym badaniu zbadano skuteczność sześciu algorytmów uczenia maszynowego i ośmiu algorytmów głębokiego uczenia się w analizie sygnałów elektroencefalogramu (EEG) w celu wykrywania napadów padaczkowych. W badaniu wykorzystano 14 kanałów w urządzeniu EMOTIV EPOC+ opartym na międzynarodowym systemie 10-20. Aby znaleźć najbardziej pouczający i wrażliwy kanał, usuwano jeden z 14 kanałów na raz. Wartości dokładności określono dla wszystkich metod przy użyciu dwóch różnych publicznie dostępnych zbiorów danych: zbioru danych dotyczących padaczki w Gwinei Bissau i zbioru danych dotyczących padaczki w Nigerii. W przypadku modeli uczenia maszynowego wydajność klasyfikatora SVM jest najlepsza przy maksymalnej dokładności 83,2% (Gwinea Bissau) i 77% (Nigeria) bez wykluczania jakichkolwiek kanałów. Nie zaobserwowano znaczącego pogorszenia wydajności w przypadku wykluczenia pojedynczego kanału tego klasyfikatora. Wśród modeli głębokiego uczenia się cztery modele o najlepszych wynikach pod względem dokładności to CNN-LSTM (92,5%), IC-RNN (91,8%), ChronoNet (91,1%) i C DRNN (88,6%). Po wykluczeniu jednego kanału na raz i zbadaniu ich wpływu na działanie czterech modeli DL zaobserwowano, że najważniejsze i najbardziej czułe kanały znajdują się w strefie czołowej i ciemieniowej. Odkrycie to będzie bardzo przydatne w praktyce, gdyż wskazuje, że elektrody w strefie czołowej i ciemieniowej powinny być umieszczone z absolutną precyzją, aby umożliwić trafną diagnostykę schorzeń. Ponadto w badaniu tym zbadano również skuteczność wybranych klasyfikatorów w wykrywaniu napadów w przypadku awarii dowolnego konkretnego kanału sygnałowego EEG.
The cognitive goal of this paper is to assess whether marker-less motion capture systems provide sufficient data to recognize human postures in the side view. The research goal is to develop a new posture classification method that allows for analysing human activities using data recorded by RGB‐D sensors. The method is insensitive to recorded activity duration and gives satisfactory results for the sagittal plane. An improved competitive Neural Network (cNN) was used. The method of pre- processing the data is first discussed. Then, a method for classifying human postures is presented. Finally, classification quality using various distance metrics is assessed. The data sets covering the selection of human activities have been created. Postures typical for these activities have been identified using the classifying neural network. The classification quality obtained using the proposed cNN network and two other popular neural networks were compared. The results confirmed the advantage of cNN network. The developed method makes it possible to recognize human postures by observing movement in the sagittal plane.
Supervised learning as a sub-discipline of machine learning enables the recognition of correlations between input variables (features) and associated outputs (classes) and the application of these to previously unknown data sets. In addition to typical areas of application such as speech and image recognition, fields of applications are also being developed in the sports and fitness sector. The purpose of this work is to implement a workflow for the automated recognition of sports exercises in the Matlab® programming environment and to carry out a comparison of different model structures. First, the acquisition of the sensor signals provided in the local network and their processing is implemented. Realised functionalities include the interpolation of lossy time series, the labelling of the activity intervals performed and, in part, the generation of sliding windows with statistical parameters. The preprocessed data are used for the training of classifiers and artificial neural networks (ANN). These are iteratively optimised in their corresponding hyper parameters for the data structure to be learned. The most reliable models are finally trained with an increased data set, validated and compared with regard to the achieved performance. In addition to the usual evaluation metrics such as F1 score and accuracy, the temporal behaviour of the assignments is also displayed graphically, allowing statements to be made about potential causes of incorrect assignments. In this context, especially the transition areas between the classes are detected as erroneous assignments as well as exercises with insufficient or clearly deviating execution. The best overall accuracy achieved with ANN and the increased dataset was 93.7 %.
Background: Fault prediction is a key problem in software engineering domain. In recent years, an increasing interest in exploiting machine learning techniques to make informed decisions to improve software quality based on available data has been observed. Aim: The study aims to build and examine the predictive capability of advanced fault prediction models based on product and process metrics by using machine learning classifiers and ensemble design. Method: Authors developed a methodological framework, consisting of three phases i.e., (i) metrics identification (ii) experimentation using base ML classifiers and ensemble design (iii) evaluating performance and cost sensitiveness. The study has been conducted on 32 projects from the PROMISE, BUG, and JIRA repositories. Result: The results shows that advanced fault prediction models built using ensemble methods show an overall median of $F$-score ranging between 76.50% and 87.34% and the ROC(AUC) between 77.09% and 84.05% with better predictive capability and cost sensitiveness. Also, non-parametric tests have been applied to test the statistical significance of the classifiers. Conclusion: The proposed advanced models have performed impressively well for inter project fault prediction for projects from PROMISE, BUG, and JIRA repositories.
The article presents the concept of using fuzzy sets methodology in modelling patientʼs disease states for preliminary medical diagnosis. The preliminary medical diagnosis is based on the identified disease symptoms. The basis of the algorithm are descriptions of the patientʼs disease status and patterns of disease entities. These patterns were defined as fuzzy sets. The paper presents simple classifiers that allow he a preliminary diagnosis based on the analysis of fuzzy sets for the use of the general practitioner.
PL
W artykule przedstawiono koncepcję wykorzystania metodologii zbiorów rozmytych w modelowaniu stanów chorobowych pacjenta w algorytmach wstępnej diagnostyki medycznej. Wstępna diagnoza lekarska opiera się na rozpoznanych objawach choroby. Podstawą algorytmu są opisy stanu chorobowego pacjenta i wzorce jednostek chorobowych. Wzorce te zostały zdefiniowane jako zbiory rozmyte. W artykule przedstawiono proste klasyfikatory, które pozwalają na wstępną diagnozę na podstawie analizy zbiorów rozmytych do użytku lekarza pierwszego kontaktu.
Pursuant to the Geodetic and Cartographic Law, the soil science based classification of land should be understood as the division of soils into valuation classes due to their productive quality, determined on the basis of soil genetic features. Pursuant to the above-mentioned Act, the task of the starosta (district administrator) is to maintain both the soil science classification of land, and the land and building records (cadastral records). The data that is the subject of the decision issued by the authority in the field of soil science classification of land constitute elements of the essential information set within land and building records, in accordance with Article 23 section 3 point 1 g of the Geodetic and Cartographic Law [PGiK]. The aim of this publication was to present the irregularities resulting from the failure to update land and building records, as well as from the lack of uniform administrative procedures in the field of soil science classification of land, which translates into the quality of the works performed. The research method used is the case study. The method was supported by the analysis of legislation in the above-mentioned subject matter.
PL
Zgodnie z ustawą Prawo Geodezyjne i Kartograficzne poprzez gleboznawczą klasyfikację gruntów należy rozumieć podział gleb na klasy bonitacyjne ze względu na ich jakość produkcyjną ustaloną na podstawie cech genetycznych gleb. Zgodnie z powyższą ustawą prowadzenie zarówno gleboznawczej klasyfikacji gruntów jak również ewidencji gruntów i budynków należy do zadań starosty. Dane będące przedmiotem wydawanej przez organ decyzji w zakresie gleboznawczej klasyfikacji gruntów są elementami zbioru informacji przedmiotowych ewidencji gruntów i budynków zgodnie z art. 23 ust.3 pkt 1 lit. g ustawy Prawo Geodezyjne i Kartograficzne [pgik]. Celem niniejszej publikacji było przedstawienie nieprawidłowości wynikających z braku aktualizacji ewidencji gruntów i budynków, a także braku jednolitych procedur administracyjnych w zakresie prac gleboznawczej klasyfikacji gruntów, które przekładają się na jakość wykonywanych prac. Stosowaną metodą badawczą jest studium przypadku. Metoda została wsparta analizą prawodawstwa w wyżej wymienionym zakresie.
From the legal point of view, the soil science classification is regulated by the Geodetic and Cartographic Law, where it is defined as the division of soils into valuation classes due to their production quality determined on the basis of the genetic characteristics of the soil. The executive act regulating the issue of soil science classification of land are the provisions included in the Regulation of the Council of Ministers of September 12, 2012 on soil science classification of land (Journal of Laws 2012, item 1246). The aim of the article was to present the problems resulting from the lack of regulation of the profession of land classifier and the lack of uniform administrative procedures regarding the selection of the classifier for the purposes of the classifications. The research method used is the case study. The method was supported by the analysis of legislation in the above-mentioned scope.
PL
Gleboznawcza klasyfikacja gruntów z punktu widzenia prawnego regulowana jest poprzez ustawę Prawo Geodezyjne i Kartograficzne, gdzie zdefiniowana jest jako podział gleb na klasy bonitacyjne ze względu na ich jakość produkcyjną ustaloną na podstawie cech genetycznych gleb. Aktem wykonawczym regulującym zagadnienie gleboznawczej klasyfikacji gruntów są przepisy zawarte w Rozporządzeniu Rady Ministrów z dnia 12 września 2012 r. w sprawie gleboznawczej klasyfikacji gruntów (Dz.U. 2012 poz.1246). Celem artykułu było przedstawienie problemów wynikających z braku uregulowania zawodu klasyfikatora gruntów oraz braku jednolitych procedur administracyjnych dotyczących wyboru klasyfikatora na potrzeby realizowanych prac klasyfikacyjnych. Stosowaną metodą badawczą jest studium przypadku. Metoda została wsparta analizą prawodawstwa w wyżej wymienionym zakresie.
11
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
In the following paper a classification problem with two multivariate normally distributed classes is considered. The problem is solved in a case of an empirical real situation (a motors data) using the Karhunen-Loeve transform and classifying functions based on estimators for unknown parameters of a multivariate normal distribution. We consider the maximum likelihood estimator and a robust one. The robust estimator bases on the Huber's functions. The corresponding classifying functions (classifiers) are compared using the Leave-One-Out metod.
PL
W artykule rozważany jest problem klasyfikacji w przypadku dwóch klas o wielowymiarowym rozkładzie normalnym. Problem ten jest rozwiązywany na podstawie przykładu empirycznego (dane dotyczące silników) z wykorzystaniem transformacji Karhunena-Loevego oraz funkcji klasyfikujących bazujących na wybranych estymatorach nieznanych parametrów wielowymiarowego rozkładu normalnego. Rozważany jest zarówno klasyczny estymator - estymator największej wiarogodności, jak również estymator odporny, który opiera się o funkcje Hubera. Uzyskane klasyfikatory są porównywane za pomocą sprawdzianu krzyżowego - metoda Leave-One-Out.
In this contribution we want to present the concept of uncertainty area of classifiers and an algorithm that uses uninorms to minimize the area of uncertainty in the pre‐ diction of new objects by complex classifiers.
Stroke is the third most common cause of death and the most common cause of long-term disability among adults around theworld. Therefore, stroke prediction and diagnosis is a very important issue. Data mining techniques come in handy to help determine the correlations between individual patient characterisation data, that is, extract from the medical information system the knowledge necessary to predict and treat various diseases. The study analysed the data of patients with stroke using eight known classification algorithms (J48 (C4.5), CART, PART, naive Bayes classifier, Random Forest, Supporting Vector Machine and neural networks Multilayer Perceptron), which allowed to build an exploration model given with an accuracy of over 88%. The potential features of patients, which may be factors that increase the risk of stroke, were also indicated.
Introduction: Software engineering continuously suffers from inadequate software testing. The automated prediction of possibly faulty fragments of source code allows developers to focus development efforts on fault-prone fragments first. Fault prediction has been a topic of many studies concentrating on C/C++ and Java programs, with little focus on such programming languages as Python. Objectives: In this study the authors want to verify whether the type of approach used in former fault prediction studies can be applied to Python. More precisely, the primary objective is conducting preliminary research using simple methods that would support (or contradict) the expectation that predicting faults in Python programs is also feasible. The secondary objective is establishing grounds for more thorough future research and publications, provided promising results are obtained during the preliminary research. Methods: It has been demonstrated that using machine learning techniques, it is possible to predict faults for C/C++ and Java projects with recall 0.71 and false positive rate 0.25. A similar approach was applied in order to find out if promising results can be obtained for Python projects. The working hypothesis is that choosing Python as a programming language does not significantly alter those results. A preliminary study is conducted and a basic machine learning technique is applied to a few sample Python projects. If these efforts succeed, it will indicate that the selected approach is worth pursuing as it is possible to obtain for Python results similar to the ones obtained for C/C++ and Java. However, if these efforts fail, it will indicate that the selected approach was not appropriate for the selected group of Python projects. Results: The research demonstrates experimental evidence that fault-prediction methods similar to those developed for C/C++ and Java programs can be successfully applied to Python programs, achieving recall up to 0.64 with false positive rate 0.23 (mean recall 0.53 with false positive rate 0.24). This indicates that more thorough research in this area is worth conducting. Conclusion: Having obtained promising results using this simple approach, the authors conclude that the research on predicting faults in Python programs using machine learning techniques is worth conducting, natural ways to enhance the future research being: using more sophisticated machine learning techniques, using additional Python-specific features and extended data sets.
15
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Automatic mass or lesion classification systems are developed to aid in distinguishing between malignant and benign lesions present in the breast DCE-MR images, the systems need to improve both the sensitivity and specificity of DCE-MR image interpretation in order to be successful for clinical use. A new classifier (a set of features together with a classification method) based on artificial neural networks trained using artificial fish swarm optimization (AFSO) algorithm is proposed in this paper. The basic idea behind the proposed classifier is to use AFSO algorithm for searching the best combination of synaptic weights for the neural network. An optimal set of features based on the statistical textural features is presented. The investigational outcomes of the proposed suspicious lesion classifier algorithm therefore confirm that the resulting classifier performs better than other such classifiers reported in the literature. Therefore this classifier demonstrates that the improvement in both the sensitivity and specificity are possible through automated image analysis.
16
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Many datasets, especially various historical medical data are incomplete. Various qualities of data can significantly hamper medical diagnosis and are bottlenecks of medical support systems. Nowadays, such systems are often used in medical diagnosis. Even great number of data can be unsuitable when data is imbalanced, missing or corrupted. In some cases these troubles can be overcome by machine learning algorithms designed for predictive modeling. Proposed approach was tested on real medical data and some benchmarks dataset form UCI repository. The liver fibrosis disease from a medical point of view is difficult to treatment and has a significant social and economic impact. Stages of liver fibrosis are diagnosed by clinical observation and evaluations, coupled with a so-called METAVIR rating scale. However, these methods may be insufficient, especially in the recognition of phase of the disease. This paper describes a newly developed algorithm to non-invasive fibrosis stage recognition using machine learning methods – a classification model based on feature projection k-NN classifier. This solution allows extracting data characteristics from the historical data which may be incomplete and may contain imbalance (unequal) sets of patients. Proposed novel solution is based on peripheral blood analysis without using any specialized biomarkers, and can be successfully included to medical diagnosis support systems and might be a powerful tool for effective estimation of liver fibrosis stages.
The aim of this work is to create a web-based system that will assist its users in the cancer diagnosis process by means of automatic classification of cytological images obtained during fine needle aspiration biopsy. This paper contains a description of the study on the quality of the various algorithms used for the segmentation and classification of breast cancer malignancy. The object of the study is to classify the degree of malignancy of breast cancer cases from fine needle aspiration biopsy images into one of the two classes of malignancy, high or intermediate. For that purpose we have compared 3 segmentation methods: k-means, fuzzy c-means and watershed, and based on these segmentations we have constructed a 25–element feature vector. The feature vector was introduced as an input to 8 classifiers and their accuracy was checked. The results show that the highest classification accuracy of 89.02 % was recorded for the multilayer perceptron. Fuzzy c–means proved to be the most accurate segmentation algorithm, but at the same time it is the most computationally intensive among the three studied segmentation methods.
In this paper, the detection of mines or other objects on the seabed from multiple side-scan sonar views is considered. Two frameworks are provided for this kind of classification. The first framework is based upon the Dempster–Shafer (DS) concept of fusion from a single-view kernel-based classifier and the second framework is based upon the concepts of multi-instance classifiers. Moreover, we consider the class imbalance problem which is always presents in sonar image recognition. Our experimental results show that both of the presented frameworks can be used in mine-like object classification and the presented methods for multi-instance class imbalanced problem are also effective in such classification.
19
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Przedstawiono rozwój konstrukcji klasyfikatorów pulsacyjnych typu KOMAG stosowanych do pozyskiwania żwiru i piasku, z jednoczesnym wydzielaniem zanieczyszczeń organicznych i mineralnych. Zamieszczono wyniki badań laboratoryjnych optymalizujących działanie klasyfikatorów. Opisano czynniki procesowe wpływające na zwiększenie skuteczności wzbogacania w zależności od charakterystyki nadawy (kruszywa).
EN
Progress in development of design of KOMAG pulsating jigs used for utilization of gravel and sand together with separation of organic and mineral impurities is presented in the paper. Results of laboratory tests aiming at optimization of classifiers operation are given. Technological factors, which have an impact on increase of beneficiation efficiency depending on feed (aggregates) characteristics, are discussed.
The estimation of the generalization error of a trained classifier by means of a test set is one of the oldest problems in pattern recognition and machine learning. Despite this problem has been addressed for several decades, it seems that the last word has not been written yet, because new proposals continue to appear in the literature. Our objective is to survey and compare old and new techniques, in terms of quality of the estimation, easiness of use, and rigorousness of the approach, so to understand if the new proposals represent an effective improvement on old ones.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.