One of the ways to reduce greenhouse gas emissions and other polluting gases caused by ships is to improve their maintenance operations through their life cycle. The maintenance manager usually does not modify the preventive intervals that the equipment manufacturer has designed to reduce the failure. Conditions of use and maintenance often change from design conditions. In these cases, continuing using the manufacturer's preventive intervals can lead to non-optimal management situations. This article proposes a new method to calculate the preventive interval when the hours of failure of the assets are unavailable. Two scenarios were created to test the effectiveness and usefulness of this new method, one without the failure hours and the other with the failure hours corresponding to a bypass valve installed in the engine of a maritime transport surveillance vessel. In an easy and fast way, the proposed method allows the maintenance manager to calculate the preventive interval of equipment that does not have installed an instrument for measuring operating hours installed.
A rockburst is a common engineering geological hazard. In order to predict rockburst potential in kimberlite at an underground diamond mine, a decision tree method was employed. Based on two fundamental premises of rockburst occurrence, σθ, σc, σt, WET are determined as indicators of rockburst, which are also partition attributes of the decision tree. 132 training samples (with 24 incomplete samples) were obtained from real rockburst cases from all over the world to build the decision tree. The decision tree based on 108 complete samples was built with an accuracy of 73% for 15 validation samples while another decision tree based on 132 samples (with 24 groups of incomplete data) shows an accuracy of 93% for validation samples. Hence, the second decision tree was employed for kimberlite burst prediction. 12 samples from lab tests and a numerical model were used as test samples. The results indicate a moderate burst liability which matches real situations at the diamond mind in question.
The use of machine learning methods in the case of incomplete data is an important task in many scientific fields, like medicine, biology, or face recognition. Typically, missing values are substituted with artificial values that are estimated from the known samples, and the classical machine learning algorithms are applied. Although this methodology is very common, it produces less informative data, because artificially generated values are treated in the same way as the known ones. In this paper, we consider a probabilistic representation of missing data, where each vector is identified with a Gaussian probability density function, modeling the uncertainty of absent attributes. This representation allows to construct an analogue of RBF kernel for incomplete data. We show that such a kernel can be successfully used in regression SVM. Experimental results confirm that our approach capture relevant information that is not captured by traditional imputation methods.
4
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
In this paper we present OvaExpert, an intelligent system for ovarian tumor diagnosis. We give an overview of its features and main design assumptions. As a theoretical framework the system uses fuzzy set theory and other soft computing techniques. This makes it possible to handle uncertainty and incompleteness of the data, which is a unique feature of the developed system. The main advantage of OvaExpert is its modular architecture which allows seamless extension of system capabilities. Three diagnostic modules are described, along with examples. The first module is based on aggregation of existing prognostic models for ovarian tumor. The second presents the novel concept of an Interval-Valued Fuzzy Classifier which is able to operate under data incompleteness and uncertainty. The third approach draws from cardinality theory of fuzzy sets and IVFSs and leads to a bipolar result that supports or rejects certain diagnoses.
In this paper we present OvaExpert, an intelligent system for ovarian tumor diagnosis. We give an overview of its features and main design assumptions. As a theoretical framework the system uses fuzzy set theory and other soft computing techniques. This makes it possible to handle uncertainty and incompleteness of the data which is an unique feature of developed system. The main advantage of OvaExpert is its modular architecture which allows seamless extension of system capabilities. Two diagnostic modules are described in the paper along with examples. First module is based on aggregation of existing prognostic models for ovarian tumor. Second, on novel concept of Interval– Valued Fuzzy Classifier which is able to operate under data incompleteness and uncertainty.
Paper describes a novel modification to a well known kNN algorithm, which enables using it for medical data, which often is a class-imbalanced data with randomly missing values. Paper presents the modified algorithm details, experiment setup, results obtained on a cross validated classification of a benchmark database with randomly removed values (missing data) and records (class imbalance), and their comparison with results of the state of the art classification algorithms.
The reliability analysis for industrial maintenance is now increasingly demanded by the industrialists in the world. Indeed, the modern manufacturing facilities are equipped by data acquisition and monitoring system, these systems generates a large volume of data. These data can be used to infer future decisions affecting the health facilities. These data can be used to infer future decisions affecting the state of the exploited equipment. However, in most practical cases the data used in reliability modelling are incomplete or not reliable. In this context, to analyze the reliability of an oil pump, this work proposes to examine and treat the incomplete, incorrect or aberrant data to the reliability modeling of an oil pump. The objective of this paper is to propose a suitable methodology for replacing the incomplete data using a regression method.
PL
Analiza niezawodności w utrzymaniu ruchu jest coraz bardziej wymagana w przemyśle. Nowoczesne urządzenia produkcyjne wyposażone są w systemy gromadzenia i monitorowania. Systemy te generują duże ilości danych. Dane te mogą być wykorzystane do podejmowania przyszłych decyzji mających wpływ na sprawność urządzeń oraz stan eksploatowanego sprzętu. Jednakże w większości praktycznych przypadków dane wykorzystywane w modelowaniu niezawodności są niekompletne lub niewiarygodne. W tym kontekście, aby poddać analizie niezawodność pompy olejowej, w artykule tym proponuje się zbadać i traktować niepełne lub nieprawidłowe dane do modelowania niezawodności pompy olejowej. Celem tego artykułu jest zaproponowanie odpowiedniej metodologii zastępowania niepełnych danych za pomocą metody regresji.
8
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
The missing values are not uncommon in real data sets. The algorithms and methods used for the data analysis of complete data sets cannot always be applied to missing value data. In order to use the existing methods for complete data, the missing value data sets are preprocessed. The other solution to this problem is creation of new algorithms dedicated to missing value data sets. The objective of our research is to compare the preprocessing techniques and specialised algorithms and to find their most advantageous usage.
Real-life data sets sometimes miss some values. The incomplete data needs specialized algorithms or preprocessing that allows the use of the algorithms for complete data. The paper presents a comparison of various techniques for handling incomplete data in the neuro-fuzzy system ANNBFIS. The crucial procedure in the creation of a fuzzy model for the neuro-fuzzy system is the partition of the input domain. The most popular approach (also used in the ANNBFIS) is clustering. The analyzed approaches for clustering incomplete data are: preprocessing (marginalization and imputation) and specialized clustering algorithms (PDS, IFCM, OCS, NPS). The objective of our research is the comparison of the preprocessing techniques and specialized clustering algorithms to find the the most-advantageous technique for handling incomplete data with a neuro-fuzzy system. This approach is also the indirect validation of clustering.
10
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
In this paper we present results of experiments on 166 incomplete data sets using three probabilistic approximations: lower, middle, and upper. Two interpretations of missing attribute values were used: lost and “do not care” conditions. Our main objective was to select the best combination of an approximation and a missing attribute interpretation. We conclude that the best approach depends on the data set. The additional objective of our research was to study the average number of distinct probabilities associated with characteristic sets for all concepts of the data set. This number is much larger for data sets with “do not care” conditions than with data sets with lost values. Therefore, for data sets with “do not care” conditions the number of probabilistic approximations is also larger.
11
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Functional dependency with degree of satisfaction (FDd) is an extended notion in data modeling, and reflects a type of integrity constraints and business rules on attributes, mainly for massive databases, in which incomplete data such as noise, null and imprecision may exist. While existing approaches are considered effective in general, attempts for further improvement in efficiency are deemed meaningful and desirable as far as knowledge discovery is concerned. This paper focuses on discovering (FDd)s as a form of useful semantic knowledge, aiming at providing an enhancement to the FDd mining process in a more efficient manner. In doing so, properties of FDd are in-depth investigated along with a measure for degree of distinctness. Subsequently, a number of optimization strategies are developed for pre-processing, which are then incorporated into the mining process, giving rise to an enhanced approach for mining functional dependency with degree of satisfaction, namely e-MFDD. Finally, data experiments revealed that e-MFDD significantly outperformed the original approach without pre-processing.
Opisano metodę wyznaczania dystrybuanty czasu funkcjonowania przy niepewnych danych. Przedstawiono zasadę metody oraz na przykładzie pokazano sposób obliczeń. Wyniki zilustrowano wykresami.
EN
The paper presents the method to calculate probability distribution of operating time on the basic incomplete observed data. The principles of this methodology are presented. Results are illustrated by the numerical example the practical application.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.