The task of underwater acoustic target recognition has significant value for both military and civilian applications involving complex marine environments. However, traditional methods often struggle to achieve robust performance due to strong noise interference, weak target signals, and insufficient complementarity among the different features. This paper therefore proposes a deep learning recognition framework based on multimodal spectrogram features, in which short-time Fourier transform magnitude spectra, mel-spectrograms, and continuous wavelet transform spectrograms are used as multi-branch inputs, and a heterogeneous convolutional neural network is employed to extract time-frequency feature representations from each modality. Tailored multi-scale convolutional kernels are designed according to the characteristics of the different features, thereby effectively capturing the energy textures, envelope characteristics, and multi-scale structural information. At the fusion stage, a structure-dependent fusion strategy selection method is introduced, which enables the optimal integration of multi-branch features through inverse-variance weighting. Experimental results demonstrate that the proposed method achieves an accuracy of 97.60% in underwater acoustic target recognition tasks, significantly outperforming single-modality and homogeneous multi-branch models, a finding that validates the effectiveness of the proposed heterogeneous multibranch architecture and its adaptive fusion strategy.
Mel spectrograms have been widely applied in music identification, often yielding successful results when combined with well-known pre-trained classification methods such as VGG16, DenseNet121, or ResNet50. However, the acquired performance may still be improved by employing fusion techniques and proposing a dataset consisting of more samples, which generally demonstrate superior results. Thus, a novel approach employing these methods with the formerly pre-trained classifiers has been introduced. The core innovation of our study is feature fusion utilizing Mel spectrograms, spectrograms, scalograms, and Mel-Frequency Cepstral Coefficients plots, created based on audio recordings from the created dataset encompassing Polish national dance music. The adaptive model is suggested as a mechanism adjusting the highly relevant features for Polish national dance music identification. Furthermore, the use of SHapley Additive exPlanations makes it possible to visualize which parts of the input feature maps are crucial to the model fusion decisions. Subsequently, the most prevalent classification metrics were employed including accuracy, precision, recall, and F1-score to compare the obtained results with state-of-the-art. Hence, the present method yields highly satisfactory results, exceeding 94% accuracy. Consequently, this study not only sets a new benchmark for Polish national dance recognition but also underscores the broader potential of multi-representation fusion as a general blueprint for next-generation audio classification systems.
Due to the lack of a standardised method for determining representative locations for measuring points, it is difficult to select sensitive data on turbocharger rotor faults. In addition, the uncertainty in the feature parameters used for diagnosis under variable rotational speeds leads to low accuracy in fault identification. To address these issues, vibration signals from a turbocharger rotor under various conditions are obtained in this study via a fault simulation test, and a fault diagnosis method for rotor faults under variable speed conditions is proposed based on a sensitivity and multi-dimensional feature analysis of measurement points. The sensitivity of the improved and traditional information entropy is evaluated using a variance analysis across the measurement points, and the most effective vibration measurement points are selected based on the improved information entropy. The effective characteristic parameters of the vibration signals at multiple measurement points are analysed and extracted, and the numer of dimensions of the feature parameters is reduced using the t-distributed stochastic neighbour embedding (t-SNE) method. Faults in the turbocharger rotor at the different speeds are classified using a one-dimensional convolutional neural network (1DCNN), and the arithmetic ability of the diagnostic algorithm is evaluated. The results demonstrate that the proposed methods of selecting measurement points and fault diagnosis can effectively identify rotor faults at different degrees and various speeds: the accuracy of fault diagnosis is 99.85%, and the arithmetic ability is markedly enhanced compared with that of traditional methods.
Accurately estimating spatio-temporal gait parameters such as stride height, stride length, stance time, swing time, and stride speed, is crucial for sports medicine and preventive healthcare. To enable users to measure their spatio-temporal gait parameters in real-life scenarios, several existing studies propose to install one inertial measurement unit (IMU) in each shoe, and design methods to estimate these gait parameters according to the readings of IMUs. Therefore, this paper proposes a novel ensemble method, NEST (standing for Novel Ensemble method for Spatio-Temporal gait parameters measurement), for the multi-task measurement of the aforementioned five spatio-temporal gait parameters. NEST consists of a K-Nearest Neighbor (KNN) regressor branch and a deep learning branch. The KNN regressor branch provides initial estimates, allowing other neural networks to learn to reduce the residual between these estimates and the ground truths. This helps NEST rapidly identify a good optimization direction during the early stage of finetuning and expedite convergence speed. The deep learning branch facilitates information sharing among multiple task-specific representations through fully-connected layers, effectively preserving the interdependencies among gait parameters. Several experiments are conducted to evaluate the performance of NEST and other prior methods. Compared to prior handcrafted-statistics-based methods, NEST demonstrates over 65.1% improvement in RMSE (Root-Mean-Square Error) when predicting spatial parameters.
Autism spectrum disorder includes symptoms like anxiety, depressive disorders, and epilepsy because of its impact on relationships, learning, and employment. Since no confirmed treatment and diagnosis are available, the emphasis is on improving an individual’s capacities through symptom mitigation. This work investigates autism screening for adults and toddlers utilizing deep learning. We investigated models for feature prediction and fused these predictions with the original dataset to be trained with deep long short-term memory (DLSTM). Features are fused from the training and testing sets and then combined with the original dataset. Data analysis is carried out to detect anomalies and outliers, and a label encoding technique is utilized to convert the categorical data into numerical values. We hyper-tuned the DLSTM model parameters to optimize and assess significant outcomes. Experimental analysis and results revealed that the proposed approach worked better than the other techniques, yielding 99.9% accuracy for toddlers and 99% for adults.
6
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Brain cancer, one of the leading causes of mortality worldwide, is caused by brain tumors. Early diagnosis of tumors and predicting their progression can help doctors to save lives. In this article, we have designed an automated approach for locating and classifying tumors from MRI images. The novelties of the research work include the following two stages: Developing an encoder-decoder type 20-Layered deep neural network (DNN) named MultiTumor Analyzer (MTA-20) with 15 down-sampling layers and 4 up-sampling layers, the segmentation is performed in the initial stage. Here, we have adhered a Leaky ReLU activation function instead of ReLU which learn a parameter with negative values that may have valuable information which is essential specifically for image segmentation. Further, a 55-layered DNN using multistage feature fusion is developed in the second stage of the work for the classification of localized tumors. The classification is performed using developed MultiTumor Analyzer (MTA-55) DNN with Softmax classifier. The efficacy of the designed network is validated using highly cited quantitative measures such as accuracy, sensitivity, specificity, dice similarity coefficient (DSC), precision, and F1-measure. It is observed that the proposed MTA-20 DNN attains the average accuracy, sensitivity, specificity, DSC, and precision of 99.2 %, 94.6 %, 99.3 %, 88 %, and 82.5 % respectively against seven state-of-the-art techniques. Also, it is found that, the proposed MTA-55 DNN provides the overall accuracy, recall, specificity, F1-measure, precision, and DSC of 99.8 %, 99.633 %, 99.844 %, 99.659 %, 99.689 %, and 99.656 % respectively as compared to thirteen state-of-the-art techniques. These results corroborate the superiority of the proposed technique.
7
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Breast cancer is a prevalent malignant tumour with high global incidence. Its diagnosis relies primarily on the analysis of pathological breast images. Owing to the complex organisation of the tumour microenvironment, neural network models are essential as efficient classification tools in the field of pathological image analysis. This study introduced spatially-aware attention swift parallel convolution network (SPA-SPCNet), a lightweight and low-latency model for classifying breast pathologies. A novel module for multi-scale feature extraction was constructed using a depthwise separable convolution method. It focuses on the multi-scale features of pathological images to alleviate recognition problems caused by similar local features in breast cancer tissues. The module concatenates the convolutions of different kernels from three branches. Second, a lightweight dynamic spatially-aware attention module was introduced to integrate the visual graph convolutional architecture in a branch. This allowed the model to capture the spatial structure and relationships in image, enabling better handling of the unique spatial distribution relationship between breast cancer tissue structures. The other branch utilises a self-attention mechanism in the transformer. The module can dynamically adjust the attention of the model to different regions in the image, allowing it to focus on the key features of the complex spatial distribution of breast cancer tissue. This feature fusion method enabled the model to capture both global semantics and local details. Compared with existing lightweight models, the proposed model has advantages in terms of tissue structure classification accuracy, parameter quantity, floating-point operations, and real-time inference speed, providing a powerful tool for computer-aided breast pathological image classification.
The need for an ensemble classifier is driven by better accuracy; reduced overfitting, increased robustness that copes with noisy data and reduced variance of individual models, combining the advantages and overcoming the drawbacks of the individual classifier. A comparison of different classifiers like Support Vector Machine (SVM), XGBoost, Random Forest (RF), Naive Bayes (NB), Convolutional Neural Network (CNN) and proposed Ensemble method used in the classification task was conducted. Among all the classifiers evaluated, CNN was found to be the most accurate having an accuracy rate of 93.7%. This indicates that CNN can identify complex data patterns that are also important for photo recognition and classification tasks. Nonetheless, NB and SVM only achieved medium results with accuracy rates of 82.66% and 85.6% respectively. These could have been due to either the complexity of data being handled or underlying assumptions made. RF and XGBoost demonstrated remarkable performances by employing ensemble learning methods as well as gradient-boosting approaches with accuracies of 83.33% and 90.7% respectively. The Ensemble method presented in this paper outperformed all individual models at an accuracy level of 95.5%, indicating that more than one technique is better when classifying correctly based on various resource allocations across techniques employed thereby improving such outcomes altogether by combining them. These results display the pros and cons of every classifier on the Plant Village dataset, giving vital data to improve plant disease classification and guide further research into precision farming and agricultural diagnostics.
In order to achieve accurate identification and segmentation of ore under complex working conditions, machine vision and neural network technology are used to carry out intelligent detection research on ore, an improved Mask RCNN instance segmentation algorithm is proposed. Aiming at the problem of misidentification of stacked ores caused by the loss of deep feature details during the feature extraction process of ore images, an improved Multipath Feature Pyramid Network (MFPN) was proposed. The network firstly adds a single bottom-up feature fusion path, and then adds with the top-down feature fusion path of the original algorithm, which can enrich the deep feature details and strengthen the fusion of the network to the feature layer, and improve the accuracy of the network to the ore recognition. The experimental results show that the algorithm proposed in this paper has a recognition accuracy of 96.5% for ore under complex working conditions, and the recall rate and recall rate function values reach 97.4% and 97.0% respectively, and the AP75 value is 6.84% higher than the original algorithm. The detection results of the ore in the actual scene show that the mask size segmented by the network is close to the actual size of the ore, indicating that the improved network model proposed in this paper has achieved a good performance in the detection of ore under different illumination, pose and background. Therefore, the method proposed in this paper has a good application prospect for stacked ore identification under complex working conditions.
PL
Aby uzyskać dokładną identyfikację i segmentację rudy w złożonych warunkach pracy, do prowadzenia inteligentnych badań wykrywania rudy wykorzystywane są technologie wizji maszynowej i sieci neuronowych, zaproponowano udoskonalony algorytm segmentacji obrazu Mask RCNN (Region Convolutional Neural Networks). Mając na celu rozwiązanie problemu błędnej identyfikacji ułożonych rud, spowodowanego utratą głębokich szczegółów cech podczas procesu ekstrakcji cech z obrazów rudy, zaproponowano ulepszoną sieć wielościeżkową piramidy cech MFPN (Multipath Feature Pyramid Network). Sieć najpierw dodaje pojedynczą ścieżkę łączenia funkcji od dołu do góry, a następnie dodaje ścieżkę łączenia funkcji od góry do dołu oryginalnego algorytmu, co może wzbogacić głębokie szczegóły funkcji i wzmocnić połączenie sieci z warstwą funkcji (obiektową) i poprawić dokładność sieci do rozpoznawania rudy. Wyniki eksperymentalne pokazują, że algorytm zaproponowany w niniejszej pracy ma dokładność rozpoznawania na poziomie 96,5% dla rudy w złożonych warunkach pracy, a wartości współczynnika czułości i współczynnika czułości funkcji osiągają odpowiednio 97,4 i 97,0%, a wartość AP75 jest wyższa o 6,84% niż oryginalny algorytm. Wyniki wykrywania rudy w rzeczywistej scenie pokazują, że rozmiar maski podzielonej na segmenty przez sieć jest zbliżony do rzeczywistego rozmiaru rudy, co wskazuje, że ulepszony model sieci zaproponowany w tym artykule osiągnął dobrą efektywność w wykrywaniu rudy przy różnym oświetleniu, ułożeniu i tle. Dlatego zaproponowana w pracy metoda ma dobre perspektywy aplikacyjne do identyfikacji usypanych rud w złożonych warunkach pracy.
10
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Automatic diagnosis of various ophthalmic diseases from ocular medical images is vital to support clinical decisions. Most current methods employ a single imaging modality, especially 2D fundus images. Considering that the diagnosis of ophthalmic diseases can greatly benefit from multiple imaging modalities, this paper further improves the accuracy of diagnosis by effectively utilizing cross-modal data. In this paper, we propose Transformerbased cross-modal multi-contrast network for efficiently fusing color fundus photograph (CFP) and optical coherence tomography (OCT) modality to diagnose ophthalmic diseases. We design multi-contrast learning strategy to extract discriminate features from crossmodal data for diagnosis. Then channel fusion head captures the semantically shared information across different modalities and the similarity features between patients of the same category. Meanwhile, we use a class-balanced training strategy to cope with the situation that medical datasets are usually class-imbalanced. Our method is evaluated on public benchmark datasets for cross-modal ophthalmic disease diagnosis. The experimental results demonstrate that our method outperforms other approaches. The codes and models are available at https://github.com/ecustyy/tcmn.
In detecting cluster targets in ports or near-shore waters, the echo amplitude is seriously disturbed by interface reverberation, which leads to the distortion of the traditional target intensity characteristics, and the appearance of multiple targets in the same or adjacent beam leads to fuzzy feature recognition. Studying and extracting spatial distribution scale and motion features that reflect the information on cluster targets physics can improve the representation accuracy of cluster target characteristics. Based on the highlight model of target acoustic scattering, the target azimuth tendency is accurately estimated by the splitting beam method to fit the spatial geometric scale formed by multiple highlights. The instantaneous frequencies of highlights are extracted from the time-frequency domain, the Doppler shift of the highlights is calculated, and the motion state of the highlights is estimated. Based on the above processing method, target highlights’ orientation, spatial scale and motion characteristics are fused, and the multiple moving highlights of typical formation distribution in the same beam are accurately identified. The features are applied to processing acoustic scattering data of multiple moving unmanned underwater vehicles (UUVs) on a lake. The results show that multiple small moving underwater targets can be effectively recognized according to the highlight scattering characteristics.
Object tracking based on Siamese networks has achieved great success in recent years, but increasingly advanced trackers are also becoming cumbersome, which will severely limit deployment on resource-constrained devices. To solve the above problems, we designed a network with the same or higher tracking performance as other lightweight models based on the SiamFC lightweight tracking model. At the same time, for the problems that the SiamFC tracking network is poor in processing similar semantic information, deformation, illumination change, and scale change, we propose a global attention module and different scale training and testing strategies to solve them. To verify the effectiveness of the proposed algorithm, this paper has done comparative experiments on the ILSVRC, OTB100, VOT2018 datasets. The experimental results show that the method proposed in this paper can significantly improve the performance of the benchmark algorithm.
13
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Scoliosis is a 3D spinal deformation where the spine takes a lateral curvature, forming an angle in the coronal plane. Diagnosis of scoliosis requires periodic detection, and frequent exposure to radiative imaging may cause cancer. A safer and more economical alternative imaging, i.e., 3D ultrasound imaging modality, is being explored. However, unlike other radiative modalities, an ultrasound image is noisy, which often suppresses the image’s useful information. Through this research, a novel hybridized CNN architecture, multi-scale feature fusion Skip-Inception U-Net (SIU-Net), is proposed for a fully automatic bony feature detection, which can be further used to assess the severity of scoliosis safely and automatically. The proposed architecture, SIU-Net, incorporates two novel features into the basic U-Net architecture: (a) an improvised Inception block and (b) newly designed decoder-side dense skip pathways. The proposed model is tested on 109 spine ultrasound image datasets. The architecture is evaluated using the popular (i) Jaccard Index (ii) Dice Coefficient and (iii) Euclidean distance, and compared with (a) the basic U-net segmentation model, (b) a more evolved UNet++ model, and (c) a newly developed MultiResUNet model. The results show that SIU-Net gives the clearest segmentation output, especially in the important regions of interest such as thoracic and lumbar bony features. The method also gives the highest average Jaccard score of 0.781 and Dice score of 0.883 and the lowest histogram Euclidean distance of 0.011 than the other three models. SIU-Net looks promising to meet the objectives of a fully automatic scoliosis detection system.
14
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
The morphological properties of retinal vessels are closely related to the diagnosis of ophthalmic diseases. However, many problems in retinal images, such as complicated directions of vessels and difficult recognition of capillaries, bring challenges to the accurate segmentation of retinal blood vessels. Thus, we propose a new retinal blood vessel segmentation method based on a dual-channel asymmetric convolutional neural network (CNN). First, we construct the thick and thin vessel extraction module based on the morphological differences in retinal vessels. A two-dimensional (2D) Gabor filter is used to perceive the scale characteristics of blood vessels after selecting the direction of blood vessels; thereby, adaptively extracting the thick vessel features characterizing the overall characteristics and the thin vessel features preserving the capillaries from fundus images. Then, considering that the single-channel network is unsuitable for the unified characterization of thick and thin vessels, we develop a dual-channel asymmetric CNN based on the U-Net model. The MainSegment-Net uses the step-by-step connection mode to achieve rapid positioning and segmentation of thick vessels; the FineSegment-Net combines dilated convolution and the skip connection to achieve the fine extraction of thin vessels. Finally, the output of the dual-channel asymmetric CNN is fused and coded to combine the segmentation results of thick and thin vessels. The performance of our method is evaluated and tested by DRIVE and CHASE_DB1. The results show that the accuracy (Acc), sensitivity (SE), and specificity (SP) of our method on the DRIVE database are 0.9630, 0.8745, and 0.9823, respectively. The evaluation indexes Acc, SE, and SP of the CHASE_DB1 database are 0.9694, 0.8916, and 0.9794, respectively. Additionally, our method combines the biological vision mechanism with deep learning to achieve rapid and automatic segmentation of retinal vessels, providing a new idea for diagnosing and analyzing subsequent medical images.
15
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
There is a close correlation between retinal vascular status and physical diseases such as eye lesions. Retinal fundus images are an important basis for diagnosing diseases such as diabetes, glaucoma, hypertension, coronary heart disease, etc. Because the thickness of the retinal blood vessels is different, the minimum diameter is only one or two pixels wide, so obtaining accurate measurement results becomes critical and challenging. In this paper, we propose a new method of retinal blood vessel segmentation that is based on a multi-path convolutional neural network, which can be used for computer-based clinical medical image analysis. First, a low-frequency image characterizing the overall characteristics of the retinal blood vessel image and a high-frequency image characterizing the local detailed features are respectively obtained by using a Gaussian low-pass filter and a Gaussian high-pass filter. Then a feature extraction path is constructed for the characteristics of the low- and high-frequency images, respectively. Finally, according to the response results of the low-frequency feature extraction path and the high-frequency feature extraction path, the whole blood vessel perception and local feature information fusion coding are realized, and the final blood vessel segmentation map is obtained. The performance of this method is evaluated and tested by DRIVE and CHASE_DB1. In the experimental results of the DRIVE database, the evaluation indexes accuracy (Acc), sensitivity (SE), and specificity (SP) are 0.9580, 0.8639, and 0.9665, respectively, and the evaluation indexes Acc, SE, and SP of the CHASE_DB1 database are 0.9601, 0.8778, and 0.9680, respectively. In addition, the method proposed in this paper could effectively suppress noise, ensure continuity after blood vessel segmentation, and provide a feasible new idea for intelligent visual perception of medical images.
16
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
This study presents a computer-aided diagnostic system for hierarchical classification of normal, fatty, and heterogeneous liver ultrasound images using feature fusion techniques. Both spatial and transform domain based features are used in the classification, since they have positive effects on the classification accuracy. After extracting gray level co-occurrence matrix and completed local binary pattern features as spatial domain features and a number of statistical features of 2-D wavelet packet transform sub-images and 2-D Gabor filter banks transformed images as transform domain features, particle swarm optimization algorithm is used to select dominant features of the parallel and serial fused feature spaces. Classification is performed in two steps: First, focal livers are classified from the diffused ones and second, normal livers are distinguished from the fatty ones. For the used database, the maximum classification accuracy of 100% and 98.86% is achieved by serial and parallel feature fusion modes, respectively, using leave-one-out cross validation (LOOCV) method and support vector machine (SVM) classifier.
Biometrics is the science of human recognition by means of their biological, chemical or behavioural traits. These systems are used in many real life applications simply from biometric based attendance system to providing security at a very sophisticated level. A biometric system deals with raw data captured using a sensor and feature template extracted from raw image. One of the challenges being faced by designers of these systems is to secure template data extracted from the biometric modalities of the user and protect the raw images. In order to minimize spoof attacks on biometric systems by unauthorised users one of the solutions is to use multi-biometric systems. Multi-modal biometric system works by using fusion technique to merge feature templates generated from different modalities of the human. In this work, a novel scheme is proposed to secure template during feature fusion level. The scheme is based on union operation of fuzzy relations of templates of modalities during fusion process of multimodal biometric systems. This approach serves dual purpose of feature fusion as well as transformation of templates into a single secured non invertible template. The proposed technique is irreversible, diverse and experimentally tested on a bimodal biometric system comprising of fingerprint and hand geometry. The given scheme results into significant improvement in the performance of the system with lower equal error rate and improvement in genuine acceptance rate.
18
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
In order to improve the recognition accuracy of the unimodal biometric system and to address the problem of the small samples recognition, a multimodal biometric recognition approach based on feature fusion level and curve tensor is proposed in this paper. The curve tensor approach is an extension of the tensor analysis method based on curvelet coefficients space. We use two kinds of biometrics: palmprint recognition and face recognition. All image features are extracted by using the curve tensor algorithm and then the normalized features are combined at the feature fusion level by using several fusion strategies. The k-nearest neighbour (KNN) classifier is used to determine the final biometric classification. The experimental results demonstrate that the proposed approach outperforms the unimodal solution and the proposed nearly Gaussian fusion (NGF) strategy has a better performance than other fusion rules.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.