Missing values frequently occur in real-world time series datasets, significantly affecting the precision and reliability of data analysis and machine learning models. This research project aims to explore the types of missing data occurrences in polish air pollution data recordings and examine various imputation methods. The imputation approaches considered range from simple statistical techniques to more complex methods such as regression models, neural networks and LSTM models. The effectiveness of these imputation techniques will be assessed using PM10 and PM2.5 atmospheric pollution data levels. The performance of each method will be evaluated based on accuracy, consistency and its impact on subsequent predictive models. The findings indicate that the LSTM models are the most effective, while the regression and MLP models, though less accurate, offer faster alternatives. Conversely, mean imputation results in the highest error values.
PL
Brakujące wartości często występują w rzeczywistych zbiorach danych szeregów czasowych, znacząco wpływając na precyzję i wiarygodność analiz danych oraz modeli uczenia maszynowego. Niniejszy projekt badawczy ma na celu zbadanie typów braków danych występujących w wybranym zbiorze oraz analizę różnych metod imputacji. Rozważane podejścia do imputacji obejmują zarówno proste techniki statystyczne, jak i bardziej złożone metody, takie jak modele regresyjne, sieci neuronowe oraz modele LSTM. Skuteczność tych technik zostanie oceniona dla danych dotyczących zanieczyszczenia powietrza, ze szczególnym uwzględnieniem poziomów PM10 i PM2.5. Wydajność każdej metody zostanie oceniona pod względem dokładności oraz spójności. Wyniki wskazują, że modele LSTM są najbardziej efektywne, podczas gdy regresja i MLP – choć mniej dokładne – stanowią szybsze alternatywy. Z kolei imputacja wartością średnią prowadzi do największych wartości błędu.
This study aims to examine the impact of green finance instruments on carbon dioxide emission intensity (CDEI), filling a methodological research gap across regions and time-series data. Specifically, it investigates how green finance instruments, namely green credit, green support, and green funds, together with gross domestic product (GDP), affect the CDEI across diverse regions from 2008 to 2021. A generalized additive mixed model (GAMM) was used to analyze panel data from 29 municipalities and provinces in China over this period. These municipalities and provinces were grouped into six administrative regions, allowing the model to capture the nonlinear relationships and interactions that vary across space and time. The results indicate that in Northern, Northeastern, and Northwestern China, GDP is associated with a higher CDEI. In contrast, green credit, green support, and green funds did not significantly reduce the CDEI during the study period. This study contributes to the discussion on the importance of developing region-specific green finance strategies. It proposes policy approaches tailored to local economic conditions to improve the effectiveness of green finance efforts, thereby supporting emission reduction and advancing environmental policies and sustainable business strategies.
Opracowano analizę statystyczną uzysku energii elektrycznej z dwóch małych systemów energetyki fotowoltaicznej: stacjonarnego i nadążnego. Podstawę przeprowadzonych badań stanowiły dane pomiarowe z lat 2019–2024 pozyskane z systemu archiwizacji danych elektrowni hybrydowej Politechniki Białostockiej. Przeprowadzono testy na stacjonarność i podano statystyki wybranych szeregów czasowych. Przeprowadzono dekompozycję szeregów czasowych na składowe za pomocą metody STL oraz wykorzystano funkcje autokorelacji ACF i autokorelacji cząstkowej PACF. Przeprowadzono mały eksperyment prognostyczny metodą SARIMA. Przeanalizowano także korelacje pomiędzy uzyskiem energii i parametrami klimatycznymi.
EN
A statistical analysis of the electricity yield from two small photovoltaic energy systems: stationary and tracking was developed. The study was based on measurement data from 2019 to 2024 that was retrieved from the hybrid power plant’s data archiving system at the Bialystok University of Technology. Stationarity tests were performed and statistics for selected time series were provided. The series were decomposed into components using STL method. The autocorrelation ACF and partial autocorrelation PACF functions were also utilized. The SARIMA method was used in a small forecasting experiment. Additionally, correlations between energy yield and climate parameters were analyzed.
This article will present the concept of creating an analytical model for inference, detection and early warning of detection of measurements, rapid measurement deviations, and the correctness of the process under study in ultrasonic tomography measurements. The research describes an algorithm's operation for inference, detection and early warning of the appearance of a stochastic trend in a time series.
PL
W tym artykule została przedstawiona koncepcja stworzenie modelu analitycznego pozwalającego na wnioskowanie, wykrywanie i wczesne ostrzeganie wykrywania błędnych pomiarów, odchyleń pomiarowych, jak również prawidłowości przebiegu badanego procesu, w pomiarach za pomocą tomografii ultradźwiękowej. Prace badawcze opisują działanie algorytmu wnioskowania, wykrywania i wczesnego ostrzegania o pojawieniu się trendu stochastycznego w szeregu czasowym.
This study focuses on the reconstruction of incomplete and spatially sparse air temperature data for the purpose of estimating evaporation from Lake Most – a large artificial reservoir in the Czech Republic with no natural inflow. The primary objective is to generate daily spatial temperature fields using spatio-temporal kriging and subsequently compute evaporation using a calibrated Hargreaves-Samani (HS) model. We utilize daily data from the years 2020-2022, collected from six low-cost microstations installed around the lake and from a nearby professional meteorological station (Kopisty, operated by the Czech Hydrometeorological Institute). Due to frequent outages, data coverage from the microstations ranges from 5% to 38%. To fill in missing values and estimate temperature over the lake surface, we apply a Gneiting covariance model. All computations are carried out in MATLAB using a in-house implementation. The reconstructed temperature fields exhibit realistic spatial structure and seasonal variability. Based on the interpolated daily mean, maximum, and minimum air temperatures, we compute daily and cumulative evaporation from the lake surface. The results show that even a sparse and unreliable sensor network can yield physically consistent inputs for evaporation estimation when combined with statistical interpolation. The proposed method is readily applicable to other reservoirs under limited measurement conditions and may support hydrological modeling and water balance analysis.
This publication presents a method for forecasting the annual amount of coal extraction for a mine. The forecast of extraction amount was based on data from an actual hard coal mine. It included an analysis of retrospective data on sales quantity, calculation of regression coefficients of the mathematical trend model and extraction amount for the next year. The planned extraction amount was corrected for the most probable forecast error. In addition, a forecast extraction plan for the analyzed mine was provided for individual months, in two variants.
PL
W niniejszej publikacji przedstawiono sposób prognozowania rocznej wielkości wydobycia węgla dla kopalni. Prognozę wielkości wydobycia przeprowadzono w oparciu o dane rzeczywistej kopalni węgla kamiennego. Obejmowała ona analizę danych retrospektywnych dotyczących wielkości sprzedaży, obliczenie współczynników regresji modelu matematycznego trendu oraz wielkość wydobycia na rok przyszły. Planowana wielkość wydobycia została skorygowana o najbardziej prawdopodobny błąd prognozy. Ponadto podano prognozowany plan produkcji dla analizowanej kopalni w odniesieniu na poszczególne miesiące, w dwóch wariantach.
Pumping systems play an important role in agriculture because they provide the necessary level of irrigation needed to increase crop yields. Pump malfunctions result in equipment downtime, reduced efficiency of agricultural production and significant financial losses. Thus, the development of an early fault detection and diagnosis system leveraging sensor analytic, filtering techniques, and machine learning (ML) technologies constitutes a critical applied research challenge. The aim of this research is to develop and validate early fault detection and classification methods for pumping systems using advanced machine learning algorithms and sensor data analysis.
Taking advantage of deep learning (DL) to extract hidden degradation signals from machinery monitoring data has led to significant advancements in predicting equipment's remaining useful life (RUL). However, existing methods that use similarity and adaptive adjacency matrices to construct graphs fail to reflect sensor relationships accurately. This article presents a cross-temporal dynamic graph convolutional network (CTDGCN) for RUL prediction to address this issue. The CTDGCN combines cross-temporal modeling with dynamic spatio-temporal graph construction, collecting multi-sensor time series signals to create dynamic graph embeddings. By constructing a cross-temporal sensor network, temporal and spatial features are extracted to design a decay graph based on temporal distance. This model utilizes decay and cross-temporal pooling layers to aggregate information and capture complicated spatio-temporal dependencies. Studies conducted on two cases indicate that the CTDGCN model significantly outperforms existing models in RUL prediction tasks.
9
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
In the era of Industry 4.0, accurate prediction of industrial process parameters is essential for optimising operations, lowering costs, and enhancing product quality. Traditional statistical methods often struggle to capture the complex temporal dependencies within industrial processes. This study explores the use of Long Short-Term Memory (LSTM), Bidirectional Long Short-Term Memory (BiLSTM), and Q-Network models to predict material quantities in an industrial dataset. The dataset was pre-processed to address missing values and outliers, and the models were evaluated based on Mean Squared Error (MSE), R2, and accuracy. The results show that the LSTM model achieved an MSE of 14.253 and an R2 of 0.700. The BiLSTM model greatly outperformed it, with an MSE of 0.714 and an R2 of 0.985. The Q-Network model produced an MSE of 0.005 and an R2 of 0.992. These findings demonstrate the Q-Network’s superior ability to capture temporal dependencies within the data.
10
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
Mild cognitive impairment (MCI) is recognized as an early stage preceding Alzheimer’s disease. Functional near-infrared spectroscopy (fNIRS) has recently been used to differentiate MCI patients from healthy controls (HCs) by analyzing their hemodynamic responses. This paper proposes a new method that uses the entire time series data from all fNIRS channels, skipping the feature extraction step. It involves a multi-scale convolutional neural network (CNN) integrated with long short-term memory (LSTM) layers to extract spatial and temporal features simultaneously. The study involves 64 participants (37 MCI patients and 27 HCs) performing three mental tasks: N-back, Stroop, and verbal fluency tests (VFT). The algorithm’s performance was assessed using 10-fold cross-validation across oxyhemoglobin (HbO), deoxyhemoglobin (HbR), and total hemoglobin (HbT). The highest classification accuracies were achieved with HbT, reaching 93.22 % for the N-back task, 91.14 % for the Stroop task, and 89.58 % for the VFT. It was found that using all types of hemodynamic signals from all channels provides better results than analyzing the region of interest data, eliminating the need for data segmentation and feature extraction procedures. Additionally, HbR (or HbT) gives better classification accuracy than HbO. The developed method can be implemented online for clinical applications and real-time monitoring of cognitive disorders.
This study employed machine learning techniques to predict time series of diffusion curves, generated in Python with NumPy library. The data was structured as a time series to enable efficient model training and evaluation. Various approaches – statistical, neural, and regression-based – were tested to model the diffusion dynamics. Results revealed notable performance differences: regression models achieved the highest accuracy with the lowest error rates, while time series models like ARIMA and TCN performed worse, likely due to difficulties in capturing the process’s complexity.
The construction sector records a significant number of occupational accidents (A) and near-misses (NM), making it one of the most dangerous in the economy. In recent years, interest in near-miss events has been growing among researchers and practicing engineers, as they are considered precursors to occupational accidents. Based on a review of the literature on the subject and their own experience, the authors of the article conclude that there is a significant gap in research on near misses in the Polish construction industry. The authors believe that such studies are necessary in the context of accident reduction. The purpose of this article is to analyze the time series of near misses and accidents at work. The data used in the study come from the system of registration of hazardous events implemented in one of the Polish construction companies, recorded in 2015-2022. Due to the specific nature of construction work and the circumstances of the event, 8 categories of hazardous events were specified. For each category, a time series was built to inform about the dynamics of the changes taking place. Box plots were developed for random variables representing the time intervals between consecutive events, informing about the statistical characteristics of a given set of events (SHEi ). This research makes it possible to predict the occurrence of specific events over time and to introduce preventive measures in construction practice.
PL
W sektorze budowlanym notuje się znaczną liczbę wypadków przy pracy (A) oraz zdarzeń potencjalnie wypadkowych (NM), co sprawia, że jest on uznawany za jeden z najbardziej niebezpiecznych w gospodarce. W ostatnich latach zainteresowanie zdarzeniami potencjalnie wypadkowymi wśród naukowców i inżynierów praktyków wzrasta, gdyż są one uważane za prekursory wypadków przy pracy. Na podstawie przeprowadzonego przeglądu literatury przedmiotu oraz doświadczeń własnych, autorzy artykułu stwierdzają, że istnieje istotna luka w badaniach dotyczących zdarzeń potencjalnie wypadkowych w polskim budownictwie. Autorzy uważają, że takie badania są niezbędne w kontekście redukcji liczby wypadków. Celem niniejszego artykułu jest analiza szeregów czasowych zdarzeń potencjalnie wypadkowych i wypadków przy pracy. Dane wykorzystane w badaniach pochodzą z systemu rejestracji zdarzeń niebezpiecznych zaimplementowanego w jednej z polskich firm budowlanych, zarejestrowanych w latach 2015-2022. Ze względu na specyfikę robót budowlanych i okoliczności zdarzenia, wyszczególniono 8 kategorii zdarzeń niebezpiecznych. Dla każdej kategorii zbudowano szereg czasowy informujący o dynamice zachodzących zmian. Opracowano wykresy skrzynkowe, dla zmiennych losowych odstępu czasu między kolejnymi zdarzeniami, informujące o charakterystykach statystycznych danego zbioru zdarzeń (SHEi ). Badania te pozwalają na prognozowanie wystąpienia określonych zdarzeń w czasie oraz wprowadzanie środków zapobiegawczych w praktyce budowlanej. Zbiór wszystkich zdarzeń niebezpiecznych (SHE) został podzielony na zbiór zdarzeń potencjalnie wypadkowych (SNM) i zbiór wypadków (SA). Każdy z tych zbiorów został poddany kategoryzacji ze względu na bezpośrednią przyczynę zdarzenia. Podzbiory zostały podzielone na 8 kategorii: SHE1 – uderzenie przedmiotami, SHE2 – najechanie / potrącenie, SHE3 – środowisko pracy, SHE4 – upadek człowieka, SHE5 – elektryczność, SHE6 – pożar / wybuch / odnalezienie niewybuchu, SHE7 – zawalenie / przysypanie / uwięzienie, SHE8 – kontakt z ruchomymi elementami maszyn. Wprowadzono nową zmienną reprezentującą czas między kolejnymi zdarzeniami. Zmienna ta określa liczbę dni pomiędzy datami następujących po sobie zdarzeń w sekwencji. Procedura wyznaczania odstępów czasowych została wykonana dla ośmiu kategorii zdarzeń, w odniesieniu do wypadków przy pracy jak i zdarzeń potencjalnie wypadkowych. Opracowano wykresy skrzynkowe dla zestawów zmiennych losowych utworzonych w każdej kategorii zdarzeń niebezpiecznych. Prezentują one kluczowe statystyki opisujące zmienne, takie jak mediany, odchylenia od nich, wartości średnie, wartości odstające oraz ekstremalne.
This study proposes a novel approach to financial time series classification by transforming numerical stock market data into candlestick chart images and analyzing them using deep convolutional neural networks (CNNs). Unlike traditional methods that rely on raw numeric sequences, our technique leverages image-based representations enriched with technical indicators (e.g., RSI, MACD, trend channels) to detect visual patterns associated with future price movements. The method is applied to daily price data from ten major publicly traded companies. A custom CNN architecture is trained to classify short-term trends (uptrend vs. downtrend) based on 30-day image windows. The model achieves a test accuracy of 92.83%, with F1-scores exceeding 92% for both classes. These results suggest that visual representations can effectively encode temporal and structural information in price data. While promising, the method’s performance may be sensitive to image resolution and labeling heuristics, which are discussed as potential limitations. Overall, this research demonstrates the feasibility and effectiveness of image-based deep learning in financial market forecasting.
The purpose of this paper is analysing the correlation between the magnitude of the annual amplitude of seasonal changes in the coordinate components of GNSS reference stations and the height of the antenna mounting above the ground. For this purpose, the daily coordinate solutions of more than 500 GNSS reference stations that are part of the IGS (International GNSS Service) network were studied due to their distribution across the globe and long operating time, for some stations dating back to the 1990s. To minimize the impact of the tectonic plate movements authors adopted coordinates of reference stations inside each of the 21 tectonic plates. The coordinates in a topocentric reference frame were detrended in accordance with a linear model, with the objective of removing first-order trends. Subsequently, the seasonal yearly functions were calculated for each North, East and Up component. Finally, the amplitude of the seasonal factor for each station was determined. As a result of the analysis, the existence of annual amplitudes of coordinate changes was demonstrated for some of the stations, but no significant correlation between this phenomenon and the height of the GNSS antenna mounting was shown. In the case of the horizontal components, the majority of the station’s time series is characterized by the amplitude of seasonal function does not exceed 2.5–3 mm, and 5 mm for the vertical component.
Purpose: The aim of the article is to use Google Trends data to examine the seasonal pattern in interest in travel insurance. Design/methodology/approach: This article examines the use of Google Trends to study the frequency of queries related to „ubezpieczenie turystyczne” [travel insurance] and „EKUZ” [European Health Insurance Card] in Poland. We applied the Holt-Winters method with multiplicative seasonality for our analysis. The statistical analysis was conducted using Statistica software. Findings: In the context of studying queries related to „ubezpieczenie turystyczne” and „EKUZ”, Google Trends offers insights into seasonal variations and emerging trends. Research limitations/implications: This research has several limitations, notably the reliance on search queries to measure interest levels, which may not accurately represent actual tourist behavior. Furthermore, the relationships and patterns identified were tested solely within a single country, thereby constraining the generalizability of the findings to other regions and cultural contexts. Practical implications: The data indicate seasonality, reflecting the public's heightened interest during peak travel periods. This seasonal pattern is crucial for policymakers and businesses in the travel and insurance sectors, enabling them to tailor their services and marketing strategies accordingly. Social implications: The extensive volume of searches conducted via Google generates trend data, which can be analyzed using Google Trends, a publicly accessible tool that compares the volume of internet search queries across different regions and time periods. Consequently, Google Trends can offer indirect estimates and has the potential to detect specific patterns earlier than traditional systems. Originality/value: This study contributes original insights by utilizing Google Trends data to analyze and understand seasonal patterns and interest level for travel insurance.
Purpose: The aim of the article was to prepare a simulation analysis of artificial neural network and XGBoost algorithm with determining which of the method was characterized by a lower level of forecast errors for time series predictions. Design/methodology/approach: The objective of the article was reached by applying, a simulation study on a sample of 1000 artificially generated time series. The analyzed XGBoost algorithm and the artificial neural network ANN model were intended to prepare forecasts for five periods ahead. These forecasts were compared with the actual implementations of the time series and proposed forecast error measures. Findings: It is possible to use simulated time series to check which of the presented algorithms were characterized by a lower forecast error. The study showed that applying of the artificial neural networks ANN to forecast future observations generated a lower level of MAPE, MAE and RMSE errors than in the case of the XGBoost algorithm. It was found that both methods generate a lower level of forecast error for time series characterized by a high level of mean value, standard deviation and variance, and levels of kurtosis and skewness close to 0. Practical implications: The research results can be used by both investors and enterprises to better adjust their business decisions to changing market prices by using a model with a lower forecast bias. Originality/value: The original contribution of this article is a comprehensive comparison of forecasts generated by the XGBoost and ANN algorithm, along with determining for which types of time series of the algorithms forecast future values with less error. Moreover, due to the use of simulated artificial time series, it was possible to test each algorithm for various market conditions.
In this research, discrete wavelet transform (DWT) is combined with MLR and ANN to develop WMLR and WANN hybrid models, respectively, for the Brahmaputra river (Pancharatna station) flow forecasting. Daily flow data for the period of 10 year were decomposed (up to fifth level) into detailed and approximation coefficients (using Daubechies wavelets db1, db2, db3, db8 and db10) which were fed as input to MLR and ANN to get the predicted discharge values two days, four days, seven days and 14 days ahead. For all lead times, the WMLR-db10 model was found to be superior as compared to WANN-db1, WANN-db2, WANN-db3, WANN-db8, WMLR-db1, WMLR-db2, WMLR-db3, WMLR-db8 and single MLR and ANN models. During testing period, the values of determination coefficient (R2) and RMSE for WMLR-db10 model for two-, four-, seven- and 14-day lead time were found to be, respectively, 0.996 (751.87 m3·s–1), 0.991 (1,174.80 m3·s–1), 0.984 (1,585.02 m3·s–1), and 0.968 (2,196.46 m3·s–1). Also, it was observed that for lower order wavelets (db1, db2, db3) WANN’s performance was better, and for higher order wavelets (db8, db10) WMLR’s performance was better. Correspondingly, it was observed that all hybrid models’ efficiency increased with increase in the decomposition level.
Light Rail Transit (LRT) plays a role in supporting the mobility of the people of a city. However, the increase in LRT use presents challenges, requiring effective solutions to anticipate changes in the number of passengers. This research aims to design and implement a prediction model using the Seasonal Autoregressive Integrated Method Moving Average to anticipate and predict the number of LRT passengers. The prediction results using the parameter model (0,1,1)(0,1,0) obtained a MAPE value of 16.69%, thus, the accuracy level obtained was 83.31%.
PL
Tranzyt koleją lekką (LRT) odgrywa rolę we wspieraniu mobilności mieszkańców miasta. Jednakże wzrost wykorzystania LRT stwarza wyzwania wymagające skutecznych rozwiązań umożliwiających przewidywanie zmian w liczbie pasażerów. Celem badania jest zaprojektowanie i wdrożenie modelu predykcyjnego wykorzystującego sezonową, zintegrowaną metodę autoregresyjną, średnią ruchomą do przewidywania i przewidywania liczby pasażerów LRT. Wyniki predykcji z wykorzystaniem modelu parametrycznego (0,1,1)(0,1,0) uzyskały wartość MAPE na poziomie 16,69%, a zatem uzyskany poziom dokładności wyniósł 83,31%.
W artykule zaprezentowano analizę wybranych aspektów pracy Krajowego Systemu Elektroenergetycznego pod kątem zapotrzebowania mocy. Przedstawiono wyniki analiz i obliczeń z wykorzystaniem programu komputerowego Statistica dla dobowej prognozy zapotrzebowania mocy i rzeczywistego zapotrzebowania mocy, a także zaprezentowano model wyrównywania wykładniczego i predykcji.
EN
The article presents an analysis of selected aspects of the operation of the National Power System in terms of power demand. The results of analysis and calculations using the Statistica computer software for daily power demand forecast and actual power demand are presented, and an exponential equalization and forecasting model is presented.
To improve forecasting accuracy, researchers employed various combination techniques for a long time. When researchers deal with time series data by using dissimilar models, the combined forecasts of these models are expected to be superior. Deriving a weighting scheme performing better than simple but hard−to−beat combining methods has always been challenging. In this study, a new weighting method based on the hybridisation of combining algorithms is proposed. Five popular datasets were utilised to demonstrate the effectiveness of the proposed method in an out-of-sample context. The results indicate that the proposed method leads to more accurate forecasts than other combining techniques used in the study.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.