The article analyses a reinforcement learning method in which the subject of learning is defined. The essence of this method is the selection of activities by a try and fail process and awarding deferred rewards. Theoretical analyses were supplemented by the practical studies, with reference to implementation of the Sarsa( Lambda) algorithm, with replacing eligibility traces and the Epsilon greedy policy.
2
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
A new "adsorption stochastic algorithm" (called ASA) is proposed for solving the unstable linear Fredholm integral equation of the first kind. The developed algorithm was applied for the calculation of the pore size distribution of activated carbons from single adsorption isotherms assuming different forms of the kernel (i.e. Dubinin and Radushkevich (DR) and/or Nguyen and Do (ND)) of a linear Fredholm integral equation of the first kind. The results obtained by ASA are compared with obtained applying, developed by Provencher, the advanced regularization CONTIN algorithm, advanced evolutionary algorithm GABI written by Arabas and modified by Kowalczyk, and simple evolutionary algorithm based on the mutation strategy labeled SASA. Additionally, the ASA results obtained by solving the integral equation with the ND kernel are compared with the results obtained by regularization solution of the integral equation with density functional theory (DFT) local isotherms as a kernel. It is shown that the developed ASA algorithm always provides stable and very similar results to the Tikhonov regularization method. Moreover, the ASA computations obtained for the ND local isotherms as a kernel are very similar to the results obtained by the most sophisticated regularization DFT software.
3
Dostęp do pełnego tekstu na zewnętrznej witrynie WWW
In this article is defined a reinforcement learning method, in which a subject of learning is analyzed. The essence of this method is the selection of activities by a try and fail process and awarding deferred rewards. If an environment is characterized by the Markov property, then step-by-step dynamics will enable forecasting of subsequent conditions and awarding subsequent rewards on the basis of the present known conditions and actions, relatively to the Markov decision making process. The relationship between the present conditions and values and the potential future conditions is defined by the Bellman equation. The article discusses also a method of temporal difference learning, mechanism of eligibility traces, as well as their algorithms TD(0) and TD(Lambda). Theoretical analyses were supplemented by the practical studies, with reference to all implementation of the Sarsa(Lambda) algorithm, with replacing eligibility traces and the Epsilon greedy policy.
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.