Uppsats

The Effect of Discount Factor in a Deep Q-learning Agent

Kandidat-uppsats

Stockholms universitet/Institutionen för data- och systemvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

The stock market is known for its unpredictability and complexity, due to the sheer number of microeconomic and macroeconomic influences. An AI trading agent could possibly adapt and thrive in these volatile and dynamic environments and, in some cases, find patterns that a human trader missed. Existing research often focuses on applying reinforcement learning agents to indices such as S&P 500 and NASDAQ. The gold market, although less common, presents unique conditions, as it is heavily influenced by macroeconomic factors such as geopolitics and interest rates. This study explores the application of model-free reinforcement learning agents on short-term gold trading. This research uses a Q-learning algorithm based on a Markov Decision Process (MDP) to boost and optimize the trading strategy over a 30-day time frame. The agent is trained on historical data from MetaTrader 5 to interact with the gold market, with the goal of maximizing profits. To evaluate the agent’s efficiency, the Sharpe ratio is used as the primary metric. Other performance metrics include: win rate, expectancy, profit factor, and maximum drawdown. The purpose of this study is to answer the research question: ”How does the discount factor affect the Sharpe Ratio and net cumulative returns of a Deep Q-learning agent optimized for trading Gold (XAUUSD) in a 30-day fixed window?” Results show that our trading agents generally outperform a B&H strategy in all types of testing and the Random Agent strategy in back testing and live testing, but not in the March test. The best exact values for the gamma rate seem to vary depending on market conditions. Under training-like conditions, the 0.6 to 0.9 range demonstrated a tendency toward stronger profits and Sharpe Ratios, though the wide confidence intervals across models limit how much weight can be placed on this finding. During March, a month with falling gold prices, the models that lost the least were 0.5, 0.6, 0.8, and 0.99, a notably different pattern from the back test. In conclusion, these results suggest that the discount factors in the 0.6 to 0.9 range may be a reasonable starting point for this type of agent in bullish or stable market conditions, but the lack of a consistent linear relationship across tests indicates that no single gamma value can be identified as universally optimal.

Information

Lärosäte / institution
Stockholms universitet/Institutionen för data- och systemvetenskap
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.