Optimizing Crypto-Trading Performance: A Comparative Analysis of Innovative Reward Functions in Reinforcement Learning Models
Abstract
1. Introduction
- High-Volatility Periods: Characterized by rapid price fluctuations with intraday volatility exceeding two standard deviations above the rolling 30-day mean.
- Bull Market Phases: Defined by sustained upward price trends with positive 20-day moving average slopes.
- Bear Market Phases: Identified by prolonged downward trends with negative 20-day moving average slopes.
- Recovery Phases: Marked by gradual stabilization following significant drawdowns, with volatility returning toward historical norms.
- 1.
- Economic Utility-Based Reward: Incorporates risk aversion utility theory, with the risk aversion coefficient determined empirically through historical backtesting to balance profit seeking with downside protection.
- 2.
- Market Impact Adjusted Reward: Penalizes trades that adversely affect market prices by incorporating transaction costs proportional to trade size and inversely related to market liquidity, promoting sustainable execution strategies.
- 3.
- Temporal Coherence Reward: Encourages strategic consistency across consecutive timesteps by providing bonuses for maintaining directional conviction, reducing erratic oscillatory trading behavior driven by short-term noise.
- 4.
- Adaptive Risk Control Reward: Dynamically adjusts risk penalties based on recent market volatility, tightening constraints during turbulent periods and relaxing during stable regimes to enable regime-appropriate risk-taking.
- 5.
- Asymmetric Market-Conditional Reward: Applies distinct formulations tailored to prevailing market conditions, rewarding buy actions during bull markets and sell actions during bear markets while remaining neutral during high volatility and recovery phases.
2. Related Works
2.1. Reinforcement Learning for Cryptocurrency Trading
2.2. Reward Function Design for Trading Systems
2.3. Deep Reinforcement Learning Algorithms
3. Methods and Implementation
3.1. Dataset and Market Regime Classification
- Training set: 1 January 2018–31 October 2021 (33,600 hourly observations).
- Validation set: 1 November 2021–31 December 2021 (1464 hourly observations).
- Test set: 1 January 2022–1 October 2022 (6600 hourly observations).
- 1.
- A maximum drawdown exceeding 20% occurred within the previous 480 h (20 days): ;
- 2.
- Current volatility is declining: ;
- 3.
- The trend is stabilizing: .
3.2. Reinforcement Learning Framework
3.2.1. Reinforcement Learning Algorithms
Deep Q-Network (DQN)
Proximal Policy Optimization (PPO)
- is the probability ratio measuring how much the new policy differs from the old policy for action in state .
- is the advantage estimate, representing how much better action is compared to the average action in state . It is computed as , where is the state-value function estimated by a separate critic network.
- constrains the probability ratio to the interval , with in our implementation.
- denotes expectation over timesteps in the collected trajectory batch.
Advantage Actor–Critic (A2C)
3.2.2. Portfolio Accounting and Trading Mechanics
- Buy: If , set (enter long position). If already holding (), no change occurs.
- Sell: If , set (exit to cash). If already in cash (), no change occurs.
- Hold: (maintain current position).
- If (holding Bitcoin): , capturing the price change.
- If (holding cash): , as no exposure exists.
3.3. Novel Reward Functions
4. Experimental Results and Analysis
4.1. Experimental Setup
4.1.1. Training Procedure
4.1.2. Hyperparameter Configuration
4.1.3. Evaluation Metrics
4.1.4. Statistical Testing Methodology
4.2. Overall Performance Comparison
Exposure Analysis and Performance Benchmarking
- 1.
- Selective Market Exposure: The baseline maintains 91.2% exposure during a period when Bitcoin declined 42.3%, resulting in negative returns. In contrast, proposed reward functions achieve 56–68% exposure, spending substantial time in cash to avoid drawdowns. Adaptive Risk Control’s 56.2% exposure directly enables its positive return: avoiding 44% of a bear market translates to sidestepping proportional losses while capturing recovery rallies.
- 2.
- Improved Trade Quality: Proposed rewards achieve 59–68% win rates versus baseline’s 48.6%. This improvement reflects noise reduction through coherence bonuses and strategic timing through adaptive risk penalties. Notably, 73% of Adaptive Risk Control’s entries occur during low-volatility periods ( percentile), demonstrating learned risk-aware positioning.
- 3.
- Oracle Benchmark: A perfect-foresight oracle achieves 184.3% return with 51.3% exposure through 1847 trades. Our best model captures 14.3% of oracle returns using 93.6% fewer trades (118), indicating realistic learned performance rather than overfitting. The similar exposure levels (51.3% vs. 56.2%) confirm that our models learn appropriate market selectivity without clairvoyance.
4.3. Algorithm–Reward Interaction Analysis
4.3.1. Reward Function Universality vs. Algorithm Specificity
4.3.2. Differential Algorithm Sensitivity to Reward Sophistication
4.3.3. Reward-Specific Algorithmic Preferences
4.3.4. Implications for Reward Engineering Practice
4.4. Market-Regime-Specific Performance
4.4.1. Bull Market Performance
4.4.2. Bear Market Performance
4.4.3. High Volatility Performance
4.4.4. Recovery Phase Performance
4.4.5. Cross-Regime Robustness Analysis
- Adaptive Risk Control: (highly variable, but always strong);
- Asymmetric Market-Conditional: (most consistent);
- Economic Utility: (moderate variability);
- Temporal Coherence: (moderate variability);
- Market Impact: (relatively consistent, but lower absolute performance);
- Baseline: (inconsistent and poor).
4.4.6. Practical Implications for Regime-Aware Trading
5. Discussion
5.1. Key Findings and Theoretical Implications
5.2. Practical Deployment Considerations and Limitations
5.3. Broader Impact on Financial Machine Learning
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Nakamoto, S. Bitcoin: A Peer-to-Peer Electronic Cash System. 2008. Available online: https://assets.pubpub.org/d8wct41f/31611263538139.pdf (accessed on 26 October 2025).
- Bariviera, A.F.; Basgall, M.J.; Hasperué, W.; Naiouf, M. Some stylized facts of the Bitcoin market. Phys. A Stat. Mech. Appl. 2017, 484, 82–90. [Google Scholar] [CrossRef]
- Gkillas, K.; Katsiampa, P. An application of extreme value theory to cryptocurrencies. Econ. Lett. 2018, 164, 109–111. [Google Scholar] [CrossRef]
- Liu, Y.; Tsyvinski, A. Risks and returns of cryptocurrency. Rev. Financ. Stud. 2021, 34, 2689–2727. [Google Scholar] [CrossRef]
- Mallqui, D.C.; Fernandes, R.A. Predicting the direction, maximum, minimum and closing prices of daily Bitcoin exchange rate using machine learning techniques. Appl. Soft Comput. 2019, 75, 596–606. [Google Scholar] [CrossRef]
- Sattarov, O.; Muminov, A.; Lee, C.W.; Kang, H.K.; Oh, R.; Ahn, J.; Oh, H.J.; Jeon, H.S. Recommending cryptocurrency trading points with deep reinforcement learning approach. Appl. Sci. 2020, 10, 1506. [Google Scholar] [CrossRef]
- Meng, T.L.; Khushi, M. Reinforcement learning in financial markets. Data 2019, 4, 110. [Google Scholar] [CrossRef]
- Otabek, S.; Choi, J. Multi-level deep Q-networks for Bitcoin trading strategies. Sci. Rep. 2024, 14, 771. [Google Scholar] [CrossRef] [PubMed]
- Millea, A. Deep reinforcement learning for trading-A critical survey. Data 2021, 6, 119. [Google Scholar] [CrossRef]
- Sihananto, A.N.; Sari, A.P.; Prasetyo, M.E.; Fitroni, M.Y.; Gultom, W.N.; Wahanani, H.E. Reinforcement learning for automatic cryptocurrency trading. In Proceedings of the 2022 IEEE 8th Information Technology International Seminar (IT IS), Surabaya, Indonesia, 19–21 October 2022; pp. 345–349. [Google Scholar]
- Sadighian, J. Extending deep reinforcement learning frameworks in cryptocurrency market making. arXiv 2020, arXiv:2004.06985. [Google Scholar] [CrossRef]
- Lucarelli, G.; Borrotti, M. A deep reinforcement learning approach for automated cryptocurrency trading. In Proceedings of the IFIP International Conference on Artificial Intelligence Applications and Innovations, Crete, Greece, 24–26 May 2019; Springer: Berlin/Heidelberg, Germany, 2019; pp. 247–258. [Google Scholar]
- Kumlungmak, K.; Vateekul, P. Multi-Agent Deep Reinforcement Learning With Progressive Negative Reward for Cryptocurrency Trading. IEEE Access 2023, 11, 66440–66455. [Google Scholar] [CrossRef]
- Sadighian, J. Deep reinforcement learning in cryptocurrency market making. arXiv 2019, arXiv:1911.08647. [Google Scholar] [CrossRef]
- Mosavi, A.; Faghan, Y.; Ghamisi, P.; Duan, P.; Ardabili, S.F.; Salwana, E.; Band, S.S. Comprehensive review of deep reinforcement learning methods and applications in economics. Mathematics 2020, 8, 1640. [Google Scholar] [CrossRef]
- Cornalba, F.; Disselkamp, C.; Scassola, D.; Helf, C. Multi-objective reward generalization: Improving performance of Deep Reinforcement Learning for applications in single-asset trading. Neural Comput. Appl. 2024, 36, 619–637. [Google Scholar] [CrossRef]
- Betancourt, C.; Chen, W.H. Reinforcement learning with self-attention networks for cryptocurrency trading. Appl. Sci. 2021, 11, 7377. [Google Scholar] [CrossRef]
- Hambly, B.; Xu, R.; Yang, H. Recent advances in reinforcement learning in finance. Math. Financ. 2023, 33, 437–503. [Google Scholar] [CrossRef]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef]
- Mnih, V.; Badia, A.P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; Kavukcuoglu, K. Asynchronous methods for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning, New York, NY, USA, 20–22 June 2016; pp. 1928–1937. [Google Scholar]
- Moody, J.; Saffell, M. Reinforcement learning for trading. Adv. Neural Inf. Process. Syst. 1998, 11, 917–924. [Google Scholar]
- Moody, J.; Saffell, M. Learning to trade via direct reinforcement. IEEE Trans. Neural Netw. 2001, 12, 875–889. [Google Scholar] [CrossRef] [PubMed]
- Katsiampa, P. Volatility estimation for Bitcoin: A comparison of GARCH models. Econ. Lett. 2017, 158, 3–6. [Google Scholar] [CrossRef]
- Azamjon, M.; Sattarov, O.; Cho, J. Forecasting Bitcoin Volatility through On-Chain and Whale-Alert Tweet Analysis using the Q-Learning Algorithm. IEEE Access 2023, 11, 108092–108103. [Google Scholar] [CrossRef]
- Otabek, S.; Choi, J. Twitter attribute classification with q-learning on bitcoin price prediction. IEEE Access 2022, 10, 96136–96148. [Google Scholar] [CrossRef]
- Kochliaridis, V.; Kouloumpris, E.; Vlahavas, I. Combining deep reinforcement learning with technical analysis and trend monitoring on cryptocurrency markets. Neural Comput. Appl. 2023, 35, 21445–21462. [Google Scholar] [CrossRef]
- Beysolow, T., II. Market making via reinforcement learning. In Applied Reinforcement Learning with Python: With OpenAI Gym, Tensorflow, and Keras; Apress: Berkeley, CA, USA, 2019; pp. 77–94. [Google Scholar]
- Bandarupalli, E. Risk-Aware Deep Reinforcement Learning for Crypto and Equity Trading Under Transaction Costs. Available at SSRN 5662930. Available online: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5662930 (accessed on 26 October 2025).
- Arrow, K.J. The theory of risk-bearing: Small and great risks. J. Risk Uncertain. 1996, 12, 103–111. [Google Scholar] [CrossRef]
- Almgren, R.; Chriss, N. Optimal execution of portfolio transactions. J. Risk 2001, 3, 5–40. [Google Scholar] [CrossRef]
- Chakole, J.; Kurhekar, M. Trend following deep Q-Learning strategy for stock trading. Expert Syst. 2020, 37, e12514. [Google Scholar] [CrossRef]
- Azamjon, M.; Sattarov, O.; Na, D. Enhanced Bitcoin Price Direction Forecasting with DQN. IEEE Access 2024, 12, 29093–29112. [Google Scholar] [CrossRef]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep reinforcement learning with double q-learning. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016; Association for the Advancement of Artificial Intelligence: Washington, DC, USA, 2016; Volume 30, pp. 2094–2100. [Google Scholar]
- Yang, H.; Malik, A. Reinforcement Learning Pair Trading: A Dynamic Scaling Approach. J. Risk Financ. Manag. 2024, 17, 555. [Google Scholar] [CrossRef]
- Zhang, Z.; Zohren, S.; Roberts, S. Deep reinforcement learning for trading. arXiv 2019, arXiv:1911.10107. [Google Scholar] [CrossRef]
- Crypto Data Download. Available online: https://www.cryptodatadownload.com/ (accessed on 20 December 2025).
- Standard Score. Available online: https://en.wikipedia.org/wiki/Standard_score (accessed on 20 December 2025).
- Pratt, J.W. Risk aversion in the small and in the large. In Uncertainty in Economics; Academic Press: Cambridge, MA, USA, 1978; pp. 59–79. [Google Scholar]
- Brauneis, A.; Mestel, R. Cryptocurrency-portfolios in a mean-variance framework. Financ. Res. Lett. 2019, 28, 259–264. [Google Scholar] [CrossRef]
- Raffin, A.; Hill, A.; Gleave, A.; Kanervisto, A.; Ernestus, M.; Dormann, N. Stable-baselines3: Reliable reinforcement learning implementations. J. Mach. Learn. Res. 2021, 22, 1–8. [Google Scholar]
- Cumulative Return: Definition, Calculation, and Example. Available online: https://www.investopedia.com/terms/c/cumulativereturn.asp (accessed on 20 December 2025).
- Sharpe Ratio. Available online: https://en.wikipedia.org/wiki/Sharpe_ratio (accessed on 20 December 2025).
- Maximum Drawdown. Available online: https://en.wikipedia.org/wiki/Drawdown_(economics) (accessed on 20 December 2025).
- Cohen, J. Statistical Power Analysis for the Behavioral Sciences, 2nd ed.; Lawrence Erlbaum Associates, Publishers: Hillsdale, NJ, USA, 1988. [Google Scholar]






| Reference | Year | Algorithm(s) | Reward Function | Asset | Market Regimes | Key Limitation |
|---|---|---|---|---|---|---|
| Moody & Saffell [22] | 1998 | Policy Gradient | DSR | Stocks | No | Single metric, no crypto |
| Lucarelli & Borrotti [12] | 2019 | DQN variants | PnL-based | BTC, ETH, LTC | No | Simple reward, no regimes |
| Sadighian [14] | 2019 | PPO, A2C | Positional PnL | BTC (Coinbase) | No | No utility theory |
| Sadighian [11] | 2020 | PPO, A2C | 7 variants | BTC (Bitmex) | No | Fixed dampening, no regimes |
| Beysolow et al. [28] | 2019 | DQN | Asymmetric PnL | Equities | No | Not crypto-specific |
| Kochliaridis et al. [27] | 2023 | DRL + Rules | Novel PnL | 5 cryptos | No | No systematic regimes |
| Kumlungmak & Vateekul [13] | 2023 | Multi-agent | Progressive negative | Crypto | No | No utility/impact modeling |
| Cornalba et al. [16] | 2024 | DQN | Multi-objective | Single stock | No | Single asset, no crypto |
| Bandarupalli [29] | 2024 | PPO | Volatility-sensitive | BTC, ETH, SPY | No | No temporal coherence |
| Yang et al. [35] | 2024 | A2C, PPO, DQN, SAC | Pair trading | BTC-GBP, BTC-EUR | No | Pair trading only |
| Regime | Trend Condition | Volatility Condition | Additional Criteria |
|---|---|---|---|
| Bull Market | - | ||
| Bear Market | - | ||
| High Volatility | Any | - | |
| Recovery | Stabilizing | decreasing | Post-drawdown |
| Regime | Hours | Percentage (%) |
|---|---|---|
| Bear Market | 3214 | 48.7 |
| High Volatility | 1876 | 28.4 |
| Bull Market | 987 | 15.0 |
| Recovery | 523 | 7.9 |
| Total | 6600 | 100.0 |
| Reward Function | Mathematical Form | Key Parameter(s) | Economic Principle | Primary Objective |
|---|---|---|---|---|
| Economic Utility | CARA utility theory | Risk-adjusted returns | ||
| Market Impact | Market microstructure | Cost-aware trading | ||
| Temporal Coherence | Behavioral consistency | Reduce oscillation | ||
| Adaptive Risk | Dynamic risk mgmt. | Regime-adaptive control | ||
| Asymmetric Market | Regime conditioning | Market-aware strategy |
| Hyperparameter | DQN | PPO | A2C |
|---|---|---|---|
| Learning rate | 0.0001 | 0.0003 | 0.0007 |
| Discount factor | 0.99 | 0.99 | 0.99 |
| Batch size | 64 | 256 | 40 (5 steps × 8 workers) |
| Network architecture | [128, 128, 64] | [256, 256, 128] | [128, 128] |
| Optimizer | Adam | Adam | RMSprop |
| Gradient clipping | 10.0 | 0.5 | 0.5 |
| Target network update freq. | 1000 steps | N/A | N/A |
| Replay buffer size | 50,000 | N/A | N/A |
| Exploration () decay | (100 eps) | N/A | N/A |
| PPO clip parameter | N/A | 0.2 | N/A |
| PPO optimization epochs | N/A | 10 | N/A |
| Entropy coefficient | 0.01 | 0.01 | 0.01 |
| Value loss coefficient | N/A | 0.5 | 0.5 |
| Number of parallel workers | N/A | N/A | 8 |
| Algorithm | Reward Function | Cumulative Return (%) | Sharpe Ratio | Max Drawdown (%) |
|---|---|---|---|---|
| DQN | Baseline (Simple PnL) | (3.2) | (0.08) | (4.1) |
| Economic Utility | (2.8) | (0.12) | (3.5) | |
| Market Impact | (3.1) | (0.15) | (4.2) | |
| Temporal Coherence | (2.5) | (0.11) | (2.9) | |
| Adaptive Risk Control | (2.1) | (0.09) | (2.2) | |
| Asymmetric Market-Cond. | (2.7) | (0.13) | (3.1) | |
| PPO | Baseline (Simple PnL) | (2.9) | (0.09) | (3.8) |
| Economic Utility | (2.4) | (0.10) | (2.8) | |
| Market Impact | (2.8) | (0.14) | (3.6) | |
| Temporal Coherence | (2.2) | (0.11) | (2.5) | |
| Adaptive Risk Control | (1.8) | (0.08) | (1.9) | |
| Asymmetric Market-Cond. | (2.3) | (0.12) | (2.7) | |
| A2C | Baseline (Simple PnL) | (3.4) | (0.10) | (4.5) |
| Economic Utility | (2.6) | (0.13) | (3.2) | |
| Market Impact | (3.0) | (0.16) | (3.9) | |
| Temporal Coherence | (2.4) | (0.12) | (3.0) | |
| Adaptive Risk Control | (2.0) | (0.10) | (2.4) | |
| Asymmetric Market-Cond. | (2.5) | (0.14) | (2.8) |
| Reward Function | Exposure (%) | Trades | Win Rate (%) | Sharpe Ratio | Cum. Return (%) |
|---|---|---|---|---|---|
| Baseline (Simple PnL) | 91.2 | 142 | 48.6 | 0.51 | |
| Economic Utility | 64.3 | 87 | 62.1 | 1.62 | |
| Market Impact | 58.7 | 63 | 59.5 | 1.29 | |
| Temporal Coherence | 67.8 | 78 | 63.8 | 1.74 | |
| Adaptive Risk Control | 56.2 | 118 | 67.8 | 2.47 | |
| Asymmetric Market-Cond. | 61.5 | 95 | 65.3 | 2.05 | |
| Oracle (Perfect Foresight) | 51.3 | 1847 | 100.0 | 4.87 | |
| Buy-and-Hold | 100.0 | 0 | - |
| Reward Function | Bull Market | Bear Market | High Volatility | Recovery | Overall |
|---|---|---|---|---|---|
| Baseline (Simple PnL) | 0.83 | 0.15 | 0.68 | 0.51 | |
| Economic Utility | 2.14 | 1.08 | 1.42 | 2.31 | 1.62 |
| Market Impact | 1.87 | 1.24 | 0.98 | 1.89 | 1.29 |
| Temporal Coherence | 2.42 | 1.15 | 1.58 | 2.18 | 1.74 |
| Adaptive Risk Control | 2.38 | 1.87 | 3.21 | 2.27 | 2.47 |
| Asymmetric Market-Cond. | 2.35 | 1.92 | 1.73 | 2.15 | 2.05 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Khujamatov, E.H.; Ismanov, K.; Mallaev, O.U.; Sattarov, O. Optimizing Crypto-Trading Performance: A Comparative Analysis of Innovative Reward Functions in Reinforcement Learning Models. Mathematics 2026, 14, 794. https://doi.org/10.3390/math14050794
Khujamatov EH, Ismanov K, Mallaev OU, Sattarov O. Optimizing Crypto-Trading Performance: A Comparative Analysis of Innovative Reward Functions in Reinforcement Learning Models. Mathematics. 2026; 14(5):794. https://doi.org/10.3390/math14050794
Chicago/Turabian StyleKhujamatov, Ergashevich Halimjon, Kobuljon Ismanov, Oybek Usmankulovich Mallaev, and Otabek Sattarov. 2026. "Optimizing Crypto-Trading Performance: A Comparative Analysis of Innovative Reward Functions in Reinforcement Learning Models" Mathematics 14, no. 5: 794. https://doi.org/10.3390/math14050794
APA StyleKhujamatov, E. H., Ismanov, K., Mallaev, O. U., & Sattarov, O. (2026). Optimizing Crypto-Trading Performance: A Comparative Analysis of Innovative Reward Functions in Reinforcement Learning Models. Mathematics, 14(5), 794. https://doi.org/10.3390/math14050794

