Robust Portfolio Optimization Considering Investor Preferences Using Maenhout Preferences and Reinforcement Learning
Keywords:
Robust Portfolio Optimization, Ambiguity Aversion, Model Uncertainty, Reinforcement Learning, Proximal Policy Optimization, Tehran Stock ExchangeAbstract
This study develops a robust and adaptive portfolio optimization framework that integrates Maenhout’s ambiguity-averse preference structure with Proximal Policy Optimization (PPO) for dynamic asset allocation in the Tehran Stock Exchange. The model is designed to address two core challenges in emerging markets: parameter estimation error and model misspecification. To do so, the reward function combines expected return with explicit penalties for tail risk and downside risk, including Conditional Value-at-Risk (CVaR) and the Sortino ratio, so that the learned policy is not only return-seeking but also resilient under adverse market conditions. Using daily data from actively traded firms over the 2011–2025 period, the empirical evaluation shows that the proposed framework consistently outperforms classical benchmarks in risk-adjusted performance, downside protection, and stability across changing regimes. In particular, the Maenhout–PPO strategy achieves higher Sharpe and Sortino ratios, lower maximum drawdown, and more controlled tail losses than competing allocation rules. The results indicate that classical mean–variance logic is not sufficiently robust for a market characterized by nonstationarity, inflation pressure, liquidity concentration, and abrupt policy shocks. By contrast, the proposed framework offers a more realistic decision architecture for portfolio management under uncertainty, demonstrating that combining robust control principles with reinforcement learning can materially improve investment performance in emerging-market settings.
References
[1] R. C. Klemkosky, R. Bharati, and S. Chen, "Asset Allocation, Market Neutral, and Portfolio Optimization," in Handbook of Financial Econometrics and Statistics: Springer, 2014, pp. 1-28.
[2] T. L. D. Huynh and V. D. Ta, "Portfolio Optimization Under Structural Uncertainty: A Systematic Review," Finance Research Letters, vol. 71, pp. 106-122, 2025.
[3] V. DeMiguel, L. Garlappi, and R. Uppal, "Optimal Versus Naive Diversification: How Inefficient Is the 1/N Portfolio Strategy?," Review of Financial Studies, vol. 22, no. 5, pp. 1915-1953, 2009, doi: 10.1093/rfs/hhm075.
[4] P. J. Maenhout, "Robust Portfolio Rules and Asset Pricing," The Review of Financial Studies, vol. 17, no. 4, pp. 951-983, 2004, doi: 10.1093/rfs/hhh003.
[5] B. Yi, F. G. Viens, B. Law, and H. Li, "Dynamic Portfolio Selection with Systemic Risk: A Robust Optimization Approach," in Innovations in Quantitative Risk Management: Springer, 2015, pp. 327-346.
[6] Y. Cheng and M. Escobar-Anel, "Robust Portfolio Choice Under a Regime-Switching Jump-Diffusion Model with Ambiguity Aversion," Quantitative Finance, vol. 23, no. 7, pp. 1091-1117, 2023.
[7] Y. Cheng and M. Escobar-Anel, "Optimal Consumption and Robust Portfolio Choice for the 3/2 and 4/2 Stochastic Volatility Models," Mathematics, vol. 11, no. 18, pp. 1-28, 2023, doi: 10.3390/math11184020.
[8] Y. Li, Z. Zheng, and S. Zhang, "Machine Learning for Portfolio Optimization: A Survey," Journal of Financial Data Science, vol. 4, no. 2, pp. 1-28, 2022.
[9] P. Panjan and S. Bhattacharya, "Adaptive Portfolio Optimization: A Survey of Machine Learning Approaches," Expert Systems with Applications, vol. 224, pp. 119-142, 2023.
[10] Y. Ye, H. Pei, B. Wang, Q. Yuan, and J. Cao, "Reinforcement-Learning-Based Portfolio Management with Augmented Liquidity Awareness," presented at the Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI), 2020.
[11] X. Li, H. Wang, and X. Chen, "Deep Reinforcement Learning for Dynamic Portfolio Optimization: A Review and New Perspectives," Engineering Applications of Artificial Intelligence, vol. 127, pp. 107-128, 2024.
[12] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, "Proximal Policy Optimization Algorithms," ed: arXiv, 2017.
[13] D. H. Bailey and M. M. Lopez de Prado, "The Sharpe Ratio Efficient Frontier," Journal of Risk, vol. 15, no. 2, pp. 3-44, 2012, doi: 10.21314/JOR.2012.255.
[14] O. Ledoit and M. Wolf, "Robust Performance Hypothesis Testing with the Sharpe Ratio," Journal of Empirical Finance, vol. 15, no. 5, pp. 850-859, 2008, doi: 10.1016/j.jempfin.2008.03.002.
[15] C. Keating and W. F. Shadwick, "A Universal Performance Measure," Journal of Performance Measurement, vol. 6, no. 3, pp. 59-84, 2002.
[16] R. T. Rockafellar and S. Uryasev, "Optimization of Conditional Value-at-Risk," Journal of Risk, vol. 2, no. 3, pp. 21-41, 2000, doi: 10.21314/JOR.2000.038.
[17] S. Nourollahi, A. Jafari, and A. Pakmaram, "The Impact of Investor Sentiment on Portfolio Optimization in the Tehran Stock Exchange and Cryptocurrency Markets," Business, Marketing, and Finance Open, pp. 1-21, 2025.
Downloads
Published
Submitted
Revised
Accepted
Issue
Section
License
Copyright (c) 2025 Mohammad Mehdi Faraghzadeh (Author); Mehrzad Minouei; Azadeh Mehrani (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.