Peningkatan Profit Melalui Strategi Diskon Dinamis Menggunakan Reinforcement Learning pada Data Transaksi Ritel Daring
Keywords:
Dynamic Discount Optimization, Reinforcement Learning, Q-Learning, E-commerce, Price ElasticityAbstract
Discount strategies play a critical role in e-commerce profitability, yet most platforms continue to rely on fixed discount rates that fail to account for heterogeneous customer behavior. This study proposes a dynamic discount optimization framework based on the Q-Learning algorithm, in which discount decisions are modeled as a Markov Decision Process (MDP) with a state space constructed from two customer behavioral indicators: transaction revenue and purchase frequency. Each customer transaction is mapped to one of six discrete states formed by combining three revenue categories (Low, Medium, High) and two frequency categories (Low, High), and the agent selects from five discount levels (0%, 5%, 10%, 15%, 20%) to maximize cumulative profit. A price elasticity mechanism is incorporated into the reward function to simulate demand responsiveness to price reductions, enabling the agent to learn non-trivial discount preferences. The framework is evaluated on the Online Retail II dataset comprising 397,885 cleaned transactions. Experimental results show that the proposed method achieves a total profit of 10,586,828.77, outperforming four of five fixed discount baselines with profit improvements of +18.80%, +11.16%, +5.60%, and +1.65% over the 0%, 5%, 10%, and 15% uniform discount scenarios, respectively. The learned policy consistently assigns 15% discounts to Low-frequency customers and 20% discounts to High-frequency customers, demonstrating that purchase frequency is the primary driver of optimal discount differentiation. These findings confirm that RL offers a more targeted and profit-competitive alternative to static promotional strategies
Downloads
References
A. N. Raji, A. O. Olawore, and D. Osahor, “The impact of E-commerce giants on SMEs: Challenges, opportunities, and the fight for survival in the digital economy,” World J. Adv. Res. Rev., vol. 20, no. 2, pp. 1412–1433, 2023.
A. Julian, C. Maheswar, R. Ramyadevi, M. S. M. Dhar, and S. Selvi, “Analysis and estimation of discounts on product prices in e-commerce sites,” in 2024 IEEE International Conference on Computing, Power and Communication Technologies (IC2PCT), 2024, pp. 759–762.
V. Pabalkar, “Price—impact on consumer behavior through e-commerce and future opportunities,” in Pandemic to Endemic, Routledge, 2024, pp. 222–239.
F. Cherbonnier and C. Gollier, “Fixing our public discounting systems,” Annu. Rev. Financ. Econ., vol. 15, no. 1, pp. 147–164, 2023.
H. Li and Zh. Chen, “How do e-commerce platforms and retailers implement discount pricing policies under consumers are strategic?,” PLoS One, vol. 19, no. 5, p. e0296654, 2024.
S. Sharma, P. Singh, A. Malik, A. Kaur, and C. Irawan, “Comprehensive Literature Review on Machine Learning--Driven Dynamic Pricing Strategies in E-Commerce,” in 2025 International Conference on Metaverse and Current Trends in Computing (ICMCTC), 2025, pp. 1–9.
A. H. Lubis, P. Sihombing, E. B. Nababan, and R. F. Rahmat, “Reformulating Q-Learning as a Hybrid Classifier for Predictive Marketing Analytics,” Int. J. Intell. Eng. Syst., vol. 18, no. 11, pp. 29–44, 2025.
A. K. Shakya, G. Pillai, and S. Chakrabarty, “Reinforcement learning algorithms: A brief survey,” Expert Syst. Appl., vol. 231, p. 120495, 2023.
G. Vivar-Estudillo, M. Estudillo-Ayala, J. Beltrán-Hernández, M. Ibarra-Manzano, and C. Lastre-Domínguez, “Industrial Applications of Q-Learning: A Systematic Review,” WSEAS Trans. Inf. Sci. Appl., vol. 22, pp. 525–539, 2025.
A. H. Lubis, E. B. Nababan, P. Sihombing, and R. F. Rahmat, “Comparison of Q-Learning and SARSA in Allocating Marketing Budget,” in 2025 International Conference on Artificial Intelligence and Technological Solutions (ICAITech), 2025, pp. 14–20.
M. Apte, P. Datar, K. Kale, and P. R. Deshmukh, “Dynamic retail pricing via q-learning-a reinforcement learning framework for enhanced revenue management,” in 2025 1st International Conference on AIML-Applications for Engineering & Technology (ICAET), 2025, pp. 1–5.
P. Mullapudi, “A Reinforcement Learning Approach to Dynamic Pricing,” IJSAT-International J. Sci. Technol., vol. 16, no. 4, 2025.
H. MuhamedAle et al., “Dynamic Pricing Algorithm Using Reinforcement Learning in Online Retail Platforms,” in 2025 3rd International Conference on Cyber Resilience (ICCR), 2025, pp. 1–7.
A. Amato and V. Di Lecce, “Data preprocessing impact on machine learning algorithm performance,” Open Comput. Sci., vol. 13, no. 1, p. 20220278, 2023.
Q. W. Khan, “Exploring Markov decision processes: a comprehensive survey of optimization applications and techniques,” Igmin Res., vol. 2, no. 7, pp. 508–517, 2024.
X. Liu, “Dynamic coupon targeting using batch deep reinforcement learning: An application to livestream shopping,” Mark. Sci., vol. 42, no. 4, pp. 637–658, 2023.
S. Mehta and S. S. Sarpal, “Maximizing Privacy in Reinforcement Learning with Federated Approaches,” in 2024 4th International Conference on Intelligent Technologies (CONIT), 2024, pp. 1–5.
T. Tan, H. Xie, and D. Lian, “Adaptive order Q-learning,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, 2024, pp. 4946–4954.
Z. Zhao, “Variants of Bellman equation on reinforcement learning problems,” in 2nd International Conference on Artificial Intelligence, Automation, and High-Performance Computing (AIAHPC 2022), 2022, pp. 470–481.
M. Hu, “Temporal difference learning,” in The Art of Reinforcement Learning: Fundamentals, Mathematics, and Implementations with Python, Springer, 2023, pp. 75–107.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Andre Hasudungan Lubis, Rizki Muliono, Nanda Novita, Susilawati Susilawati, Muhathir Muhathir, Sayuti Rahman

This work is licensed under a Creative Commons Attribution 4.0 International License.
Universitas Harapan Medan






.png)


.png)


.png)

.png)


.png)



