Enhancing Stable Behavioral Imitation through Adaptive Reward Weighting in TD3-SAC-GAIL

Authors

  • Mehran Ali Department of Computer Science, Gomal University, D.I.K, Pakistan Author
  • Zia Ullah Medical Lab Technology, Sarhad University of Science and Information Technology, Pakistan Author
  • Aliza Ashfaq Mathematical Sciences, Fatima Jinnah Women University, Pakistan Author

DOI:

https://doi.org/10.63544/tvffe878

Keywords:

Imitation Learning, Generative Adversarial Imitation Learning (GAIL), TD3-SAC Hybrid Reinforcement Learning, Adaptive Reward Weighting, Stable Policy Optimization

Abstract

Imitation learning enables reinforcement learning agents to acquire complex behaviors from expert demonstrations, but its performance remains strongly influenced by the quality of expert data and the balance between imitation and exploration during policy optimization. In particular, TD3-SAC-GAIL enhances the exploration capability of Generative Adversarial Imitation Learning (GAIL) by combining deterministic policy smoothing from Twin Delayed Deep Deterministic Policy Gradient (TD3) with entropy-driven exploration from Soft Actor-Critic (SAC). However, the use of fixed reward or exploration weighting can lead to an inappropriate exploration–exploitation balance across different training stages and environments. To address this limitation, this paper proposes an adaptive reward weighting mechanism to enhance imitation learning stability within the TD3-SAC-GAIL framework. The proposed mechanism dynamically adjusts the contribution of exploration and imitation signals according to the current learning condition, encouraging exploration when learning progress is limited while placing greater emphasis on policy quality as the training process becomes stable. The proposed framework is evaluated in four continuous-control environments, namely Half Cheetah, Walker2d, Hopper, and Lunar Lander Continuous. Experimental evaluation considers expert-surpassing performance, training behavior, reward stability, and the effect of adaptive weighting. The results demonstrate the potential of adaptive reward weighting to provide a systematic mechanism for controlling the exploration–imitation trade-off and enhancing the stability and robustness of GAIL-based policy learning while retaining the exploration advantages of the TD3-SAC hybrid framework.

REFERENCES

[1] M. Ayyildiz and Ö. Polat, “ES-SAC: A hybrid evolution strategy and reinforcement learning approach for humanoid locomotion control,” Bulletin of the Polish Academy of Sciences Technical Sciences, vol. 74, no. 4, pp. e158974, 2026.

[2] M. Zare, P. M. Kebria, A. Khosravi, and S. Nahavandi, “A survey of imitation learning: Algorithms, recent developments, and challenges,” IEEE Transactions on Cybernetics, vol. 54, no. 12, pp. 7173–7186, 2024.

[3] T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, and J. Peters, “An algorithmic perspective on imitation learning,” Foundations and Trends in Robotics, vol. 7, nos. 1–2, pp. 1–179, 2018.

[4] U. Imtiaz, “Ghost Signals: Ethical RF adversarial testing of consumer alarm ecosystems,” in Proc. Int. Conf. Data Science, Computation and Security, Cham, Switzerland: Springer Nature Switzerland, Nov. 2025, pp. 240–253.

[5] A. Khan, F. Amin, and U. Imtiaz, “SENTINEL-WHEEL: Entropy-compressed edge intelligence for explainable self-healing cyber defense in connected vehicles,” International Journal of Innovative Research, vol. 4, no. 1, pp. 227–237, 2026.

[6] J. Ho and S. Ermon, “Generative adversarial imitation learning,” in Advances in Neural Information Processing Systems, vol. 29, 2016.

[7] S. B. Amarat, C. Shi, and Y. Wang, “Generative adversarial imitation learning method based on TD3-SAC hybrid algorithm for robot motion control,” in Proc. 2025 8th Int. Conf. Artificial Intelligence and Big Data (ICAIBD), May 2025, pp. 846–852.

[8] H. Zhang, B. Li, J. Huang, C. Song, P. He, and E. Neretin, “A parallel multi-demonstrations generative adversarial imitation learning approach on UAV target tracking decision,” Chinese Journal of Electronics, vol. 34, no. 4, pp. 1185–1198, 2025.

[9] D. Patel, “Time aware intelligence for efficient and resilient control,” 2025.

[10] N. Bunzeck, P. Dayan, R. J. Dolan, and E. Duzel, “A common mechanism for adaptive scaling of reward and novelty,” Human Brain Mapping, vol. 31, no. 9, pp. 1380–1394, 2010.

[11] M. Danaei, M. Akbarpour Shirazi, and A. Sheikh, “An ensemble deep reinforcement learning framework for multi-channel advertising budget optimization: A practical AI approach,” Applied Artificial Intelligence, vol. 40, no. 1, Art. no. 2684155, 2026.

[12] Z. Shang, R. Li, C. Zheng, H. Li, and Y. Cui, “Relative entropy regularized sample-efficient reinforcement learning with continuous actions,” IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 1, pp. 475–485, 2025.

[13] M. I. K. Jabed, M. P. Ahmed, F. M. Tofa, M. F. Islam, C. A. Gomes, and R. M. Sirazy, “Federated intrusion detection for Internet of Medical Things networks: Differential privacy, non-IID robustness, and cross-device generalization,” Journal of Computer Science and Technology Studies, vol. 8, no. 8, pp. 303–315, 2026.

[14] S. Liu, “An evaluation of DDPG, TD3, SAC, and PPO: Deep reinforcement learning algorithms for controlling continuous system,” in Proc. 2023 Int. Conf. Data Science, Advanced Algorithm and Intelligent Computing (DAI 2023), Feb. 2024, pp. 15–24.

[15] M. I. K. Jabed, M. A. Manzoor, F. M. Tofa, and M. H. Khan, “Interpretable ensemble learning approach for breast cancer diagnosis using SHAP-based explainable AI,” Journal of Computer Science and Technology Studies, vol. 8, no. 8, pp. 244–255, 2026.

[16] U. Iqbal and Y. Bhutto, “Digital transformation through artificial intelligence and advance business analytic in American operational management,” Journal of Theoretical and Applied Econometrics, vol. 3, no. 1, pp. 37–50, 2026.

[17] U. Iqbal, S. Bekmez, and F. A. Qurashi, “Operational risk management through machine learning and business intelligence in U.S. businesses,” Spanish Journal of Innovation and Integrity, vol. 54, pp. 239–253, 2026. [Online]. Available: https://www.sjii.es/index.php/journal/article/view/1140

[18] M. A. Rahman, R. K. Devnath, S. B. Niloy, C. M. Mehedi, T. H. Chowdhury, and M. I. K. Jabed, “A stacking ensemble framework for predicting employee turnover: Explainable AI with SHAP,” in Proc. 2025 IEEE 2nd Int. Conf. Computing, Applications and Systems (COMPAS), Oct. 2025, pp. 1–6.

[19] M. I. K. Jabed, M. R. M. Sirazy, S. Mandal, S. A. Akter, A. Hassan, and H. Esa, “Developing AI-based financial forecasting and cybersecurity systems for the U.S. digital economy,” Frontiers in Computer Science and Artificial Intelligence, vol. 5, no. 5, pp. 30–38, 2026.

[20] M. I. K. Jabed, M. Imran, A. A. Khan, M. Mehedi, A. Islam, and R. Pervez, “Explainable machine learning framework for early heart disease detection using SMOTE and SHAP,” Vascular and Endovascular Review, vol. 9, no. 1, pp. 316–324, 2026.

[21] J. Huang, H. Chen, J. Ren, S. Peng, and L. Deng, “A general adaptive dual-level weighting mechanism for remote sensing pansharpening,” in Proc. 2025 IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), Jun. 2025, pp. 7447–7456.

[22] R. M. Kretchmar, P. M. Young, C. W. Anderson, D. C. Hittle, M. L. Anderson, and C. C. Delnero, “Robust reinforcement learning control with static and dynamic stability,” International Journal of Robust and Nonlinear Control, vol. 11, no. 15, pp. 1469–1500, 2001.

[23] U. Iqbal, “AI-powered supplier risk intelligence: Predicting financial and geopolitical supply chain disruptions in U.S. critical industries,” Journal of Engineering and Computational Intelligence Review, vol. 3, no. 2, pp. 173–193, 2025.

[24] F. Bertolotti, “Practical evaluation of DDPG, TD3, and SAC for HVAC control: A comparative study of training methods and deployment strategies,” Ph.D. dissertation, Politecnico di Torino, 2025.

[25] P. Probst, M. N. Wright, and A. L. Boulesteix, “Hyperparameters and tuning strategies for random forest,” Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, vol. 9, no. 3, Art. no. e1301, 2019.

[26] K. Dutta, P. Gupta, and D. Bajaj, “Robo-Net: A novel reinforced walking biped design using an augmented random search approach,” in Proc. 2024 Int. Conf. Augmented Reality, Intelligent Systems, and Industrial Automation (ARIIA), Dec. 2024, pp. 1–9.

[27] S. Dey, P. Dasgupta, and S. Dey, “Safe reinforcement learning through phasic safety-oriented policy optimization,” in Proc. SafeAI@AAAI, 2023.

[28] U. Imtiaz, “Dynamic security certification framework for evolving distributed architectures,” in Proc. Int. Conf. Data Science, Computation and Security, Cham, Switzerland: Springer Nature Switzerland, Nov. 2025, pp. 347–363.

[29] F. Amin, U. Imtiaz, and A. Khan, “FALCON-Guard: A lightweight explainable framework for real-time cyber threat detection and adaptive risk mitigation in intelligent driving networks,” Multidisciplinary Research in Computing Information Systems, vol. 5, no. 12, pp. 1223–1235, 2025.

[30] U. Iqbal, “AI-driven predictive maintenance for U.S. smart manufacturing: Deep learning models for equipment failure prediction and operational resilience,” Journal of Engineering and Computational Intelligence Review, vol. 3, no. 1, 2025.

[31] U. Iqbal, “AI-enhanced network optimization for electric vehicle charging infrastructure expansion in the United States using graph theory and demand analytics,” Journal of Engineering and Computational Intelligence Review, vol. 2, no. 2, pp. 112–129, 2024.

Author Biographies

Downloads

Published

25-08-2026

How to Cite

Enhancing Stable Behavioral Imitation through Adaptive Reward Weighting in TD3-SAC-GAIL. (2026). Journal of Engineering and Computational Intelligence Review, 4(2), 57-74. https://doi.org/10.63544/tvffe878

Share

Similar Articles

11-20 of 57

You may also start an advanced similarity search for this article.