Abstract
Background: Reinforcement learning (RL) has achieved substantial success in automated negotiation. However, training RL-based agents typically requires costly online interaction, which can be unsafe in real-world settings. Furthermore, these agents often struggle when opponents change their strategies or preferences unexpectedly. Method: To address these challenges, we introduce ANAgent, a novel Automated Negotiating Agent. Our approach is two-fold: first, it pre-trains a general negotiation strategy on offline datasets collected by a variety of unknown strategies. Second, it employs a safe and efficient offline-to-online fine-tuning mechanism, allowing the pre-trained strategy to quickly adapt to new opponents' preferences and strategies during deployment. Results: We evaluate ANAgent's performance through extensive experiments against a wide range of state-of-theart negotiating agents. The results demonstrate that ANAgent successfully learns high-performing strategies from offline data that surpass the performance of the strategies that generated the data. Furthermore, it robustly finetunes its strategy online in response to changes in opponent behavior. Beyond traditional average score metrics, a comprehensive empirical game-theoretic analysis confirms the robustness and strategic strength of our agent. Conclusion: This work presents a viable pathway for developing negotiation agents that can safely learn from historical data and adapt efficiently in real-time interactions, moving closer to robust and deployable automated negotiation systems.
| Original language | English |
|---|---|
| Article number | 114374 |
| Number of pages | 25 |
| Journal | Applied Soft Computing |
| Volume | 188 |
| DOIs | |
| Publication status | Published - 1 Feb 2026 |
Keywords
- Automated negotiation
- Deep learning
- E-commerce
- Offline reinforcement learning
- Empirical game theory
Fingerprint
Dive into the research topics of 'Building automated negotiating agent with offline-to-online deep reinforcement learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver