Extracting tactics learned from self-play in general games

Dennis J. N. J. Soemers; Spyridon Samothrakis; Eric Piette; Matthew Stephenson

doi:10.1016/j.ins.2022.12.080

Extracting tactics learned from self-play in general games

Dennis J. N. J. Soemers^*, Spyridon Samothrakis, Eric Piette, Matthew Stephenson

^*Corresponding author for this work

Research output: Contribution to journal › Article › Academic › peer-review

Abstract

Local, spatial state-action features can be used to effectively train linear policies from self-play in a wide variety of board games. Such policies can play games directly, or be used to bias tree search agents. However, the resulting feature sets can be large, with a significant amount of overlap and redundancies between features. This is a problem for two reasons. Firstly, large feature sets can be computationally expensive, which reduces the playing strength of agents based on them. Secondly, redundancies and correlations between fea -tures impair the ability for humans to analyse, interpret, or understand tactics learned by the policies. We look towards decision trees for their ability to perform feature selection, and serve as interpretable models. Previous work on distilling policies into decision trees uses states as inputs, and distributions over the complete action space as outputs. In con -trast, we propose and evaluate a variety of decision tree types, which take state-action pairs as inputs, and provide various different types of outputs on a per-action basis. An empirical evaluation over 43 different board games is presented, and two of those games are used as case studies where we attempt to interpret the discovered features.(c) 2022 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).

Original language	English
Pages (from-to)	277-298
Number of pages	22
Journal	Information Sciences
Volume	624
DOIs	https://doi.org/10.1016/j.ins.2022.12.080
Publication status	Published - May 2023

Access to Document

10.1016/j.ins.2022.12.080Licence: CC BY

Cite this

@article{389773e78f9f4f599ac8336219c9c755,

title = "Extracting tactics learned from self-play in general games",

abstract = "Local, spatial state-action features can be used to effectively train linear policies from self-play in a wide variety of board games. Such policies can play games directly, or be used to bias tree search agents. However, the resulting feature sets can be large, with a significant amount of overlap and redundancies between features. This is a problem for two reasons. Firstly, large feature sets can be computationally expensive, which reduces the playing strength of agents based on them. Secondly, redundancies and correlations between fea -tures impair the ability for humans to analyse, interpret, or understand tactics learned by the policies. We look towards decision trees for their ability to perform feature selection, and serve as interpretable models. Previous work on distilling policies into decision trees uses states as inputs, and distributions over the complete action space as outputs. In con -trast, we propose and evaluate a variety of decision tree types, which take state-action pairs as inputs, and provide various different types of outputs on a per-action basis. An empirical evaluation over 43 different board games is presented, and two of those games are used as case studies where we attempt to interpret the discovered features.(c) 2022 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).",

author = "Soemers, {Dennis J. N. J.} and Spyridon Samothrakis and Eric Piette and Matthew Stephenson",

note = "Funding Information: This research is funded by the European Research Council as part of the Digital Ludeme Project (ERC Consolidator Grant #771292) led by Cameron Browne at Maastricht University's Department of Advanced Computing Sciences. This work was carried out on the Dutch national e-infrastructure with the support of SURF Cooperative. We wish to thank Cameron Browne and Mark Winands for feedback on earlier drafts of the paper. We thank the anonymous reviewers for helpful feedback on the paper. Funding Information: This research is funded by the European Research Council as part of the Digital Ludeme Project (ERC Consolidator Grant #771292) led by Cameron Browne at Maastricht University{\textquoteright}s Department of Advanced Computing Sciences. This work was carried out on the Dutch national e-infrastructure with the support of SURF Cooperative. We wish to thank Cameron Browne and Mark Winands for feedback on earlier drafts of the paper. We thank the anonymous reviewers for helpful feedback on the paper. Publisher Copyright: {\textcopyright} 2022 The Author(s)",

year = "2023",

month = may,

doi = "10.1016/j.ins.2022.12.080",

language = "English",

volume = "624",

pages = "277--298",

journal = "Information Sciences",

issn = "0020-0255",

publisher = "Elsevier Inc.",

}

TY - JOUR

T1 - Extracting tactics learned from self-play in general games

AU - Soemers, Dennis J. N. J.

AU - Samothrakis, Spyridon

AU - Piette, Eric

AU - Stephenson, Matthew

N1 - Funding Information: This research is funded by the European Research Council as part of the Digital Ludeme Project (ERC Consolidator Grant #771292) led by Cameron Browne at Maastricht University's Department of Advanced Computing Sciences. This work was carried out on the Dutch national e-infrastructure with the support of SURF Cooperative. We wish to thank Cameron Browne and Mark Winands for feedback on earlier drafts of the paper. We thank the anonymous reviewers for helpful feedback on the paper. Funding Information: This research is funded by the European Research Council as part of the Digital Ludeme Project (ERC Consolidator Grant #771292) led by Cameron Browne at Maastricht University’s Department of Advanced Computing Sciences. This work was carried out on the Dutch national e-infrastructure with the support of SURF Cooperative. We wish to thank Cameron Browne and Mark Winands for feedback on earlier drafts of the paper. We thank the anonymous reviewers for helpful feedback on the paper. Publisher Copyright: © 2022 The Author(s)

PY - 2023/5

Y1 - 2023/5

N2 - Local, spatial state-action features can be used to effectively train linear policies from self-play in a wide variety of board games. Such policies can play games directly, or be used to bias tree search agents. However, the resulting feature sets can be large, with a significant amount of overlap and redundancies between features. This is a problem for two reasons. Firstly, large feature sets can be computationally expensive, which reduces the playing strength of agents based on them. Secondly, redundancies and correlations between fea -tures impair the ability for humans to analyse, interpret, or understand tactics learned by the policies. We look towards decision trees for their ability to perform feature selection, and serve as interpretable models. Previous work on distilling policies into decision trees uses states as inputs, and distributions over the complete action space as outputs. In con -trast, we propose and evaluate a variety of decision tree types, which take state-action pairs as inputs, and provide various different types of outputs on a per-action basis. An empirical evaluation over 43 different board games is presented, and two of those games are used as case studies where we attempt to interpret the discovered features.(c) 2022 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).

AB - Local, spatial state-action features can be used to effectively train linear policies from self-play in a wide variety of board games. Such policies can play games directly, or be used to bias tree search agents. However, the resulting feature sets can be large, with a significant amount of overlap and redundancies between features. This is a problem for two reasons. Firstly, large feature sets can be computationally expensive, which reduces the playing strength of agents based on them. Secondly, redundancies and correlations between fea -tures impair the ability for humans to analyse, interpret, or understand tactics learned by the policies. We look towards decision trees for their ability to perform feature selection, and serve as interpretable models. Previous work on distilling policies into decision trees uses states as inputs, and distributions over the complete action space as outputs. In con -trast, we propose and evaluate a variety of decision tree types, which take state-action pairs as inputs, and provide various different types of outputs on a per-action basis. An empirical evaluation over 43 different board games is presented, and two of those games are used as case studies where we attempt to interpret the discovered features.(c) 2022 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).

U2 - 10.1016/j.ins.2022.12.080

DO - 10.1016/j.ins.2022.12.080

M3 - Article

SN - 0020-0255

VL - 624

SP - 277

EP - 298

JO - Information Sciences

JF - Information Sciences

ER -