О книге
Рассмотрены современные и классические алгоритмы одновременного машинного обучения множества агентов, основанные на теории игр, табличных, нейросетевых, эволюционных и роевых технологиях. Представлено последовательное развитие теоретической модели алгоритмов, базирующееся на марковских процессах принятия решений. Реализация алгоритмов выполнена на языке программирования Python с использованием библиотеки глубокого обучения PyTorch. Средой машинного обучения является компьютерная игра StarCraft II с интерфейсом кооперативного мультиагентного обучения SMAC.
Для магистрантов и аспирантов направления подготовки «Информатика и вычислительная техника».
Список литературы
- Девятков В.В. Системы искусственного интеллекта: учеб. пособие для вузов. М.: Изд-во МГТУ им. Н.Э. Баумана, 2001. 352 с.
- Тарасов В.Б. От многоагентных систем к интеллектуальным организациям. М.: Едиториал УРСС, 2002. 353 с.
- Buşoniu L., Babuška R., De Schutter B. Multiagent reinforcement learning: An overview // Innovations in multi-agent systems and applications. Berlin: Springer, 2010. P. 183–221.
- Claus C., Boutilier C. The dynamics of reinforcement learning in cooperative multiagent systems // AAAI/IAAI. 1998. Vol. 2. P. 746–752.
- Greenwald A., Hall K., Serrano R. Correlated Q-learning // ICML. 2003. Vol. 20. No. 1. P. 242.
- Hernandez-Leal P. et al. A survey of learning in multiagent environments: Dealing with non-stationarity // URL: https://arxiv.org/abs/1707.09183 (дата об-ращения 01.11.2020).
- Klein D., Abbeel P. CS 188: Artificial Intelligence // URL: https://inst.eecs.berkeley.edu/~cs188/sp20 (дата обращения 01.11.2020).
- Könönen V. Asymmetric multiagent reinforcement learning // Web Intelligence and Agent Systems: An international journal. 2004. Vol. 2. No. 2. P. 105–121.
- Lauer M., Riedmiller M. An algorithm for distributed reinforcement learning in cooperative multiagent systems // Proceedings of the Seventeenth International Conference on Machine Learning. 2000.
- Laurent G.J. et al. The world of independent learners is not Markovian // International Journal of Knowledge-based and Intelligent Engineering Systems. 2011. Vol. 15. No. 1. P. 55–64.
- Leike J. et al. AI safety gridworlds. URL: https://arxiv.org/abs/1711.09883 (дата обращения 01.11.2020).
- Littman M.L. Value-function reinforcement learning in Markov games // Cognitive systems research. 2001. Vol. 2. No. 1. P. 55–66.
- Matignon L., Laurent G.J., Le Fort-Piat N. Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems // Knowledge Engineering Review. 2012. Vol. 27, No. 1. P. 1–31.
- Pollack M.E., Ringuette M. Introducing the Tileworld: Experimentally evaluating agent architectures // AAAI. 1990. Vol. 90. P. 183–189.
- Schwartz H.M. Multi-agent machine learning: A reinforcement approach. NY: John Wiley & Sons, 2014. 264 p.
- Sutton R.S., Barto A.G. Reinforcement learning: An introduction. Cambridge: MIT press, 2018. 548 p.
- Tan M. Multi-agent reinforcement learning: Independent vs. cooperative agents // Proceedings of the tenth international conference on machine learning. 1993. P. 330–337.
- Tesauro G. Extending Q-learning to general adaptive multiagent systems // Advances in neural information processing systems. 2003. Vol. 16. P. 871–878.
- Tuyls K., Weiss G. Multiagent learning: Basics, challenges, and prospects // AI Magazine. 2012. Vol. 33. No. 3. P. 41–41.
- Verma T., Varakantham P., Lau H.C. Entropy based independent learning in anonymous multi-agent settings // Proceedings of the International Conference on Automated Planning and Scheduling. 2019. Vol. 29. No. 1. P. 655–663.
- Whiteson S. Learning with Opponent — Learning Awareness // Proceedings of the 17th International Conference on Autonomous Agents and Multiagent Systems. 2018.
- Zhao Y. et al. Winning Isn’t Everything: Enhancing Game Development with Intelligent Agents // IEEE Transactions on Games. 2020. Vol. 12. No. 2. P. 199—212.
- Zinkevich M. Online convex programming and generalized infinitesimal gradient ascent // Proceedings of the 20th international conference on machine learning. 2003. P. 928–936.
- Abouheaf M. Optimization and reinforcement learning techniques in multiagent graphical games and economic dispatch: PhD thesis. The University of Texas at Arlington, 2012. 213 p.
- Beck-Courcelle D. et al. Study of Multiple Multiagent Reinforcement Learning Algorithms in Grid Games: MASc thesis. Carleton University, 2013. 109 p.
- Berger U. Brown’s original fictitious play // Journal of Economic Theory. 2007. Vol. 135. No. 1. P. 9.
- Bowling M. Multiagent learning in the presence of agents with limitations: PhD thesis. Carnegie Mellon University, 2003. P. 172.
- Bowling M., Veloso M. Multiagent learning using a variable learning rate // Artificial Intelligence. 2002. Vol. 136. No. 2. P. 215–250.
- Bowling M., Veloso M. Rational and Convergent Learning in Stochastic Games // International joint conference on artificial intelligence. 2001. Vol. 17. No. 1. P. 1021–1026.
- Brau B.C. An Exploration of Multiagent Learning Within the Game of Sheephead: MSc thesis. Minnesota: State University, 2011. 79 p.
- Busoniu L. et al. Multiagent reinforcement learning: A survey // 9th International Conference on Control, Automation, Robotics and Vision. 2006. P. 7.
- Busoniu L. et al. Multiagent Reinforcement Learning: An Overview // Innovations in MultiAgent Systems and Applications. 2010. No. 1. P. 183–221.
- Casgrain P. et al. Deep Q-Learning for Nash Equilibria: Nash-DQN. URL: https://arxiv.org/abs/1904.10554 (дата обращения 01.11.2020).
- Claus C., Boutilier C. The dynamics of reinforcement learning in cooperative multiagent systems // AAAI/IAAI. 1998. Vol. 1. P. 7
- Conitzer V., Sandholm T. AWESOME: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents // Machine Learning. 2007. No. 67. P. 23–43.
- Hernandez-Leal P. et al. A survey of learning in multiagent environments: Dealing with non-stationarity. URL: https://arxiv.org/abs/1707.09183. 2019 (дата обращения 01.11.2020).
- Hu J., Wellman M.P. Multiagent reinforcement learning: Theoretical framework and an algorithm // ICML. 1998. Vol. 98. P. 242–250.
- Hu J., Wellman M.P. Nash Q-Learning for General-Sum Stochastic Games // Journal of machine learning research. 2003. No. 4. P. 1039–1069.
- Kapetanakis S., Kudenko D. Reinforcement Learning of Coordination in Cooperative Multiagent Systems // AAAI/IAAI. 2002. Vol. 1. P. 6.
- Lemke C.E., Howson J.T. Equilibrium Points of Bimatrix Games // Journal of the Society for Industrial and Applied Mathematics. 1964. Vol. 12. No. 2. P. 413–423.
- Littman M.L. Markov games as a framework for multiagent reinforcement learning // Machine learning proceedings. 1994. P. 157–163.
- Littman M.L., Stone P. Implicit Negotiation in Repeated Games // International Workshop on Agent Theories, Architectures, and Languages. 2001. P. 393–404.
- Liu T. et al. Game Theoretic Control of Multiagent Systems // SIAM Journal on Control and Optimization. 2019. Vol. 57. No. 3. P. 1691–1709.
- Lu X. MultiAgent Reinforcement Learning in Games: PhD thesis. Carleton University, 2012. 188 p.
- Nowé A. et al. Game Theory and Multiagent Reinforcement Learning // Reinforcement Learning. 2012. P. 441–470.
- Park Y.J. et al. Multiagent reinforcement learning with approximate model learning for competitive games // PLoS ONE. 2019. Vol. 14. No. 9. P. 21.
- Powers R., Shoham Y. New Criteria and a New Algorithm for Learning in MultiAgent Systems // Advances in Neural Information Processing Systems. 2004. P. 8.
- Qiao H. Multiagent learning with bargaining — a game theoretic approach: PhD thesis. University of Arizona, 2007. P. 112.
- Roberson B. The Colonel Blotto game // Economic Theory. Vol. 29. No. 1. 2006. P. 24.
- Rokhlin D.B. Q-learning in a stochastic stackelberg game between an uninformed leader and a naïve follower // Theory Probability Application. 2019. Vol. 64. No. 1. P. 41–58
- Roughgarden T. et al. Algorithmic game theory // Communications of the ACM. 2010. Vol. 53. No. 7. P. 775
- Schwartz H. MultiAgent Machine Learning: A Reinforcement Approach. NY: John Wiley & Sons, 2014. 264 p.
- Sheppard J.W. Multiagent reinforcement learning in markov games: PhD thesis. The Johns Hopkins University, 1998. 268 p.
- Shwartz A., Makowski M.A. Comparing Policies in Markov Decision Processes: Mandl’s Lemma Revisited // Mathematics of Operations Research. 1990. Vol. 15. No. 1. P. 155–174.
- Singh S.P. et al. Nash Convergence of Gradient Dynamics in General-Sum Games // Uncertainty in artificial intelligence proceedings. 2000. P. 541—548.
- Tadelis S. Game Theory An Introduction. Princeton: Princeton University Press, 2013. 416 p.
- Tuyls K., Weiss G. Multiagent Learning: Basics, Challenges, and Prospects // AI Magazine. 2012. Vol. 33. No. 3. P. 41–52
- Uther W., Veloso M. Adversarial Reinforcement Learning // Proceedings of the AAAI Fall Symposium on Model Directed Autonomous Systems. 1997. P. 22.
- Wei E. Learning to play cooperative games via reinforcement learning: PhD thesis. George Mason University. 2018. 164 p.
- Wheeler Jr., R. Narendra K. Decentralized Learning in Finite Markov Chains // IEEE Transactions on Automatic Control. 1986. Vol. 31. No. 6. P. 519–526
- Xu D. An integrated simulation, learning and game-theoretic framework for supply chain competition: PhD thesis. The University of Arizona, 2014. P. 233.
- Yang Z. et al. A theoretical analysis of deep Q-learning. URL: https:// arxiv.org/abs/1901.00137v3 (дата обращения 01.11.2020).
- Baker B., Kanitscheider I., Markov T., Wu Y. Emergent Tool Use From Multi-Agent Autocurricula. URL: https://arxiv.org/abs/1909.07528 (дата обращения 01.11.2020).
- Cassandra A. A Survey of POMDP Applications // AAAI. 1998. Vol. 1724. P. 1–9.
- Foerster J. Deep multi-agent reinforcement learning DeepMARL: PhD thesis. University of Oxford. 2018. URL: https://ora.ox.ac.uk/objects/uuid:a55621b3-53c0-4e1b-ad1c-92438b57ffa4 (дата обращения 01.11.2020).
- Foerster J., Assael I., Freitas N., Whiteson S. Learning to Communicate with Deep Multi-Agent Reinforcement Learning // NIPS. 2017. Vol. 29. P. 1—13.
- Foerster J., Chen R., Al-Shedivat M., Whiteson S. Learning with Opponent-Learning Awareness. URL: https://arxiv.org/abs/1709.04326 (дата обращения 01.11.2020).
- Foerster J., Farquhar G.‚ Afouras T.‚ Nardelli N. Counterfactual multiagent policy gradients. URL: https://arxiv.org/abs/1705.08926 (дата обращения 01.11.2020).
- Foerster J., Nardelli N., Farquhar G., Afouras T. Stabilising experience replay for deep multiagent reinforcement learning. URL: https://arxiv.org/abs/1702.08887 (дата обращения 01.11.2020)
- Gupta J., Egorov M., Kochenderfer M. Cooperative multiagent control using deep reinforcement learning // International Conference on Autonomous Agents and Multiagent Systems. 2017. P. 66–83.
- Hausknecht M.J. Cooperation and communication in multiagent deep reinforcement learning: PhD thesis. The University of Texas at Austin, 2016. 169 p.
- Hausknecht M., Stone P. Deep Recurrent Q-Learning for Partially Observable MDPs. URL: https://arxiv.org/abs/1507.06527 (дата обращения 01.11.2020).
- He H., Boyd-Graber J., Kwok K., Daume H. Opponent Modeling in Deep Reinforcement Learning. URL: https://arxiv.org/abs/1609.05559 (дата обращения 01.11.2020).
- Hernandez-Leal P., Kartal B., Taylor M. A survey and critique of multiagent deep reinforcement learning // Autonomous Agents and Multi-Agent Systems. 2019. Vol. 33. P. 750–797.
- Hernandez-Leal P., Kartal B., Taylor M. Is multiagent deep reinforcement learning the answer or the question? A brief survey. URL: https://arxiv.org/abs/1810.05587v2 (дата обращения 01.11.2020).
- Hessel M., Modayil J., Hasselt H.V., Schaul T. Rainbow: Combining Improvements in Deep Reinforcement Learning. URL: https://arxiv.org/abs/1710.02298 (дата обращения 01.11.2020).
- Hong Z., Su S., Shann T., Chang Y. A Deep Policy Inference Q-Network for Multi-Agent Systems. URL: https://arxiv.org/abs/1712.07893 (дата обращения 01.11.2020)
- Jaakkola T., Singh S., Jordan M. Reinforcement learning algorithm for partially observable Markov decision problems // NIPS. 1994. Vol. 4. P. 345–352.
- Konda V., Tsitsiklis J. Actor-Critic Algorithms // NIPS. 2000. Vol. 13. P. 1008–1014.
- Lapan M. Deep Reinforcement Learning Hands-On: Apply modern RL methods, with deep Q-networks, value iteration, policy gradients, TRPO, AlphaGo Zero and more. Birmingham: Packt Publishing Ltd, 2020. 827 p.
- Lazaridou A., Peysakhovich A., Baroni M. Multi-Agent Cooperation and the Emergence of (Natural) Language. URL: https://arxiv.org/abs/1612.07182 (дата обращения 01.11.2020)
- Lerer A., Peysakhovich A. Maintaining cooperation in complex social dilemmas using deep reinforcement learning. URL: https://arxiv.org/abs/1707.01068 (дата обращения 01.11.2020).
- Lillicrap T., Hunt J., Pritzel A., Heess N. Continuous control with deep reinforcement learning. URL: https://arxiv.org/abs/1509.02971 (дата обращения 01.11.2020).
- Liu Y., Wang W., Hu Y., Hao J. Multi-Agent Game Abstraction via Graph Attention Neural Network. URL: https://arxiv.org/abs/1911.10715v1 (дата обра-щения 01.11.2020).
- Lowe R., Wu Y., Tamar A., Harb J. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. URL: https://arxiv.org/abs/1706.02275 (дата обращения 01.11.2020).
- Matignon L., Laurent G.J., Le Fort-Piat N. Independent reinforcement learners in cooperative Markov games a survey regarding coordination problems // CUP. 2012. Vol. 27. No. 1. P. 1–31.
- Mnih V., Kavukcuoglu K., Silver D., Graves A. Playing Atari with Deep Reinforcement Learning. URL: https://arxiv.org/abs/1312.5602 (дата обращения 01.11.2020).
- Mnih V., Kavukcuoglu K., Silver D., Rusu A. Human-Level Control Through Deep Reinforcement Learning // Nature. 2015. Vol. 518. P. 529–533.
- Oliehoek F., Amato C. A concise introduction to decentralized POMDPs. Berlin: Springer, 2016. 141 p.
- Omidshafiei S., Pazis J., Amato C., How J. Deep decentralized multi-task multi-agent reinforcement learning under partial observability. URL: https://arxiv.org/abs/1703.06182 (дата обращения 01.11.2020).
- Oroojlooy Jadid A., Hajinezhad D. A Review of Cooperative Multi-Agent Deep Reinforcement Learning. URL: https://arxiv.org/abs/1908.03963 (дата обращения 01.11.2020).
- Palmer G., Tuyls K., Bloembergen D., Savani R. Lenient multi-agent deep reinforcement learning. URL: https://arxiv.org/abs/1707.04402 (дата обращения 01.11.2020).
- Panait L., Sullivan K., Luke S. Lenient learners in cooperative multiagent systems // Proceedings of the fifth international joint conference on Autonomous agents and multiagent systems. 2006. P. 801–803
- Pham H.N.A., Triantaphyllou E. The Impact of Overfitting and Over-generalization on the Classification Accuracy in Data Mining: PhD thesis. Louisiana State University, 2011. 127 p.
- Puigdomènech A., Piot B., Kapturowski S., Sprechmann P. Agent57: Outper-forming the Atari Human Benchmark. URL: https://arxiv.org/abs/2003.13350 (дата обращения 01.11.2020).
- Rashid T., Samvelyan M., Witt C., Farquhar G. QMIX: Monotonic Value Function Factorisation for DeepMultiAgent Reinforcement Learning. URL: https://arxiv.org/abs/1803.11485 (дата обращения 01.11.2020).
- Schmidhuber J., Hochreiter S. Long short-term memory // Neural Pomputation. 1997. Vol. 9. No. 8. P. 1735–1780.
- Schrittwieser J., Antonoglou I., Hubert T., Simonyan K. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model. URL: https://arxiv.org/abs/1911.08265 (дата обращения 01.11.2020).
- Schulman J., Levine S., Abbeel P., Jordan M. Trust region policy optimization // PMLR. 2015. Vol. 37. P. 1889–1897.
- Silver D., Hubert T., Schrittwieser J., Antonoglou I. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play // Science. 2018. Vol. 362. P. 1140–1144.
- Silver D., Lever G., Heess N., Degris T. Deterministic policy gradient algorithms // ICML. 2014. Vol. 32. P. 387–395.
- Silver D., Schrittwieser J., Simonyan K., Antonoglou I. Mastering the game of Go without human knowledge // Nature. 2017. Vol. 550. P. 354–359.
- Son K., Kim D., Kang W.J., Hostallero D.E., Yi Y. QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning. URL: https://arxiv.org/abs/1905.05408 (дата обращения 01.11.2020).
- Sukhbaatar S., Szlam A., Fergus R. Learning Multiagent Communication with Backpropagation. URL: https://arxiv.org/abs/1605.07736 (дата обращения 01.11.2020).
- Sunehag P., Lever G., Gruslys A., Czarnecki W.M. Value-decomposition networks for cooperative multi-agent learning based on team reward // AAMAS. 2018. Vol. 17. P. 2085–2087.
- Tampuu A., Matiisen T., Kodelja D., Kuzovkin I. Multiagent cooperation and competition with deep reinforcement learning // PLoS One. 2017. Vol. 12. No. 4.
- Usunier N., Synnaeve G., Lin Z., Chintala S. Episodic Exploration for Deep Deterministic Policies: An Application to StarCraft Micromanagement Tasks. URL: https://arxiv.org/abs/1609.02993 (дата обращения 01.11.2020).
- 215Sukhbaatar S., Szlam A., Fergus R. Learning Multiagent Communication with Backpropagation. URL: https://arxiv.org/abs/1605.07736 (дата обращения 01.11.2020).Sunehag P., Lever G., Gruslys A., Czarnecki W.M. Value-decomposition networks for cooperative multi-agent learning based on team reward // AAMAS. 2018. Vol. 17. P. 2085–2087.Tampuu A., Matiisen T., Kodelja D., Kuzovkin I. Multiagent cooperation and competition with deep reinforcement learning // PLoS One. 2017. Vol. 12. No. 4.Usunier N., Synnaeve G., Lin Z., Chintala S. Episodic Exploration for Deep Deterministic Policies: An Application to StarCraft Micromanagement Tasks. URL: https://arxiv.org/abs/1609.02993 (дата обращения 01.11.2020).
- Vinyals O., Babuschkin I., Czarnecki W., Mathieu M. Grandmaster level in StarCraft II using multi-agent reinforcement learning // Nature. 2019. Vol. 575. P. 350–354.
- Wang W., Yang T., Liu L., Hao J. From Few to More: Large-Scale Dynamic Multiagent Curriculum Learning. URL: https://arxiv.org/abs/1909.02790v2 (дата обращения 01.11.2020)
- Wei E., Luke S. Lenient learning in independent-learner stochastic cooperative games // The Journal of MLR. 2016. Vol. 17. No. 1.
- Werbos P.J. Backpropagation through time: what it does and how to do it // IEEE. 1990. Vol. 78. No. 10. P. 1550–1560.
- Zheng Y., Meng Z., Hao J., Zhang Z., Yang T. A deep bayesian policy reuse approach against non-stationary agents // NIPS. 2018. Vol. 1. P. 962–972.
- Карпенко А.П. Современные алгоритмы поисковой оптимизации. Ал-горитмы, вдохновленные природой: учебное пособие. М.: Изд-во МГТУ им. Н.Э. Баумана, 2017. 446 с.
- Almufti S., Marqas R., Ashqi V. Taxonomy of bio-inspired optimization algorithms // Journal of Advanced Computer Science & Technology. 2019. Vol. 8. No. 2. P. 23–31.
- Baar W., Bauso D. Networked Bio-Inspired Evolutionary Dynamics on a Multi-Population // 18th European Control Conference. 2019. P. 1023–1028.
- De Jong K. Evolutionary computation: a unified approach // Proceedings of the Genetic and Evolutionary Computation Conference Companion. 2020. P. 327–342.
- Ficici S.G., Pollack J.B. Pareto optimality in coevolutionary learning // European Conference on Artificial Life. Berlin: Springer, 2001. P. 316–325.
- Gill S.S., Buyya R. Bio-inspired algorithms for big data analytics: a survey, taxonomy, and open challenges // Big Data Analytics for Intelligent Healthcare Management. NY: Academic Press, 2019. P. 1–17.
- Gomes J., Mariano P., Christensen A.L. Cooperative coevolution of partially heterogeneous multiagent systems // Proceedings of the International Conference on Autonomous Agents and Multiagent Systems. 2015. P. 297–305.
- Gomes J., Mariano P., Christensen A.L. Dynamic team heterogeneity in cooperative coevolutionary algorithms // IEEE Transactions on Evolutionary Computation. 2017. Vol. 22. No. 6. P. 934–948
- Jacob C. et al. Illustrating evolutionary computation with Mathematica. Voltem: Morgan Kaufmann, 2001. 547 p.
- Jaderberg M. et al. Population based training of neural networks. URL: https://arxiv.org/abs/1711.09846 (дата обращения 01.11.2020).
- Khalid M.A., Yusof U., Aman K.K. A survey on bio-inspired multi-agent system for versatile manufacturing assembly line // ICIC Express Letters. 2016. Vol. 10. No. 1. P. 1–7.
- Knudson M., Tumer K. Coevolution of heterogeneous multi-robot teams // Proceedings of the 12th annual conference on Genetic and evolutionary computation. 2010. P. 127–134.
- Leibo J.Z. et al. Malthusian reinforcement learning. URL: https://arxiv.org/abs/1812.07019 (дата обращения 01.11.2020).
- Levin S.A., Udovic J.D. A mathematical model of coevolving populations // The American Naturalist. 1977. Vol. 111. No. 980. P. 657–675.
- Li Z., Liu J. A multi-agent genetic algorithm for community detection in complex networks // Physica A: Statistical Mechanics and its Applications. 2016. Vol. 449. P. 336–347.
- Liekens A.M.L., ten Eikelder H.M.M., Hilbers P.A.J. Finite population models of co-evolution and their application to haploidy versus diploidy // Genetic and Evolutionary Computation Conference. Berlin: Springer, 2003. P. 344–355.
- Liu S. et al. Emergent coordination through competition. URL: https:// arxiv.org/abs/1902.07151(дата обращения 01.11.2020).
- Miikkulainen R. et al. Multiagent learning through neuroevolution // IEEE World Congress on Computational Intelligence. Berlin: Springer, 2012. P. 24–46.
- Moriarty D.E., Miikkulainen R. Forming neural networks through efficient and adaptive coevolution // Evolutionary computation. 1997. Vol. 5. No. 4. P. 373–399
- Omidshafiei S. et al. α-rank: Multi-agent evaluation by evolution // Scientific reports. 2019. Vol. 9. No. 1. P. 1–29.
- Panait L., Luke S. Cooperative multi-agent learning: The state of the art // Autonomous agents and multi-agent systems. 2005. Vol. 11. No. 3. P. 387–434.
- Panait L., Sullivan K., Luke S. Lenience towards teammates helps in cooperative multiagent learning // Proceedings of the Fifth International Joint Conference on Autonomous Agents and Multi Agent Systems. 2006. P. 1–10
- Paredis J. Coevolutionary computation // Artificial life. 1995. Vol. 2. No. 4. P. 355–375.
- Peng Z., Wu J., Chen J. Three-dimensional multi-constraint route planning of unmanned aerial vehicle low-altitude penetration based on coevolutionary multi-agent genetic algorithm // Journal of Central South University of Technology. 2011. Vol. 18. No. 5. P. 1502
- Rockefeller G., Khadka S., Tumer K. Multi-level Fitness Critics for Cooperative Coevolution // Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems. 2020. P. 1143–1151.
- Rossi F. et al. Review of multi-agent algorithms for collective behavior: a structural taxonomy // IFAC-PapersOnLine. 2018. Vol. 51. No. 12. P. 112–117.
- Schmitt L.M. Theory of coevolutionary genetic algorithms // International Symposium on Parallel and Distributed Processing and Applications. Berlin: Springer, 2003. P. 285–293.
- Schmitt L.M. Theory of genetic algorithms // Theoretical Computer Science. 2001. Vol. 259. No. 1–2. P. 1–61
- Seredynski F. Coevolutionary multi-agent systems: the application to mapping and scheduling problems // Proceedings of the IEEE International Conference on Industrial Technology. 1996. P. 431–435.
- Seredynski F. Competitive coevolutionary multi-agent systems: The application to mapping and scheduling problems // Journal of Parallel and Distributed Computing. 1997. Vol. 47. No. 1. P. 39–57.
- Singh S. et al. Intrinsically motivated reinforcement learning: An evolutionary perspective // IEEE Transactions on Autonomous Mental Development. 2010. Vol. 2. No. 2. P. 70–82.
- Singh S., Lewis R.L., Barto A.G. Where do rewards come from // Proceedings of the annual conference of the cognitive science society. 2009. P. 2601–2606
- Srivastava V., Leonard N.E. Bio-inspired decision-making and control: From honeybees and neurons to network design // American Control Conference. 2017. P. 2026–2039.
- Stella L. Bio-Inspired Collective Decision-Making in Game Theoretic Models and Multi-Agent Systems: PhD thesis. University of Sheffield, 2019. 141 p.
- Stone P., Veloso M. Multiagent systems: A survey from a machine learning perspective // Autonomous Robots. 2000. Vol. 8. No. 3. P. 345–383
- Uriot T., Izzo D. Safe Crossover of Neural Networks Through Neuron Alignment. URL: https://arxiv.org/abs/2003.10306 (дата обращения 01.11.2020).
- Wang Y., Qi Y., Li Y. Memory-based multiagent coevolution modeling for robust moving object tracking // The Scientific World Journal. 2013. Vol. 2013. P. 1—13.
- Whiteson S., Stone P. Evolutionary function approximation for reinforcement learning // Journal of Machine Learning Research. 2006. Vol. 7. P. 877–917.
- Wiegand R.P., Liles W.C., De Jong K.A. Modeling Variation in Cooperative Coevolution Using Evolutionary Game Theory // 7th International Workshop Foundations of Genetic Algorithms. 2002. P. 203–220.
- Yang Z., Tang K., Yao X. Large scale evolutionary optimization using cooperative coevolution // Information sciences. 2008. Vol. 178. No. 15. P. 2985–2999.
- Yong C.H., Miikkulainen R. Coevolution of role-based cooperation in multiagent systems // IEEE Transactions on Autonomous Mental Development. 2009. Vol. 1. No. 3. P. 170–186.
- Aubret A., Matignon L., Hassas S. A survey on intrinsic motivation in reinforcement learning. URL: https://arxiv.org/abs/1908.06976(дата обращения 01.11.2020).
- Bellemare M. et al. Unifying count-based exploration and intrinsic motivation // NIPS. 2016. Vol. 29. P. 1471–1479.
- Chakraborty A., Kar A.K. Swarm intelligence: A review of algorithms // Nature-Inspired Computing and Optimization. Berlin: Springer, 2017. P. 475–494.
- Coppola M. et al. A Survey on Swarming With Micro Air Vehicles: Fundamental Challenges and Constraints // Frontiers in Robotics and AI. 2020. Vol. 7. P. 18.
- Dorigo M., Blum C. Ant colony optimization theory: A survey // Theoretical computer science. 2005. Vol. 344. No. 2–3. P. 243–278.
- Dorigo M., Stützle T. Ant colony optimization: overview and recent advances // Handbook of metaheuristics. Berlin: Springer, 2019. P. 311–351
- Gambardella L.M., Dorigo M. Ant-Q: A reinforcement learning approach to the traveling salesman problem // Machine Learning Proceedings. 1995. P. 252–260.
- GhasemAghaei R. et al. Ant colony-based reinforcement learning algorithm for routing in wireless sensor networks // Instrumentation & Measurement Technology Conference. 2007. P. 1–6.
- Hüttenrauch M. et al. Deep reinforcement learning for swarm systems // Journal of Machine Learning Research. 2019. Vol. 20. No. 54. P. 1–31.
- Hüttenrauch M., Šošić A., Neumann G. Local communication protocols for learning complex swarm behaviors with deep reinforcement learning // International Conference on Swarm Intelligence. Berlin: Springer, 2018. P. 71–83.
- Khan A. et al. Collaborative multiagent reinforcement learning in homogeneous swarms. URL: https://openreview.net/pdf?id=ByeDojRcYQ (дата обращения 01.11.2020).
- Kho L.C. et al. Ant colony optimization for 2 satisfiability in restricted neural symbolic integration // AIP Conference Proceedings. 2020. Vol. 2266. No. 1. P. 050006.
- Matta M. et al. Q-RTS: a real-time swarm intelligence based on multi-agent Q-learning // Electronics Letters. 2019. Vol. 55. No. 10. P. 589–591.
- Oh K.K., Park M.C., Ahn H.S. A survey of multi-agent formation control // Automatica. 2015. Vol. 53. P. 424–440.
- Ostrovski G. et al. Count-based exploration with neural density models. URL: https://arxiv.org/abs/1703.01310 (дата обращения 01.11.2020).
- Schranz M. et al. Swarm Robotic Behaviors and Current Applications // Frontiers in Robotics and AI. 2020. Vol. 7. P. 36.
- Singh P. et al. Swarm Intelligence Algorithms: A Tutorial. Boca Raton: CRC press, 2020. 363 p.
- Socha K., Blum C. An ant colony optimization algorithm for continuous optimization: application to feed-forward neural network training // Neural Computing and Applications. 2007. Vol. 16. No. 3. P. 235–247.
- Šošic A. et al. Inverse reinforcement learning in swarm systems // Proceedings of the 1st Workshop on Transferin Reinforcement Learning at the 16th International Conference on Autonomous Agents and Multiagent Systems. Sao Paulo, Brazil, 2017. P. 17.
- Rizk Y., Awad M., Tunstel E.W. Decision making in multiagent systems: A survey // IEEE Transactions on Cognitive and Developmental Systems. 2018. Vol. 10. No. 3. P. 514–529.
- Tarassov V.B., Gapanyuk Y.E. Complex Graphs in the Modeling of Multi-agent Systems: From Goal-Resource Networks to Fuzzy Metagraphs // Russian Conference on Artificial Intelligence. Berlin: Springer, 2020. P. 177–198.
- Yang X.S. (ed.). Nature-Inspired Computation and Swarm Intelligence: Algorithms, Theory and Applications. NY: Academic Press, 2020.
- Zhou S., Yan S. The stability analysis for a class of multi-agent group formation with input saturation constraints // Proceedings of Chinese Guidance, Navigation and Control Conference. 2014. P. 1618–1623.