Multi-Labelled Value Networks for Computer Go
This paper proposes a new approach to a novel value network architecture for the game Go, called a multi-labelled (ML) value network. In the ML value network, different values (win rates) are trained simultaneously for different settings of komi, a compensation given to balance the initiative of playing first. The ML value network has three advantages, (a) it outputs values for different komi, (b) it supports dynamic komi, and (c) it lowers the mean squared error (MSE). This paper also proposes a new dynamic komi method to improve game-playing strength. This paper also performs experiments to demonstrate the merits of the architecture. First, the MSE of the ML value network is generally lower than the value network alone. Second, the program based on the ML value network wins by a rate of 67.6% against the program based on the value network alone. Third, the program with the proposed dynamic komi method significantly improves the playing strength over the baseline that does not use dynamic komi, especially for handicap games. To our knowledge, up to date, no handicap games have been played openly by programs using value networks. This paper provides these programs with a useful approach to playing handicap games.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Model-free inference of unseen attractors: Reconstructing phase space features from a single noisy trajectory using reservoir computing
Reservoir computers are powerful tools for chaotic time series prediction. They can be trained to approximate phase space flows and can thus both predict future values to a high accuracy, as well as reconstruct the gener…
Time SeriesTime Series AnalysisTime Series PredictionAnomaly Detection in Automated Fibre Placement: Learning with Data Limitations
Conventional defect detection systems in Automated Fibre Placement (AFP) typically rely on end-to-end supervised learning, necessitating a substantial number of labelled defective samples for effective training. However,…
Anomaly DetectionBinary ClassificationDefect DetectionNoise-Tolerant Few-Shot Unsupervised Adapter for Vision-Language Models
Recent advances in large-scale vision-language models have achieved impressive performance in various zero-shot image classification tasks. While prior studies have demonstrated significant improvements by introducing fe…
image-classificationImage ClassificationKnowledge DistillationPseudo Label+1GANsfer Learning: Combining labelled and unlabelled data for GAN based data augmentation
Medical imaging is a domain which suffers from a paucity of manually annotated data for the training of learning algorithms. Manually delineating pathological regions at a pixel level is a time consuming process, especia…
Data AugmentationSegmentationApplications and Effect Evaluation of Generative Adversarial Networks in Semi-Supervised Learning
In recent years, image classification, as a core task in computer vision, relies on high-quality labelled data, which restricts the wide application of deep learning models in practical scenarios. To alleviate the proble…
Classificationimage-classificationImage ClassificationImage Generation+1