Search arXiv⌕ Search

arXiv subjects

Kazuyuki Hara

Publications and source records attributed to Kazuyuki Hara.

13 recordsLinked to original sources

Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task Performance

Knowledge distillation trains a small student model to reproduce the outputs of a large teacher model, and its progress is typically monitored through the teacher--student discrepancy. The quantity of ultimate interest, however, is the student's error with respect to the true task. We study the relation between these two objectives in a minimal three-party model, a true teacher (generative model), a teacher, and a student, all soft committee machines, in which the true teacher contains a shared latent factor that the teacher cannot represent, with mismatch strength controlled by a single scalar $\dmiss$. Within an order-parameter description of online distillation, and exploiting closed-form (arcsine-type) expressions for all errors under error-function activations, we prove that the learning dynamics and the distillation error $\Ets$ are exactly invariant to $\dmiss$, whereas the true error $\Etzs$ and the gap $Δ=\Etzs-\Ets$ are strictly increasing in $\dmiss$, with a rate that is amplified linearly by the complexity $M_0$ of the true teacher. Numerical phase diagrams over the plane spanned by true-teacher complexity and student capacity confirm the predicted deformation: the contours of $\Ets$ do not move while the landscape of $\Etzs$ rises systematically, and a teacher-miss regime, where mimicry succeeds but the task fails, expands with $\dmiss$. The results give a quantitative warning against evaluating distillation solely through teacher-mimicry metrics and identify the gap $Δ$ as a minimal diagnostic for distinguishing teacher-miss from capacity-limited failure.

cs.LG↗

Cardiomyopathy Diagnosis Model from Endomyocardial Biopsy Specimens: Appropriate Feature Space and Class Boundary in Small Sample Size Data

As the number of patients with heart failure increases, machine learning (ML) has garnered attention in cardiomyopathy diagnosis, driven by the shortage of pathologists. However, endomyocardial biopsy specimens are often small sample size and require techniques such as feature extraction and dimensionality reduction. This study aims to determine whether texture features are effective for feature extraction in the pathological diagnosis of cardiomyopathy. Furthermore, model designs that contribute toward improving generalization performance are examined by applying feature selection (FS) and dimensional compression (DC) to several ML models. The obtained results were verified by visualizing the inter-class distribution differences and conducting statistical hypothesis testing based on texture features. Additionally, they were evaluated using predictive performance across different model designs with varying combinations of FS and DC (applied or not) and decision boundaries. The obtained results confirmed that texture features may be effective for the pathological diagnosis of cardiomyopathy. Moreover, when the ratio of features to the sample size is high, a multi-step process involving FS and DC improved the generalization performance, with the linear kernel support vector machine achieving the best results. This process was demonstrated to be potentially effective for models with reduced complexity, regardless of whether the decision boundaries were linear, curved, perpendicular, or parallel to the axes. These findings are expected to facilitate the development of an effective cardiomyopathy diagnostic model for its rapid adoption in medical practice.

cs.LG↗

Theoretical Analysis of SIRVVD Model to Provide Insight on the Target Rate of COVID-19/SARS-CoV-2 Vaccination in Japan

The effectiveness of the first and second dose vaccinations are different for COVID-19; therefore, a susceptible-infected-recovered-vaccination1-vaccination2-death (SIRVVD) model that can represent the states of the first and second vaccination doses has been proposed. By the previous study, we can carry out simulating the spread of infectious disease considering the effects of the first and second doses of the vaccination based on the SIRVVD model. However, theoretical analysis of the SIRVVD Model is insufficient. Therefore, we obtained an analytical expression of the infectious number, by treating the numbers of susceptible persons and vaccinated persons as parameters. We used the solution to determine the target rate of the vaccination for decreasing the infection numbers of the COVID-19 Delta variant (B.1.617) in Japan. Further, we investigated the target vaccination rates for cases with strong or weak variants by comparison with the COVID-19 Delta variant (B.1.617). This study contributes to the mathematical development of the SIRVVD model and provides insight into the target rate of the vaccination to decrease the number of infections.

q-bio.PE↗

Effectiveness of the COVID-19 Contact-Confirming Application (COCOA) based on a Multi Agent Simulation

As of Aug. 2020, coronavirus disease 2019 (COVID-19) is still spreading in the world. In Japan, the Ministry of Health, Labor, and Welfare developed "COVID-19 Contact-Confirming Application (COCOA)," which was released on Jun. 19, 2020. By utilizing COCOA, users can know whether or not they had contact with infected persons. If those who had contact with infectors keep staying at home, they may not infect those outside. However, effectiveness decreasing the number of infectors depending on the app's various usage parameters is not clear. If it is clear, we could set the objective value of the app's usage parameters (e.g., the usage rate of the total populations) and call for installation of the app. Therefore, we develop a multi-agent simulator that can express COVID-19 spreading and usage of the apps, such as COCOA. In this study, we describe the simulator and the effectiveness of the app in various scenarios. The result obtained in this study supports those of previously conducted studies.

cs.CY↗

Analysis of Dropout in Online Learning

Deep learning is the state-of-the-art in fields such as visual object recognition and speech recognition. This learning uses a large number of layers and a huge number of units and connections. Therefore, overfitting is a serious problem with it, and the dropout which is a kind of regularization tool is used. However, in online learning, the effect of dropout is not well known. This paper presents our investigation on the effect of dropout in online learning. We analyzed the effect of dropout on convergence speed near the singular point. Our results indicated that dropout is effective in online learning. Dropout tends to avoid the singular point for convergence speed near that point.

cs.LG↗

Analysis of dropout learning regarded as ensemble learning

Deep learning is the state-of-the-art in fields such as visual object recognition and speech recognition. This learning uses a large number of layers, huge number of units, and connections. Therefore, overfitting is a serious problem. To avoid this problem, dropout learning is proposed. Dropout learning neglects some inputs and hidden units in the learning process with a probability, p, and then, the neglected inputs and hidden units are combined with the learned network to express the final output. We find that the process of combining the neglected hidden units with the learned network can be regarded as ensemble learning, so we analyze dropout learning from this point of view.

cs.LG↗

Statistical Mechanics of Node-perturbation Learning with Noisy Baseline

Node-perturbation learning is a type of statistical gradient descent algorithm that can be applied to problems where the objective function is not explicitly formulated, including reinforcement learning. It estimates the gradient of an objective function by using the change in the object function in response to the perturbation. The value of the objective function for an unperturbed output is called a baseline. Cho et al. proposed node-perturbation learning with a noisy baseline. In this paper, we report on building the statistical mechanics of Cho's model and on deriving coupled differential equations of order parameters that depict learning dynamics. We also show how to derive the generalization error by solving the differential equations of order parameters. On the basis of the results, we show that Cho's results are also apply in general cases and show some general performances of Cho's model.

stat.ML↗

Statistical Mechanics of On-line Ensemble Teacher Learning through a Novel Perceptron Learning Rule

In ensemble teacher learning, ensemble teachers have only uncertain information about the true teacher, and this information is given by an ensemble consisting of an infinite number of ensemble teachers whose variety is sufficiently rich. In this learning, a student learns from an ensemble teacher that is iteratively selected randomly from a pool of many ensemble teachers. An interesting point of ensemble teacher learning is the asymptotic behavior of the student to approach the true teacher by learning from ensemble teachers. The student performance is improved by using the Hebbian learning rule in the learning. However, the perceptron learning rule cannot improve the student performance. On the other hand, we proposed a perceptron learning rule with a margin. This learning rule is identical to the perceptron learning rule when the margin is zero and identical to the Hebbian learning rule when the margin is infinity. Thus, this rule connects the perceptron learning rule and the Hebbian learning rule continuously through the size of the margin. Using this rule, we study changes in the learning behavior from the perceptron learning rule to the Hebbian learning rule by considering several margin sizes. From the results, we show that by setting a margin of kappa > 0, the effect of an ensemble appears and becomes significant when a larger margin kappa is used.

cond-mat.dis-nn↗

Optimization of the Asymptotic Property of Mutual Learning Involving an Integration Mechanism of Ensemble Learning

We propose an optimization method of mutual learning which converges into the identical state of optimum ensemble learning within the framework of on-line learning, and have analyzed its asymptotic property through the statistical mechanics method.The proposed model consists of two learning steps: two students independently learn from a teacher, and then the students learn from each other through the mutual learning. In mutual learning, students learn from each other and the generalization error is improved even if the teacher has not taken part in the mutual learning. However, in the case of different initial overlaps(direction cosine) between teacher and students, a student with a larger initial overlap tends to have a larger generalization error than that of before the mutual learning. To overcome this problem, our proposed optimization method of mutual learning optimizes the step sizes of two students to minimize the asymptotic property of the generalization error. Consequently, the optimized mutual learning converges to a generalization error identical to that of the optimal ensemble learning. In addition, we show the relationship between the optimum step size of the mutual learning and the integration mechanism of the ensemble learning.

cond-mat.dis-nn↗

Analysis of ensemble learning using simple perceptrons based on online learning theory

Ensemble learning of $K$ nonlinear perceptrons, which determine their outputs by sign functions, is discussed within the framework of online learning and statistical mechanics. One purpose of statistical learning theory is to theoretically obtain the generalization error. This paper shows that ensemble generalization error can be calculated by using two order parameters, that is, the similarity between a teacher and a student, and the similarity among students. The differential equations that describe the dynamical behaviors of these order parameters are derived in the case of general learning rules. The concrete forms of these differential equations are derived analytically in the cases of three well-known rules: Hebbian learning, perceptron learning and AdaTron learning. Ensemble generalization errors of these three rules are calculated by using the results determined by solving their differential equations. As a result, these three rules show different characteristics in their affinity for ensemble learning, that is ``maintaining variety among students." Results show that AdaTron learning is superior to the other two rules with respect to that affinity.

cond-mat.dis-nn↗

Ensemble learning of linear perceptron; Online learning theory

Within the framework of on-line learning, we study the generalization error of an ensemble learning machine learning from a linear teacher perceptron. The generalization error achieved by an ensemble of linear perceptrons having homogeneous or inhomogeneous initial weight vectors is precisely calculated at the thermodynamic limit of a large number of input elements and shows rich behavior. Our main findings are as follows. For learning with homogeneous initial weight vectors, the generalization error using an infinite number of linear student perceptrons is equal to only half that of a single linear perceptron, and converges with that of the infinite case with O(1/K) for a finite number of K linear perceptrons. For learning with inhomogeneous initial weight vectors, it is advantageous to use an approach of weighted averaging over the output of the linear perceptrons, and we show the conditions under which the optimal weights are constant during the learning process. The optimal weights depend on only correlation of the initial weight vectors.

cond-mat.dis-nn↗

On-line learning through simple perceptron with a margin

We analyze a learning method that uses a margin $κ$ {\it a la} Gardner for simple perceptron learning. This method corresponds to the perceptron learning when $κ=0$, and to the Hebbian learning when $κ\to \infty$. Nevertheless, we found that the generalization ability of the method was superior to that of the perceptron and the Hebbian methods at an early stage of learning. We analyzed the asymptotic property of the learning curve of this method through computer simulation and found that it was the same as for perceptron learning. We also investigated an adaptive margin control method.

cond-mat.dis-nn↗