Search arXiv⌕ Search

arXiv subjects

Yu Pan

Publications and source records attributed to Yu Pan.

At least 145 records · Page 8Linked to original sources

Semantically Proportional Patchmix for Few-Shot Learning

Few-shot learning aims to classify unseen classes with only a limited number of labeled data. Recent works have demonstrated that training models with a simple transfer learning strategy can achieve competitive results in few-shot classification. Although excelling at distinguishing training data, these models are not well generalized to unseen data, probably due to insufficient feature representations on evaluation. To tackle this issue, we propose Semantically Proportional Patchmix (SePPMix), in which patches are cut and pasted among training images and the ground truth labels are mixed proportionally to the semantic information of the patches. In this way, we can improve the generalization ability of the model by regional dropout effect without introducing severe label noise. To learn more robust representations of data, we further take rotate transformation on the mixed images and predict rotations as a rule-based regularizer. Extensive experiments on prevalent few-shot benchmarks have shown the effectiveness of our proposed method.

cs.CV↗

High precision measurement of cosmic curvature: from gravitational waves and cosmic chronometer

Although the spatial curvature has been measured with very high precision, it still suffers from the well known cosmic curvature tension. In this paper, we propose an improved method to determine the cosmic curvature, by using the simulated data of binary neutron star mergers observed by the second generation space-based DECi-hertz Interferometer Gravitational-wave Observatory (DECIGO). By applying the Hubble parameter observations of cosmic chronometers to the DECIGO standard sirens, we explore different possibilities of making measurements of the cosmic curvature referring to a distant past: one is to reconstruct the Hubble parameters through the Gaussian process without the influence of hypothetical models, and the other is deriving constraints on $Ω_K$ in the framework of non-flat $Λ$ cold dark matter model. It is shown that in the improved method DECIGO could provide a reliable and stringent constraint on the cosmic curvature ($Ω_{K} = -0.007\pm0.016$), while we could only expect the zero cosmic curvature to be established at the precision of $ΔΩ_K=0.12$ in the second model-dependent method. Therefore, our results indicate that in the framework of methodology proposed in this paper, the increasing number of well-measured standard sirens in DECIGO could significantly reduce the bias of estimations for cosmic curvature. Such constraint is also comparable to the precision of Planck 2018 results with the newest cosmic microwave background (CMB) observations ($ΔΩ_{K} \approx 0.018$), based on the concordance $Λ$CDM model.

astro-ph.CO↗

Quantum Language Model with Entanglement Embedding for Question Answering

Quantum Language Models (QLMs) in which words are modelled as quantum superposition of sememes have demonstrated a high level of model transparency and good post-hoc interpretability. Nevertheless, in the current literature word sequences are basically modelled as a classical mixture of word states, which cannot fully exploit the potential of a quantum probabilistic description. A full quantum model is yet to be developed to explicitly capture the non-classical correlations within the word sequences. We propose a neural network model with a novel Entanglement Embedding (EE) module, whose function is to transform the word sequences into entangled pure states of many-body quantum systems. Strong quantum entanglement, which is the central concept of quantum information and an indication of parallelized correlations among the words, is observed within the word sequences. Numerical experiments show that the proposed QLM with EE (QLM-EE) achieves superior performance compared with the classical deep neural network models and other QLMs on Question Answering (QA) datasets. In addition, the post-hoc interpretability of the model can be improved by quantizing the degree of entanglement among the words.

cs.CL↗

Making Adversarial Examples More Transferable and Indistinguishable

Fast gradient sign attack series are popular methods that are used to generate adversarial examples. However, most of the approaches based on fast gradient sign attack series cannot balance the indistinguishability and transferability due to the limitations of the basic sign structure. To address this problem, we propose a method, called Adam Iterative Fast Gradient Tanh Method (AI-FGTM), to generate indistinguishable adversarial examples with high transferability. Besides, smaller kernels and dynamic step size are also applied to generate adversarial examples for further increasing the attack success rates. Extensive experiments on an ImageNet-compatible dataset show that our method generates more indistinguishable adversarial examples and achieves higher attack success rates without extra running time and resource. Our best transfer-based attack NI-TI-DI-AITM can fool six classic defense models with an average success rate of 89.3% and three advanced defense models with an average success rate of 82.7%, which are higher than the state-of-the-art gradient-based attacks. Additionally, our method can also reduce nearly 20% mean perturbation. We expect that our method will serve as a new baseline for generating adversarial examples with better transferability and indistinguishability.

cs.CV↗

Augmented Legendrian cobordism in $J^1S^1$

We consider Legendrian links and tangles in $J^1S^1$ and $J^1[0,1]$ equipped with Morse complex families over a field $\mathbb{F}$ and classify them up to Legendrian cobordism. When the coefficient field is $\mathbb{F}_2$ this provides a cobordism classification for Legendrians equipped with augmentations of the Legendrian contact homology DG-algebras. A complete set of invariants, for which arbitrary values may be obtained, is provided by the fiber cohomology, a graded monodromy matrix, and a mod $2$ spin number. We apply the classification to construct augmented Legendrian surfaces in $J^1M$ with $\dim M = 2$ realizing any prescribed monodromy representation, $Φ:π_1(M,x_0) \rightarrow \mathit{GL}(\mathbf{n}, \mathbb{F})$.

math.SG↗

TedNet: A Pytorch Toolkit for Tensor Decomposition Networks

Tensor Decomposition Networks (TDNs) prevail for their inherent compact architectures. To give more researchers a flexible way to exploit TDNs, we present a Pytorch toolkit named TedNet. TedNet implements 5 kinds of tensor decomposition(i.e., CANDECOMP/PARAFAC (CP), Block-Term Tucker (BTT), Tucker-2, Tensor Train (TT) and Tensor Ring (TR) on traditional deep neural layers, the convolutional layer and the fully-connected layer. By utilizing the basic layers, it is simple to construct a variety of TDNs. TedNet is available at https://github.com/tnbar/tednet.

cs.LG↗

Anomalous thermoelectric effects and quantum oscillations in the kagome metal CsV$_3$Sb$_5$

The kagome metal compounds $A$V$_3$Sb$_5$ ($A$ = K, Rb, and Cs) feature a wealth of phenomena including nontrivial band topology, charge density wave (CDW), and superconductivity. One intriguing property is the time-reversal symmetry breaking in the CDW state without local moments, which leads to anomalous transport responses. Here, we report the investigation of magneto-thermoelectric effects on high-quality CsV$_3$Sb$_5$ single crystals. A large anomalous Nernst effect is observed at temperatures below 30 K. Multiple Fermi surfaces with small effective masses are revealed by quantum oscillations in Nernst and Seebeck signals under high magnetic field. Furthermore, we find an unknown frequency, and attribute it to the magnetic breakdown across two smaller Fermi surfaces. A gap around 20 meV can be resolved from the breakdown threshold field, which we propose to be introduced by the CDW. These results shed new light on the CDW-related phenomena, particularly in $A$V$_3$Sb$_5$ compounds.

cond-mat.mtrl-sci↗

Detecting quantum entanglement with unsupervised learning

Quantum properties, such as entanglement and coherence, are indispensable resources in various quantum information processing tasks. However, there still lacks an efficient and scalable way to detecting these useful features, especially for high-dimensional and multipartite quantum systems. In this work, we exploit the convexity of samples without the desired quantum features and design an unsupervised machine learning method to detect the presence of such features as anomalies. Particularly, in the context of entanglement detection, we propose a complex-valued neural network composed of pseudo-siamese network and generative adversarial net, and then train it with only separable states to construct non-linear witnesses for entanglement. It is shown via numerical examples, ranging from two-qubit to ten-qubit systems, that our network is able to achieve high detection accuracy which is above 97.5% on average.Moreover, it is capable of revealing rich structures of entanglement, such as partial entanglement among subsystems. Our results are readily applicable to the detection of other quantum resources such as Bell nonlocality and steerability, and thus our work could provide a powerful tool to extract quantum features hidden in multipartite quantum data.

quant-ph↗

Giant anomalous Nernst signal in the antiferromagnet YbMnBi2

Searching for a high anomalous Nernst effect (ANE) is crucial for thermoelectric energy conversion applications because the associated unique transverse geometry facilitates module fabrication. Topological ferromagnets with large Berry curvatures show high ANEs; however, they face drawbacks such as strong magnetic disturbances and low mobility due to high magnetization. Herein, we demonstrate that YbMnBi2, a canted antiferromagnet, has a large ANE conductivity of ~10 Am-1K-1 that surpasses the common high values (i.e. 3-5 Am-1K-1) observed so far in ferromagnets. The canted spin structure of Mn guarantees a nonzero Berry curvature but generates only a weak magnetization three orders of magnitude lower than that of general ferromagnets. The heavy Bi with a large spin-orbit coupling enables a high ANE and low thermal conductivity, whereas its highly dispersive px/y orbitals ensure low resistivity. The high anomalous transverse thermoelectric performance and extremely small magnetization makes YbMnBi2 an excellent candidate for transverse thermoelectrics.

cond-mat.mtrl-sci↗

Fast query-by-example speech search using separable model

Traditional Query-by-Example (QbE) speech search approaches usually use methods based on frame-level features, while state-of-the-art approaches tend to use models based on acoustic word embeddings (AWEs) to transform variable length audio signals into fixed length feature vector representations. However, these approaches cannot meet the requirements of the search quality as well as speed at the same time. In this paper, we propose a novel fast QbE speech search method based on separable models to fix this problem. First, a QbE speech search training framework is introduced. Second, we design a novel model inference scheme based on RepVGG which can efficiently improve the QbE search quality. Third, we modify and improve our QbE speech search model according to the proposed model inference scheme. Experiments on keywords dataset shows that our proposed method can improve the GPU Real-time Factor (RTF) from 1/150 to 1/2300 by just applying separable model scheme and outperforms other state-of-the-art methods.

eess.AS↗

A new way to test the WIMP dark matter models

In this paper, we investigate the possibility of testing the weakly interacting massive particle (WIMP) dark matter (DM) models by applying the simplest phenomenological model which introduces an interaction term between dark energy (DE) and WIMP DM, i.e., $Q = 3γ_{DM} Hρ_{DM}$. In general, the coupling strength $γ_{DE}$ is close to $0$ as the interaction between DE and WIMP DM is very weak, thus the effect of $γ_ {DE}$ on the evolution of $Y$ associated with DM energy density can be safely neglected. Meanwhile, our numerical calculation also indicates that $x_f\approx20$ is associated with DM freeze-out temperature, which is the same as the vanishing interaction scenario. As for DM relic density, it will be magnified by $\frac{2-3γ_{DM}}{2}[{2πg_* m_{DM}^3}/{(45 s_0 x_f^3})]^{γ_{DM}}$ times, which provides a new way to test WIMP DM models. As an example, we analyze the case in which WIMP DM is a scalar DM. (SGL+SNe+Hz) and (CMB+BAO+SNe) cosmological observations will give $γ_{DM}=0.134^{+0.17}_{-0.069}$ and $γ_{DM}=-0.0008\pm0.0016$, respectively. After further considering the constraints from DM direct detection experiment, DM indirect detection experiment, and DM relic density, we find that the allowed parameter space of the scalar DM model will be completely excluded for the former cosmological observations, while it will increase for the latter ones. Those two cosmological observations lead to an almost paradoxical conclusion. Therefore, one could expect more stringent constraints on the WMIP DM models, with the accumulation of more accurate cosmological observations in the near future.

astro-ph.CO↗

Large topological Hall effect in an easy-cone ferromagnet (Cr0.9B0.1)Te

The Berry phase understanding of electronic properties has attracted special interest in condensed matter physics, leading to phenomena such as the anomalous Hall effect and the topological Hall effect. A non-vanishing Berry phase, induced in momentum space by the band structure or in real space by a non-coplanar spin structure, is the origin of both effects. Here, we report a sign conversion of the anomalous Hall effect and a large topological Hall effect in (Cr0.9B0.1)Te single crystals. The spin reorientation from an easy-axis structure at high temperature to an easy-cone structure below 140 K leads to conversion of the Berry curvature, which influences both, anomalous and topological, Hall effects in the presence of an applied magnetic field and current. We compare and summarize the topological Hall effect in four categories with different mechanisms and have a discussion into the possible artificial fake effect of topological Hall effect in polycrystalline samples, which provides a deep understanding of the relation between spin structure and Hall properties.

cond-mat.mtrl-sci↗

Heuristic Rank Selection with Progressively Searching Tensor Ring Network

Recently, Tensor Ring Networks (TRNs) have been applied in deep networks, achieving remarkable successes in compression ratio and accuracy. Although highly related to the performance of TRNs, rank selection is seldom studied in previous works and usually set to equal in experiments. Meanwhile, there is not any heuristic method to choose the rank, and an enumerating way to find appropriate rank is extremely time-consuming. Interestingly, we discover that part of the rank elements is sensitive and usually aggregate in a narrow region, namely an interest region. Therefore, based on the above phenomenon, we propose a novel progressive genetic algorithm named Progressively Searching Tensor Ring Network Search (PSTRN), which has the ability to find optimal rank precisely and efficiently. Through the evolutionary phase and progressive phase, PSTRN can converge to the interest region quickly and harvest good performance. Experimental results show that PSTRN can significantly reduce the complexity of seeking rank, compared with the enumerating method. Furthermore, our method is validated on public benchmarks like MNIST, CIFAR10/100, UCF11 and HMDB51, achieving the state-of-the-art performance.

cs.CV↗

AFINet: Attentive Feature Integration Networks for Image Classification

Convolutional Neural Networks (CNNs) have achieved tremendous success in a number of learning tasks including image classification. Recent advanced models in CNNs, such as ResNets, mainly focus on the skip connection to avoid gradient vanishing. DenseNet designs suggest creating additional bypasses to transfer features as an alternative strategy in network design. In this paper, we design Attentive Feature Integration (AFI) modules, which are widely applicable to most recent network architectures, leading to new architectures named AFI-Nets. AFI-Nets explicitly model the correlations among different levels of features and selectively transfer features with a little overhead.AFI-ResNet-152 obtains a 1.24% relative improvement on the ImageNet dataset while decreases the FLOPs by about 10% and the number of parameters by about 9.2% compared to ResNet-152.

cs.CV↗

Testing f(R) gravity with the simulated data of gravitational waves from the Einstein Telescope

In this paper we analyze the implications of gravitational waves (GWs) as standard sirens on the modified gravity models by using the third-generation gravitational wave detector, i.e., the Einstein Telescope. Two viable models in $f(R)$ theories within the Palatini formalism are considered in our analysis ($f_{1}(\mathcal{R})=\mathcal{R}-\fracβ{\mathcal{R}^{n}}$ and $f_{2}(\mathcal{R})=\mathcal{R}+α\ln{\mathcal{R}}-β$), with the combination of simulated GW data and the latest electromagnetic (EM) observational data (including the recently released Pantheon type Ia supernovae sample, the cosmic chronometer data, and baryon acoustic oscillation distance measurements). Our analysis reveals that the standard sirens GWs, which provide an independent and complementary alternative to current experiments, could effectively eliminate the degeneracies among parameters in the two modified gravity models. In addition, we thoroughly investigate the nature of geometrical dark energy in the modified gravity theories with the assistance of $Om(z)$ and statefinder diagnostic analysis. The present analysis makes it clear-cut that the simplest cosmological constant model is still the most preferred by the current data. However, the combination of future naturally improved GW data most recent EM observations will reveal the consistency or acknowledge the tension between the $Λ$CDM model and modified gravity theories.

astro-ph.CO↗

Spectrum Attention Mechanism for Time Series Classification

Time series classification(TSC) has always been an important and challenging research task. With the wide application of deep learning, more and more researchers use deep learning models to solve TSC problems. Since time series always contains a lot of noise, which has a negative impact on network training, people usually filter the original data before training the network. The existing schemes are to treat the filtering and training as two stages, and the design of the filter requires expert experience, which increases the design difficulty of the algorithm and is not universal. We note that the essence of filtering is to filter out the insignificant frequency components and highlight the important ones, which is similar to the attention mechanism. In this paper, we propose an attention mechanism that acts on spectrum (SAM). The network can assign appropriate weights to each frequency component to achieve adaptive filtering. We use L1 regularization to further enhance the frequency screening capability of SAM. We also propose a segmented-SAM (SSAM) to avoid the loss of time domain information caused by using the spectrum of the whole sequence. In which, a tumbling window is introduced to segment the original data. Then SAM is applied to each segment to generate new features. We propose a heuristic strategy to search for the appropriate number of segments. Experimental results show that SSAM can produce better feature representations, make the network converge faster, and improve the robustness and classification accuracy.

cs.LG↗

RegNet: Self-Regulated Network for Image Classification

The ResNet and its variants have achieved remarkable successes in various computer vision tasks. Despite its success in making gradient flow through building blocks, the simple shortcut connection mechanism limits the ability of re-exploring new potentially complementary features due to the additive function. To address this issue, in this paper, we propose to introduce a regulator module as a memory mechanism to extract complementary features, which are further fed to the ResNet. In particular, the regulator module is composed of convolutional RNNs (e.g., Convolutional LSTMs or Convolutional GRUs), which are shown to be good at extracting Spatio-temporal information. We named the new regulated networks as RegNet. The regulator module can be easily implemented and appended to any ResNet architecture. We also apply the regulator module for improving the Squeeze-and-Excitation ResNet to show the generalization ability of our method. Experimental results on three image classification datasets have demonstrated the promising performance of the proposed architecture compared with the standard ResNet, SE-ResNet, and other state-of-the-art architectures.

eess.IV↗

Constructions of Lagrangian cobordisms

Lagrangian cobordisms between Legendrian knots arise in Symplectic Field Theory and impose an interesting and not well-understood relation on Legendrian knots. There are some known "elementary" building blocks for Lagrangian cobordisms that are smoothly the attachment of $0$- and $1$-handles. An important question is whether every pair of non-empty Legendrians that are related by a connected Lagrangian cobordism can be related by a ribbon Lagrangian cobordism, in particular one that is "decomposable" into a composition of these elementary building blocks. We will describe these and other combinatorial building blocks as well as some geometric methods, involving the theory of satellites, to construct Lagrangian cobordisms. We will then survey some known results, derived through Heegaard Floer Homology and contact surgery, that may provide a pathway to proving the existence of nondecomposable (nonribbon) Lagrangian cobordisms.

math.SG↗