Search arXiv⌕ Search

arXiv subjects

Yang Yu

Publications and source records attributed to Yang Yu.

At least 343 records · Page 19Linked to original sources

Interaction expansion inchworm Monte Carlo solver for lattice and impurity models

Multi-orbital quantum impurity models with general interaction and hybridization terms appear in a wide range of applications including embedding, quantum transport, and nanoscience. However, most quantum impurity solvers are restricted to a few impurity orbitals, discretized baths, diagonal hybridizations, or density-density interactions. Here, we generalize the inchworm quantum Monte Carlo method to the interaction expansion and explore its application to typical single- and multi-orbital problems encountered in investigations of impurity and lattice models. Our implementation generically outperforms bare and bold-line quantum Monte Carlo algorithms in the interaction expansion. So far, for the systems studied here, it remains inferior to the more specialized hybridization expansion and auxiliary field algorithms. The problem of convergence to unphysical fixed points, which hampers so-called bold-line methods, is not encountered in inchworm Monte Carlo.

cond-mat.str-el↗

Long-term Trends of Regolith Movement on the Surface of Small Bodies

This paper studies the long-term migration of disturbed regolith materials on the surface of Solar System small bodies from the viewpoint of nonlinear dynamics. We propose an approximation model for secular mass movement, which combines the complex topography and irregular gravitational field. Choosing asteroid 101955 Bennu as a representative, the global change of the dynamical environment is examined, which presents a division of the creeping-sliding-shedding regions for a spun-up asteroid. In the creeping region, the dynamical equation of disturbed regolith grains is established based on the assumption of "trigger-slide" motion mode. The equilibrium points, local manifolds and large-scale trajectories of the system are calculated to clarify the dynamical characteristics of long-term regolith movement. Generally, we find for a low spin rate, the surface regolith grains flow toward the middle latitudes from the polar/equatorial regions, which is dominated by the gradient of the geopotential. While spun up to a high rate, regolith grains tend to migrate toward the equator, which happens in parallel with a topological shift of the local equilibria at low latitudes. From a long-term perspective, we find the equilibrium points dominate the global trends of regolith movements. Using the methodology developed in this paper, we give a prospect or retrospect to the secular motion of regolith materials during the spin-up process, and the results reveal a significant regulatory role of the equilibria. Through a detailed look at the dynamical scheme under different spin rates, we achieve a macro forecast of the global trends of regolith motion during the spin-up process, which explains the global geologic evolution driven by the long-term movements of regolith materials.

astro-ph.EP↗

Progressive Multi-stage Interactive Training in Mobile Network for Fine-grained Recognition

Fine-grained Visual Classification (FGVC) aims to identify objects from subcategories. It is a very challenging task because of the subtle inter-class differences. Existing research applies large-scale convolutional neural networks or visual transformers as the feature extractor, which is extremely computationally expensive. In fact, real-world scenarios of fine-grained recognition often require a more lightweight mobile network that can be utilized offline. However, the fundamental mobile network feature extraction capability is weaker than large-scale models. In this paper, based on the lightweight MobilenetV2, we propose a Progressive Multi-Stage Interactive training method with a Recursive Mosaic Generator (RMG-PMSI). First, we propose a Recursive Mosaic Generator (RMG) that generates images with different granularities in different phases. Then, the features of different stages pass through a Multi-Stage Interaction (MSI) module, which strengthens and complements the corresponding features of different stages. Finally, using the progressive training (P), the features extracted by the model in different stages can be fully utilized and fused with each other. Experiments on three prestigious fine-grained benchmarks show that RMG-PMSI can significantly improve the performance with good robustness and transferability.

cs.CV↗

A catalogue of 323 cataclysmic variables from LAMOST DR6

In this work, we present a catalog of cataclysmic variables (CVs) identified from the Sixth Data Release (DR6) of the Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST). To single out the CV spectra, we introduce a novel machine-learning algorithm called UMAP to screen out a total of 169,509 H$α$-emission spectra, and obtain a classification accuracy of the algorithm of over 99.6$\%$ from the cross-validation set. We then apply the template matching program PyHammer v2.0 to the LAMOST spectra to obtain the optimal spectral type with metallicity, which helps us identify the chromospherically active stars and potential binary stars from the 169,509 spectra. After visually inspecting all the spectra, we identify 323 CV candidates from the LAMOST database, among them 52 objects are new. We further discuss the new CV candidates in subtypes based on their spectral features, including five DN subtype during outbursts, five NL subtype and four magnetic CVs (three AM Her type and one IP type). We also find two CVs that have been previously identified by photometry, and confirm their previous classification by the LAMOST spectra.

astro-ph.SR↗

Stochastic optimal scheduling of demand response-enabled microgrids with renewable generations: An analytical-heuristic approach

In the context of transition towards cleaner and sustainable energy production, microgrids have become an effective way for tackling environmental pollution and energy crisis issues. With the increasing penetration of renewables, how to coordinate demand response and renewable generations is a critical and challenging issue in the field of microgrid scheduling. To this end, a bi-level scheduling model is put forward for isolated microgrids with consideration of multi-stakeholders in this paper, where the lower- and upper-level models respectively aim to the minimization of user cost and microgrid operational cost under real-time electricity pricing environments. In order to solve this model, this research combines Jaya algorithm and interior point method (IPM) to develop a hybrid analysis-heuristic solution method called Jaya-IPM, where the lower- and upper- levels are respectively addressed by the IPM and the Jaya, and the scheduling scheme is obtained via iterations between the two levels. After that, the real-time prices updated by the upper-level model and the electricity plans determined by the lower-level model will be alternately iterated between the upper- and lower- levels through the real-time pricing mechanism to obtain an optimal scheduling plan. The test results show that the proposed method can coordinate the uncertainty of renewable generations with demand response strategies, thereby achieving a balance between the interests of microgrid and users; and that by leveraging demand response, the flexibility of the load side can be fully exploited to achieve peak load shaving while maintaining the balance of supply and demand. In addition, the Jaya-IPM algorithm is proven to be superior to the traditional hybrid intelligent algorithm (HIA) and the CPLEX solver in terms of optimization results and calculation efficiency.

eess.SY↗

Optimal control of stimulated Raman adiabatic passage in a superconducting qudit

Stimulated Raman adiabatic passage (STIRAP) is a widely used protocol to realize high-fidelity and robust quantum control in various quantum systems. However, further application of this protocol in superconducting qubits is limited by population leakage caused by the only weak anharmonicity. Here we introduce an optimally controlled shortcut-to-adiabatic (STA) technique to speed up the STIRAP protocol in a superconducting qudit. By modifying the shapes of the STIRAP pulses, we experimentally realize a fast (32 ns) and high-fidelity quantum state transfer. In addition, we demonstrate that our protocol is robust against control parameter perturbations. Our stimulated Raman shortcut-to-adiabatic passage transition provides an efficient and practical approach for quantum information processing.

quant-ph↗

Calculus of Consent via MARL: Legitimating the Collaborative Governance Supplying Public Goods

Public policies that supply public goods, especially those involve collaboration by limiting individual liberty, always give rise to controversies over governance legitimacy. Multi-Agent Reinforcement Learning (MARL) methods are appropriate for supporting the legitimacy of the public policies that supply public goods at the cost of individual interests. Among these policies, the inter-regional collaborative pandemic control is a prominent example, which has become much more important for an increasingly inter-connected world facing a global pandemic like COVID-19. Different patterns of collaborative strategies have been observed among different systems of regions, yet it lacks an analytical process to reason for the legitimacy of those strategies. In this paper, we use the inter-regional collaboration for pandemic control as an example to demonstrate the necessity of MARL in reasoning, and thereby legitimizing policies enforcing such inter-regional collaboration. Experimental results in an exemplary environment show that our MARL approach is able to demonstrate the effectiveness and necessity of restrictions on individual liberty for collaborative supply of public goods. Different optimal policies are learned by our MARL agents under different collaboration levels, which change in an interpretable pattern of collaboration that helps to balance the losses suffered by regions of different types, and consequently promotes the overall welfare. Meanwhile, policies learned with higher collaboration levels yield higher global rewards, which illustrates the benefit of, and thus provides a novel justification for the legitimacy of, promoting inter-regional collaboration. Therefore, our method shows the capability of MARL in computationally modeling and supporting the theory of calculus of consent, developed by Nobel Prize winner J. M. Buchanan.

cs.AI↗

An Introduction of mini-AlphaStar

StarCraft II (SC2) is a real-time strategy game in which players produce and control multiple units to fight against opponent's units. Due to its difficulties, such as huge state space, various action space, a long time horizon, and imperfect information, SC2 has been a research hotspot in reinforcement learning. Recently, an agent called AlphaStar (AS) has been proposed, which shows good performance, obtaining a high win rate of 99.8% against human players. We implemented a mini-scaled version of it called mini-AlphaStar (mAS) based on AS's paper and pseudocode. The difference between AS and mAS is that we substituted the hyper-parameters of AS with smaller ones for mini-scale training. Codes of mAS are all open-sourced (https://github.com/liuruoze/mini-AlphaStar) for future research.

cs.AI↗

Regret Minimization Experience Replay in Off-Policy Reinforcement Learning

In reinforcement learning, experience replay stores past samples for further reuse. Prioritized sampling is a promising technique to better utilize these samples. Previous criteria of prioritization include TD error, recentness and corrective feedback, which are mostly heuristically designed. In this work, we start from the regret minimization objective, and obtain an optimal prioritization strategy for Bellman update that can directly maximize the return of the policy. The theory suggests that data with higher hindsight TD error, better on-policiness and more accurate Q value should be assigned with higher weights during sampling. Thus most previous criteria only consider this strategy partially. We not only provide theoretical justifications for previous criteria, but also propose two new methods to compute the prioritization weight, namely ReMERN and ReMERT. ReMERN learns an error network, while ReMERT exploits the temporal ordering of states. Both methods outperform previous prioritized sampling algorithms in challenging RL benchmarks, including MuJoCo, Atari and Meta-World.

cs.LG↗

Efficient Reinforcement Learning for StarCraft by Abstract Forward Models and Transfer Learning

Injecting human knowledge is an effective way to accelerate reinforcement learning (RL). However, these methods are underexplored. This paper presents our discovery that an abstract forward model (thought-game (TG)) combined with transfer learning (TL) is an effective way. We take StarCraft II as our study environment. With the help of a designed TG, the agent can learn a 99% win-rate on a 64x64 map against the Level-7 built-in AI, using only 1.08 hours in a single commercial machine. We also show that the TG method is not as restrictive as it was thought to be. It can work with roughly designed TGs, and can also be useful when the environment changes. Comparing with previous model-based RL, we show TG is more effective. We also present a TG hypothesis that gives the influence of different fidelity levels of TG. For real games that have unequal state and action spaces, we proposed a novel XfrNet of which usefulness is validated while achieving a 90% win-rate against the cheating Level-10 AI. We argue that the TG method might shed light on further studies of efficient RL with human knowledge.

cs.LG↗

A general framework for scintillation in nanophotonics

Bombardment of materials by high-energy particles (e.g., electrons, nuclei, X- and $γ$-ray photons) often leads to light emission, known generally as scintillation. Scintillation is ubiquitous and enjoys widespread applications in many areas such as medical imaging, X-ray non-destructive inspection, night vision, electron microscopy, and high-energy particle detectors. A large body of research focuses on finding new materials optimized for brighter, faster, and more controlled scintillation. Here, we develop a fundamentally different approach based on integrating nanophotonic structures into scintillators to enhance their emission. To start, we develop a unified and ab initio theory of nanophotonic scintillators that accounts for the key aspects of scintillation: the energy loss by high-energy particles, as well as the light emission by non-equilibrium electrons in arbitrary nanostructured optical systems. This theoretical framework allows us, for the first time, to experimentally demonstrate nearly an order-of-magnitude enhancement of scintillation, in both electron-induced, and X-ray-induced scintillation. Our theory also allows the discovery of structures that could eventually achieve several orders-of-magnitude scintillation enhancement. The framework and results shown here should enable the development of a new class of brighter, faster, and higher-resolution scintillators with tailored and optimized performances - with many potential applications where scintillators are used.

physics.optics↗

QPLEX: Duplex Dueling Multi-Agent Q-Learning

We explore value-based multi-agent reinforcement learning (MARL) in the popular paradigm of centralized training with decentralized execution (CTDE). CTDE has an important concept, Individual-Global-Max (IGM) principle, which requires the consistency between joint and local action selections to support efficient local decision-making. However, in order to achieve scalability, existing MARL methods either limit representation expressiveness of their value function classes or relax the IGM consistency, which may suffer from instability risk or may not perform well in complex domains. This paper presents a novel MARL approach, called duPLEX dueling multi-agent Q-learning (QPLEX), which takes a duplex dueling network architecture to factorize the joint value function. This duplex dueling structure encodes the IGM principle into the neural network architecture and thus enables efficient value function learning. Theoretical analysis shows that QPLEX achieves a complete IGM function class. Empirical experiments on StarCraft II micromanagement tasks demonstrate that QPLEX significantly outperforms state-of-the-art baselines in both online and offline data collection settings, and also reveal that QPLEX achieves high sample efficiency and can benefit from offline datasets without additional online exploration.

cs.LG↗

Labor-right Protecting Dispatch of Meal Delivery Platforms

The boom in the meal delivery industry brings growing concern about the labor rights of riders. Current dispatch policies of meal-delivery platforms focus mainly on satisfying consumers or minimizing the number of riders for cost savings. There are few discussions on improving the working conditions of riders by algorithm design. The lack of concerns on labor rights in mechanism and dispatch design has resulted in a very large time waste for riders and their risky driving. In this research, we propose a queuing-model-based framework to discuss optimal dispatch policy with the goal of labor rights protection. We apply our framework to develop an algorithm minimizing the waiting time of food delivery riders with guaranteed user experience. Our framework also allows us to manifest the value of restaurants' data about their offline-order numbers on improving the benefits of riders.

cs.PF↗

UserBERT: Contrastive User Model Pre-training

User modeling is critical for personalized web applications. Existing user modeling methods usually train user models from user behaviors with task-specific labeled data. However, labeled data in a target task may be insufficient for training accurate user models. Fortunately, there are usually rich unlabeled user behavior data which encode rich information of user characteristics and interests. Thus, pre-training user models on unlabeled user behavior data has the potential to improve user modeling for many downstream tasks. In this paper, we propose a contrastive user model pre-training method named UserBERT. Two self-supervision tasks are incorporated in UserBERT for user model pre-training on unlabeled user behavior data to empower user modeling. The first one is masked behavior prediction, which aims to model the relatedness between user behaviors. The second one is behavior sequence matching, which aims to capture the inherent user interests that are consistent in different periods. In addition, we propose a medium-hard negative sampling framework to select informative negative samples for better contrastive pre-training. We maintain a synchronously updated candidate behavior pool and an asynchronously updated candidate behavior sequence pool to select the locally hardest negative behaviors and behavior sequences in an efficient way. Extensive experiments on two real-world datasets in different tasks show that UserBERT can effectively improve various user models.

cs.IR↗

NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application

Pre-trained language models (PLMs) like BERT have made great progress in NLP. News articles usually contain rich textual information, and PLMs have the potentials to enhance news text modeling for various intelligent news applications like news recommendation and retrieval. However, most existing PLMs are in huge size with hundreds of millions of parameters. Many online news applications need to serve millions of users with low latency tolerance, which poses huge challenges to incorporating PLMs in these scenarios. Knowledge distillation techniques can compress a large PLM into a much smaller one and meanwhile keeps good performance. However, existing language models are pre-trained and distilled on general corpus like Wikipedia, which has some gaps with the news domain and may be suboptimal for news intelligence. In this paper, we propose NewsBERT, which can distill PLMs for efficient and effective news intelligence. In our approach, we design a teacher-student joint learning and distillation framework to collaboratively learn both teacher and student models, where the student model can learn from the learning experience of the teacher model. In addition, we propose a momentum distillation method by incorporating the gradients of teacher model into the update of student model to better transfer useful knowledge learned by the teacher model. Extensive experiments on two real-world datasets with three tasks show that NewsBERT can effectively improve the model performance in various intelligent news applications with much smaller models.

cs.CL↗

The Flare and Warp of the Young Stellar Disk traced with LAMOST DR5 OB-type stars

We present analysis of the spatial density structure for the outer disk from 8$-$14 \,kpc with the LAMOST DR5 13534 OB-type stars and observe similar flaring on north and south sides of the disk implying that the flaring structure is symmetrical about the Galactic plane, for which the scale height at different Galactocentric distance is from 0.14 to 0.5 \,kpc. By using the average slope to characterize the flaring strength we find that the thickness of the OB stellar disk is similar but flaring is slightly stronger compared to the thin disk as traced by red giant branch stars, possibly implying that secular evolution is not the main contributor to the flaring but perturbation scenarios such as interactions with passing dwarf galaxies should be more possible. When comparing the scale height of OB stellar disk of the north and south sides with the gas disk, the former one is slightly thicker than the later one by $\approx$ 33 and 9 \,pc, meaning that one could tentatively use young OB-type stars to trace the gas properties. Meanwhile, we unravel that the radial scale length of the young OB stellar disk is 1.17 $\pm$ 0.05 \,kpc, which is shorter than that of the gas disk, confirming that the gas disk is more extended than stellar disk. What is more, by considering the mid-plane displacements ($Z_{0}$) in our density model we find that almost all of $Z_{0}$ are within 100 \,pc with the increasing trend as Galactocentric distance increases.

astro-ph.GA↗

Neural-to-Tree Policy Distillation with Policy Improvement Criterion

While deep reinforcement learning has achieved promising results in challenging decision-making tasks, the main bones of its success --- deep neural networks are mostly black-boxes. A feasible way to gain insight into a black-box model is to distill it into an interpretable model such as a decision tree, which consists of if-then rules and is easy to grasp and be verified. However, the traditional model distillation is usually a supervised learning task under a stationary data distribution assumption, which is violated in reinforcement learning. Therefore, a typical policy distillation that clones model behaviors with even a small error could bring a data distribution shift, resulting in an unsatisfied distilled policy model with low fidelity or low performance. In this paper, we propose to address this issue by changing the distillation objective from behavior cloning to maximizing an advantage evaluation. The novel distillation objective maximizes an approximated cumulative reward and focuses more on disastrous behaviors in critical states, which controls the data shift effect. We evaluate our method on several Gym tasks, a commercial fight game, and a self-driving car simulator. The empirical results show that the proposed method can preserve a higher cumulative reward than behavior cloning and learn a more consistent policy to the original one. Moreover, by examining the extracted rules from the distilled decision trees, we demonstrate that the proposed method delivers reasonable and robust decisions.

cs.LG↗

Imitate TheWorld: A Search Engine Simulation Platform

Recent E-commerce applications benefit from the growth of deep learning techniques. However, we notice that many works attempt to maximize business objectives by closely matching offline labels which follow the supervised learning paradigm. This results in models obtain high offline performance in terms of Area Under Curve (AUC) and Normalized Discounted Cumulative Gain (NDCG), but cannot consistently increase the revenue metrics such as purchases amount of users. Towards the issues, we build a simulated search engine AESim that can properly give feedback by a well-trained discriminator for generated pages, as a dynamic dataset. Different from previous simulation platforms which lose connection with the real world, ours depends on the real data in AliExpress Search: we use adversarial learning to generate virtual users and use Generative Adversarial Imitation Learning (GAIL) to capture behavior patterns of users. Our experiments also show AESim can better reflect the online performance of ranking models than classic ranking metrics, implying AESim can play a surrogate of AliExpress Search and evaluate models without going online.

cs.AI↗