Search arXiv⌕ Search

arXiv subjects

Lei Yu

Publications and source records attributed to Lei Yu.

At least 271 records · Page 15Linked to original sources

Sentence Encoding with Tree-constrained Relation Networks

The meaning of a sentence is a function of the relations that hold between its words. We instantiate this relational view of semantics in a series of neural models based on variants of relation networks (RNs) which represent a set of objects (for us, words forming a sentence) in terms of representations of pairs of objects. We propose two extensions to the basic RN model for natural language. First, building on the intuition that not all word pairs are equally informative about the meaning of a sentence, we use constraints based on both supervised and unsupervised dependency syntax to control which relations influence the representation. Second, since higher-order relations are poorly captured by a sum of pairwise relations, we use a recurrent extension of RNs to propagate information so as to form representations of higher order relations. Experiments on sentence classification, sentence pair classification, and machine translation reveal that, while basic RNs are only modestly effective for sentence representation, recurrent RNs with latent syntax are a reliably powerful representational device.

cs.CL↗

Tackling Sequence to Sequence Mapping Problems with Neural Networks

In Natural Language Processing (NLP), it is important to detect the relationship between two sequences or to generate a sequence of tokens given another observed sequence. We call the type of problems on modelling sequence pairs as sequence to sequence (seq2seq) mapping problems. A lot of research has been devoted to finding ways of tackling these problems, with traditional approaches relying on a combination of hand-crafted features, alignment models, segmentation heuristics, and external linguistic resources. Although great progress has been made, these traditional approaches suffer from various drawbacks, such as complicated pipeline, laborious feature engineering, and the difficulty for domain adaptation. Recently, neural networks emerged as a promising solution to many problems in NLP, speech recognition, and computer vision. Neural models are powerful because they can be trained end to end, generalise well to unseen examples, and the same framework can be easily adapted to a new domain. The aim of this thesis is to advance the state-of-the-art in seq2seq mapping problems with neural networks. We explore solutions from three major aspects: investigating neural models for representing sequences, modelling interactions between sequences, and using unpaired data to boost the performance of neural models. For each aspect, we propose novel models and evaluate their efficacy on various tasks of seq2seq mapping.

cs.CL↗

Analog-to-digital conversion revolutionized by deep learning

As the bridge between the analog world and digital computers, analog-to-digital converters are generally used in modern information systems such as radar, surveillance, and communications. For the configuration of analog-to-digital converters in future high-frequency broadband systems, we introduce a revolutionary architecture that adopts deep learning technology to overcome tradeoffs between bandwidth, sampling rate, and accuracy. A photonic front-end provides broadband capability for direct sampling and speed multiplication. Trained deep neural networks learn the patterns of system defects, maintaining high accuracy of quantized data in a succinct and adaptive manner. Based on numerical and experimental demonstrations, we show that the proposed architecture outperforms state-of-the-art analog-to-digital converters, confirming the potential of our approach in future analog-to-digital converter design and performance enhancement of future information systems.

eess.SP↗

Codegree threshold for tiling $k$-graphs with two edges sharing exactly $\ell$ vertices

Given integer $k$ and a $k$-graph $F$, let $t_{k-1}(n,F)$ be the minimum integer $t$ such that every $k$-graph $H$ on $n$ vertices with codegree at least $t$ contains an $F$-factor. For integers $k\geq3$ and $0\leq\ell\leq k-1$, let $\mathcal{Y}_{k,\ell}$ be a $k$-graph with two edges that shares exactly $\ell$ vertices. Han and Zhao (JCTA, 2015) asked the following question: For all $k\ge 3$, $0\le \ell\le k-1$ and sufficiently large $n$ divisible by $2k-\ell$, determine the exact value of $t_{k-1}(n,\mathcal{Y}_{k,\ell})$. In this paper, we show that $t_{k-1}(n,\mathcal{Y}_{k,\ell})=\frac{n}{2k-\ell}$ for $k\geq3$ and $1\leq\ell\leq k-2$, combining with two previously known results of Rödl, Ruciński and Szemerédi {(JCTA, 2009)} and Gao, Han and Zhao (arXiv, 2016), the question of Han and Zhao is solved completely.

math.CO↗

Distortion Bounds for Source Broadcast Problems

This paper investigates the joint source-channel coding problem of sending a memoryless source over a memoryless broadcast channel. An inner bound and several outer bounds on the admissible distortion region are derived, which respectively generalize and unify several existing bounds. As a consequence, we also obtain an inner bound and an outer bound for the degraded broadcast channel case. When specialized to the Gaussian or binary source broadcast, the inner bound and outer bound not only recover the best known inner bound and outer bound in the literature, but also generate some new results. Besides, we also extend the inner bound and outer bounds to the Wyner-Ziv source broadcast problem, i.e., source broadcast with side information available at decoders. Some new bounds are obtained when specialized to the Wyner-Ziv Gaussian and Wyner-Ziv binary cases.

cs.IT↗

Diverse Exploration for Fast and Safe Policy Improvement

We study an important yet under-addressed problem of quickly and safely improving policies in online reinforcement learning domains. As its solution, we propose a novel exploration strategy - diverse exploration (DE), which learns and deploys a diverse set of safe policies to explore the environment. We provide DE theory explaining why diversity in behavior policies enables effective exploration without sacrificing exploitation. Our empirical study shows that an online policy improvement algorithm framework implementing the DE strategy can achieve both fast policy improvement and safe online performance.

cs.LG↗

Exponential Strong Converse for Content Identification with Lossy Recovery

We revisit the high-dimensional content identification with lossy recovery problem (Tuncel and Gündüz, 2014) and establish an exponential strong converse theorem. As a corollary of the exponential strong converse theorem, we derive an upper bound on the joint identification-error and excess-distortion exponent for the problem. Our main results can be specialized to the biometrical identification problem~(Willems, 2003) and the content identification problem~(Tuncel, 2009) since these two problems are both special cases of the content identification with lossy recovery problem. We leverage the information spectrum method introduced by Oohama and adapt the strong converse techniques therein to be applicable to the problem at hand.

cs.IT↗

Local exact one-sided boundary null controllability of entropy solutions to a class of hyperbolic systems of balance laws

We consider nxn hyperbolic systems of balance laws in one-space dimension under the assumption that all negative (resp. positive) characteristics are linearly degenerate. We prove the local exact one-sided boundary null controllability of entropy solutions to this class of systems, which generalizes the corresponding results from the case without source terms to that with source terms. In order to apply the strategy used for conservation law, we essentially modify the constructive method by introducing two different kinds of approximate solutions to system in the forward sense and to the system in the rightward (resp. leftward) sense, respectively, and we prove that their limit solutions are equivalent to some extend.

math.AP↗

Automatic Streaming Segmentation of Stereo Video Using Bilateral Space

In this paper, we take advantage of binocular camera and propose an unsupervised algorithm based on semi-supervised segmentation algorithm and extracting foreground part efficiently. We creatively embed depth information into bilateral grid in the graph cut model and achieve considerable segmenting accuracy in the case of no user input. The experi- ment approves the high precision, time efficiency of our algorithm and its adaptation to complex natural scenario which is significant for practical application.

cs.CV↗

The size of $3$-uniform hypergraphs with given matching number and codegree

Determine the size of $r$-graphs with given graph parameters is an interesting problem. Chvátal and Hanson (JCTB, 1976) gave a tight upper bound of the size of 2-graphs with restricted maximum degree and matching number; Khare (DM, 2014) studied the same problem for linear $3$-graphs with restricted matching number and maximum degree. In this paper, we give a tight upper bound of the size of $3$-graphs with bounded codegree and matching number.

math.CO↗

Generalized Common Informations: Measuring Commonness by the Conditional Maximal Correlation

In literature, different common informations were defined by Gács and Körner, by Wyner, and by Kumar, Li, and Gamal, respectively. In this paper, we define two generalized versions of common informations, named approximate and exact information-correlation functions, by exploiting the conditional maximal correlation as a commonness or privacy measure. These two generalized common informations encompass the notions of Gács-Körner's, Wyner's, and Kumar-Li-Gamal's common informations as special cases. Furthermore, to give operational characterizations of these two generalized common informations, we also study the problems of private sources synthesis and common information extraction, and show that the information-correlation functions are equal to the minimum rates of commonness needed to ensure that some conditional maximal correlation constraints are satisfied for the centralized setting versions of these problems. As a byproduct, the conditional maximal correlation has been studied as well.

cs.IT↗

Joint Source-Channel Secrecy Using Uncoded Schemes: Towards Secure Source Broadcast

This paper investigates a joint source-channel secrecy problem for the Shannon cipher broadcast system. We suppose list secrecy is applied, i.e., a wiretapper is allowed to produce a list of reconstruction sequences and the secrecy is measured by the minimum distortion over the entire list. For discrete communication cases, we propose a permutation-based uncoded scheme, which cascades a random permutation with a symbol-by-symbol mapping. Using this scheme, we derive an inner bound for the admissible region of secret key rate, list rate, wiretapper distortion, and distortions of legitimate users. For the converse part, we easily obtain an outer bound for the admissible region from an existing result. Comparing the outer bound with the inner bound shows that the proposed scheme is optimal under certain conditions. Besides, we extend the proposed scheme to the scalar and vector Gaussian communication scenarios, and characterize the corresponding performance as well. For these two cases, we also propose another uncoded scheme, orthogonal-transform-based scheme, which achieves the same performance as the permutation-based scheme. Interestingly, by introducing the random permutation or the random orthogonal transform into the traditional uncoded scheme, the proposed uncoded schemes, on one hand, provide a certain level of secrecy, and on the other hand, do not lose any performance in terms of the distortions for legitimate users.

cs.IT↗

Odd induced subgraphs in graphs with treewidth at most two

A long-standing conjecture asserts that there exists a constant $c>0$ such that every graph of order $n$ without isolated vertices contains an induced subgraph of order at least $cn$ with all degrees odd. Scott (1992) proved that every graph $G$ has an induced subgraph of order at least $|V(G)|/(2χ(G))$ with all degrees odd, where $χ(G)$ is the chromatic number of $G$, this implies the conjecture for graphs with { bounded} chromatic number. But the factor $1/(2χ(G))$ seems to be not best possible, for example, Radcliffe and Scott (1995) proved $c=\frac 23$ for trees, Berman, Wang and Wargo (1997) showed that $c=\frac 25$ for graphs with maximum degree $3$, so it is interesting to determine the exact value of $c$ for special family of graphs. In this paper, we further confirm the conjecture for graphs with treewidth at most 2 with $c=\frac{2}{5}$, and the bound is best possible.

math.CO↗

An algorithm of frequency estimation for multi-channel coprime sampling

In some applications of frequency estimation, it is challenging to sample at as high as the Nyquist rate due to hardware limitations. An effective solution is to use multiple sub-Nyquist channels with coprime undersampling ratios to jointly sample. In this paper, an algorithm suitable for any number of channels is proposed, which is based on subspace techniques. Numerical simulations show that the proposed algorithm has high accuracy and good robustness.

cs.IT↗

Frequency Estimation of Multiple Sinusoids with Three Sub-Nyquist Channels

Frequency estimation of multiple sinusoids is significant in both theory and application. In some application scenarios, only sub-Nyquist samples are available to estimate the frequencies. A conventional approach is to sample the signals at several lower rates. In this paper, we address frequency estimation of the signals in the time domain through undersampled data. We analyze the impact of undersampling and demonstrate that three sub-Nyquist channels are generally enough to estimate the frequencies provided the undersampling ratios are pairwise coprime. We deduce the condition that leads to the failure of resolving frequency ambiguity when two coprime undersampling channels are utilized. When three-channel sub-Nyquist samples are used jointly, the frequencies can be determined uniquely and the correct frequencies are estimated. Numerical experiments verify the correctness of our analysis and conclusion.

cs.IT↗

The Neural Noisy Channel

We formulate sequence to sequence transduction as a noisy channel decoding problem and use recurrent neural networks to parameterise the source and channel models. Unlike direct models which can suffer from explaining-away effects during training, noisy channel models must produce outputs that explain their inputs, and their component models can be trained with not only paired training samples but also unpaired samples from the marginal output distribution. Using a latent variable to control how much of the conditioning sequence the channel model needs to read in order to generate a subsequent symbol, we obtain a tractable and effective beam search decoder. Experimental results on abstractive sentence summarisation, morphological inflection, and machine translation show that noisy channel models outperform direct models, and that they significantly benefit from increased amounts of unpaired output data that direct models cannot easily use.

cs.CL↗

Source-Channel Secrecy for Shannon Cipher System

Recently, a secrecy measure based on list-reconstruction has been proposed [2], in which a wiretapper is allowed to produce a list of $2^{mR_{L}}$ reconstruction sequences and the secrecy is measured by the minimum distortion over the entire list. In this paper, we show that this list secrecy problem is equivalent to the one with secrecy measured by a new quantity \emph{lossy-equivocation}, which is proven to be the minimum optimistic 1-achievable source coding rate (the minimum coding rate needed to reconstruct the source within target distortion with positive probability for \emph{infinitely many blocklengths}) of the source with the wiretapped signal as two-sided information, and also can be seen as a lossy extension of conventional equivocation. Upon this (or list) secrecy measure, we study source-channel secrecy problem in the discrete memoryless Shannon cipher system with \emph{noisy} wiretap channel. Two inner bounds and an outer bound on the achievable region of secret key rate, list rate, wiretapper distortion, and distortion of legitimate user are given. The inner bounds are derived by using uncoded scheme and (operationally) separate scheme, respectively. Thanks to the equivalence between lossy-equivocation secrecy and list secrecy, information spectrum method is leveraged to prove the outer bound. As special cases, the admissible region for the case of degraded wiretap channel or lossless communication for legitimate user has been characterized completely. For both these two cases, separate scheme is proven to be optimal. Interestingly, however, separation indeed suffers performance loss for other certain cases. Besides, we also extend our results to characterize the achievable region for Gaussian communication case. As a side product optimistic lossy source coding has also been addressed.

cs.IT↗

Distortion Bounds for Transmitting Correlated Sources with Common Part over MAC

This paper investigates the joint source-channel coding problem of sending two correlated memoryless sources with common part over a memoryless multiple access channel (MAC). An inner bound and two outer bounds on the achievable distortion region are derived. In particular, they respectively recover the existing bounds for several special cases, such as communication without common part, lossless communication, and noiseless communication. When specialized to quadratic Gaussian communication case, transmitting Gaussian sources with Gaussian common part over Gaussian MAC, the inner bound and outer bound are used to generate two new bounds. Numerical result shows that common part improves the distortion of such distributed source-channel coding problem.

cs.IT↗