Search arXiv⌕ Search

arXiv · cond-mat/0402143

Personal Email Networks: An Effective Anti-Spam Tool

Abstract

We provide an automated graph theoretic method for identifying individual users' trusted networks of friends in cyberspace. We routinely use our social networks to judge the trustworthiness of outsiders, i.e., to decide where to buy our next car, or to find a good mechanic for it. In this work, we show that an email user may similarly use his email network, constructed solely from sender and recipient information available in the email headers, to distinguish between unsolicited commercial emails, commonly called "spam", and emails associated with his circles of friends. We exploit the properties of social networks to construct an automated anti-spam tool which processes an individual user's personal email network to simultaneously identify the user's core trusted networks of friends, as well as subnetworks generated by spams. In our empirical studies of individual mail boxes, our algorithm classified approximately 53% of all emails as spam or non-spam, with 100% accuracy. Some of the emails are left unclassified by this network analysis tool. However, one can exploit two of the following useful features. First, it requires no user intervention or supervised training; second, it results in no false negatives i.e., spam being misclassified as non-spam, or vice versa. We demonstrate that these two features suggest that our algorithm may be used as a platform for a comprehensive solution to the spam problem when used in concert with more sophisticated, but more cumbersome, content-based filters.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

P. Oscar Boykin, Vwani Roychowdhury. 2004-02-04. Personal Email Networks: An Effective Anti-Spam Tool. https://arxiv.org/abs/cond-mat/0402143

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Perfect resonance and fragile localization suppression in correlated disordered chains

Spatial correlations can suppress scattering in disordered chains and produce perfectly transmitting resonances. The practical value of this protection, however, depends not on the resonance peak itself but on the width of the surrounding transmission window and its sensitivity to local ordering errors. We show that a resonance can remain exactly transparent at its center while arbitrarily rare adjacent-swap errors restore an inverse localization length proportional to the square of the energy detuning. For lossless single-channel chains assembled from independent blocks of fixed length and composition, this positive quadratic term is guaranteed by a local scattering invariant and holds for every arrangement and every fixed swap probability between zero and one. With exact tuning and matched contacts, the central transmission remains unity. A microscopic quantum chain exhibits both higher-order suppression of scattering in the ideal recursive arrangement and the predicted response to local exchanges. These results reveal a limitation of spatial ordering that is invisible to a measurement at the resonance alone: spatial ordering protects the resonance peak, not the transport around it. The effect can therefore be tested experimentally by measuring transmission spectra before and after exchanges, without identifying microscopic defects.

cond-mat.dis-nn↗

Unifying Physical Backpropagation

Physical computing systems exploit device dynamics for computation, but their gradient-based optimization is challenging: backpropagation through a digital twin suffers from a model-reality gap. On-device gradient computation could resolve this issue, and a handful of theoretical and experimental studies have proposed ways to achieve it. Yet a unifying theory identifying when a physical system can compute the gradient of its own performance has been missing. Here we develop such a unification based on the adjoint method: we identify sufficient conditions under which the adjoint field required for formally exact gradients can be generated on the same hardware that performs the computation. Linear and nonlinear systems obey fundamentally different conditions: for linear systems, damping or gain is admissible provided reciprocity is preserved. For nonlinear trajectory systems, the sufficient conditions are reciprocity of the linearized system and the existence of a time-reversal mirror. Algorithmically, the nonlinear case requires infinitesimal nudging, whereas linear systems admit a finite-amplitude experiment. We recover (quantum) Equilibrium Propagation, Hamiltonian echo backpropagation, fully forward mode training and in situ gradient methods in integrated-photonic and free-space-optical systems. Finally, we show that reciprocity is a special case of more general intertwining conditions. For linear systems, these permit exact on-device gradients in a class of non-Hermitian, non-reciprocal systems. For nonlinear trajectories, they combine with generalized time-reversal mirrors to cover, e.g., PT-symmetric equations. The framework also includes time-dependent parameters and Onsager-reciprocal dynamics, providing a unified basis for formally exact physical learning.

cond-mat.dis-nn↗

Rank-One Signal Recovery in Sparse Wishart Noise

We study the high-dimensional recovery of a signal vector $\mathbf{x}$ in the presence of sparse Wishart-like noise. We define an $N \times N$ matrix $A = J+(θ/N)\mathbf{xx}^{\top}$, where $\mathbf{xx}^{\top}$ is the rank-one deformation of the random noise matrix $J$. We consider a Wishart-like matrix $J={X}^{\top} X$, where $X$ is a sparse $M \times N$ random matrix with entries $X_{ij} = c_{ij}W_{ij}$, with $c_{ij}$ regulating the density of non-zero elements, and $W_{ij}$ the bond weights. Using the replica method, we compute analytically the top eigenpair statistics of $A$, and their dependence on the signal strength $θ$, the rectangularity ratio $α=\sqrt{M/N}$, and the average connectivity of the noise. The spectral observables are expressed in terms of a system of Recursive Distributional Equations, which are efficiently solved via a Population Dynamics algorithm. They allow us to compute the average largest eigenvalue $\langleλ_1\rangle_{A}$, the average top eigenvector component density, and the average overlap between the top eigenvector of $A$ and $\mathbf{x}$. We identify a critical threshold $θ_{\mathrm{crit}}$--depending on the average connectivity of the noise--that marks a BBP-like phase transition: below this value, $\langleλ_1\rangle_{A}$ is unaffected by the signal, and the overlap vanishes. Thus, the signal is not recoverable from the top eigenvector of $A$. For $θ>θ_{\mathrm{crit}}$, the signal-related outlier eigenvalue becomes $\langleλ_1\rangle_{A}$ and the overlap is nonzero, allowing for recovery of the signal. The results are in excellent agreement with numerical diagonalisation. We show that in the dense limit, the recovery threshold and eigen-statistics converge to the results predicted by the classical BBP transition for additive rank-one deformations of dense Wishart matrices.

cond-mat.dis-nn↗