Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 523 records · Page 29Linked to original sources

Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models

Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users' geographic locations from casually shared photographs, by performing structured reasoning over subtle visual cues such as architectural styles, vegetation, and lighting conditions. In this work, we present a systematic study of MLRM-driven geolocation privacy leakage. We first reveal that refusal-based safeguards are critically insufficient, as carefully crafted jailbreak prompts can raise model response rates to 100%. We further identify that existing defenses, which inject imperceptible perturbations into shared images, suffer from structural limitations intrinsic to their pixel-space optimization, resulting in degraded black-box transferability and pronounced visual artifacts. Motivated by these findings, we propose a diffusion-based framework that provides targeted, proactive defense against geolocation privacy leakage. By injecting perturbations into the latent space of a diffusion model during reverse sampling, our method operates directly on high-level semantic representations, thereby resolving the effectiveness-utility bottlenecks by construction. We further ground our optimization with GeoCLIP, a model explicitly aligned with GPS coordinates, as a surrogate to pinpoint and disrupt the geographic signals that MLRMs exploit for location inference. This targeted semantic disruption yields significantly stronger black-box transferability while preserving perceptual image quality, offering a seamless integration on social media platforms. Code is available at https://github.com/RachelWolowitz/Hiding_in_plain_sight.

cs.CV↗

Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources

Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central challenge of nonlinear Independent Component Analysis (nICA). We prove identifiability (exact recovery) up to trivial ambiguities for real analytic generating functions when source probability density functions have a finite number of discontinuities in the first derivative. The Laplace distribution is the most prominent example satisfying this assumption. Our proof relies on the contrast between kinks in the source distribution and the smoothness of real analytic functions. Real analytic functions comprise a broad class of generating mechanisms, and can be approximated with Normalizing Flows or Variational Autoencoders with standard activation functions (e.g., tanh, softplus, GELU), so our result applies with minimal changes to existing training pipelines. We perform experiments on real and synthetic data with both Normalizing Flows and Variational Auto-Encoders demonstrating their identifiability properties. In experiments on CelebA data we recover several interpretable latent factors controlling unique attributes across the dataset.

cs.LG↗

The Three-Dimensional Erdős Box Problem Has Exponent $11/4$

Let $z(n)$ be the maximum number of edges in a tripartite $3$-uniform hypergraph with $n$ vertices in each part and no copy of $K_{2,2,2}^{(3)}$ (a ``box''). Erdős (1964) proved that $z(n) = O(n^{11/4})$, whereas the best previous lower bound, due to Katz, Krop, and Maggioni (2002), was $Ω(n^{8/3})$. For each $q = 2^m$, we construct a box-free hypergraph with $q^4$ vertices in each part and $q^{11}$ edges, showing that $z(n) = Θ(n^{11/4})$. The construction uses the power map $τ(s) = s^{q^2-q+1}$ on $\F_{q^3}$, which sends the fibers of $s \mapsto τ(s+1) + τ(s)$ to pairwise skew affine lines over $\F_q$.

math.CO↗

Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts

We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore reconstructs missing spans without requiring oracle knowledge of their length. To our knowledge, it is the first large-scale decoder model for historical Greek, and the first for any ancient Mediterranean language. Evaluated as in prior work, on short gaps of up to ten characters, Apollo Restore places the correct restoration among its top twenty candidates for 80.6%/54.6%/61.0% of documentary-papyrus, literary-papyrus, and stone-inscription lacunae, exceeding the strongest published models by $1.6\times$/$2.6\times$/$1.4\times$. Prior evaluation protocols, however, inflate scores through a bias toward trivially short gaps; under a length-balanced metric Apollo Restore's advantage over the strongest published models grows to $2.3\times$/$3.5\times$/$1.6\times$ and degrades gracefully, even given incorrect length hints. In a blind study, 20 expert papyrologists, epigraphists, and philologists strongly preferred Apollo Restore to the strongest baseline and judged its performance at least as good as human restorations in 77% of cases. Apollo Restore also improves the published reading of PHerc. 1667---a papyrus roll carbonised in the eruption of Vesuvius in 79 CE and digitally unrolled and edited after Apollo Restore's training data was compiled. Apollo Restore is an output of the Decoding Antiquity initiative to build specialized LLMs for historical languages and manuscripts, led by the Austrian Academy of Sciences.

cs.CL↗

SC Derandomization for Regular ROBPs and Models Beyond BPL

We study SC derandomizations for regular read-once branching programs (ROBPs) and computation models beyond BPL. For regular ROBPs with length $n$, width $w$, and multiple accept nodes, we attain the following results. 1. When $n \le w$, we show an SC derandomization with space $O(\log^2 n+\log w)$ and two-sided error $1/\text{poly}(nw)$. 2. When $n \ge w$, we show an SC derandomization with space $O(\log n \log w)$ and two-sided error $1/\text{poly}(w)$. In addition, when $w=O(\log n)$, we attain an optimal $O(\log n)$ space derandomization with two-sided error $1/\text{poly}(w)$. 3. When $w \le 2^{O(\sqrt{\log n})}$, we show that reachability of regular ROBPs (i.e. derandmization of one-sided unbounded small error ROBPs) can be computed in SC. We also show that when regular ROBPs are powering, optimal deterministic logspace can be attained for computing reachability and for bounded two-sided error derandomization, when $w = n^{1/c}$ for some constant $c$. We further show that two super sets of BPL can be computed in SC. 1. For probabilistic logspace TMs with a two-way access random tape, we show that it can be approximated in SC if each entry of the random tape is accessed for at most a constant number of times. 2. For probabilistic logspace TMs with a polynomial size stack, i.e. probabilistic logspace Auxiliary Push-down Machines (AuxPDMs), we show that it can be approximated in SC if the timings of push/pop/idle stack operations do not depend on the randomness. The first model is the read-multiplicity model considered by Impagliazzo, Nisan, Wigderson (STOC'94), in which they show that their INW generator can fool such computations. For the second model, we indicate that it contains candidate languages separating BQL from BPL considered by Apers and Edenhofer (CCC'25).

cs.CC↗

Do Language Models Need Music Supervision? Verifiable Rewards for Multi-Constraint Symbolic Music Generation

Language models now generate symbolic music from text, and research has focused on musicality. However, many applications require a score that meets explicit constraints, which models struggle to satisfy jointly: on MusicConstraintBench, our benchmark of 2,180 items over eight families of programmatically verifiable constraints, Llama-3.1-70B satisfies 0.630 of single-constraint items but only 0.044 of four-constraint ones. As a remedy, we introduce MusicRLVR, which trains a language model with group relative policy optimisation (GRPO) on verifier rewards alone, needing no human annotation, reward model or music-domain supervised fine-tuning. MusicRLVR incorporates (1) a hard validation gate that rejects malformed scores, (2) graded per-family credit that, unlike a binary reward, separates partially correct outputs, and (3) an all-satisfied bonus for meeting every constraint at once. Extensive experiments show that, in under four hours of training, MusicRLVR raises Qwen3-4B-Instruct-2507 from 0.160 to 0.797 on mixed constraints, outperforming Llama-3.1-70B, and generalises to unseen property combinations, out-of-range parameters and more constraints than any training prompt. The recipe transfers to Qwen3-8B, and neither trained model loses significant accuracy on general benchmarks.

cs.SD↗

THz Spectroscopy of Urine Vapors from Patients with Prostate Cancer and Benign Prostatic Hyperplasia: A Critical Review and Statistical Reanalysis

The cancer specificity of serum prostate-specific antigen (PSA) motivates complementary, non-invasive approaches to prostate disease assessment. Building on reported high-resolution terahertz (THz) analysis of urine-derived products, this review integrates spectroscopic evidence with a statistical reanalysis of clinical data from 24 patients with prostate cancer (PC) and 14 with histologically confirmed benign prostatic hyperplasia (BPH). Nonparametric statistical comparison shows higher PSA levels in PC, while overlapping patient distributions and receiver-operating-characteristic analysis with bootstrap uncertainty estimation demonstrate incomplete discrimination from BPH. Fast frequency-sweep spectroscopy characterizes urine-derived volatile and thermal-decomposition products, distinguishing compounds shared between groups from candidate PC-associated assignments, including glycolaldehyde, butyronitrile, pentanenitrile and methyl isocyanate. Glycolaldehyde has been detected by urinary carbonyl profiling, while benzaldehyde and phenol have been reported in urinary volatile studies. Related aldehydes identified by gas chromatography-mass spectrometry and glycine-associated metabolites reported in proton nuclear magnetic resonance studies provide complementary chemical and metabolic context. These correspondences support plausibility of THz findings without establishing PC specificity. Together, the dataset, statistical assessment and comparative literature synthesis support THz urine spectroscopy as a platform for molecular discovery and candidate prioritization. Standardized sample preparation, separation of native urinary constituents from processing-derived products, and validation in independent cohorts remain necessary.

cond-mat.other↗

On the Born rule in a new quantum approach

In the context of a new approach towards quantum foundation, the Born rule is proved under weak conditions through several steps. Some of the steps turn out to have connections to statistical theory. One step gives a generalized likelihood principle, a result of independent interest. The sole additional assumption behind the Born rule is a consequence of the likelihood principle of statistics. This principle is discussed from several points of view. The whole approach towards quantum foundation advocated here, leads to a simpler and more intuitive pure state concept than the traditional one in quantum mechanics. It also provides links to statistical theory and to Andrei Khrennikov's macroscopical quantum-like models, in particular, models connected to cognition and decisions. Finally, the approach has connections to relativity theory and to quantum field theory. It is a hope that this approach now can be seen as a good alternative to the traditional formal foundation of quantum theory. The approach is here discussed in the case of finite-values variables, but continuous variables and variables taking a countable set of values can be included in the theory by approximating them with finite-valued ones.

quant-ph↗

Vacancy aggregation enhances NV- spin coherence in diamond: a cluster-correlation-expansion study of multi-vacancy spin baths in semiconductors

Annealing a semiconductor makes its vacancies mobile; they aggregate into multi-vacancy complexes that often carry spin. Such centres are a magnetic-noise source for any spin qubit among them, in silicon and silicon carbide as well as in diamond, and they are accordingly blamed for NV-decoherence in irradiated and implanted diamond, with two coherence records credited to removing them. Cluster-correlation-expansion simulations driven by published electron-paramagnetic-resonance parameters invert that attribution. At a fixed paramagnetic spin density a multi-vacancy bath gives a Hahn-echo coherence time 3.1-7.1 times longer than a bath of isolated negative vacancies, over three decades of concentration. A bath dephases the qubit because its spins exchange spin projections with one another, and two can exchange only if their transition frequencies match. A fine-structure splitting shifts those frequencies. Were every defect on the same crystallographic site, all would shift alike and still match: worth only a factor 1.15. Real defects occupy symmetry-equivalent sites pointing in different directions, so neighbours land at different frequencies and stop exchanging: a further 2.7. What governs the coherence time is therefore the fraction of bath pairs sharing a transition frequency, not the zero-field splitting. Vacancy aggregation extends NV- spin coherence rather than shortening it.

cond-mat.mtrl-sci↗

Spectroscopic parameters of $B_c$ meson

We study the spectroscopy and decays of the $B_c$ meson in a nonrelativistic quark model in which all observables follow from a single set of wave functions, with spin-dependent effects included directly in the bound-state equation. With parameters calibrated to the measured $B_c(1S)$ and $B_c(2S)$ masses, we predict the $S$-, $P$-, and $D$-wave spectrum up to $n=5$, the pseudoscalar and vector decay constants with one-loop QCD corrections, the E1 and M1 radiative widths, and the radial Regge trajectories of ten spin-parity families. The $1P$ multiplet is predicted at 6692--6746~MeV with the normal fine-structure ordering, compatible with lattice QCD and located in the region of the two structures observed by LHCb in the $B_c^+γ$ spectrum. The $1D$ states lie near 7008~MeV, below the $BD$ threshold, and all four $D$-wave families are provided. The decay constant $f_{B_c}=507.1$~MeV is consistent with recent QCD sum-rule results and corresponds to $\mathcal{B}(B_c\toτν_τ)\approx3\%$. The E1 widths follow the expected angular weights, while their magnitudes are governed mainly by the radial overlap. The $1P\to1S$ (81--112~keV) and $1D\to1P$ (71--93~keV) widths suggest the $1D\to1P\to1S$ cascade as a promising route toward the first $D$-wave $B_c$ state. The Regge trajectories are nearly linear, with slopes decreasing from about 5.8~GeV$^2$ for the $S$-wave to 4.7--4.8~GeV$^2$ for the $D$-wave states. These results may guide forthcoming measurements at the LHC and future $e^+e^-$ facilities.

hep-ph↗

How Children Design and Reason about Trustworthy AI Chatbots

Children increasingly interact with AI chatbots, making trust calibration essential to AI literacy. Prior research has examined children's trust in AI mainly as users evaluating systems built by others, rather than as designers of their own chatbots. We developed a chatbot-building environment with adjustable trust-relevant traits (e.g., confidence, transparency, formality, assertiveness), rules, and persona. We conducted mixed-methods study with 115 learners (ages 8-18) who made 119 chatbots. We examined how children configured their chatbots, reasoned about trustworthiness, and how closely chatbot behavior aligned with their designs. Younger students (age 10-13) set significantly higher confidence than older students (age 14-18), and some deliberately built chatbots that gave wrong answers on purpose, yet still called them trustworthy, arguing that a chatbot does what it was built to do. Younger students equated trust with purpose-fulfillment, while older students linked it to transparent, calibrated design. Students also calibrated academic chatbots to be more transparent and formal than hobby chatbots. We identify seven design dimensions describing what children believe makes a chatbot trustworthy, and discuss implications for AI literacy tools.

cs.HC↗

Toward Service-Balanced ISAC: From Coupled RAN to De-Coupled RAN

Sixth-generation (6G) applications require radio access networks (RANs) to support reliable communication and seamless sensing across their operating regions. In coupled RAN deployments, shared downlink transmitting and uplink receiving sites constrain the network's ability to accommodate asymmetric links and different sensing geometries. De-Coupled RAN (DC-RAN) separates these functions, allowing independently deployed and coordinated base stations to extend uplink and downlink communication and sensing coverage. How this flexibility translates into balanced communication and sensing services, however, remains insufficiently explored. This article revisits the evolution from coupled to DC-RAN from the perspective of service-balanced integrated sensing and communication (ISAC). It examines how architectural choices affect the availability of both services, with communication-sensing coverage symmetry capturing their spatial alignment under application-specific quality requirements. Practical challenges include preserving communication consistency, maintaining sensing continuity, and coordinating distributed resources. Two case studies illustrate how DC-RAN can support coverage symmetry alongside consistent communication, and how complementary observations can sustain continuous and accurate sensing. These examples inform a discussion of future research toward service-balanced ISAC.

eess.SP↗

LUNA: Luneburg-Lens-Aided Reconfigurable Array for 6G-and-Advanced Wireless Networks

This article introduces the LUneburg-lens-aided recoNfigurable Array (LUNA), an antenna architecture that unifies multiple-input multiple-output (MIMO) and network-controlled repeater (NCR) functionalities in the Luneburg lens-enabled hardware platform. A Luneburg lens, fabricated from graded-index dielectric materials, passively converts the radiation of a low-gain feed into a highly directional beam without active phase shifting, while a dense passive feed bank and a reconfigurable feed-selection network electronically switch the beam directions with minimal hardware complexity and power consumption. We commence by reviewing the basic principles and application history of Luneburg lenses in radar and wireless communications, which motivates their role in 6G-and-advanced networks. Then, we highlight how a Luneburg lens and a reconfigurable feed array construct both LUNA-MIMO and LUNA-NCR, where the lens and feed bank can be reused across functions and frequency bands. Case studies demonstrate that LUNA achieves the satisfactory spectral and energy efficiency with a few radio-frequency chains, and it also improves positioning performance for sensing tasks. Finally, some open problems and research directions are provided to inspire follow-up research on LUNA.

eess.SP↗

Beyond expressiveness in pairwise and higher-order models

The debate over pairwise and higher-order models is often framed as a contest of expressive power. This framing is misleading. A graph with arbitrary multivariate node functions can reproduce the node-level dynamics of many hypergraph models, but the hypergraph's grouping information does not thereby disappear: it may simply move from the structure to the rule. We separate four notions frequently conflated here: structural representation, functional representability, statistical identifiability, and mechanistic adequacy. We then prove that interaction order, the number of variables that must act jointly in some term of any additive decomposition of a node's update, is the same in every exact representation: no change of structural language can lower it. As a description-length problem, at the unrestricted algorithmic level a fixed compiler redistributes information between structure and rule at constant overhead, so expressiveness alone cannot privilege either language. Preferences arise only relative to explicit model classes, code families, and data. This yields an operational minimum-description-length criterion combining structural cost, conditional rule cost, and imperfect fit. Within it, for $M$ disjoint groups of size $k$, the clique projection's edge list is asymptotically $k-1$ times longer than the hyperedge list it replaces: the projection encodes the same grouping at higher cost. Examples from diffusion, Boolean dynamics, ecology, and ambiguous projections give graph-preferred, hypergraph-preferred, and unresolved cases; for bipartite and multilayer lifts, cost, fit, and identifiability tie, and the choice turns on which entities are posited as primitive. The position is symmetric: higher-order structure should not be inferred from phenomenology alone, nor does a graph's ability to emulate a system make it the most parsimonious or adequate description.

physics.soc-ph↗

Quaternionic LVM Manifolds

An admissible configuration $Λ\subset\Hh^m\simeq\Rr^{4m}$ defines a compact intersection of quadrics $Z_{\Hh}(Λ)$ and a free diagonal $\Sp(1)$ quotient $N_{\Hh}(Λ)$. We realize this quotient as the leaf space of a real dilation action on an open subset of quaternionic projective space. The total space is a quaternionic moment-angle polyhedral product. Every quotient is $2$-connected; when there are no indispensable coordinates, it is $3$-connected and its fourth integral cohomology is generated by the Euler class $e$ of the associated oriented rank-four bundle. We compute $p_1(TN_{\Hh})=2(n-2)e$ and prove that no LVMQ manifold admits a quaternionic Kähler metric. An explicit infinite family consists of products of spheres admitting both homogeneous positive Einstein metrics and quaternionic toric structures in the sense of Gentili--Gori--Sarfatti. We also describe the trace foliations of the real action in the Poincaré domain. For the separate rank-one distribution $X\Hh$, an invertible quaternionic-linear field is integrable only in the real scalar case. A nonlinear field with positive real scalar derivative at the origin and involutive $X\Hh$ has exactly the standard quaternionic Hopf foliation on every sufficiently small centered sphere.

math.DG↗

Rainbow spanning configurations in uniformly coloured pseudorandom graphs

We prove a quantitative palette-transference principle for rainbow spanning configurations in uniformly edge-coloured pseudorandom graphs. The input to our transference principle is an embedding result of a spanning configuration in an appropriately bijumbled graph $H$ with sufficiently large minimum degree. The output of our transference principle is the asymptotically almost sure existence of a rainbow copy of the same configuration in a uniformly edge-coloured graph $G$ whose bijumbledness and minimum are comparable and sometimes coincide with those of $H$. We then apply our transference principle in order to asymptotically almost surely obtain $K_k$-factors, including perfect matchings, Hamilton cycles, and a prescribed bounded-degree spanning tree in bijumbled graphs with appropriate parameters. In all of our results, the palette size exceeds the size of the target configuration by $\varepsilon n$, where $\varepsilon > 0$ is arbitrarily small yet fixed, and $n$ is the order of the configuration.

math.CO↗

A Szemerédi-Trotter Theorem in Arbitrary Fields

Let $k$ be a field of characteristic $p\ge0$. We prove that $m$ points and $n$ lines in $k^2$ determine at most $3(mn)^{2/3}+m+n+2mn/p$ incidences, the last term being omitted in characteristic zero. Over the prime field $\mathbb{F}_p$ the coefficient of $mn/p$ can be replaced by $1$. The proof uses the polynomial method, and for $m=n$ the bound is sharp up to an absolute constant over prime fields. As applications, over prime fields in which $-1$ is not a square we obtain the $L^2\to L^r$ extension estimate for the paraboloid in $\mathbb{F}_p^3$ for $r>10/3$. Over every odd prime field, we show that a two-source extractor construction of Bourgain has exponentially small error at every min-entropy rate greater than $1/3$. We also improve sum-product estimates for small sets in positive characteristic and obtain projection and Furstenberg estimates over prime fields. The incidence inequalities with exact constants have been formalized in Lean.

math.CO↗