Search arXivSearch

arXiv · 2508.07115

Sensory robustness through top-down feedback and neural stochasticity in recurrent vision models

Abstract

Biological systems leverage top-down feedback for visual processing, yet most artificial vision models succeed in image classification using purely feedforward or recurrent architectures, calling into question the functional significance of descending cortical pathways. Here, we trained convolutional recurrent neural networks (ConvRNN) on image classification in the presence or absence of top-down feedback projections to elucidate the specific computational contributions of those feedback pathways. We found that ConvRNNs with top-down feedback exhibited remarkable speed-accuracy trade-off and robustness to noise perturbations and adversarial attacks, but only when they were trained with stochastic neural variability, simulated by randomly silencing single units via dropout. By performing detailed analyses to identify the reasons for such benefits, we observed that feedback information substantially shaped the representational geometry of the post-integration layer, combining the bottom-up and top-down streams, and this effect was amplified by dropout. Moreover, feedback signals coupled with dropout optimally constrained network activity onto a low-dimensional manifold and encoded object information more efficiently in out-of-distribution regimes, with top-down information stabilizing the representational dynamics at the population level. Together, these findings uncover a dual mechanism for resilient sensory coding. On the one hand, neural stochasticity prevents unit-level co-adaptation albeit at the cost of more chaotic dynamics. On the other hand, top-down feedback harnesses high-level information to stabilize network activity on compact low-dimensional manifolds.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Antonino Greco, Marco D'Alessandro, Karl J. Friston, Giovanni Pezzulo, Markus Siegel. 2025-08-09. Sensory robustness through top-down feedback and neural stochasticity in recurrent vision models. https://arxiv.org/abs/2508.07115

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Differences in Neurovascular Coupling in Patients with Major Depressive Disorder: Evidence from Simultaneous Resting-State EEG-fNIRS

Neurovascular coupling (NVC), the relationship between neural activity and cerebral hemodynamic responses, remains poorly understood in major depressive disorder (MDD). To investigate alterations in NVC associated with depressive symptom severity, we simultaneously recorded resting-state electroencephalography (rsEEG) and functional near-infrared spectroscopy (fNIRS) in 206 participants, including 134 patients with MDD and 72 healthy controls stratified by age to disentangle disease-related alterations from age-associated effects. NVC in the prefrontal cortex (PFC) was characterized by the consistency and temporal lag between spontaneous electrophysiological peaks and corresponding hemodynamic responses. We found that age significantly enhances the NVC consistency (p < 0.05), whereas this relationship was altered in patients with MDD during both the oxygen-consumption phase and the subsequent hemodynamic compensation phase. Furthermore, among patients with MDD, the strength of this consistency decreased more significantly with greater illness severity (partial r = -0.336, p = 0.0601). These findings suggest that changes in neurovascular coupling effects are strongly associated with the severity of depression. By leveraging wearable neuroimaging techniques, this study provides multimodal evidence for altered neurovascular coupling in depression and highlights its potential as a biomarker for disease monitoring and recovery trajectories.

q-bio.NC

Critical Flicker Fusion Frequency As An Experience-Restricted Constraint On Visual Temporal Resolution: What Does And Does Not Change It

Experience-dependent plasticity is fundamental to adaptive behaviour, yet the conditions under which basic sensory timing can be modified in adulthood remain poorly specified. Critical flicker fusion frequency (CFFF), the threshold at which flicker is perceived as continuous, is unusually informative here: what fails to change it is as well documented as what does. Reviewing the stability and training literature, we argue CFFF is best characterised neither as non-plastic nor as held near a physiological ceiling, but as modifiable only by a specific class of experience, not by amount. Repeated testing, cognitive training without temporal content, and incoherent-flicker exposure leave the threshold unchanged regardless of duration. By contrast, a narrow class of perceptual-learning paradigms -- pairing coherent directional motion with a task-relevant target -- reportedly raises the threshold substantially, with gains retained at one year in a small subsample. Gains are largest below typical values, as in amblyopia, and absent in normally sighted observers under the same protocol. Three qualifications apply: the evidence rests on small samples; the threshold-raising paradigms have been assessed almost exclusively with heterochromatic flicker photometry rather than luminance-defined CFFF, leaving construct equivalence unestablished; and a minority of untrained controls show comparable changes. We examine whether the restriction originates at peripheral, thalamocortical, or cortical levels; available data do not adjudicate between them. Functional arguments for why such a restriction might be adaptive are offered as rationales, not evidence. Proposed links to working-memory precision and metacognition are stated as predictions; the first test of them was largely negative.

q-bio.NC

AI-Driven Neural Surrogates for In Silico Design of Cognitive-Affective Neuromodulation Targets

In neuropsychiatry, the primary goal is often not only to decode brain activity but to change it, for example to lessen a negative affective bias or an overly salient memory. Motivated by control theory, we develop an AI-driven neural-surrogate framework that proposes candidate representational changes and tests their predicted perceptual effects from snapshots of stimulus-evoked fMRI activity, without physical stimulation. The framework combines fMRI decoding, deep generative modeling, and constrained latent-space steering. Valence and memorability are used only as worked examples. Using more than 36,000 image-fMRI observations from four deeply sampled Natural Scenes Dataset participants, subject-specific models recovered coarse generative structure from visually responsive cortex (two-way identification, 0.79-0.88; chance, 0.5). Graded perturbations were reconstructed as images and evaluated with automated scorers and human ratings from 7,200 trials by 18 participants. In the primary VDVAE model, valence shifted from -0.61 to +1.03 SD and memorability from -1.34 to +1.45 SD; a later Versatile Diffusion refinement reduced or altered these effects. Across five perturbation levels, human valence ratings moved in the predicted direction under the linear time-correction model (mean slope, 0.038 SD per unit of alpha; 95 percent CI, 0.003-0.074; positive in 16 of 18 participants). Perceived memorability did not change reliably. Baseline agreement with the automated assessor was suggestive for valence (r = 0.30) and weak for memorability (r = 0.10). Extreme perturbations drifted from the original stimulus, so intended change must be weighed against loss of fidelity. These findings provide a falsifiable upstream method for designing and behaviorally testing candidate representational targets for future neuromodulation in psychiatry, while marking the limits of the present static approximation.

q-bio.NC