Search arXiv⌕ Search

arXiv · 1606.04698

Invariant recognition drives neural representations of action sequences

Abstract

Recognizing the actions of others from visual stimuli is a crucial aspect of human visual perception that allows individuals to respond to social cues. Humans are able to identify similar behaviors and discriminate between distinct actions despite transformations, like changes in viewpoint or actor, that substantially alter the visual appearance of a scene. This ability to generalize across complex transformations is a hallmark of human visual intelligence. Advances in understanding motion perception at the neural level have not always translated in precise accounts of the computational principles underlying what representation our visual cortex evolved or learned to compute. Here we test the hypothesis that invariant action discrimination might fill this gap. Recently, the study of artificial systems for static object perception has produced models, CNNs, that achieve human level performance in complex discriminative tasks. Within this class of models, architectures that better support invariant object recognition also produce image representations that match those implied by human and primate neural data. However, whether these models produce representations of action sequences that support recognition across complex transformations and closely follow neural representations remains unknown. Here we show that spatiotemporal CNNs appropriately categorize video stimuli into actions, and that deliberate model modifications that improve performance on an invariant action recognition task lead to data representations that better match human neural recordings. Our results support our hypothesis that performance on invariant discrimination dictates the neural representations of actions computed by human visual cortex.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Andrea Tacchetti, Leyla Isik, Tomaso Poggio. 2017-04-20. Invariant recognition drives neural representations of action sequences. https://doi.org/10.1371/journal.pcbi.1005859

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Differences in Neurovascular Coupling in Patients with Major Depressive Disorder: Evidence from Simultaneous Resting-State EEG-fNIRS

Neurovascular coupling (NVC), the relationship between neural activity and cerebral hemodynamic responses, remains poorly understood in major depressive disorder (MDD). To investigate alterations in NVC associated with depressive symptom severity, we simultaneously recorded resting-state electroencephalography (rsEEG) and functional near-infrared spectroscopy (fNIRS) in 206 participants, including 134 patients with MDD and 72 healthy controls stratified by age to disentangle disease-related alterations from age-associated effects. NVC in the prefrontal cortex (PFC) was characterized by the consistency and temporal lag between spontaneous electrophysiological peaks and corresponding hemodynamic responses. We found that age significantly enhances the NVC consistency (p < 0.05), whereas this relationship was altered in patients with MDD during both the oxygen-consumption phase and the subsequent hemodynamic compensation phase. Furthermore, among patients with MDD, the strength of this consistency decreased more significantly with greater illness severity (partial r = -0.336, p = 0.0601). These findings suggest that changes in neurovascular coupling effects are strongly associated with the severity of depression. By leveraging wearable neuroimaging techniques, this study provides multimodal evidence for altered neurovascular coupling in depression and highlights its potential as a biomarker for disease monitoring and recovery trajectories.

q-bio.NC↗

Critical Flicker Fusion Frequency As An Experience-Restricted Constraint On Visual Temporal Resolution: What Does And Does Not Change It

Experience-dependent plasticity is fundamental to adaptive behaviour, yet the conditions under which basic sensory timing can be modified in adulthood remain poorly specified. Critical flicker fusion frequency (CFFF), the threshold at which flicker is perceived as continuous, is unusually informative here: what fails to change it is as well documented as what does. Reviewing the stability and training literature, we argue CFFF is best characterised neither as non-plastic nor as held near a physiological ceiling, but as modifiable only by a specific class of experience, not by amount. Repeated testing, cognitive training without temporal content, and incoherent-flicker exposure leave the threshold unchanged regardless of duration. By contrast, a narrow class of perceptual-learning paradigms -- pairing coherent directional motion with a task-relevant target -- reportedly raises the threshold substantially, with gains retained at one year in a small subsample. Gains are largest below typical values, as in amblyopia, and absent in normally sighted observers under the same protocol. Three qualifications apply: the evidence rests on small samples; the threshold-raising paradigms have been assessed almost exclusively with heterochromatic flicker photometry rather than luminance-defined CFFF, leaving construct equivalence unestablished; and a minority of untrained controls show comparable changes. We examine whether the restriction originates at peripheral, thalamocortical, or cortical levels; available data do not adjudicate between them. Functional arguments for why such a restriction might be adaptive are offered as rationales, not evidence. Proposed links to working-memory precision and metacognition are stated as predictions; the first test of them was largely negative.

q-bio.NC↗

AI-Driven Neural Surrogates for In Silico Design of Cognitive-Affective Neuromodulation Targets

In neuropsychiatry, the primary goal is often not only to decode brain activity but to change it, for example to lessen a negative affective bias or an overly salient memory. Motivated by control theory, we develop an AI-driven neural-surrogate framework that proposes candidate representational changes and tests their predicted perceptual effects from snapshots of stimulus-evoked fMRI activity, without physical stimulation. The framework combines fMRI decoding, deep generative modeling, and constrained latent-space steering. Valence and memorability are used only as worked examples. Using more than 36,000 image-fMRI observations from four deeply sampled Natural Scenes Dataset participants, subject-specific models recovered coarse generative structure from visually responsive cortex (two-way identification, 0.79-0.88; chance, 0.5). Graded perturbations were reconstructed as images and evaluated with automated scorers and human ratings from 7,200 trials by 18 participants. In the primary VDVAE model, valence shifted from -0.61 to +1.03 SD and memorability from -1.34 to +1.45 SD; a later Versatile Diffusion refinement reduced or altered these effects. Across five perturbation levels, human valence ratings moved in the predicted direction under the linear time-correction model (mean slope, 0.038 SD per unit of alpha; 95 percent CI, 0.003-0.074; positive in 16 of 18 participants). Perceived memorability did not change reliably. Baseline agreement with the automated assessor was suggestive for valence (r = 0.30) and weak for memorability (r = 0.10). Extreme perturbations drifted from the original stimulus, so intended change must be weighed against loss of fidelity. These findings provide a falsifiable upstream method for designing and behaviorally testing candidate representational targets for future neuromodulation in psychiatry, while marking the limits of the present static approximation.

q-bio.NC↗