Search arXiv⌕ Search

arXiv · 2009.04083

Generalizing Complex/Hyper-complex Convolutions to Vector Map Convolutions

Abstract

We show that the core reasons that complex and hypercomplex valued neural networks offer improvements over their real-valued counterparts is the weight sharing mechanism and treating multidimensional data as a single entity. Their algebra linearly combines the dimensions, making each dimension related to the others. However, both are constrained to a set number of dimensions, two for complex and four for quaternions. Here we introduce novel vector map convolutions which capture both of these properties provided by complex/hypercomplex convolutions, while dropping the unnatural dimensionality constraints they impose. This is achieved by introducing a system that mimics the unique linear combination of input dimensions, such as the Hamilton product for quaternions. We perform three experiments to show that these novel vector map convolutions seem to capture all the benefits of complex and hyper-complex networks, such as their ability to capture internal latent relations, while avoiding the dimensionality restriction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chase J Gaudet, Anthony S Maida. 2020-09-09. Generalizing Complex/Hyper-complex Convolutions to Vector Map Convolutions. https://arxiv.org/abs/2009.04083

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Structured Bayesian Modeling of Dynamic Receptive1 Fields in Salamander Retinal Ganglion Cells

Neurons in the visual system are selective for specific spatial and temporal stimulus features, described by their \emph{receptive field}. Estimating one means a coefficient per pixel per time bin from few trials -- a high-dimensional problem requiring regularization. Sparse regularizers such as the LASSO handle the dimension but select pixels independently at each time point, with nothing to keep the region coherent in space or smooth in time; it can fragment or reorganize discontinuously even when the true response evolves smoothly, a failure since this evolving pattern is what a receptive-field estimate should capture. We formulate dynamic receptive-field estimation as a high-dimensional Bayesian problem: a Poisson model combining a Gaussian Markov random field in space with an autoregressive process in time, so the estimated field is smooth and coherent across space and time. On recordings from $155$ salamander retinal ganglion cells, fitting this model independently per neuron recovers a coherent surface, where a pixel-level Poisson-LASSO comparison instead returns a fragmented one. Summarizing each neuron's surface by its space-averaged temporal response and clustering these curves with a model-based functional-clustering procedure, BIC selects three balanced temporal-response phenotypes ($85$, $32$, $38$ neurons), against a degenerate grouping from clustering the raw surfaces. A simulation study with known ground truth confirms the same pattern, with the model beating an unregularized Poisson GLM, LASSO, and the elastic net on recovery and estimation accuracy, though LASSO controls false positives better. The per-neuron field identification, its contrast with LASSO, and the functional-clustering population typing constitute this paper's contribution.

cs.NE↗

Landscape Limits of Quantum-Inspired Evolutionary Optimization across 256 continuous functions

Quantum-inspired evolutionary optimization (QIEO) represents design variables as a set of qubits and searches a continuous, multi-dimensional landscape through rotation of the qubit's amplitude pair. Every generation rotates those amplitudes toward a single elite, which corresponds to that generation's best. The update is cheap, almost parameter-free, and well-suited for massive parallel implementation, which has encouraged its adoption in engineering, design, and planning applications. However, there are critical issues with this formulation, principally, the treatment of design variables as independent probability components which make it incapable of exploiting local curvature, anisotropy, or variable coupling. Despite this, QIEO is believed to hold promise, and has been used extensively to solve real-world problems, with significant qualitative and computational advantage over its classical counterpart, Genetic Algorithm (GA). A collection of 256 (actually 508; 256 unshifted + 252 shifted, 4 could not be shifted) continuous function are selected from the prior works, in such a way that they represent eleven landscape characteristics, namely continuity, differentiability, separability, scalability, modality, convexity, conditioning, symmetry, maximum dimensionality, dimension dependency, and the coupling pattern of the design variables. These functions are then solved by three QIEO variants, two GA encodings and Hansen's Covariance Matrix Adaptation Evolution Strategy (CMA-ES). The results are evaluated in terms of computational cost, solution precision, and specialization across landscape characteristics. They identify the conditions under which QIEO provides competitive performance, clarify where its independent-variable representation becomes limiting, and establish whether particular QIEO variants offer advantages for specific landscape characteristics.

cs.NE↗

Context-sensitive neocortical neurons transform the effectiveness and efficiency of neural information processing

Deep learning (DL) has big-data processing capabilities that are as good, or even better, than those of humans in many real-world domains, but at the cost of high energy requirements that may be unsustainable in some applications and of errors, that, though infrequent, can be large. We hypothesise that a fundamental weakness of DL lies in its intrinsic dependence on integrate-and-fire point neurons that maximise information transmission irrespective of whether it is relevant in the current context or not. This leads to unnecessary neural firing and to the feedforward transmission of conflicting messages, which makes learning difficult and processing energy inefficient. Here we show how to circumvent these limitations by mimicking the capabilities of context-sensitive neocortical neurons that receive input from diverse sources as a context to amplify and attenuate the transmission of relevant and irrelevant information, respectively. We demonstrate that a deep network composed of such local processors seeks to maximise agreement between the active neurons, thus restricting the transmission of conflicting information to higher levels and reducing the neural activity required to process large amounts of heterogeneous real-world data. As shown to be far more effective and efficient than current forms of DL, this two-point neuron study offers a possible step-change in transforming the cellular foundations of deep network architectures.

cs.NE↗