Search arXiv⌕ Search

arXiv · 0709.3587

Self-organizing maps and symbolic data

Abstract

In data analysis new forms of complex data have to be considered like for example (symbolic data, functional data, web data, trees, SQL query and multimedia data, ...). In this context classical data analysis for knowledge discovery based on calculating the center of gravity can not be used because input are not $\mathbb{R}^p$ vectors. In this paper, we present an application on real world symbolic data using the self-organizing map. To this end, we propose an extension of the self-organizing map that can handle symbolic data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aïcha El Golli, Brieuc Conan-Guez, Fabrice Rossi. 2007-09-22. Self-organizing maps and symbolic data. https://arxiv.org/abs/0709.3587

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Structured Bayesian Modeling of Dynamic Receptive1 Fields in Salamander Retinal Ganglion Cells

Neurons in the visual system are selective for specific spatial and temporal stimulus features, described by their \emph{receptive field}. Estimating one means a coefficient per pixel per time bin from few trials -- a high-dimensional problem requiring regularization. Sparse regularizers such as the LASSO handle the dimension but select pixels independently at each time point, with nothing to keep the region coherent in space or smooth in time; it can fragment or reorganize discontinuously even when the true response evolves smoothly, a failure since this evolving pattern is what a receptive-field estimate should capture. We formulate dynamic receptive-field estimation as a high-dimensional Bayesian problem: a Poisson model combining a Gaussian Markov random field in space with an autoregressive process in time, so the estimated field is smooth and coherent across space and time. On recordings from $155$ salamander retinal ganglion cells, fitting this model independently per neuron recovers a coherent surface, where a pixel-level Poisson-LASSO comparison instead returns a fragmented one. Summarizing each neuron's surface by its space-averaged temporal response and clustering these curves with a model-based functional-clustering procedure, BIC selects three balanced temporal-response phenotypes ($85$, $32$, $38$ neurons), against a degenerate grouping from clustering the raw surfaces. A simulation study with known ground truth confirms the same pattern, with the model beating an unregularized Poisson GLM, LASSO, and the elastic net on recovery and estimation accuracy, though LASSO controls false positives better. The per-neuron field identification, its contrast with LASSO, and the functional-clustering population typing constitute this paper's contribution.

cs.NE↗

Landscape Limits of Quantum-Inspired Evolutionary Optimization across 256 continuous functions

Quantum-inspired evolutionary optimization (QIEO) represents design variables as a set of qubits and searches a continuous, multi-dimensional landscape through rotation of the qubit's amplitude pair. Every generation rotates those amplitudes toward a single elite, which corresponds to that generation's best. The update is cheap, almost parameter-free, and well-suited for massive parallel implementation, which has encouraged its adoption in engineering, design, and planning applications. However, there are critical issues with this formulation, principally, the treatment of design variables as independent probability components which make it incapable of exploiting local curvature, anisotropy, or variable coupling. Despite this, QIEO is believed to hold promise, and has been used extensively to solve real-world problems, with significant qualitative and computational advantage over its classical counterpart, Genetic Algorithm (GA). A collection of 256 (actually 508; 256 unshifted + 252 shifted, 4 could not be shifted) continuous function are selected from the prior works, in such a way that they represent eleven landscape characteristics, namely continuity, differentiability, separability, scalability, modality, convexity, conditioning, symmetry, maximum dimensionality, dimension dependency, and the coupling pattern of the design variables. These functions are then solved by three QIEO variants, two GA encodings and Hansen's Covariance Matrix Adaptation Evolution Strategy (CMA-ES). The results are evaluated in terms of computational cost, solution precision, and specialization across landscape characteristics. They identify the conditions under which QIEO provides competitive performance, clarify where its independent-variable representation becomes limiting, and establish whether particular QIEO variants offer advantages for specific landscape characteristics.

cs.NE↗

Context-sensitive neocortical neurons transform the effectiveness and efficiency of neural information processing

Deep learning (DL) has big-data processing capabilities that are as good, or even better, than those of humans in many real-world domains, but at the cost of high energy requirements that may be unsustainable in some applications and of errors, that, though infrequent, can be large. We hypothesise that a fundamental weakness of DL lies in its intrinsic dependence on integrate-and-fire point neurons that maximise information transmission irrespective of whether it is relevant in the current context or not. This leads to unnecessary neural firing and to the feedforward transmission of conflicting messages, which makes learning difficult and processing energy inefficient. Here we show how to circumvent these limitations by mimicking the capabilities of context-sensitive neocortical neurons that receive input from diverse sources as a context to amplify and attenuate the transmission of relevant and irrelevant information, respectively. We demonstrate that a deep network composed of such local processors seeks to maximise agreement between the active neurons, thus restricting the transmission of conflicting information to higher levels and reducing the neural activity required to process large amounts of heterogeneous real-world data. As shown to be far more effective and efficient than current forms of DL, this two-point neuron study offers a possible step-change in transforming the cellular foundations of deep network architectures.

cs.NE↗