Search arXiv⌕ Search

arXiv subjects

Max Tegmark

Publications and source records attributed to Max Tegmark.

At least 91 records · Page 5Linked to original sources

The HERA-19 Commissioning Array: Direction Dependent Effects

Foreground power dominates the measurements of interferometers that seek a statistical detection of highly-redshifted HI emission from the Epoch of Reionization (EoR). The chromaticity of the instrument creates a boundary in the Fourier transform of frequency (proportional to $k_\parallel$) between spectrally smooth emission, characteristic of the strong synchrotron foreground (the "wedge"), and the spectrally structured emission from HI in the EoR (the "EoR window"). Faraday rotation can inject spectral structure into otherwise smooth polarized foreground emission, which through instrument effects or miscalibration could possibly pollute the EoR window. Using data from the HERA 19-element commissioning array, we investigate the polarization response of this new instrument in the power spectrum domain. We perform a simple image-based calibration based on the unpolarized diffuse emission of the Global Sky Model, and show that it achieves qualitative redundancy between the nominally-redundant baselines of the array and reasonable amplitude accuracy. We construct power spectra of all fully polarized coherencies in all pseudo-Stokes parameters. We compare to simulations based on an unpolarized diffuse sky model and detailed electromagnetic simulations of the dish and feed, confirming that in Stokes I, the calibration does not add significant spectral structure beyond the expected level. Further, this calibration is stable over the 8 days of observations considered. Excess power is seen in the power spectra of the linear polarization Stokes parameters which is not easily attributable to leakage via the primary beam, and results from some combination of residual calibration errors and actual polarized emission. Stokes V is found to be highly discrepant from the expectation of zero power, strongly pointing to the need for more accurate polarized calibration.

astro-ph.IM↗

The role of artificial intelligence in achieving the Sustainable Development Goals

The emergence of artificial intelligence (AI) and its progressively wider impact on many sectors across the society requires an assessment of its effect on sustainable development. Here we analyze published evidence of positive or negative impacts of AI on the achievement of each of the 17 goals and 169 targets of the 2030 Agenda for Sustainable Development. We find that AI can support the achievement of 128 targets across all SDGs, but it may also inhibit 58 targets. Notably, AI enables new technologies that improve efficiency and productivity, but it may also lead to increased inequalities among and within countries, thus hindering the achievement of the 2030 Agenda. The fast development of AI needs to be supported by appropriate policy and regulation. Otherwise, it would lead to gaps in transparency, accountability, safety and ethical standards of AI-based technology, which could be detrimental towards the development and sustainable use of AI. Finally, there is a lack of research assessing the medium- and long-term impacts of AI. It is therefore essential to reinforce the global debate regarding the use of AI and to develop the necessary regulatory insight and oversight for AI-based technologies.

cs.CY↗

Latent Representations of Dynamical Systems: When Two is Better Than One

A popular approach for predicting the future of dynamical systems involves mapping them into a lower-dimensional "latent space" where prediction is easier. We show that the information-theoretically optimal approach uses different mappings for present and future, in contrast to state-of-the-art machine-learning approaches where both mappings are the same. We illustrate this dichotomy by predicting the time-evolution of coupled harmonic oscillators with dissipation and thermal noise, showing how the optimal 2-mapping method significantly outperforms principal component analysis and all other approaches that use a single latent representation, and discuss the intuitive reason why two representations are better than one. We conjecture that a single latent representation is optimal only for time-reversible processes, not for e.g. text, speech, music or out-of-equilibrium physical systems.

physics.data-an↗

Ensemble Inhibition and Excitation in the Human Cortex: an Ising Model Analysis with Uncertainties

The pairwise maximum entropy model, also known as the Ising model, has been widely used to analyze the collective activity of neurons. However, controversy persists in the literature about seemingly inconsistent findings, whose significance is unclear due to lack of reliable error estimates. We therefore develop a method for accurately estimating parameter uncertainty based on random walks in parameter space using adaptive Markov Chain Monte Carlo after the convergence of the main optimization algorithm. We apply our method to the spiking patterns of excitatory and inhibitory neurons recorded with multielectrode arrays in the human temporal cortex during the wake-sleep cycle. Our analysis shows that the Ising model captures neuronal collective behavior much better than the independent model during wakefulness, light sleep, and deep sleep when both excitatory (E) and inhibitory (I) neurons are modeled; ignoring the inhibitory effects of I-neurons dramatically overestimates synchrony among E-neurons. Furthermore, information-theoretic measures reveal that the Ising model explains about 80%-95% of the correlations, depending on sleep state and neuron type. Thermodynamic measures show signatures of criticality, although we take this with a grain of salt as it may be merely a reflection of long-range neural correlations.

cond-mat.dis-nn↗

Meta-learning autoencoders for few-shot prediction

Compared to humans, machine learning models generally require significantly more training examples and fail to extrapolate from experience to solve previously unseen challenges. To help close this performance gap, we augment single-task neural networks with a meta-recognition model which learns a succinct model code via its autoencoder structure, using just a few informative examples. The model code is then employed by a meta-generative model to construct parameters for the task-specific model. We demonstrate that for previously unseen tasks, without additional training, this Meta-Learning Autoencoder (MeLA) framework can build models that closely match the true underlying models, with loss significantly lower than given by fine-tuned baseline networks, and performance that compares favorably with state-of-the-art meta-learning algorithms. MeLA also adds the ability to identify influential training examples and predict which additional data will be most valuable to acquire to improve model prediction.

cs.LG↗

The power of deeper networks for expressing natural functions

It is well-known that neural networks are universal approximators, but that deeper networks tend in practice to be more powerful than shallower ones. We shed light on this by proving that the total number of neurons $m$ required to approximate natural classes of multivariate polynomials of $n$ variables grows only linearly with $n$ for deep neural networks, but grows exponentially when merely a single hidden layer is allowed. We also provide evidence that when the number of hidden layers is increased from $1$ to $k$, the neuron requirement grows exponentially not with $n$ but with $n^{1/k}$, suggesting that the minimum number of layers required for practical expressibility grows only logarithmically with $n$.

cs.LG↗

Gated Orthogonal Recurrent Units: On Learning to Forget

We present a novel recurrent neural network (RNN) based model that combines the remembering ability of unitary RNNs with the ability of gated RNNs to effectively forget redundant/irrelevant information in its memory. We achieve this by extending unitary RNNs with a gating mechanism. Our model is able to outperform LSTMs, GRUs and Unitary RNNs on several long-term dependency benchmark tasks. We empirically both show the orthogonal/unitary RNNs lack the ability to forget and also the ability of GORU to simultaneously remember long term dependencies while forgetting irrelevant information. This plays an important role in recurrent neural networks. We provide competitive results along with an analysis of our model on many natural sequential tasks including the bAbI Question Answering, TIMIT speech spectrum prediction, Penn TreeBank, and synthetic tasks that involve long-term dependencies such as algorithmic, parenthesis, denoising and copying tasks.

cs.LG↗

Nanophotonic Particle Simulation and Inverse Design Using Artificial Neural Networks

We propose a method to use artificial neural networks to approximate light scattering by multilayer nanoparticles. We find the network needs to be trained on only a small sampling of the data in order to approximate the simulation to high precision. Once the neural network is trained, it can simulate such optical processes orders of magnitude faster than conventional simulations. Furthermore, the trained neural network can be used solve nanophotonic inverse design problems by using back- propogation - where the gradient is analytical, not numerical.

physics.comp-ph↗

Criticality in Formal Languages and Statistical Physics

We show that the mutual information between two symbols, as a function of the number of symbols between the two, decays exponentially in any probabilistic regular grammar, but can decay like a power law for a context-free grammar. This result about formal languages is closely related to a well-known result in classical statistical mechanics that there are no phase transitions in dimensions fewer than two. It is also related to the emergence of power-law correlations in turbulence and cosmological inflation through recursive generative processes. We elucidate these physics connections and comment on potential applications of our results to machine learning tasks like training artificial recurrent neural networks. Along the way, we introduce a useful quantity which we dub the rational mutual information and discuss generalizations of our claims involving more complicated Bayesian networks.

cond-mat.dis-nn↗

Why does deep and cheap learning work so well?

We show how the success of deep learning could depend not only on mathematics but also on physics: although well-known mathematical theorems guarantee that neural networks can approximate arbitrary functions well, the class of functions of practical interest can frequently be approximated through "cheap learning" with exponentially fewer parameters than generic ones. We explore how properties frequently encountered in physics such as symmetry, locality, compositionality, and polynomial log-probability translate into exceptionally simple neural networks. We further argue that when the statistical process generating the data is of a certain hierarchical form prevalent in physics and machine-learning, a deep neural network can be more efficient than a shallow one. We formalize these claims using information theory and discuss the relation to the renormalization group. We prove various "no-flattening theorems" showing when efficient linear deep networks cannot be accurately approximated by shallow ones without efficiency loss, for example, we show that $n$ variables cannot be multiplied using fewer than 2^n neurons in a single hidden layer.

cond-mat.dis-nn↗

Tunable Efficient Unitary Neural Networks (EUNN) and their application to RNNs

Using unitary (instead of general) matrices in artificial neural networks (ANNs) is a promising way to solve the gradient explosion/vanishing problem, as well as to enable ANNs to learn long-term correlations in the data. This approach appears particularly promising for Recurrent Neural Networks (RNNs). In this work, we present a new architecture for implementing an Efficient Unitary Neural Network (EUNNs); its main advantages can be summarized as follows. Firstly, the representation capacity of the unitary space in an EUNN is fully tunable, ranging from a subspace of SU(N) to the entire unitary space. Secondly, the computational complexity for training an EUNN is merely $\mathcal{O}(1)$ per parameter. Finally, we test the performance of EUNNs on the standard copying task, the pixel-permuted MNIST digit recognition benchmark as well as the Speech Prediction Test (TIMIT). We find that our architecture significantly outperforms both other state-of-the-art unitary RNNs and the LSTM architecture, in terms of the final performance and/or the wall-clock training speed. EUNNs are thus promising alternatives to RNNs and LSTMs for a wide variety of applications.

cs.LG↗

On the Impossibility of Supersized Machines

In recent years, a number of prominent computer scientists, along with academics in fields such as philosophy and physics, have lent credence to the notion that machines may one day become as large as humans. Many have further argued that machines could even come to exceed human size by a significant margin. However, there are at least seven distinct arguments that preclude this outcome. We show that it is not only implausible that machines will ever exceed human size, but in fact impossible.

cs.CY↗

The Hydrogen Epoch of Reionization Array Dish III: Measuring Chromaticity of Prototype Element with Reflectometry

The experimental efforts to detect the redshifted 21 cm signal from the Epoch of Reionization (EoR) are limited predominantly by the chromatic instrumental systematic effect. The delay spectrum methodology for 21 cm power spectrum measurements brought new attention to the critical impact of an antenna's chromaticity on the viability of making this measurement. This methodology established a straightforward relationship between time-domain response of an instrument and the power spectrum modes accessible to a 21 cm EoR experiment. We examine the performance of a prototype of the Hydrogen Epoch of Reionization Array (HERA) array element that is currently observing in Karoo desert, South Africa. We present a mathematical framework to derive the beam integrated frequency response of a HERA prototype element in reception from the return loss measurements between 100-200 MHz and determined the extent of additional foreground contamination in the delay space. The measurement reveals excess spectral structures in comparison to the simulation studies of the HERA element. Combined with the HERA data analysis pipeline that incorporates inverse covariance weighting in optimal quadratic estimation of power spectrum, we find that in spite of its departure from the simulated response, HERA prototype element satisfies the necessary criteria posed by the foreground attenuation limits and potentially can measure the power spectrum at spatial modes as low as $k_{\parallel} > 0.1h$~Mpc$^{-1}$. The work highlights a straightforward method for directly measuring an instrument response and assessing its impact on 21 cm EoR power spectrum measurements for future experiments that will use reflector-type antenna.

astro-ph.IM↗

The Hydrogen Epoch of Reionization Array Dish I: Beam Pattern Measurements and Science Implications

The Hydrogen Epoch of Reionization Array (HERA) is a radio interferometer aiming to detect the power spectrum of 21 cm fluctuations from neutral hydrogen from the Epoch of Reionization (EOR). Drawing on lessons from the Murchison Widefield Array (MWA) and the Precision Array for Probing the Epoch of Reionization (PAPER), HERA is a hexagonal array of large (14 m diameter) dishes with suspended dipole feeds. Not only does the dish determine overall sensitivity, it affects the observed frequency structure of foregrounds in the interferometer. This is the first of a series of four papers characterizing the frequency and angular response of the dish with simulations and measurements. We focus in this paper on the angular response (i.e., power pattern), which sets the relative weighting between sky regions of high and low delay, and thus, apparent source frequency structure. We measure the angular response at 137 MHz using the ORBCOMM beam mapping system of Neben et al. We measure a collecting area of 93 m^2 in the optimal dish/feed configuration, implying HERA-320 should detect the EOR power spectrum at z~9 with a signal-to-noise ratio of 12.7 using a foreground avoidance approach with a single season of observations, and 74.3 using a foreground subtraction approach. Lastly we study the impact of these beam measurements on the distribution of foregrounds in Fourier space.

astro-ph.IM↗

Improved Measures of Integrated Information

Although there is growing interest in measuring integrated information in computational and cognitive systems, current methods for doing so in practice are computationally unfeasible. Existing and novel integration measures are investigated and classified by various desirable properties. A simple taxonomy of Phi-measures is presented where they are each characterized by their choice of factorization method (5 options), choice of probability distributions to compare (3x4 options) and choice of measure for comparing probability distributions (7 options). When requiring the Phi-measures to satisfy a minimum of attractive properties, these hundreds of options reduce to a mere handful, some of which turn out to be identical. Useful exact and approximate formulas are derived that can be applied to real-world data from laboratory experiments without posing unreasonable computational demands.

q-bio.NC↗

An Improved Model of Diffuse Galactic Radio Emission from 10 MHz to 5 THz

We present an improved Global Sky Model (GSM) of diffuse galactic radio emission from 10 MHz to 5 THz, whose uses include foreground modeling for CMB and 21 cm cosmology. Our model improves on past work both algorithmically and by adding new data sets such as the Planck maps and the enhanced Haslam map. Our method generalizes the Principal Component Analysis approach to handle non-overlapping regions, enabling the inclusion of 29 sky maps with no region of the sky common to all. We also perform a blind separation of our GSM into physical components with a method that makes no assumptions about physical emission mechanisms (synchrotron, free-free, dust, etc). Remarkably, this blind method automatically finds five components that have previously only been found "by hand", which we identify with synchrotron, free-free, cold dust, warm dust, and the CMB anisotropy, with maps and spectra agreeing with previous work but in many cases with smaller error bars. The improved GSM is available online at github.com/jeffzhen/gsm2016.

astro-ph.CO↗

The Hydrogen Epoch of Reionization Array Dish II: Characterization of Spectral Structure with Electromagnetic Simulations and its science Implications

We use time-domain electromagnetic simulations to determine the spectral characteristics of the Hydrogen Epoch of Reionization Arrays (HERA) antenna. These simulations are part of a multi-faceted campaign to determine the effectiveness of the dish's design for obtaining a detection of redshifted 21 cm emission from the epoch of reionization. Our simulations show the existence of reflections between HERA's suspended feed and its parabolic dish reflector that fall below -40 dB at 150 ns and, for reasonable impedance matches, have a negligible impact on HERA's ability to constrain EoR parameters. It follows that despite the reflections they introduce, dishes are effective for increasing the sensitivity of EoR experiments at relatively low cost. We find that electromagnetic resonances in the HERA feed's cylindrical skirt, which is intended to reduce cross coupling and beam ellipticity, introduces significant power at large delays ($-40$ dB at 200 ns) which can lead to some loss of measurable Fourier modes and a modest reduction in sensitivity. Even in the presence of this structure, we find that the spectral response of the antenna is sufficiently smooth for delay filtering to contain foreground emission at line-of-sight wave numbers below $k_\parallel \lesssim 0.2$ $h$Mpc$^{-1}$, in the region where the current PAPER experiment operates. Incorporating these results into a Fisher Matrix analysis, we find that the spectral structure observed in our simulations has only a small effect on the tight constraints HERA can achieve on parameters associated with the astrophysics of reionization.

astro-ph.CO↗

Hydrogen Epoch of Reionization Array (HERA)

The Hydrogen Epoch of Reionization Array (HERA) is a staged experiment to measure 21 cm emission from the primordial intergalactic medium (IGM) throughout cosmic reionization ($z=6-12$), and to explore earlier epochs of our Cosmic Dawn ($z\sim30$). During these epochs, early stars and black holes heated and ionized the IGM, introducing fluctuations in 21 cm emission. HERA is designed to characterize the evolution of the 21 cm power spectrum to constrain the timing and morphology of reionization, the properties of the first galaxies, the evolution of large-scale structure, and the early sources of heating. The full HERA instrument will be a 350-element interferometer in South Africa consisting of 14-m parabolic dishes observing from 50 to 250 MHz. Currently, 19 dishes have been deployed on site and the next 18 are under construction. HERA has been designated as an SKA Precursor instrument. In this paper, we summarize HERA's scientific context and provide forecasts for its key science results. After reviewing the current state of the art in foreground mitigation, we use the delay-spectrum technique to motivate high-level performance requirements for the HERA instrument. Next, we present the HERA instrument design, along with the subsystem specifications that ensure that HERA meets its performance requirements. Finally, we summarize the schedule and status of the project. We conclude by suggesting that, given the realities of foreground contamination, current-generation 21 cm instruments are approaching their sensitivity limits. HERA is designed to bring both the sensitivity and the precision to deliver its primary science on the basis of proven foreground filtering techniques, while developing new subtraction techniques to unlock new capabilities. The result will be a major step toward realizing the widely recognized scientific potential of 21 cm cosmology.

astro-ph.IM↗