Search arXivSearch

arXiv subjects

Jonathan Baxter

Publications and source records attributed to Jonathan Baxter.

At least 19 recordsLinked to original sources

The Future-Record Model of Everettian Probability

I recently showed how objective Everettian probability can be defined over the records borne by an observer's later continuations, without unique actualization or additional premeasurement subjects \citep{Baxter2026}. Here I extend that model to all decoherent physical record states, including records that support no such continuation. The observer-indexed model is then a restriction of an exhaustive physical-record model.

quant-ph

Does Probability Require a Single History?

Everettian quantum mechanics is often said to lack coherent probability because every possible measurement result occurs. If all of them exist globally, what is probability supposed to be a probability of? This paper defends an axiomatic answer: normalized Born measure is postulated as objective probability over the records borne by an observer's future continuations. A collapse law assigns the same probabilities to possible global outcomes. Neither theory derives the probabilities from its ontology. I relate this minimal proposal to earlier axiomatic accounts, separate it from their accompanying theories of personal persistence, and defend it against recent arguments that the axiomatic route cannot succeed. At the level of recorded results, the Everettian law supports the same predictions and empirical inferences as the collapse law, despite the theories' different ontologies.

quant-ph

Does the Actuality of Life Favor Many Actual Histories?

This essay develops a simple conditional argument for many-history ontologies from the fact that life exists. The question is not why our exact evolutionary sequence occurred, but why reality contains the broad class of life-bearing histories at all. If such histories are sufficiently rare, then a single-history ontology makes the existence of life surprising, while an ontology containing many histories can make it unsurprising. In the limiting case, life may be vanishingly unlikely in any one history while almost certain to occur somewhere if many histories are actualized. Everettian quantum mechanics is a natural candidate for this kind of ontology because its multiplicity is not introduced for anthropic purposes, but arises from taking universal unitary quantum mechanics seriously. The Fermi paradox is discussed as pressure against the claim that life is common within a single history.

physics.hist-ph

Point Charges in Classical Electrodynamics

LaTeX transcription (2025) of a 1989 honours thesis (University of Adelaide) on point charges in classical electrodynamics and the Lorentz-Dirac radiation-reaction equation. The thesis reviews the retarded field of an arbitrarily moving charge, energy-momentum conservation, and derives the Lorentz-Dirac equation via momentum balance. It discusses self-interaction and mass renormalization, and presents world-line self-field definitions including retarded averaging and an analytic continuation approach. Appendices include Mathematica listings used to obtain near-world-line expansions of the field and related quantities.

physics.class-ph

Reinforcement Learning From State and Temporal Differences

TD($\lambda$) with function approximation has proved empirically successful for some complex reinforcement learning problems. For linear approximation, TD($\lambda$) has been shown to minimise the squared error between the approximate value of each state and the true value. However, as far as policy is concerned, it is error in the relative ordering of states that is critical, rather than error in the state values. We illustrate this point, both in simple two-state and three-state systems in which TD($\lambda$)--starting from an optimal policy--converges to a sub-optimal policy, and also in backgammon. We then present a modified form of TD($\lambda$), called STD($\lambda$), in which function approximators are trained with respect to relative state values on binary decision problems. A theoretical analysis, including a proof of monotonic policy improvement for STD($\lambda$) in the context of the two-state system, is presented, along with a comparison with Bertsekas' differential training method [1]. This is followed by successful demonstrations of STD($\lambda$) on the two-state system and a variation on the well known acrobot problem.

cs.LG

A result relating convex n-widths to covering numbers with some applications to neural networks

In general, approximating classes of functions defined over high-dimensional input spaces by linear combinations of a fixed set of basis functions or ``features'' is known to be hard. Typically, the worst-case error of the best basis set decays only as fast as $\Theta\(n^{-1/d}\)$, where $n$ is the number of basis functions and $d$ is the input dimension. However, there are many examples of high-dimensional pattern recognition problems (such as face recognition) where linear combinations of small sets of features do solve the problem well. Hence these function classes do not suffer from the ``curse of dimensionality'' associated with more general classes. It is natural then, to look for characterizations of high-dimensional function classes that nevertheless are approximated well by linear combinations of small sets of features. In this paper we give a general result relating the error of approximation of a function class to the covering number of its ``convex core''. For one-hidden-layer neural networks, covering numbers of the class of functions computed by a single hidden node upper bound the covering numbers of the convex core. Hence, using standard results we obtain upper bounds on the approximation rate of neural network classes.

cs.LG

Reinforcement Learning in POMDP's via Direct Gradient Ascent

This paper discusses theoretical and experimental aspects of gradient-based approaches to the direct optimization of policy performance in controlled POMDPs. We introduce GPOMDP, a REINFORCE-like algorithm for estimating an approximation to the gradient of the average reward as a function of the parameters of a stochastic policy. The algorithm's chief advantages are that it requires only a single sample path of the underlying Markov chain, it uses only one free parameter $\beta\in [0,1)$, which has a natural interpretation in terms of bias-variance trade-off, and it requires no knowledge of the underlying state. We prove convergence of GPOMDP and show how the gradient estimates produced by GPOMDP can be used in a conjugate-gradient procedure to find local optima of the average reward.

cs.LG

Scaling Internal-State Policy-Gradient Methods for POMDPs

Policy-gradient methods have received increased attention recently as a mechanism for learning to act in partially observable environments. They have shown promise for problems admitting memoryless policies but have been less successful when memory is required. In this paper we develop several improved algorithms for learning policies with memory in an infinite-horizon setting -- directly when a known model of the environment is available, and via simulation otherwise. We compare these algorithms on some large POMDPs, including noisy robot navigation and multi-agent problems.

cs.LG

A Multi-Agent, Policy-Gradient approach to Network Routing

Network routing is a distributed decision problem which naturally admits numerical performance measures, such as the average time for a packet to travel from source to destination. OLPOMDP, a policy-gradient reinforcement learning algorithm, was successfully applied to simulated network routing under a number of network models. Multiple distributed agents (routers) learned co-operative behavior without explicit inter-agent communication, and they avoided behavior which was individually desirable, but detrimental to the group's overall performance. Furthermore, shaping the reward signal by explicitly penalizing certain patterns of sub-optimal behavior was found to dramatically improve the convergence rate.

cs.LG

The Evolution of Learning Algorithms for Artificial Neural Networks

In this paper we investigate a neural network model in which weights between computational nodes are modified according to a local learning rule. To determine whether local learning rules are sufficient for learning, we encode the network architectures and learning dynamics genetically and then apply selection pressure to evolve networks capable of learning the four boolean functions of one variable. The successful networks are analysed and we show how learning behaviour emerges as a distributed property of the entire network. Finally the utility of genetic algorithms as a tool of discovery is discussed.

cs.NE

Bayesian Beta-Binomial Prevalence Estimation Using an Imperfect Test

Following [Diggle 2011, Greenland 1995], we give a simple formula for the Bayesian posterior density of a prevalence parameter based on unreliable testing of a population. This problem is of particular importance when the false positive test rate is close to the prevalence in the population being tested. An efficient Monte Carlo algorithm for approximating the posterior density is presented, and applied to estimating the Covid-19 infection rate in Santa Clara county, CA using the data reported in [Bendavid 2020]. We show that the true Bayesian posterior places considerably more mass near zero, resulting in a prevalence estimate of 5,000--70,000 infections (median: 42,000) (2.17% (95CI 0.27%--3.63%)), compared to the estimate of 48,000--81,000 infections derived in [Bendavid 2020] using the delta method. A demonstration, with code and additional examples, is available at testprev.com.

stat.AP

Theoretical Models of Learning to Learn

A Machine can only learn if it is biased in some way. Typically the bias is supplied by hand, for example through the choice of an appropriate set of features. However, if the learning machine is embedded within an {\em environment} of related tasks, then it can {\em learn} its own bias by learning sufficiently many tasks from the environment. In this paper two models of bias learning (or equivalently, learning to learn) are introduced and the main theoretical results presented. The first model is a PAC-type model based on empirical process theory, while the second is a hierarchical Bayes model.

cs.LG

Some observations concerning Off Training Set (OTS) error

A form of generalisation error known as Off Training Set (OTS) error was recently introduced in [Wolpert, 1996b], along with a theorem showing that small training set error does not guarantee small OTS error, unless assumptions are made about the target function. Here it is shown that the applicability of this theorem is limited to models in which the distribution generating training data has no overlap with the distribution generating test data. It is argued that such a scenario is of limited relevance to machine learning.

cs.LG

General Matrix-Matrix Multiplication Using SIMD features of the PIII

Generalised matrix-matrix multiplication forms the kernel of many mathematical algorithms. A faster matrix-matrix multiply immediately benefits these algorithms. In this paper we implement efficient matrix multiplication for large matrices using the floating point Intel Pentium SIMD (Single Instruction Multiple Data) architecture. A description of the issues and our solution is presented, paying attention to all levels of the memory hierarchy. Our results demonstrate an average performance of 2.09 times faster than the leading public domain matrix-matrix multiply routines.

cs.PF

Hebbian Synaptic Modifications in Spiking Neurons that Learn

In this paper, we derive a new model of synaptic plasticity, based on recent algorithms for reinforcement learning (in which an agent attempts to learn appropriate actions to maximize its long-term average reward). We show that these direct reinforcement learning algorithms also give locally optimal performance for the problem of reinforcement learning with multiple agents, without any explicit communication between agents. By considering a network of spiking neurons as a collection of agents attempting to maximize the long-term average of a reward signal, we derive a synaptic update rule that is qualitatively similar to Hebb's postulate. This rule requires only simple computations, such as addition and leaky integration, and involves only quantities that are available in the vicinity of the synapse. Furthermore, it leads to synaptic connection strengths that give locally optimal values of the long term average reward. The reinforcement learning paradigm is sufficiently broad to encompass many learning problems that are solved by the brain. We illustrate, with simulations, that the approach is effective for simple pattern classification and motor learning tasks.

cs.LG

Learning Model Bias

In this paper the problem of {\em learning} appropriate domain-specific bias is addressed. It is shown that this can be achieved by learning many related tasks from the same domain, and a theorem is given bounding the number tasks that must be learnt. A corollary of the theorem is that if the tasks are known to possess a common {\em internal representation} or {\em preprocessing} then the number of examples required per task for good generalisation when learning $n$ tasks simultaneously scales like $O(a + \frac{b}{n})$, where $O(a)$ is a bound on the minimum number of examples required to learn a single task, and $O(a + b)$ is a bound on the number of examples required to learn each task independently. An experiment providing strong qualitative support for the theoretical results is reported.

cs.LG

The Canonical Distortion Measure for Vector Quantization and Function Approximation

To measure the quality of a set of vector quantization points a means of measuring the distance between a random point and its quantization is required. Common metrics such as the {\em Hamming} and {\em Euclidean} metrics, while mathematically simple, are inappropriate for comparing natural signals such as speech or images. In this paper it is shown how an {\em environment} of functions on an input space $X$ induces a {\em canonical distortion measure} (CDM) on X. The depiction 'canonical" is justified because it is shown that optimizing the reconstruction error of X with respect to the CDM gives rise to optimal piecewise constant approximations of the functions in the environment. The CDM is calculated in closed form for several different function classes. An algorithm for training neural networks to implement the CDM is presented along with some encouraging experimental results.

cs.LG

A Bayesian/Information Theoretic Model of Bias Learning

In this paper the problem of learning appropriate bias for an environment of related tasks is examined from a Bayesian perspective. The environment of related tasks is shown to be naturally modelled by the concept of an {\em objective} prior distribution. Sampling from the objective prior corresponds to sampling different learning tasks from the environment. It is argued that for many common machine learning problems, although we don't know the true (objective) prior for the problem, we do have some idea of a set of possible priors to which the true prior belongs. It is shown that under these circumstances a learner can use Bayesian inference to learn the true prior by sampling from the objective prior. Bounds are given on the amount of information required to learn a task when it is simultaneously learnt with several other tasks. The bounds show that if the learner has little knowledge of the true prior, and the dimensionality of the true prior is small, then sampling multiple tasks is highly advantageous.

cs.LG