Search arXiv⌕ Search

arXiv subjects

Johnny Jingze Li

Publications and source records attributed to Johnny Jingze Li.

5 recordsLinked to original sources

Emergence in Network Systems is Bounded by Boundary-Crossing Paths

Emergence, features of a whole that no part shows alone, is central to complex systems but rarely measured so as to locate its source. Viewing a system structurally, as a network of interacting parts, we evaluate emergence as what observing the coupled whole reveals beyond, or erases from, the combined observations of the parts. For observations that build what they see route by route, we show that every emergent feature is produced by its own route across the interface between parts, a route that traces it to the components and interactions responsible. We show the number of all such boundary-crossing paths bounds the total emergence discrepancy, and equals the emergence of measures that count routes. The bound is computed locally around the interface, validated through network examples, and in a nervous system the routes show that interneurons relay most of the emergence. Our work associates emergence with path structure in the network system that can be anticipated, attributed, and engineered.

physics.soc-ph↗

How Language Models Differ in Redistributing Attention-Head Activity Under Serial Demand

The way a model distributes activity over each layer's attention heads offers a coarse view of how it routes information through depth; how this changes with the task is part of what a mechanistic account must explain. Holding prompt length fixed, we vary how many serial steps a task demands and measure, in every layer of 17 open-weight models, whether activity concentrates on a few heads or spreads across many as demand rises. Both occur: in most models, layers just before mid-depth concentrate activity and later layers spread it. Models differ in where and how strongly this happens. The Qwen2.5 base models from 0.5B to 7B, for example, spread less than the average model in every task and concentrate activity in parts of their second half, where Llama models from 1B to 8B and OLMo-2 spread; the contrast largely holds between Llama-3.1-70B and Qwen2.5-72B, which have the same number of layers and heads. These differences are reproducible, and post-trained models keep much of their base model's pattern. An ablation study suggests that, within a task, models whose activity is more concentrated on their top heads also depend more on those heads for the answer. Concentration and spreading across layers thus offer a new way to compare models, by how they route information through depth. Code is available at https://github.com/johnnyjli/serial-demand-heads.

cs.LG↗

A Categorical Framework for Quantifying Emergent Effects in Network Topology

Emergent effect is crucial to understanding the properties of complex systems that do not appear in their basic units, but there has been a lack of theories to measure and understand its mechanisms. In this paper, we consider emergence as a kind of structural nonlinearity, discuss a framework based on homological algebra that encodes emergence as the mathematical structure of cohomologies, and then apply it to network models to develop a computational measure of emergence. This framework ties the potential for emergent effects of a system to its network topology and local structures, paving the way to predict and understand the cause of emergent effects. We show in our numerical experiment that our measure of emergence correlates with the existing information-theoretic measure of emergence.

nlin.AO↗

Advancing Neural Network Performance through Emergence-Promoting Initialization Scheme

Emergence in machine learning refers to the spontaneous appearance of complex behaviors or capabilities that arise from the scale and structure of training data and model architectures, despite not being explicitly programmed. We introduce a novel yet straightforward neural network initialization scheme that aims at achieving greater potential for emergence. Measuring emergence as a kind of structural nonlinearity, our method adjusts the layer-wise weight scaling factors to achieve higher emergence values. This enhancement is easy to implement, requiring no additional optimization steps for initialization compared to GradInit. We evaluate our approach across various architectures, including MLP and convolutional architectures for image recognition and transformers for machine translation. We demonstrate substantial improvements in both model accuracy and training speed, with and without batch normalization. The simplicity, theoretical innovation, and demonstrable empirical advantages of our method make it a potent enhancement to neural network initialization practices. These results suggest a promising direction for leveraging emergence to improve neural network training methodologies. Code is available at: https://github.com/johnnyjingzeli/EmergenceInit.

cs.LG↗

Quantifying Emergence in Neural Networks: Insights from Pruning and Training Dynamics

Emergence, where complex behaviors develop from the interactions of simpler components within a network, plays a crucial role in enhancing neural network capabilities. We introduce a quantitative framework to measure emergence during the training process and examine its impact on network performance, particularly in relation to pruning and training dynamics. Our hypothesis posits that the degree of emergence, defined by the connectivity between active and inactive nodes, can predict the development of emergent behaviors in the network. Through experiments with feedforward and convolutional architectures on benchmark datasets, we demonstrate that higher emergence correlates with improved trainability and performance. We further explore the relationship between network complexity and the loss landscape, suggesting that higher emergence indicates a greater concentration of local minima and a more rugged loss landscape. Pruning, which reduces network complexity by removing redundant nodes and connections, is shown to enhance training efficiency and convergence speed, though it may lead to a reduction in final accuracy. These findings provide new insights into the interplay between emergence, complexity, and performance in neural networks, offering valuable implications for the design and optimization of more efficient architectures.

cs.LG↗