Search arXivSearch

arXiv · 2510.16183

Forward-Backward Binarization

Abstract

Binarization of gene expression data is a \textbf{critical prerequisite} for the synthesis of Boolean gene regulatory network (GRN) models from omics datasets. Because Boolean networks encode gene activity as binary variables, the accuracy of binarization directly conditions whether the inferred models can faithfully reproduce biological experiments, capture regulatory dynamics, and support downstream analyses such as controllability and therapeutic strategy design. In practice, binarization is most often performed using thresholding methods that partition expression values into two discrete levels, representing the absence or presence of gene expression. However, such approaches oversimplify the underlying biology: gene-specific functional roles, measurement uncertainty, and the scarcity of time-resolved experimental data render thresholding alone insufficient. To overcome these limitations, we propose a novel \textbf{regulation-based binarization method} tailored to snapshot data. Our approach combines thresholding with functional binary value completion guided by the regulatory graph, propagating values between regulators and targets according to Boolean regulation rules. This strategy enables the inference of missing or uncertain values and ensures that binarization remains biologically consistent with both regulatory interactions and Boolean modeling principles of the gene regulation. Validation against ODE simulations of artificial and established Boolean GRNs demonstrates that the method achieves accurate and robust binarization, thereby strengthening the reliability of Boolean network synthesis.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ismail Belgacem, Franck Delaplace. 2025-10-17. Forward-Backward Binarization. https://arxiv.org/abs/2510.16183

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Flip Dynamics for Sampling Colorings: Improving $(11/6-ε)$ Using a Simple Metric

We present improved bounds for randomly sampling $k$-colorings of graphs with maximum degree $Δ$; our results hold without any further structural assumptions on the graph. The Glauber dynamics is a simple single-site update Markov chain. Jerrum (1995) proved an optimal $O(n\log{n})$ mixing-time bound for Glauber dynamics whenever $k>2Δ$ where $Δ$ is the maximum degree of the input graph. This bound was improved by Vigoda (1999) to $k>(11/6)Δ$ using a "flip" dynamics which recolors (small) maximal two-colored components in each step. Vigoda's result was the best known for general graphs for 20 years until Chen et al. (2019) established optimal mixing of the flip dynamics for $k>(11/6-\varepsilon)Δ$ where $\varepsilon\approx 10^{-5}$. We present the first substantial improvement over these results. We prove an optimal mixing-time bound of $O(n\log{n})$ for the flip dynamics when $Δ\geq125$ and $k\geq1.809Δ$. This yields, through recent spectral independence results, an optimal $O(n\log{n})$ mixing time for the Glauber dynamics for every fixed $Δ\geq125$ in the same range of $k/Δ$. Our proof utilizes path coupling with a simple weighted Hamming distance for "unblocked" neighbors.

cs.DM

Factorisability of Low Dimensional Non-Negative Integer Matrices

We consider the problem of determining if a given two-dimensional nonnegative integer matrix $M$ is the product of two such matrices, excluding trivial units. A matrix $M$ with no such factorisation is called prime and therefore belongs to the minimal (infinite rank) generator of $2 \times 2$ matrices over the natural numbers, otherwise it is called composite. We also consider the problem of finding a (non-unique) factorisation of a composite matrix. Our results have applications in computational group theory and the theory of codes, where such matrices are called incidence matrices. We analyse the complexity of primality and finding a factorisation for a composite matrix, providing a first efficient algorithm.

cs.DM