Search arXivSearch

arXiv subjects

Song Feng

Publications and source records attributed to Song Feng.

At least 19 recordsLinked to original sources

DigiPhen: a new paradigm for building predictive models of biological systems

Reengineered biological systems have the potential to revolutionize chemical and material production, enhance critical mineral recovery, serve as threat sensors and improve human health. Unfortunately, the extreme complexity of organisms has made it difficult to achieve this potential in all but the simplest cases. Recent technological advances, however, have provided a foundation for solving this problem. Here, we describe a capability for accelerating the reengineering of cells by providing accurate predictions of the impact of genetic or environmental changes on cell phenotype. This digital phenome platform (DigiPhen) consists of integrated experimental, analytical and modeling workflows for building a digital representation of microbial or plant systems. It is designed around an expanding set of interchangeable, interconnecting software and experimental modules that can accurately represent the mechanistic determinants of phenotype. The DigiPhen platform will systematically collect data on cell composition, spatial organization, metabolic pathways and regulatory networks in a semi-autonomous fashion and use this information to build modular, multi-scale models of biological systems. These models will be used to predict molecular and environmental changes needed for producing desired biological outcomes. DigiPhen is intended to be the heart of community research campaigns that will meet the immediate needs of individual researchers while fulfilling long-term goals of the scientific community. Altogether, the DigiPhen platform represents a new paradigm for building predictive models of biological systems.

q-bio.MN

Phase-drifting with emitting plasma temperature in the quasi-periodic pulsations of an X-class solar flare

Recent multi-wavelength observations of solar flares have provided new constraints on the physical origin of quasi-periodic pulsations (QPPs). In an X-class flare, we detect a short-lived $\sim$5-minute QPP simultaneously in hard X-rays, extreme-ultraviolet (EUV), and soft X-ray emissions, exhibiting a clear phase-drifting behavior with emitting plasma temperature. Based on phase-resolved timing analysis, it is found that (i) the QPPs in all diagnostics share nearly identical oscillation periods, (ii) a systematic temperature-dependent phase drifting is present, with the phase delay relative to the hard X-ray emission increases systematically from the hottest to cooler EUV channels, and (iii) the QPP persists for only a few cycles during the impulsive phase. These properties imply that periodic magnetic reconnection, possibly triggered by the leakage of 5-minute oscillations from the lower atmosphere, modulates the non-thermal electrons responsible for the leading Hard X-ray QPPs. Subsequently, plasma heating and cooling processes manifest sequentially across passbands with different temperature responses, resulting in the observed temperature-dependent phase drifting. These results provide novel observational evidence supporting the use of multi-temperature, multi-wavelength phase relationships to constrain the temporal evolution of flare energy release and the origins of QPPs.

astro-ph.SR

A Markov-Chain-Monte-Carlo-based Hybrid Noise Inference for Continuous Wavelet Power Spectra: with Applications to Solar and Stellar Oscillatory Signals

Detecting oscillations in solar and stellar time series is complicated by non-stationary red noise and evolving background emission. Methods based on detrending and AR(1)-based wavelet analysis can introduce spurious periodicities and do not adequately describe time-dependent backgrounds. We develop a Bayesian approach that combines the continuous wavelet transform with MCMC sampling to infer a time-dependent background spectrum. The background is represented by a power-law plus white-noise component, with parameters allowed to vary smoothly in time, so that significance levels can be evaluated locally without explicit detrending. Tests with synthetic data show that injected oscillations are recovered reliably, while false detections are suppressed in pure-noise cases. Using a frequency-domain signal-to-noise ratio (S/N), we find that oscillations can be identified robustly when the S/N is greater than or equal to 2 under mixed noise conditions. The detectable period range is limited by wavelet resolution, from about 3-4 sampling intervals up to roughly one-quarter of the total duration. Application to GOES soft X-ray flare observations shows that the method isolates quasi-periodic oscillations with improved temporal localization compared to standard wavelet and Fourier-based approaches. Meanwhile, this behavior is consistent across a range of noise conditions and signal morphologies.

astro-ph.IM

Chromospheric resonator model for sunspot revealed by multi-height observation of umbral wave

Sunspots are transient, magnetically intense features that host oscillations linked to magnetohydrodynamic (MHD) waves. These waves may contribute to plasma heating and drive mass flows in the solar wind. Beyond their energetic role, they serve as diagnostic tools for probing sunspot structure. In this study, we investigated chromospheric wave propagation in a sunspot using high-resolution, multi-wavelength observations from the Goode Solar Telescope at Big Bear Solar Observatory. Spectral analysis shows that the intensity at H$\alpha$ line core and its wings exhibited oscillatory signal at about 3 min. We performed a cross-wavelet analysis to examine the phase relationship between the wing-integrated and line-core intensity oscillations of the H$\alpha$ line and the centroid-derived H$\alpha$ Doppler velocity. We also analyze the phase relationships between intensity pairs from different passband combinations of the H$\alpha$ line. The results indicate the presence of slow magnetoacoustic modes manifesting standing waves along with upward propagating waves. The observed phase patterns suggest that umbral waves are confined within a non-ideal acoustic resonator, providing measurable wave properties that could serve as input for sunspot seismology and refine models of sunspot atmospheric structure.

astro-ph.SR

Effect of Solar Flares on Decayless Kink Oscillations in Nearby Coronal Loops

We present a statistical study of 130 solar flares (B to X class) that lack soft X-ray quasi-periodic pulsations and show no kink oscillations of nearby coronal loops visible in SDO/AIA 171~\AA~images. The aim is to investigate whether decayless kink oscillations of coronal loops respond to nearby flaring activity. Using the Fractional Anisotropy-based Video Motion Magnification technique, we detected low-amplitude decayless oscillations in all 130 loops before, during, and after each flare, confirming their ubiquitous nature. Oscillation periods are found to range from 122~s to 268~s, and the projected displacement amplitudes are 0.023--0.111~Mm. No amplitude--period correlation is found. For each event, we estimated the amplitude before, during, and after the flare. Across all flare classes, the average amplitude remains unchanged. However, in some specific cases, the oscillation amplitude may exhibit minor changes. For B-, C-, and M-class flares, the fraction of events with an amplitude change exceeding 10% is approximately 23%, 41%, and 36%, respectively. In M-class flares, such minor amplitude increases occur four times more often than decreases; in X-class flares (only six events), decreases dominate by a factor of three. The fraction of events that exhibit an increase in the amplitude of more than 20% appears to be highest when the loop centre is located at a distance of 100--120~Mm from the flare site, reaching 33% (6 out of 18 events). Overall, the amplitude of decayless kink oscillations does not undergo a major change in response to nearby flares, especially for less powerful classes, suggesting that flare-related processes such as blast waves and reconnection inflows have little effect on the energy supply to oscillating loops.

astro-ph.SR

Thiol post-translational modifications modulate allosteric regulation of the OpcA-G6PDH complex through conformational gate control

Cyanobacteria require ultra-fast metabolic switching to maintain reducing power balance during environmental fluctuations. Glucose-6-phosphate dehydrogenase (G6PDH), catalyzing the rate-limiting step of the oxidative pentose phosphate pathway (OPPP), provides essential NADPH and metabolic intermediates for biosynthetic processes and redox homeostasis. In cyanobacteria, the unique redox-sensitive protein OpcA acts as a metabolic switch for G6PDH, enabling rapid adjustment of reducing power generation from glycogen catabolism and resulting in precise regulation of carbon flux between anabolic and catabolic pathways. While the redox-sensitive cysteine structures of OpcA are known to regulate G6PDH, the detailed mechanisms of how redox post-translational modifications (PTMs) influence OpcA's allosteric effects on G6PDH structures and function remain elusive. To investigate this mechanism, we utilized computational modeling combined with experimental redox proteomics using Synechococcus elongatus PCC 7942 as a model system. Redox proteomics captured modified cysteine residues under light/dark or circadian shifts. Computational simulation revealed that thiol PTMs near the OpcA-G6PDH interface are crucial to allosteric regulation of regions affecting the G6PDH activity, including a potential gate region for substrate ingress and product egress, as well as critical hydrogen bond networks within the active site. These PTMs promote rapid metabolic switching by enhancing G6PDH catalytic activity when OpcA is oxidized. This study provides evidence for novel molecular mechanisms that elucidate the importance of thiol PTMs of OpcA in modulating G6PDH structure and function in an allosteric manner, demonstrating how PTM-level regulation provides a critical control mechanism that enables cyanobacteria to rapidly adapt to environmental fluctuations through precise metabolic fine-tuning.

physics.bio-ph

ProCaliper: functional and structural analysis, visualization, and annotation of proteins

Understanding protein function at the molecular level requires connecting residue-level annotations with physical and structural properties. This can be cumbersome and error-prone when functional annotation, computation of physico-chemical properties, and structure visualization are separated. To address this, we introduce ProCaliper, an open-source Python library for computing and visualizing physico-chemical properties of proteins. It can retrieve annotation and structure data from UniProt and AlphaFold databases, compute residue-level properties such as charge, solvent accessibility, and protonation state, and interactively visualize the results of these computations along with user-supplied residue-level data. Additionally, ProCaliper incorporates functional and structural information to construct and optionally sparsify networks that encode the distance between residues and/or annotated functional sites or regions. The package ProCaliper and its source code, along with the code used to generate the figures in this manuscript, are freely available at https://github.com/PNNL-Predictive-Phenomics/ProCaliper.

q-bio.BM

From FAIR to CURE: Guidelines for Computational Models of Biological Systems

Guidelines for managing scientific data have been established under the FAIR principles requiring that data be Findable, Accessible, Interoperable, and Reusable. In many scientific disciplines, especially computational biology, both data and models are key to progress. For this reason, and recognizing that such models are a very special type of 'data', we argue that computational models, especially mechanistic models prevalent in medicine, physiology and systems biology, deserve a complementary set of guidelines. We propose the CURE principles, emphasizing that models should be Credible, Understandable, Reproducible, and Extensible. We delve into each principle, discussing verification, validation, and uncertainty quantification for model credibility; the clarity of model descriptions and annotations for understandability; adherence to standards and open science practices for reproducibility; and the use of open standards and modular code for extensibility and reuse. We outline recommended and baseline requirements for each aspect of CURE, aiming to enhance the impact and trustworthiness of computational models, particularly in biomedical applications where credibility is paramount. Our perspective underscores the need for a more disciplined approach to modeling, aligning with emerging trends such as Digital Twins and emphasizing the importance of data and modeling standards for interoperability and reuse. Finally, we emphasize that given the non-trivial effort required to implement the guidelines, the community moves to automate as many of the guidelines as possible.

q-bio.OT

Graph-GIC: A Smart and Parallelized Geomagnetically Induced Current Modelling Algorithm Based on Graph Theory for Space Weather Applications

Geomagnetically Induced Current (GIC) refers to the electromagnetic response of the Earth and its conductive modern infrastructures to space weather and would pose a significant threat to high-voltage power grids designed for the alternative current operation. To assess the impact of space weather on the power grid, one needs to calculate the GIC on a national or continental scale. In this study, we developed a smart and parallelized GIC modelling algorithm, Graph GIC. This algorithm deploys a graph representing a power grid in a single-line diagram, in which substations/transformers act as nodes and transmission lines as edges. With these denotations, a power grid and its electric parameters are mathematically represented with an adjacency matrix and an admittance matrix. We used sparse matrix and parallelisation techniques to expedite the intensive computation in cases of large-scale power grids. The Graph GIC was validated with a benchmark grid, applied to the GIC calculation of the 500 kV power grid of Guangdong, China, and conducted preliminary analysis on the grid's susceptibility to geomagnetic storms. The Graph GIC algorithm has the advantage of an intuitive and highly scalable graph representation of a power grid at any scale. It achieves high-accuracy calculation and a speedup of about 18 times after parallelisation. This algorithm could be applied to assess the impact of space weather on a power grid up to continental scales and could be incorporated into global space weather modelling frameworks.

physics.space-ph

DFlow: Diverse Dialogue Flow Simulation with Large Language Models

Developing language model-based dialogue agents requires effective data to train models that can follow specific task logic. However, most existing data simulation methods focus on increasing diversity in language, topics, or dialogue acts at the utterance level, largely neglecting a critical aspect of task logic diversity at the dialogue level. This paper proposes a novel data simulation method designed to enhance the diversity of synthetic dialogues by focusing on task execution logic. Our method uses LLMs to generate decision tree-structured task plans, which enables the derivation of diverse dialogue trajectories for a given task. Each trajectory, referred to as a "dialog flow", guides the generation of a multi-turn dialogue that follows a unique trajectory. We apply this method to generate a task-oriented dialogue dataset comprising 3,886 dialogue flows across 15 different domains. We validate the effectiveness of this dataset using the next action prediction task, where models fine-tuned on our dataset outperform strong baselines, including GPT-4. Upon acceptance of this paper, we plan to release the code and data publicly.

cs.CL

Transcriptome and Redox Proteome Reveal Temporal Scales of Carbon Metabolism Regulation in Model Cyanobacteria Under Light Disturbance

We develop a systems approach based on an energy-landscape concept to differentiate interactions involving redox activities and conformational changes of proteins and nucleic acids interactions in multi-layered protein-DNA regulatory networks under light disturbance. Our approach is a data-driven modeling workflow using a physics-informed machine learning algorithm to train a non-linear mathematical model for interpreting gene expression dynamics and to lead discovery for protein regulators using redox proteome analysis. We distinguish light-responsive elements within central carbon metabolism pathways from independent variables like circadian time using the publicly available transcriptome datasets of Synechococcus elongatus over diel cycles responding to light perturbations. Our approach provides interpretable de novo models for elucidating events of reactions in complex regulatory pathways in response to stressful disturbance from the environment. We discovered protein regulators in response to light disturbance in the proteome analysis involving shifts in protein abundance as well as cysteine redox states under constant illumination and after two hours of darkness. We discovered significant shifts in cysteine redox states in regulatory proteins such as transcription sigma factors and metabolic enzymes in the oxidative pentose phosphate pathway and the Calvin-Benson cycle, while the changes in their protein abundance were minimal. These results indicate that regulatory dynamics in reductant generation link photo-induced electron transport pathways and redox metabolic pathways with circadian rhythms through fast redox-induced conformational changes or slow expression regulations across networks.

q-bio.MN

Structured List-Grounded Question Answering

Document-grounded dialogue systems aim to answer user queries by leveraging external information. Previous studies have mainly focused on handling free-form documents, often overlooking structured data such as lists, which can represent a range of nuanced semantic relations. Motivated by the observation that even advanced language models like GPT-3.5 often miss semantic cues from lists, this paper aims to enhance question answering (QA) systems for better interpretation and use of structured lists. To this end, we introduce the LIST2QA dataset, a novel benchmark to evaluate the ability of QA systems to respond effectively using list information. This dataset is created from unlabeled customer service documents using language models and model-based filtering processes to enhance data quality, and can be used to fine-tune and evaluate QA models. Apart from directly generating responses through fine-tuned models, we further explore the explicit use of Intermediate Steps for Lists (ISL), aligning list items with user backgrounds to better reflect how humans interpret list items before generating responses. Our experimental results demonstrate that models trained on LIST2QA with our ISL approach outperform baselines across various metrics. Specifically, our fine-tuned Flan-T5-XL model shows increases of 3.1% in ROUGE-L, 4.6% in correctness, 4.5% in faithfulness, and 20.6% in completeness compared to models without applying filtering and the proposed ISL method.

cs.CL

Radio Frequency Interference Detection Using Efficient Multi-Scale Convolutional Attention UNet

Studying the universe through radio telescope observation is crucial. However, radio telescopes capture not only signals from the universe but also various interfering signals, known as Radio Frequency Interference (RFI). The presence of RFI can significantly impact data analysis. Ensuring the accuracy, reliability, and scientific integrity of research findings by detecting and mitigating or eliminating RFI in observational data, presents a persistent challenge in radio astronomy. In this study, we proposed a novel deep learning model called EMSCA-UNet for RFI detection. The model employs multi-scale convolutional operations to extract RFI features of various scale sizes. Additionally, an attention mechanism is utilized to assign different weights to the extracted RFI feature maps, enabling the model to focus on vital features for RFI detection. We evaluated the performance of the model using real data observed from the 40-meter radio telescope at Yunnan Observatory. Furthermore, we compared our results to other models, including U-Net, RFI-Net, and R-Net, using four commonly employed evaluation metrics: precision, recall, F1 score, and IoU. The results demonstrate that our model outperforms the other models on all evaluation metrics, achieving an average improvement of approximately 5\% compared to U-Net. Our model not only enhances the accuracy and comprehensiveness of RFI detection but also provides more detailed edge detection while minimizing the loss of useful signals.

astro-ph.IM

TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization

Single document news summarization has seen substantial progress on faithfulness in recent years, driven by research on the evaluation of factual consistency, or hallucinations. We ask whether these advances carry over to other text summarization domains. We propose a new evaluation benchmark on topic-focused dialogue summarization, generated by LLMs of varying sizes. We provide binary sentence-level human annotations of the factual consistency of these summaries along with detailed explanations of factually inconsistent sentences. Our analysis shows that existing LLMs hallucinate significant amounts of factual errors in the dialogue domain, regardless of the model's size. On the other hand, when LLMs, including GPT-4, serve as binary factual evaluators, they perform poorly and can be outperformed by prevailing state-of-the-art specialized factuality evaluation metrics. Finally, we conducted an analysis of hallucination types with a curated error taxonomy. We find that there are diverse errors and error distributions in model-generated summaries and that non-LLM based metrics can capture all error types better than LLM-based evaluators.

cs.CL

DFEE: Interactive DataFlow Execution and Evaluation Kit

DataFlow has been emerging as a new paradigm for building task-oriented chatbots due to its expressive semantic representations of the dialogue tasks. Despite the availability of a large dataset SMCalFlow and a simplified syntax, the development and evaluation of DataFlow-based chatbots remain challenging due to the system complexity and the lack of downstream toolchains. In this demonstration, we present DFEE, an interactive DataFlow Execution and Evaluation toolkit that supports execution, visualization and benchmarking of semantic parsers given dialogue input and backend database. We demonstrate the system via a complex dialog task: event scheduling that involves temporal reasoning. It also supports diagnosing the parsing results via a friendly interface that allows developers to examine dynamic DataFlow and the corresponding execution results. To illustrate how to benchmark SoTA models, we propose a novel benchmark that covers more sophisticated event scheduling scenarios and a new metric on task success evaluation. The codes of DFEE have been released on https://github.com/amazonscience/dataflow-evaluation-toolkit.

cs.CL

Inter-Correlation Between Sunspot Oscillations and Their Internal Structures

Three- and five-minute oscillations are commonly found in any sunspot. As they are modulated by the internal thermal and magnetic structures of a sunspot, therehence, they could be used as an effective tool for sunspot seismology. In this paper, we investigate the properties of oscillations in sunspot groups with varying size and magnetic field, and aim to establish the relationships between sunspot oscillations and its internal structure comparatively. We selected three groups of unipolar sunspot with approximately axial-symmetric magnetic field and calculated their Fourier spectra based on the Ultraviolet(UV)/Extreme ultraviolet(EUV) emission intensity variations recorded by the Solar Dynamics Observatory/Atmospheric Imaging Assembly (SDO/AIA). We found that the distribution of three minute oscillation is defined by the joint effect of diverging magnetic field and the stratification of sunspot atmosphere. Its distribution could be modified by any invading magnetic structures in the umbra. Whereas the five minute oscillations are more prominent in small spots, it implies that five minute oscillation is very closely connected with umbral dynamics.

astro-ph.SR

SBbadger: Biochemical Reaction Networks with Definable Degree Distributions

Motivation: An essential step in developing computational tools for the inference, optimization, and simulation of biochemical reaction networks is gauging tool performance against earlier efforts using an appropriate set of benchmarks. General strategies for the assembly of benchmark models include collection from the literature, creation via subnetwork extraction and de novo generation. However, with respect to biochemical reaction networks, these approaches and their associated tools are either poorly suited to generate models that reflect the wide range of properties found in natural biochemical networks or to do so in numbers that enable rigorous statistical analysis. Results: In this work we present SBbadger, a python-based software tool for the generation of synthetic biochemical reaction or metabolic networks with user-defined degree distributions, multiple available kinetic formalisms, and a host of other definable properties. SBbadger thus enables the creation of benchmark model sets that reflect properties of biological systems and generate the kinetics and model structures typically targeted by computational analysis and inference software. Here we detail the computational and algorithmic workflow of SBbadger, demonstrate its performance under various settings, provide samples outputs, and compare it to currently available biochemical reaction network generation software.

q-bio.MN

DG2: Data Augmentation Through Document Grounded Dialogue Generation

Collecting data for training dialog systems can be extremely expensive due to the involvement of human participants and need for extensive annotation. Especially in document-grounded dialog systems, human experts need to carefully read the unstructured documents to answer the users' questions. As a result, existing document-grounded dialog datasets are relatively small-scale and obstruct the effective training of dialogue systems. In this paper, we propose an automatic data augmentation technique grounded on documents through a generative dialogue model. The dialogue model consists of a user bot and agent bot that can synthesize diverse dialogues given an input document, which are then used to train a downstream model. When supplementing the original dataset, our method achieves significant improvement over traditional data augmentation methods. We also achieve great performance in the low-resource setting.

cs.CL