Search arXivSearch

arXiv subjects

Daniel Rodriguez

Publications and source records attributed to Daniel Rodriguez.

18 recordsLinked to original sources

Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level. We present an AI-assisted approach that generates candidate hazard scenarios from NASA's Aviation Safety Reporting System (ASRS). Given a target adverse outcome, it produces a structured hypothesis as categorical factors and a narrative scenario describing an operational event sequence consistent with the structure. Each scenario includes by a plausibility score from historical co-occurrence evidence and traceability to the most similar held-out ASRS reports. We then propose a hybrid variant, conditioning narrative generation on a structured hypothesis produced via evolutionary abduction, improving correctness and reducing variability. We evaluate multiple large language models, zero-shot versus few-shot prompting, and optional fine-tuning, measuring how prompting and model choice affect the validity and realism of the generated structures and narratives.

cs.AI

Laser Spectroscopy of Thulium Isotopes Near the (N=82) Shell Closure: Nuclear Moment and Charge Radius of ${}^{152\mathrm{m}}\mathrm{Tm}$

We report on resonance ionization laser spectroscopy measurements performed on both neutron-deficient and neutron-rich thulium ($\mathrm{Tm}, Z=69$) isotopes. Isotope shifts were determined for three atomic ground-state transitions at wavelengths of $389.8\,\mathrm{nm}$, $388.4\,\mathrm{nm}$, and $388.8\,\mathrm{nm}$ in the isotopes ${}^{152\mathrm{m}}\mathrm{Tm}$, ${}^{153}\mathrm{Tm}$, ${}^{154\mathrm{m}}\mathrm{Tm}$, and ${}^{169}\mathrm{Tm}$. In addition, for the $389.8\,\mathrm{nm}$ transition, measurements were extended to the isotope ${}^{170}\mathrm{Tm}$, and the hyperfine structure was partially resolved for all five isotopes. For this transition, the isotope shift could be determined for one more isotope, ${}^{154\mathrm{m}}\mathrm{Tm}$. From the extracted hyperfine coupling constants, the nuclear magnetic dipole moment for ${}^{152\mathrm{m}}\mathrm{Tm}$ was determined for the first time, resulting in $\mu\left({}^{152\mathrm{m}}\mathrm{Tm}\right) = 5.8(3) \mu_\mathrm{N}$. Furthermore, the mean-square nuclear charge radius $\delta\langle r^2\rangle^{152\mathrm{m},169} = -1.86(25)\,\mathrm{fm}^2$ for ${}^{152\mathrm{m}}\mathrm{Tm}$ was extracted from the measured isotope shifts.

nucl-ex

A lower bound for the distance between CM points on Shimura curves

In this paper, we establish a quantitative Diophantine approximation result for complex multiplication (CM) points on Shimura curves. Specifically, we prove a lower bound for the distance between a sequence of CM points $P_n$ converging to a fixed CM point $P$ on a Shimura curve $X(D,1)$ in terms of the discriminant of the endomorphism rings of $P_n$. The proof exploits the complex geometry of the Fuchsian uniformization, the explicit matrix representation of the underlying quaternion algebra, and Liouville's inequality. We show that the distance between the corresponding fixed points $\tau_n$ and $\tau$ in the upper half-plane is bounded below by a positive constant times a negative power of the discriminant. This result provides a Shimura curve analogue of a result of Habegger on singular moduli and modular curves.

math.NT

An Explainable Multi-Task Similarity Measure: Integrating Accumulated Local Effects and Weighted Fr\'echet Distance

In many machine learning contexts, tasks are often treated as interconnected components with the goal of leveraging knowledge transfer between them, which is the central aim of Multi-Task Learning (MTL). Consequently, this multi-task scenario requires addressing critical questions: which tasks are similar, and how and why do they exhibit similarity? In this work, we propose a multi-task similarity measure based on Explainable Artificial Intelligence (XAI) techniques, specifically Accumulated Local Effects (ALE) curves. ALE curves are compared using the Fr\'echet distance, weighted by the data distribution, and the resulting similarity measure incorporates the importance of each feature. The measure is applicable in both single-task learning scenarios, where each task is trained separately, and multi-task learning scenarios, where all tasks are learned simultaneously. The measure is model-agnostic, allowing the use of different machine learning models across tasks. A scaling factor is introduced to account for differences in predictive performance across tasks, and several recommendations are provided for applying the measure in complex scenarios. We validate this measure using four datasets, one synthetic dataset and three real-world datasets. The real-world datasets include a well-known Parkinson's dataset and a bike-sharing usage dataset -- both structured in tabular format -- as well as the CelebA dataset, which is used to evaluate the application of concept bottleneck encoders in a multitask learning setting. The results demonstrate that the measure aligns with intuitive expectations of task similarity across both tabular and non-tabular data, making it a valuable tool for exploring relationships between tasks and supporting informed decision-making.

cs.LG

The challenge of generating and evolving real-life like synthetic test data without accessing real-world raw data -- a Systematic Review

Background: High-level system testing of applications that use data from e-Government services as input requires test data that is real-life-like but where the privacy of personal information is guaranteed. Applications with such strong requirement include information exchange between countries, medicine, banking, etc. This review aims to synthesize the current state-of-the-practice in this domain. Objectives: The objective of this Systematic Review is to identify existing approaches for creating and evolving synthetic test data without using real-life raw data. Methods: We followed well-known methodologies for conducting systematic literature reviews, including the ones from Kitchenham as well as guidelines for analysing the limitations of our review and its threats to validity. Results: A variety of methods and tools exist for creating privacy-preserving test data. Our search found 1,013 publications in IEEE Xplore, ACM Digital Library, and SCOPUS. We extracted data from 75 of those publications and identified 37 approaches that answer our research question partly. A common prerequisite for using these methods and tools is direct access to real-life data for data anonymization or synthetic test data generation. Nine existing synthetic test data generation approaches were identified that were closest to answering our research question. Nevertheless, further work would be needed to add the ability to evolve synthetic test data to the existing approaches. Conclusions: None of the publications really covered our requirements completely, only partially. Synthetic test data evolution is a field that has not received much attention from researchers but needs to be explored in Digital Government Solutions, especially since new legal regulations are being placed in force in many countries.

cs.LG

Black hole entropy from the quantum atmosphere of bound gravitational fluctuations

Black hole entropy is identified with the counting of the dynamical degrees of freedom of trapped gravitational modes continually sourced by the Hawking-Unruh process. In the context of linear perturbations of Schwarzschild spacetime the density of states is derived from the orthogonality of states in the solution space of the Regge-Wheeler-Zerilli equation. The otherwise divergent energy and entropy is cutoff by the Planck scale closest approach of constantly accelerating observers near the horizon. The thermal distribution of the trapped modes, which represent shape fluctuations in the near horizon geometry, store a significant fraction of the spacetime mass as observed from far away. Unlike quasi-normal modes the modes are not directly observable outside of $\sim 3 M$ but, being external to the horizon, they affect the propagation of null rays near the black hole. The characteristic frequencies, around 100 Hz for solar mass black holes, are discussed in relation to possible observations.

gr-qc

Android in the Wild: A Large-Scale Dataset for Android Device Control

There is a growing interest in device-control systems that can interpret human natural language instructions and execute them on a digital device by directly controlling its user interface. We present a dataset for device-control research, Android in the Wild (AITW), which is orders of magnitude larger than current datasets. The dataset contains human demonstrations of device interactions, including the screens and actions, and corresponding natural language instructions. It consists of 715k episodes spanning 30k unique instructions, four versions of Android (v10-13),and eight device types (Pixel 2 XL to Pixel 6) with varying screen resolutions. It contains multi-step tasks that require semantic understanding of language and visual context. This dataset poses a new challenge: actions available through the user interface must be inferred from their visual appearance. And, instead of simple UI element-based actions, the action space consists of precise gestures (e.g., horizontal scrolls to operate carousel widgets). We organize our dataset to encourage robustness analysis of device-control systems, i.e., how well a system performs in the presence of new task descriptions, new applications, or new platform versions. We develop two agents and report performance across the dataset. The dataset is available at https://github.com/google-research/google-research/tree/master/android_in_the_wild.

cs.LG

On the preferred flapping motion of round twin jets

Linear stability theory (LST) is often used to model the large-scale flow structures in the turbulent mixing region and near pressure field of high-speed jets. For perfectly-expanded single round jets, these models predict the dominance of $m=0$ and $m = 1$ helical modes for the lower frequency range, in agreement with empirical data. When LST is applied to twin-jet systems, four solution families appear following the odd/even behaviour of the pressure field about the symmetry planes. The interaction between the unsteady pressure fields of the two jets also results in their coupling. The individual modes of the different solution families no longer correspond to helical motions, but to flapping oscillations of the jet plumes. In the limit of large jet separations, when the jet coupling vanishes, the eigenvalues corresponding to the $m=1$ mode in each family are identical, and a linear combination of them recovers the helical motion. Conversely, as the jet separation decreases, the eigenvalues for the $m=1$ modes of each family diverge, thus favouring a particular flapping oscillation over the others and preventing the appearance of helical motions. The dominant mode of oscillation for a given jet Mach number $M_j$ and temperature ratio $T_R$ depends on the Strouhal number $St$ and jet separation $s$. Increasing both $M_j$ and $T_R$ independently is found to augment the jet coupling and modify the $(St,s)$ map of the preferred oscillation mode. Present results predict the preference of two modes when the jet interaction is relevant, namely varicose and especially sinuous flapping oscillations on the nozzles plane.

physics.flu-dyn

Linear recurrent cryptography: golden-like cryptography for higher order linear recurrences

We develop matrix cryptography based on linear recurrent sequences of any order that allows securing encryption against brute force and chosen plaintext attacks. In particular, we solve the problem of generalizing error detection and correction algorithms of golden cryptography previously known only for recurrences of a special form. They are based on proving the checking relations (inequalities satisfied by the ciphertext) under the condition that the analog of the golden $Q$-matrix has the strong Perron-Frobenius property. These algorithms are proved to be especially efficient when the characteristic polynomial of the recurrence is a Pisot polynomial. Finally, we outline algorithms for generating recurrences that satisfy our conditions.

cs.CR

Software defect prediction with zero-inflated Poisson models

In this work we apply several Poisson and zero-inflated models for software defect prediction. We apply different functions from several R packages such as pscl, MASS, R2Jags and the recent glmmTMB. We test the functions using the Equinox dataset. The results show that Zero-inflated models, fitted with either maximum likelihood estimation or with Bayesian approach, are slightly better than other models, using the AIC as selection criterion.

cs.SE

Distributed Correlation-Based Feature Selection in Spark

CFS (Correlation-Based Feature Selection) is an FS algorithm that has been successfully applied to classification problems in many domains. We describe Distributed CFS (DiCFS) as a completely redesigned, scalable, parallel and distributed version of the CFS algorithm, capable of dealing with the large volumes of data typical of big data applications. Two versions of the algorithm were implemented and compared using the Apache Spark cluster computing model, currently gaining popularity due to its much faster processing times than Hadoop's MapReduce model. We tested our algorithms on four publicly available datasets, each consisting of a large number of instances and two also consisting of a large number of features. The results show that our algorithms were superior in terms of both time-efficiency and scalability. In leveraging a computer cluster, they were able to handle larger datasets than the non-distributed WEKA version while maintaining the quality of the results, i.e., exactly the same features were returned by our algorithms when compared to the original algorithm available in WEKA.

cs.LG

Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale

Labeling training data is one of the most costly bottlenecks in developing machine learning-based applications. We present a first-of-its-kind study showing how existing knowledge resources from across an organization can be used as weak supervision in order to bring development time and cost down by an order of magnitude, and introduce Snorkel DryBell, a new weak supervision management system for this setting. Snorkel DryBell builds on the Snorkel framework, extending it in three critical aspects: flexible, template-based ingestion of diverse organizational knowledge, cross-feature production serving, and scalable, sampling-free execution. On three classification tasks at Google, we find that Snorkel DryBell creates classifiers of comparable quality to ones trained with tens of thousands of hand-labeled examples, converts non-servable organizational resources to servable models for an average 52% performance improvement, and executes over millions of data points in tens of minutes.

cs.LG

Distributed ReliefF based Feature Selection in Spark

Feature selection (FS) is a key research area in the machine learning and data mining fields, removing irrelevant and redundant features usually helps to reduce the effort required to process a dataset while maintaining or even improving the processing algorithm's accuracy. However, traditional algorithms designed for executing on a single machine lack scalability to deal with the increasing amount of data that has become available in the current Big Data era. ReliefF is one of the most important algorithms successfully implemented in many FS applications. In this paper, we present a completely redesigned distributed version of the popular ReliefF algorithm based on the novel Spark cluster computing model that we have called DiReliefF. Spark is increasing its popularity due to its much faster processing times compared with Hadoop's MapReduce model implementation. The effectiveness of our proposal is tested on four publicly available datasets, all of them with a large number of instances and two of them with also a large number of features. Subsets of these datasets were also used to compare the results to a non-distributed implementation of the algorithm. The results show that the non-distributed implementation is unable to handle such large volumes of data without specialized hardware, while our design can process them in a scalable way with much better processing times and memory usage.

cs.LG

Merge Non-Dominated Sorting Algorithm for Many-Objective Optimization

Many Pareto-based multi-objective evolutionary algorithms require to rank the solutions of the population in each iteration according to the dominance principle, what can become a costly operation particularly in the case of dealing with many-objective optimization problems. In this paper, we present a new efficient algorithm for computing the non-dominated sorting procedure, called Merge Non-Dominated Sorting (MNDS), which has a best computational complexity of $\Theta(NlogN)$ and a worst computational complexity of $\Theta(MN^2)$. Our approach is based on the computation of the dominance set of each solution by taking advantage of the characteristics of the merge sort algorithm. We compare the MNDS against four well-known techniques that can be considered as the state-of-the-art. The results indicate that the MNDS algorithm outperforms the other techniques in terms of number of comparisons as well as the total running time.

cs.NE

The HOD Dichotomy

This paper provides an accessible introduction to some of the work of Woodin on suitable extender models. We define the HOD conjecture, prove it is equivalent to a formulation in terms of weak extender models for supercompactness, and give some of its consequences.

math.LO

Polytope conditioning and linear convergence of the Frank-Wolfe algorithm

It is known that the gradient descent algorithm converges linearly when applied to a strongly convex function with Lipschitz gradient. In this case the algorithm's rate of convergence is determined by the condition number of the function. In a similar vein, it has been shown that a variant of the Frank-Wolfe algorithm with away steps converges linearly when applied to a strongly convex function with Lipschitz gradient over a polytope. In a nice extension of the unconstrained case, the algorithm's rate of convergence is determined by the product of the condition number of the function and a certain condition number of the polytope. We shed new light into the latter type of polytope conditioning. In particular, we show that previous and seemingly different approaches to define a suitable condition measure for the polytope are essentially equivalent to each other. Perhaps more interesting, they can all be unified via a parameter of the polytope that formalizes a key premise linked to the algorithm's linear convergence. We also give new insight into the linear convergence property. For a convex quadratic objective, we show that the rate of convergence is determined by a condition number of a suitably scaled polytope.

math.OC

On the von Neumann and Frank-Wolfe Algorithms with Away Steps

The von Neumann algorithm is a simple coordinate-descent algorithm to determine whether the origin belongs to a polytope generated by a finite set of points. When the origin is in the of the polytope, the algorithm generates a sequence of points in the polytope that converges linearly to zero. The algorithm's rate of convergence depends on the radius of the largest ball around the origin contained in the polytope. We show that under the weaker condition that the origin is in the polytope, possibly on its boundary, a variant of the von Neumann algorithm that includes generates a sequence of points in the polytope that converges linearly to zero. The new algorithm's rate of convergence depends on a certain geometric parameter of the polytope that extends the above radius but is always positive. Our linear convergence result and geometric insights also extend to a variant of the Frank-Wolfe algorithm with away steps for minimizing a strongly convex function over a polytope.

math.OC

On the biological and cultural evolution of shame: Using internet search tools to weight values in many cultures

Shame has clear biological roots and its precise form of expression affects social cohesion and cultural characteristics. Here we explore the relative importance between shame and guilt by using Google Translate to produce translation for the words shame, guilt, pain, embarrassment and fear to the 64 languages covered. We also explore the meanings of these concepts among the Yanomami, a horticulturist hunter-gatherer tribe in the Orinoquia. Results show that societies previously described as 'guilt societies' have more words for guilt than for shame, but the large majority, including the societies previously described as 'shame societies', have more words for shame than for guilt. Results are consistent with evolutionary models of shame which predict a wide scatter in the relative importance between guilt and shame, suggesting that cultural evolution of shame has continued the work of biological evolution, and that neither provides a strong adaptive advantage to either shame or guilt. We propose that the study of shame will improve our understanding of the interaction between biological and cultural evolution in the evolution of cognition and emotions.

cs.CY