Search arXivSearch

arXiv · 2511.09805

Pervasive Label Errors in Seismological Machine Learning Datasets

Abstract

The recent boom in artificial intelligence and machine learning has been powered by large datasets with accurate labels, combined with algorithmic advances and efficient computing. The quality of data can be a major factor in determining model performance. Here, we detail observations of commonly occurring errors in popular seismological machine learning datasets. We used an ensemble of available deep learning models PhaseNet and EQTransformer to evaluate the dataset labels and found four types of errors ranked from most prevalent to least prevalent: (1) unlabeled earthquakes; (2) noise samples that contain earthquakes; (3) inaccurately labeled arrival times, and (4) absent earthquake signals. We checked a total of 8.6 million examples from the following datasets: Iquique, ETHZ, PNW, TXED, STEAD, INSTANCE, AQ2009, and CEED. The average error rate across all datasets is 3.9 %, ranging from nearly zero to 8 % for individual datasets. These faulty data and labels are likely to degrade model training and performance. By flagging these errors, we aim to increase the quality of the data used to train machine learning models, especially for the measurement of arrival times, and thereby to improve the reliability of the models. We present a companion list of examples that contain problems, aiming to integrate them into training routines so that only the reliable data is used for training.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Albert Leonardo Aguilar Suarez, Gregory Beroza. 2025-11-12. Pervasive Label Errors in Seismological Machine Learning Datasets. https://arxiv.org/abs/2511.09805

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Direction-Aware Masked Pretraining for 3D Seismic Representation Learning and Transfer to Cross-Area Acoustic Impedance Inversion

Large archives of unlabeled three-dimensional seismic data offer opportunities for self-supervised representation learning and subsequent transfer to acoustic impedance inversion. However, conventional masked pretraining often treats three axes equivalently, overlooking differences between lateral reflector structure and vertical waveform characteristics. We propose a direction-aware masked autoencoder for three-dimensional post-stack seismic data, combining anisotropic tokenization, direction-aware representation, trace-aligned tube masking, and reconstruction constraints designed for reflector continuity and waveform characteristics. We evaluate reconstruction quality and downstream transferability using field data through masked reconstruction and cross-area acoustic impedance inversion. Reconstruction is more sensitive to lateral token resolution than to moderate changes in vertical patch length. Within the evaluated configurations, increasing encoder capacity does not fully compensate for reconstruction fidelity loss associated with coarser tokenization. Preferred token scales and masking strategies differ between reconstruction and inversion, indicating that reconstruction fidelity alone is not a reliable indicator of transferability. For cross-area inversion, the pretrained model is fine-tuned in the source area and applied to the target area without further parameter updates. With limited target-area well control, the proposed framework reduces normalized root-mean-square error by 20.8% relative to a pretrained conventional masked autoencoder across eight target-area test wells under matched tokenization and downstream settings. These results demonstrate the value of direction-aware masked pretraining for field seismic inversion and show that token scale and masking strategy should be selected according to downstream-task requirements rather than reconstruction accuracy alone.

physics.geo-ph

Time Distribution of Heavy Rainfall in Brazil: Empirical Huff Curves from 290,164 Sub-Daily Storm Events

Temporal rainfall distributions are widely used in design storm construction; however, many countries with limited sub-daily observations, such as Brazil, still rely on frameworks developed under different hydroclimatic conditions. This mismatch may introduce bias in hydrological design and water resources estimation. Here, we develop the first national-scale empirical Huff curves for Brazil using sub-daily rainfall observations. We compiled data from 3,164 stations and, after quality control, retained 290,164 storm events from 1,045 stations spanning 2010 to 2025. Empirical cumulative rainfall mass curves were constructed and fitted using seventh-degree polynomials at station, biome, state, and municipality scales. Our findings show a dominance of first-quartile (Q1; front-loaded) storm patterns, occurring at 94.4% of stations nationally, increasing to 99.2% in the Amazon and Cerrado biomes and decreasing to 88.7% in the Atlantic Forest. The national Q1 median curve closely matches the Huff (1967) reference (MAE = 0.045; Dmax = 0.097), with narrow bootstrap uncertainty. Q1 dominance is robust to inter-event time definition, with 84.7% of stations showing consistent classification across 2 to 12 h thresholds. A Soil Conservation Service Curve Number experiment across 579 headwater catchments shows that Brazilian curves increase design peak discharge by a median of 8% and up to 11% in the Cerrado relative to the Illinois reference, indicating potential underestimation when using non-local distributions. Biome-, state-, and municipality-scale parameters are provided as open data and through an interactive platform, offering locally calibrated design-storm alternatives for Brazil.

physics.geo-ph

Predicting the Elastic Properties of a Cemented Granular Material during Chemical Damage (Debonding)

While underground reservoirs emerge as essential elements to face global warming, these systems represent complex multi-physical and multiscale problems. The considered injection of fluids during hydrogen storage, carbon dioxide sequestration, or geothermal energy recovery involves a modification of the chemical equilibrium of the fluid in the porous reservoir. Chemical reactions can induce microstructural changes of the rock matrix, leading to a reduction of elastic properties of the material, and to potential settlement or stress redistribution. Consequently, it becomes pivotal to establish predictive behavior laws to describe the effect of chemical damage on elastic properties. Facing the difficulties to estimate experimentally the impact of chemical damage on mechanical properties, a Digital Rock Physics approach is proposed in this contribution. This numerical homogenization scheme is used to compare two distinct types of microstructure models: the first one consists in a Discrete Element Model, while the second one employs a continuous description. This continuous formulation is based on a Phase-Field description to predict the evolution of the microstructure subjected to chemical alterations and on the Fast Fourier Transform to estimate the macroscopic properties of the material. Finally, these frameworks establish different softening laws that can be used as constitutive ingredients for a cemented material during its weathering.

physics.geo-ph