Search arXiv⌕ Search

arXiv subjects

Ahmed El-Kotory

Publications and source records attributed to Ahmed El-Kotory.

5 recordsLinked to original sources

Three-group variance-ratio tests for heteroscedasticity in linear regression: an exact null distribution and an outlier-resistant version

Tests for heteroscedasticity in linear regression lose their level or power in three situations: when the data contain outliers, when the variance is not monotone, and when it changes along a variable outside the regressors, such as time. The proposed tests sort the observations by any chosen variable, split them into three equal parts, fit the regression in each part and compare the largest with the smallest error scale. With least squares fits and normal errors, the ratio follows Hartley's maximum F distribution with three groups exactly. With least trimmed squares fits, the ratio resists outliers spread along the ordering. Its square is approximately a maximum F ratio with effective degrees of freedom, whose limit we derive in closed form from the influence function of the trimmed variance. A factorial simulation covered 72 settings of variance shape, ordering variable, number of regressors, contamination and sample size. The robust test had the highest mean size-adjusted power, 67%, against 45% for the next test; when the variance changed along a regressor, which all tests were given, it led White's test after an outlier screen, 67% against 56%. With heavy-tailed t3 errors its lead was similar, 64% against 35%. Unlike the tests built on the regressors, it can follow a variance changing along time, and it identifies the part of the data where the variance changes. Two published robust versions of the Goldfeld-Quandt test, as implemented from their descriptions, did not hold their nominal level. The R package KOTORY implements the methods.

stat.ME↗

Foundation or Formula? A Simulation-Based Comparison of Google's TimesFM-3 and Classical Time-Series Models

Time-series foundation models such as Google's TimesFM-3 forecast series they were never trained on, but whether they should replace classical methods such as ARIMA and exponential smoothing is hard to settle on public benchmarks, because few benchmark datasets are absent from every model's pre-training corpus. This exploratory study compares TimesFM-3 with classical methods on data generated from nine known processes, which the model cannot have seen. Across 7200 series and a pre-registered protocol with a 12-step horizon, neither family dominates. TimesFM-3's error is never more than 1.31 times that of the best method in a scenario, whereas every classical method's is at least 2.2 times somewhere; the largest classical failures occur with only two seasonal cycles, where the automatic methods fall back to non-seasonal models. With four or more cycles automatic ARIMA is 21-24% more accurate than TimesFM-3 on seasonal ARIMA data. Among five foundation models this robustness is specific to TimesFM-3, and it holds only at short horizons: at steps 25 to 48 the three foundation models tested there forecast a decline on a saturating process whose level stays flat, and automatic ARIMA has the smaller worst case elsewhere. On intermittent demand TimesFM-3 shows no advantage over simple benchmarks built from the history. Randomised parameters, heavy-tailed noise and outliers leave these patterns qualitatively unchanged; on the official test period of 1000 M4 monthly series TimesFM-3 is level with the fifth- to seventh-ranked M4 entries, behind the four best, and on 101 macroeconomic series observed after the models' documented training data it is statistically level with the classical methods. The study characterises the behaviour of a black-box model, not the reasons for it, and its findings are conditional on the processes and horizons examined.

stat.ME↗

EDGE: a closed-form directed test for the calibration of probabilistic binary classifiers

A probabilistic binary classifier is judged almost everywhere by discrimination - accuracy, the ROC curve, the area under it. Every such criterion is invariant to a monotone distortion of the predicted probabilities, so a classifier can rank perfectly and still return probabilities that are badly wrong. Calibration is the property decisions need, and the field's instrument for it, the binned expected calibration error with its reliability diagram, is descriptive: it has no null distribution, so it cannot say whether the miscalibration it displays is real or noise, and it depends on the binning. We propose EDGE, a calibration test for the canonical probabilistic classifier, logistic regression. EDGE reads the same binned predicted-versus-observed table a reliability diagram plots, and projects its standardized bin residuals onto a small pre-specified basis of smooth calibration-distortion shapes. Its null distribution is a weighted sum of chi-square variables in closed form, costing one pass over the data and one small eigendecomposition: no refit, no resampling, no tuning, so it can run inside cross-validation loops. Binning also makes it robust to the sparsity continuous features create. Across link and feature misspecification the pre-specified default led or tied every rival binned test on the fitted index in 19 of 22 detectable scenarios, and stayed computable where the refit-based Stukel score test separates in 20% to 28% of sparse samples. Its honest limit is rough, high-frequency miscalibration, where omnibus statistics win - a limit an elementary resolution argument shows is shared by every binned instrument, the calibration error included.

stat.ME↗

A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression

Goodness-of-fit assessment for the binary logistic regression model is difficult when covariates are continuous: the data are effectively sparse, the classical Pearson and deviance tests fail, and practitioners rely on partition-based tests, such as the Hosmer-Lemeshow test, that group observations before comparing observed and expected counts. We study a partition test that modifies the Hosmer-Lemeshow statistic with a single directional correction term, weighted by $(1-2\barπ_g)$ and referred to a $χ^2_{G-2}$ distribution. The correction is the grouped form of the Osius-Rojek/Farrington standardization; grouping makes it well defined in the sparse regime, and it targets the asymmetric over- and under-prediction that a misspecified link induces. A single alignment functional captures its effect, predicting where the test gains power (asymmetric-link misspecification) and where it does not (symmetric departures, and covariate-space structure that no probability-grouping test can see). In simulations the test holds its size; no well-calibrated partition test is more sensitive to asymmetric-link misfit, and it clearly exceeds Hosmer-Lemeshow there, most so for the complementary log-log link -- a modest gain that fades as $n$ grows; it ties Hosmer-Lemeshow on an omitted interaction and is less powerful on an omitted quadratic (by about ten percentage points at $n=1000$). A real-data application illustrates its use, and the test is implemented in the R package ebrahim.gof.

stat.ME↗

Benchmarking Goodness-of-Fit and Calibration Algorithms for Logistic Regression Classifiers: A Large-Scale Simulation Study under Sparse Data

Binary logistic regression is among the most widely used classification algorithms, yet a classifier is only trustworthy if its predicted probabilities are well calibrated. The classical checks -- the Pearson chi-square and deviance statistics -- break down precisely in the modern setting where predictors are continuous and the data are sparse (one covariate pattern per observation). Four decades of research have produced dozens of alternative goodness-of-fit and calibration algorithms, yet practitioners still default to the Hosmer-Lemeshow test because it ships with their software. This paper provides a unified taxonomy and a large-scale, reproducible simulation benchmark; more than twenty tests are implemented in the open-source R package ebrahim.gof. We evaluate them across five covariate distributions and four misspecification scenarios, with 10,000 replications each, measuring both Type I error and power. Several classical tests prove liberal, rejecting correct models far too often, while others have little power. A compact core -- McCullagh, Osius-Rojek, le Cessie-van Houwelingen, Stute-Zhu, and the GiViTI calibration test -- delivers the best balance of correct size and high power, and is consistently more powerful than the ubiquitous Hosmer-Lemeshow test. A low-birth-weight application reinforces the point: a model with omitted interactions slips past nearly every test, exposed only by pairing sensitive tests with a calibration (reliability) curve. We translate these findings into practical, evidence-based guidance for assessing logistic regression fit.

stat.ME↗