Search arXiv⌕ Search

arXiv subjects

Usman Ahmed

Publications and source records attributed to Usman Ahmed.

4 recordsLinked to original sources

Adapting a Large Language Model Crash-Severity Pipeline to Tennessee: Performance Across Sampling Strategies

State crash databases differ in structure, coding, and injury-severity distributions, limiting direct reuse of predictive workflows across jurisdictions. This study adapts the SafeTraffic Copilot large language model (LLM) crash-severity workflow to a three-year Tennessee inventory of 624,392 crashes. Tennessee crash, roadway, vehicle, and person attributes were harmonized and converted into textual prompts while unavailable values were preserved rather than inferred. Llama 3.1 8B was fine-tuned using low-rank adaptation to classify five injury-severity categories. Random, county-based, and severity-balanced sampling strategies were evaluated using separate in-sample and unseen-test experiments; unseen-test experiments used 70/15/15 training, validation, and test splits. On their respective unseen test sets, random and county-based models achieved weighted F1-scores near 80%, but macro F1 remained below 47% and fatal-crash F1 below 27%, showing that strong aggregate performance can mask weak recognition of rare outcomes. Within its balanced evaluation population, the severity-balanced model achieved weighted and macro F1-scores of 57.7% and a fatalcrash F1-score of 65.9%, yielding more even class-level performance. Because each sampling strategy used a different test subset, crossstrategy differences are descriptive rather than controlled rankings. The results highlight the importance of interpreting LLM crashseverity performance together with sampling design, class balance, class-level metrics, and evaluation-population composition.

cs.CE↗

Residential Price Modeling using Spatially Validated Machine Learning Methods: A Comparison Across Geographic Contexts

Land use and transportation infrastructure are tightly linked systems, and residential property prices play a central role in transportation planning. Transportation investments also influence property values, making accurate forecasting of both systems an interconnected research challenge. This study examines these interactions and evaluates machine learning methods for modeling residential real estate prices across contrasting urban contexts. We use XGBoost and Random Forest models to assess how land use and transportation infrastructure shape dwelling prices, comparing results between the Rawalpindi and Islamabad Metropolitan Area in Pakistan and the City of Toronto in Canada. We also examine the performance of machine learning models on spatial data by comparing nonspatial and spatial cross validation, and use SHAP values to interpret feature impacts. Despite differences in demographics and economic development, both cities show similar effects of transportation infrastructure and local amenities on prices. Proximity to major urban cores increases sale price, while high quality transit such as subway and bus rapid transit raises prices and conventional bus stop proximity lowers them. Nonspatial cross validation overestimates predictive accuracy for spatial datasets. XGBoost performs slightly better than Random Forest. We recommend careful application of machine learning methods when modeling spatially dependent land price data.

econ.EM↗

A Pseudo Panel Difference-in-Differences (DiD) Analysis of Online Shopping Behavior in the Puget Sound Regional Council (PSRC) Region

Online shopping is a growing trend, particularly following the COVID-19 lock-downs enacted by many cities. Understanding these trends requires robust panel dat methods. However, panel data (i.e., with repeated measurements of the same observational units) are often unavailable. In this study, we use a propensity score weighting (PSW) approach to adjust repeated cross-sectional travel diary surveys collected in the Puget Sound Regional Council (PSRC) region. The most common home delivery is packages (2.09 days per week in 2023), followed by food (0.36 days per week in 2023). We find that single-family detached and townhouse residents tend to receive more home deliveries than those living in apartments. We also find a difference in several patterns pre- and post-COVID pandemic. Vehicle deficient households did not exhibit the same increase in home delivery frequency as other households pre-COVID, but the pattern reversed in the post-COVID period - i.e., vehicle deficient households saw a relative increase in delivery frequency. Overall, this study demonstrates the pseudo-panel approach to causal inference by leveraging differences in PSW-weighted in delivery frequency across treatment groups.

econ.EM↗

Density functional benchmark for quadruple hydrogen bonds

Hydrogen bonding is an important non-covalent interaction that plays a major role in molecular self-organization and supramolecular structures. It can be described accurately with ab initio quantum chemical wave function methods, which become computationally expensive for large molecular assemblies. Density functional theory (DFT) offers a better balance between accuracy and computational cost, and can be routinely applied to large systems. A large number of density functional approximations (DFAs) has been developed, but their accuracy depend on the application, necessitating benchmark studies to guide their selection for use in applications. Some of us have recently determined highly accurate hydrogen bonding energies of 14 quadruply hydrogen-bonded dimers by extrapolating coupled-cluster energies to the complete basis set limit as well as extrapolating electron correlation contributions with a continued-fraction approach [U. Ahmed et al, Phys. Chem. Chem. Phys., 2024, 26, 24470]. In this work, we study the reproduction of these bonding energies at the DFT level using 152 DFAs. The top ten functionals are composed of eight variants of the Berkeley functionals both with and without dispersion corrections, and two Minnesota 2011 functionals augmented with a further dispersion correction. We find the B97M-V functional with the non-local correlation functional replaced by an empirical D3BJ dispersion correction to be the best functional, while changes to the dispersion part in other Berkeley functionals lead to poorer performance in our study.

physics.chem-ph↗