Search arXivSearch

arXiv · 2210.15785

Supply Chain Characteristics as Predictors of Cyber Risk: A Machine-Learning Assessment

Abstract

This paper provides the first large-scale data-driven analysis to evaluate the predictive power of different attributes for assessing risk of cyberattack data breaches. Furthermore, motivated by rapid increase in third party enabled cyberattacks, the paper provides the first quantitative empirical evidence that digital supply-chain attributes are significant predictors of enterprise cyber risk. The paper leverages outside-in cyber risk scores that aim to capture the quality of the enterprise internal cybersecurity management, but augment these with supply chain features that are inspired by observed third party cyberattack scenarios, as well as concepts from network science research. The main quantitative result of the paper is to show that supply chain network features add significant detection power to predicting enterprise cyber risk, relative to merely using enterprise-only attributes. Particularly, compared to a base model that relies only on internal enterprise features, the supply chain network features improve the out-of-sample AUC by 2.3\%. Given that each cyber data breach is a low probability high impact risk event, these improvements in the prediction power have significant value. Additionally, the model highlights several cybersecurity risk drivers related to third party cyberattack and breach mechanisms and provides important insights as to what interventions might be effective to mitigate these risks.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kevin Hu, Retsef Levi, Raphael Yahalom, El Ghali Zerhouni. 2023-11-13. Supply Chain Characteristics as Predictors of Cyber Risk: A Machine-Learning Assessment. https://arxiv.org/abs/2210.15785

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Quantifying the 2027 Solvency II Risk Margin Reform

The 2027 Solvency II reform recalibrates the Risk Margin by reducing the prescribed cost-of-capital rate from 6% to 4.75% and introducing a time-dependent attenuation of future Solvency Capital Requirements. This paper develops an analytical and numerical framework for characterizing the effect of the final regulatory calibration. By normalizing discounted projected capital requirements into a probability distribution over run-off time, we obtain an exact representation of the ratio between the revised and previous Risk Margins. The framework yields sharp universal and horizon-specific bounds, characterizes the effect of later capital timing through stochastic dominance, and shows that mean run-off time alone does not determine the reform effect when temporal dispersion varies. Additional bounds are derived conditional on mean run-off time and horizon. Conditional on a given projected SCR path, the mechanical reduction lies between 20.83% and 60.42%. For proportional Best Estimate projections, an exact covariance decomposition identifies how departures from proportionality affect the relative reform ratio. Holding the projected SCR path fixed, the revised formula has a lower direct interest-rate semi-elasticity than the previous calibration. A reduced-form stochastic extension further quantifies convexity effects arising from uncertainty in capital persistence. Numerical applications reconstruct Risk Margin calculations from published actuarial run-off profiles and complement them with controlled long-horizon, interest-rate, and persistence experiments. The results show that the reform effect is governed by the temporal structure of future capital and provide tractable tools for assessing projected capital profiles, Risk Margin simplifications, and direct discount-rate sensitivity.

q-fin.RM

Towards foundation models for insurance risk modelling

Claim narratives, images and sensor data contain information about insured risks that is difficult to use through existing actuarial models. Foundation models learn patterns from large datasets before being adapted to particular tasks. By turning these high-dimensional sources into variables or numerical representations, they could help insurers use more of the information they already collect, potentially reducing the experience needed to develop each application. For example, a language model could identify a worsening injury in a new claim note, allowing a reserving model to recognise the change in expected cost before the payments reveal the deterioration. In this paper, we review language, vision, geospatial, time series, tabular and scientific models, explaining existing insurance applications and potential future uses. Scientific models extend this approach to future weather and climate conditions: their simulations can inform loss estimates once local hazards are linked to asset damage, repair costs and insurance coverage. We propose a process to connect these model outputs to actuarial calculations and to assess their predictive contribution, stability and compliance with rules on information use. Evaluating these applications is difficult when final claim costs become known only after long delays, large losses are rare or patterns learned elsewhere fail to transfer to the target portfolio. Richer data can reveal private information and support finer risk classification, which can change access to insurance. Reusing the same models across insurers also creates dependence on shared predictions and providers.

q-fin.RM

Forecasting Liquidity Withdraw with Machine Learning Models

Liquidity withdrawal is a critical indicator of market fragility. In this project, I test a framework for forecasting liquidity withdrawal at the individual-stock level, ranging from less liquid stocks to highly liquid large-cap tickers, and evaluate the relative performance of competing model classes in predicting short-horizon order book stress. We introduce the Liquidity Withdrawal Index (LWI) -- defined as the ratio of order cancellations to the sum of standing depth and new additions at the best quotes -- as a bounded, interpretable measure of transient liquidity removal. Using Nasdaq market-by-order (MBO) data, we compare a spectrum of approaches: linear benchmarks (AR, HAR), and non-linear tree ensembles (XGBoost), across horizons ranging from 250\,ms to 5\,s. Beyond predictive accuracy, our results provide insights into order placement and cancellation dynamics, identify regimes where linear versus non-linear signals dominate, and highlight how early-warning indicators of liquidity withdrawal can inform both market surveillance and execution.

q-fin.RM