Search arXivSearch

arXiv · 1907.01800

P2P Loan acceptance and default prediction with Artificial Intelligence

Abstract

Logistic Regression and Support Vector Machine algorithms, together with Linear and Non-Linear Deep Neural Networks, are applied to lending data in order to replicate lender acceptance of loans and predict the likelihood of default of issued loans. A two phase model is proposed; the first phase predicts loan rejection, while the second one predicts default risk for approved loans. Logistic Regression was found to be the best performer for the first phase, with test set recall macro score of $77.4 \%$. Deep Neural Networks were applied to the second phase only, were they achieved best performance, with validation set recall score of $72 \%$, for defaults. This shows that AI can improve current credit risk models reducing the default risk of issued loans by as much as $70 \%$. The models were also applied to loans taken for small businesses alone. The first phase of the model performs significantly better when trained on the whole dataset. Instead, the second phase performs significantly better when trained on the small business subset. This suggests a potential discrepancy between how these loans are screened and how they should be analysed in terms of default prediction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jeremy D. Turiel, Tomaso Aste. 2019-07-03. P2P Loan acceptance and default prediction with Artificial Intelligence. https://arxiv.org/abs/1907.01800

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Quantifying the 2027 Solvency II Risk Margin Reform

The 2027 Solvency II reform recalibrates the Risk Margin by reducing the prescribed cost-of-capital rate from 6% to 4.75% and introducing a time-dependent attenuation of future Solvency Capital Requirements. This paper develops an analytical and numerical framework for characterizing the effect of the final regulatory calibration. By normalizing discounted projected capital requirements into a probability distribution over run-off time, we obtain an exact representation of the ratio between the revised and previous Risk Margins. The framework yields sharp universal and horizon-specific bounds, characterizes the effect of later capital timing through stochastic dominance, and shows that mean run-off time alone does not determine the reform effect when temporal dispersion varies. Additional bounds are derived conditional on mean run-off time and horizon. Conditional on a given projected SCR path, the mechanical reduction lies between 20.83% and 60.42%. For proportional Best Estimate projections, an exact covariance decomposition identifies how departures from proportionality affect the relative reform ratio. Holding the projected SCR path fixed, the revised formula has a lower direct interest-rate semi-elasticity than the previous calibration. A reduced-form stochastic extension further quantifies convexity effects arising from uncertainty in capital persistence. Numerical applications reconstruct Risk Margin calculations from published actuarial run-off profiles and complement them with controlled long-horizon, interest-rate, and persistence experiments. The results show that the reform effect is governed by the temporal structure of future capital and provide tractable tools for assessing projected capital profiles, Risk Margin simplifications, and direct discount-rate sensitivity.

q-fin.RM

Towards foundation models for insurance risk modelling

Claim narratives, images and sensor data contain information about insured risks that is difficult to use through existing actuarial models. Foundation models learn patterns from large datasets before being adapted to particular tasks. By turning these high-dimensional sources into variables or numerical representations, they could help insurers use more of the information they already collect, potentially reducing the experience needed to develop each application. For example, a language model could identify a worsening injury in a new claim note, allowing a reserving model to recognise the change in expected cost before the payments reveal the deterioration. In this paper, we review language, vision, geospatial, time series, tabular and scientific models, explaining existing insurance applications and potential future uses. Scientific models extend this approach to future weather and climate conditions: their simulations can inform loss estimates once local hazards are linked to asset damage, repair costs and insurance coverage. We propose a process to connect these model outputs to actuarial calculations and to assess their predictive contribution, stability and compliance with rules on information use. Evaluating these applications is difficult when final claim costs become known only after long delays, large losses are rare or patterns learned elsewhere fail to transfer to the target portfolio. Richer data can reveal private information and support finer risk classification, which can change access to insurance. Reusing the same models across insurers also creates dependence on shared predictions and providers.

q-fin.RM

Forecasting Liquidity Withdraw with Machine Learning Models

Liquidity withdrawal is a critical indicator of market fragility. In this project, I test a framework for forecasting liquidity withdrawal at the individual-stock level, ranging from less liquid stocks to highly liquid large-cap tickers, and evaluate the relative performance of competing model classes in predicting short-horizon order book stress. We introduce the Liquidity Withdrawal Index (LWI) -- defined as the ratio of order cancellations to the sum of standing depth and new additions at the best quotes -- as a bounded, interpretable measure of transient liquidity removal. Using Nasdaq market-by-order (MBO) data, we compare a spectrum of approaches: linear benchmarks (AR, HAR), and non-linear tree ensembles (XGBoost), across horizons ranging from 250\,ms to 5\,s. Beyond predictive accuracy, our results provide insights into order placement and cancellation dynamics, identify regimes where linear versus non-linear signals dominate, and highlight how early-warning indicators of liquidity withdrawal can inform both market surveillance and execution.

q-fin.RM