Search arXivSearch

arXiv subjects

Martin Neil

Publications and source records attributed to Martin Neil.

At least 19 recordsLinked to original sources

The biased interaction game: Its dynamics and application in modelling social systems

The biased interaction game described the operation of systems rooted in boundedly rational interactions under conditions of scarcity. The game explored the influence of bias and demonstrated how hierarchy and inequality are emergent system properties when sources of bias, such as power and scarcity, affect the outcome of interactions in an environment. Bias also impacts the likelihood of the emergence of cooperation. This paper serves as a companion piece to the paper introducing the biased interaction game. It investigates the general applicability of the game and demonstrates how the consideration of bias can modify and improve upon prior systems thinking. In particular, it shows how social systems can be successfully modelled using the biased interaction game and confirms its suitability for modelling extreme examples such as hyper-capitalism and social egalitarianism. It also reveals how biased systems can demonstrate non-linear behaviour, where long periods of system stability are punctuated by short bursts of rapid hierarchical transitions, mimicking real-world observations of social mobility. The paper concludes with a simplified real-world application, modelling the merits of two competing wealth redistribution philosophies: social welfare and a universal basic income

cs.GT

Stacking Factorizing Partitioned Expressions in Hybrid Bayesian Network Models

Hybrid Bayesian networks (HBN) contain complex conditional probabilistic distributions (CPD) specified as partitioned expressions over discrete and continuous variables. The size of these CPDs grows exponentially with the number of parent nodes when using discrete inference, resulting in significant inefficiency. Normally, an effective way to reduce the CPD size is to use a binary factorization (BF) algorithm to decompose the statistical or arithmetic functions in the CPD by factorizing the number of connected parent nodes to sets of size two. However, the BF algorithm was not designed to handle partitioned expressions. Hence, we propose a new algorithm called stacking factorization (SF) to decompose the partitioned expressions. The SF algorithm creates intermediate nodes to incrementally reconstruct the densities in the original partitioned expression, allowing no more than two continuous parent nodes to be connected to each child node in the resulting HBN. SF can be either used independently or combined with the BF algorithm. We show that the SF+BF algorithm significantly reduces the CPD size and contributes to lowering the tree-width of a model, thus improving efficiency.

cs.AI

A hybrid Bayesian network for medical device risk assessment and management

ISO 14971 is the primary standard used for medical device risk management. While it specifies the requirements for medical device risk management, it does not specify a particular method for performing risk management. Hence, medical device manufacturers are free to develop or use any appropriate methods for managing the risk of medical devices. The most commonly used methods, such as Fault Tree Analysis (FTA), are unable to provide a reasonable basis for computing risk estimates when there are limited or no historical data available or where there is second-order uncertainty about the data. In this paper, we present a novel method for medical device risk management using hybrid Bayesian networks (BNs) that resolves the limitations of classical methods such as FTA and incorporates relevant factors affecting the risk of medical devices. The proposed BN method is generic but can be instantiated on a system-by-system basis, and we apply it to a Defibrillator device to demonstrate the process involved for medical device risk management during production and post-production. The example is validated against real-world data.

cs.LG

Product safety idioms: a method for building causal Bayesian networks for product safety and risk assessment

Idioms are small, reusable Bayesian network (BN) fragments that represent generic types of uncertain reasoning. This paper shows how idioms can be used to build causal BNs for product safety and risk assessment that use a combination of data and knowledge. We show that the specific product safety idioms that we introduce are sufficient to build full BN models to evaluate safety and risk for a wide range of products. The resulting models can be used by safety regulators and product manufacturers even when there are limited (or no) product testing data.

cs.AI

Smart Automotive Technology Adherence to the Law: (De)Constructing Road Rules for Autonomous System Development, Verification and Safety

Driving is an intuitive task that requires skills, constant alertness and vigilance for unexpected events. The driving task also requires long concentration spans focusing on the entire task for prolonged periods, and sophisticated negotiation skills with other road users, including wild animals. These requirements are particularly important when approaching intersections, overtaking, giving way, merging, turning and while adhering to the vast body of road rules. Modern motor vehicles now include an array of smart assistive and autonomous driving systems capable of subsuming some, most, or in limited cases, all of the driving task. The UK Department of Transport's response to the Safe Use of Automated Lane Keeping System consultation proposes that these systems are tested for compliance with relevant traffic rules. Building these smart automotive systems requires software developers with highly technical software engineering skills, and now a lawyer's in-depth knowledge of traffic legislation as well. These skills are required to ensure the systems are able to safely perform their tasks while being observant of the law. This paper presents an approach for deconstructing the complicated legalese of traffic law and representing its requirements and flow. The approach (de)constructs road rules in legal terminology and specifies them in structured English logic that is expressed as Boolean logic for automation and Lawmaps for visualisation. We demonstrate an example using these tools leading to the construction and validation of a Bayesian Network model. We strongly believe these tools to be approachable by programmers and the general public, and capable of use in developing Artificial Intelligence to underpin motor vehicle smart systems, and in validation to ensure these systems are considerate of the law when making decisions.

cs.AI

Bayesian hypothesis testing and hierarchical modelling of ivermectin effectiveness in treating Covid-19

Ivermectin is an antiparasitic drug that some have claimed is an effective treatment for reducing Covid-19 deaths. To test this claim, two recent peer reviewed papers both conducted a meta-analysis on a similar set of randomized controlled trials data, applying the same classical statistical approach. Although the statistical results were similar, one of the papers (Bryant et al, 2021) concluded that ivermectin was effective for reducing Covid-19 deaths, while the other (Roman et al, 2021) concluded that there was insufficient quality of evidence to support the conclusion Ivermectin was effective. This paper applies a Bayesian approach, to a subset of the same trial data, to test several causal hypotheses linking Covid-19 severity and ivermectin to mortality and produce an alternative analysis to the classical approach. Applying diverse alternative analysis methods which reach the same conclusions should increase overall confidence in the result. We show that there is strong evidence to support a causal link between ivermectin, Covid-19 severity and mortality, and: i) for severe Covid-19 there is a 90.7% probability the risk ratio favours ivermectin; ii) for mild/moderate Covid-19 there is an 84.1% probability the risk ratio favours ivermectin. To address concerns expressed about the veracity of some of the studies we evaluate the sensitivity of the conclusions to any single study by removing one study at a time. In the worst case, where (Elgazzar 2020) is removed, the results remain robust, for both severe and mild to moderate Covid-19. The paper also highlights advantages of using Bayesian methods over classical statistical methods for meta-analysis. All studies included in the analysis were prior to data on the delta variant.

q-bio.QM

Calculating the Likelihood Ratio for Multiple Pieces of Evidence

When presenting forensic evidence, such as a DNA match, experts often use the Likelihood ratio (LR) to explain the impact of evidence . The LR measures the probative value of the evidence with respect to a single hypothesis such as 'DNA comes from the suspect', and is defined as the probability of the evidence if the hypothesis is true divided by the probability of the evidence if the hypothesis is false. The LR is a valid measure of probative value because, by Bayes Theorem, the higher the LR is, the more our belief in the probability the hypothesis is true increases after observing the evidence. The LR is popular because it measures the probative value of evidence without having to make any explicit assumptions about the prior probability of the hypothesis. However, whereas the LR can in principle be easily calculated for a distinct single piece of evidence that relates directly to a specific hypothesis, in most realistic situations 'the evidence' is made up of multiple dependent components that impact multiple different hypotheses. In such situations the LR cannot be calculated . However, once the multiple pieces of evidence and hypotheses are modelled as a causal Bayesian network (BN), any relevant LR can be automatically derived using any BN software application.

stat.AP

A Bayesian-network-based cybersecurity adversarial risk analysis framework with numerical examples

Cybersecurity risk analysis plays an essential role in supporting organizations make effective decision about how to manage and control cybersecurity risk. Cybersecurity risk is a function of the interplay between the defender, i.e., the organisation, and the attacker: decisions and actions made by the defender second guess the decisions and actions taken by the attacker and vice versa. Insight into this game between these two agents provides a means for the defender to identify and make optimal decisions. To date, the adversarial risk analysis framework has provided a decision-analytical approach to solve such game problems in the presence of uncertainty and uses Monte Carlo simulation to calculate and identify optimal decisions. We propose an alternative framework to construct and solve a serial of sequential Defend-Attack models, that incorporates the adversarial risk analysis approach, but uses a new class of influence diagrams algorithm, called hybrid Bayesian network inference, to identify optimal decision strategies. Compared to Monte Carlo simulation the proposed hybrid Bayesian network inference is more versatile because it provides an automated way to compute hybrid Defend-Attack models and extends their use to involve mixtures of continuous and discrete variables, of any kind. More importantly, the hybrid Bayesian network approach is novel in that it supports dynamic decision making whereby new real-time observations can update the Defend-Attack model in practice. We also extend the Defend-Attack model to support cases involving extra variables and longer decision sequence. Examples are presented, illustrating how the proposed framework can be adjusted for more complicated scenarios, including dynamic decision making.

cs.GT

Positive results from UK single gene testing for SARS-COV-2 may be inconclusive, negative or detecting past infections

The UK Office for National Statistics (ONS) publish a regular infection survey that reports data on positive RT-PCR test results for SARS-COV-2 virus. This survey reports that a large proportion of positive test results may be based on the detection of a single target gene rather than on two or more target genes as required in the manufacturer instructions for use, and by the WHO in their emergency use assessment. Without diagnostic validation, for both the original virus and any variants, it is not clear what can be concluded from a positive test resulting from a single target gene call, especially if there was no confirmatory testing. Given this, many of the reported positive results may be inconclusive, negative or from people who suffered past infection for SARS-COV-2.

stat.AP

How do some Bayesian Network machine learned graphs compare to causal knowledge?

The graph of a Bayesian Network (BN) can be machine learned, determined by causal knowledge, or a combination of both. In disciplines like bioinformatics, applying BN structure learning algorithms can reveal new insights that would otherwise remain unknown. However, these algorithms are less effective when the input data are limited in terms of sample size, which is often the case when working with real data. This paper focuses on purely machine learned and purely knowledge-based BNs and investigates their differences in terms of graphical structure and how well the implied statistical models explain the data. The tests are based on four previous case studies whose BN structure was determined by domain knowledge. Using various metrics, we compare the knowledge-based graphs to the machine learned graphs generated from various algorithms implemented in TETRAD spanning all three classes of learning. The results show that, while the algorithms produce graphs with much higher model selection score, the knowledge-based graphs are more accurate predictors of variables of interest. Maximising score fitting is ineffective in the presence of limited sample size because the fitting becomes increasingly distorted with limited data, guiding algorithms towards graphical patterns that share higher fitting scores and yet deviate considerably from the true graph. This highlights the value of causal knowledge in these cases, as well as the need for more appropriate fitting scores suitable for limited data. Lastly, the experiments also provide new evidence that support the notion that results from simulated data tell us little about actual real-world performance.

cs.AI

Product risk assessment: a Bayesian network approach

Product risk assessment is the overall process of determining whether a product, which could be anything from a type of washing machine to a type of teddy bear, is judged safe for consumers to use. There are several methods used for product risk assessment, including RAPEX, which is the primary method used by regulators in the UK and EU. However, despite its widespread use, we identify several limitations of RAPEX including a limited approach to handling uncertainty and the inability to incorporate causal explanations for using and interpreting test data. In contrast, Bayesian Networks (BNs) are a rigorous, normative method for modelling uncertainty and causality which are already used for risk assessment in domains such as medicine and finance, as well as critical systems generally. This article proposes a BN model that provides an improved systematic method for product risk assessment that resolves the identified limitations with RAPEX. We use our proposed method to demonstrate risk assessments for a teddy bear and a new uncertified kettle for which there is no testing data and the number of product instances is unknown. We show that, while we can replicate the results of the RAPEX method, the BN approach is more powerful and flexible.

stat.OT

On the limitations of probabilistic claims about the probative value of mixed DNA profile evidence

The likelihood ratio (LR) is a commonly used measure for determining the strength of forensic match evidence. When a forensic expert determines a high LR for DNA found at a crime scene matching the DNA profile of a suspect they typically report that 'this provides strong support for the prosecution hypothesis that the DNA comes from the suspect'. However, even with a high LR, the evidence might not support the prosecution hypothesis if the defence hypothesis used to determine the LR is not the negation of the prosecution hypothesis (such as when the alternative is 'DNA comes from a person unrelated to the defendant' instead of 'DNA does not come from the suspect'). For DNA mixture profiles, especially low template DNA (LTDNA), the value of a high LR for a 'match' - typically computed from probabilistic genotyping software - can be especially questionable. But this is not just because of the use of non-exhaustive hypotheses in such cases. In contrast to single profile DNA 'matches', where the only residual uncertainty is whether a person other than the suspect has the same matching DNA profile, it is possible for all the genotypes of the suspect's DNA profile to appear at each locus of a DNA mixture, even though none of the contributors has that DNA profile. In fact, in the absence of other evidence, we show it is possible to have a very high LR for the hypothesis 'suspect is included in the mixture' even though the posterior probability that the suspect is included is very low. Yet, in such cases a forensic expert will generally still report a high LR as 'strong support for the suspect being a contributor'. Our observations suggest that, in certain circumstances, the use of the LR may have led lawyers and jurors into grossly overestimating the probative value of a LTDNA mixed profile 'match'

stat.AP

The role of collider bias in understanding statistics on racially biased policing

Contradictory conclusions have been made about whether unarmed blacks are more likely to be shot by police than unarmed whites using the same data. The problem is that, by relying only on data of 'police encounters', there is the possibility that genuine bias can be hidden. We provide a causal Bayesian network model to explain this bias, which is called collider bias or Berkson's paradox, and show how the different conclusions arise from the same model and data. We also show that causal Bayesian networks provide the ideal formalism for considering alternative hypotheses and explanations of bias.

stat.AP

Medical idioms for clinical Bayesian network development

Bayesian Networks (BNs) are graphical probabilistic models that have proven popular in medical applications. While numerous medical BNs have been published, most are presented fait accompli without explanation of how the network structure was developed or justification of why it represents the correct structure for the given medical application. This means that the process of building medical BNs from experts is typically ad hoc and offers little opportunity for methodological improvement. This paper proposes generally applicable and reusable medical reasoning patterns to aid those developing medical BNs. The proposed method complements and extends the idiom-based approach introduced by Neil, Fenton, and Nielsen in 2000. We propose instances of their generic idioms that are specific to medical BNs. We refer to the proposed medical reasoning patterns as medical idioms. In addition, we extend the use of idioms to represent interventional and counterfactual reasoning. We believe that the proposed medical idioms are logical reasoning patterns that can be combined, reused and applied generically to help develop medical BNs. All proposed medical idioms have been illustrated using medical examples on coronary artery disease. The method has also been applied to other ongoing BNs being developed with medical experts. Finally, we show that applying the proposed medical idioms to published BN models results in models with a clearer structure.

cs.AI

Bluetooth Smartphone Apps: Are they the most private and effective solution for COVID-19 contact tracing?

Many digital solutions mainly involving Bluetooth technology are being proposed for Contact Tracing Apps (CTA) to reduce the spread of COVID-19. Concerns have been raised regarding privacy, consent, uptake required in a given population, and the degree to which use of CTAs can impact individual behaviours. However, very few groups have taken a holistic approach and presented a combined solution. None has presented their CTA in such a way as to ensure that even the most suggestible member of our community does not become complacent and assume that CTA operates as an invisible shield, making us and our families impenetrable or immune to the disease. We propose to build on some of the digital solutions already under development that, with addition of a Bayesian model that predicts likelihood for infection supplemented by traditional symptom and contact tracing, that can enable us to reach 90% of a population. When combined with an effective communication strategy and social distancing, we believe solutions like the one proposed here can have a very beneficial effect on containing the spread of this pandemic.

cs.CY

Simpson's Paradox and the implications for medical trials

This paper describes Simpson's paradox, and explains its serious implications for randomised control trials. In particular, we show that for any number of variables we can simulate the result of a controlled trial which uniformly points to one conclusion (such as 'drug is effective') for every possible combination of the variable states, but when a previously unobserved confounding variable is included every possible combination of the variables state points to the opposite conclusion ('drug is not effective'). In other words no matter how many variables are considered, and no matter how 'conclusive' the result, one cannot conclude the result is truly 'valid' since there is theoretically an unobserved confounding variable that could completely reverse the result.

stat.ME

Modelling Competing Legal Arguments using Bayesian Model Comparison and Averaging

Bayesian models of legal arguments generally aim to produce a single integrated model, combining each of the legal arguments under consideration. This combined approach implicitly assumes that variables and their relationships can be represented without any contradiction or misalignment, and in a way that makes sense with respect to the competing argument narratives. This paper describes a novel approach to compare and 'average' Bayesian models of legal arguments that have been built independently and with no attempt to make them consistent in terms of variables, causal assumptions or parametrisation. The approach involves assessing whether competing models of legal arguments are explained or predict facts uncovered before or during the trial process. Those models that are more heavily disconfirmed by the facts are given lower weight, as model plausibility measures, in the Bayesian model comparison and averaging framework adopted. In this way a plurality of arguments is allowed yet a single judgement based on all arguments is possible and rational.

stat.AP

Region Based Approximation for High Dimensional Bayesian Network Models

Performing efficient inference on Bayesian Networks (BNs), with large numbers of densely connected variables is challenging. With exact inference methods, such as the Junction Tree algorithm, clustering complexity can grow exponentially with the number of nodes and so computation becomes intractable. This paper presents a general purpose approximate inference algorithm called Triplet Region Construction (TRC) that reduces the clustering complexity for factorized models from worst case exponential to polynomial. We employ graph factorization to reduce connection complexity and produce clusters of limited size. Unlike MCMC algorithms TRC is guaranteed to converge and we present experiments that show that TRC achieves accurate results when compared with exact solutions.

cs.AI