Search arXiv⌕ Search

arXiv subjects

Ahmad Shahi

Publications and source records attributed to Ahmad Shahi.

2 recordsLinked to original sources

Conformal Adversarial Generative Ensemble

Accurate time series forecasting is critical across various domains, yet traditional ensemble methods often suffer from the disproportionate influence of extreme forecasts. We introduce the Conformal Adversarial Generative Ensemble (CAGE), a novel framework that combines generative modeling, adversarial discrimination, and conformal prediction to enhance forecast reliability and accuracy. CAGE employs multiple generative models to produce initial forecasts, which are then evaluated by a discriminative component using conformal prediction techniques. P-values derived from nonconformity scores help dynamically adjust model weights, minimizing the impact of unreliable forecasts. This approach ensures that only the most credible predictions contribute to the final ensemble output. Our empirical and statistical analyses of time series data from New Zealand's milk collection and the global health data from the public owid-monkeypox dataset show that the CAGE outperforms traditional ensemble methods, especially in handling outliers and noisy data. By incorporating conformal prediction, CAGE delivers accurate and statistically rigorous forecasts, enhancing decision-making. We have demonstrated performance on two different datasets deliberately to showcase that the proposed method offers a versatile solution potentially applicable across finance, weather, and supply chain management.

cs.LG↗

Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models

Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, particularly regarding bias and hallucination. In this work, we evaluate the robustness of LLMs to perturbed variations of the original inquiry in decision-making tasks. We show that contrary to previous studies, perturbations can mitigate bias and hallucination in some LLMs over other models. It's found that Claude 3 is more effective for the tasks represented in most datasets, whereas models like GPT3.5 exhibit varying levels of adequacy, performing comparably in some cases but falling significantly behind in others. These insights are crucial for understanding the practical implications of deploying LLM-based assistants as effective decision-support tools in real-world applications, emphasising the need for rigorous testing and validation to ensure reliability and effectiveness. This study contributes to the growing body of research on LLM evaluation and provides insights for developing more robust and trustworthy AI assistants in critical decision-making contexts.

cs.CL↗