Search arXivSearch

arXiv subjects

Anna Simoni

Publications and source records attributed to Anna Simoni.

12 recordsLinked to original sources

Testing for Endogeneity: A Moment-Based Bayesian Approach

A standard assumption in the Bayesian estimation of linear regression models is that the regressors are exogenous in the sense that they are uncorrelated with the model error term. In practice, however, this assumption can be invalid. In this paper, using the exponentially tilted empirical likelihood framework, we develop a Bayes factor test for endogeneity that compares a base model that is correctly specified under exogeneity but misspecified under endogeneity against an extended model that is correctly specified in either case. We provide a comprehensive study of the log-marginal exponentially tilted empirical likelihood. We demonstrate that our testing procedure is consistent from a frequentist point of view: as the sample grows, it almost surely selects the base model if and only if the regressors are exogenous, and the extended model if and only if the regressors are endogenous. The methods are illustrated with simulated data, and problems concerning the causal effect of automobile prices on automobile demand and the causal effect of potentially endogenous airplane ticket prices on passenger volume.

econ.EM

Panel data models with randomly generated groups

We develop a structural framework for modeling and inferring unobserved heterogeneity in dynamic panel-data models. Unlike methods treating clustering as a descriptive device, we model heterogeneity as arising from a latent clustering mechanism, where the number of clusters is unknown and estimated. Building on the mixture of finite mixtures (MFM) approach, our method avoids the clustering inconsistency issues of Dirichlet process mixtures and provides an interpretable representation of the population clustering structure. We extend the Telescoping Sampler of Fruhwirth-Schnatter et al. (2021) to dynamic panels with covariates, yielding an efficient MCMC algorithm that delivers full Bayesian inference and credible sets. We show that asymptotically the posterior distribution of the mixing measure contracts around the truth at parametric rates in Wasserstein distance, ensuring recovery of clustering and structural parameters. Simulations demonstrate strong finite-sample performance. Finally, an application to the income-democracy relationship reveals latent heterogeneity only when controlling for additional covariates.

econ.EM

Bayesian Bi-level Sparse Group Regressions for Macroeconomic Density Forecasting

We propose a Machine Learning approach for optimal macroeconomic density forecasting in a high-dimensional setting where the underlying model exhibits a known group structure. Our approach is general enough to encompass specific forecasting models featuring either many covariates, or unknown nonlinearities, or series sampled at different frequencies. By relying on the novel concept of bi-level sparsity in time-series econometrics, we construct density forecasts based on a prior that induces sparsity both at the group level and within groups. We demonstrate the consistency of both posterior and predictive distributions. We show that the posterior distribution contracts at the minimax-optimal rate and, asymptotically, puts mass on a set that includes the support of the model. Our theory allows for correlation between groups, while predictors in the same group can be characterized by strong covariation as well as common characteristics and patterns. Finite sample performance is illustrated through comprehensive Monte Carlo experiments and a real-data nowcasting exercise of the US GDP growth rate.

econ.EM

Bayesian Estimation and Comparison of Conditional Moment Models

We consider the Bayesian analysis of models in which the unknown distribution of the outcomes is specified up to a set of conditional moment restrictions. The nonparametric exponentially tilted empirical likelihood function is constructed to satisfy a sequence of unconditional moments based on an increasing (in sample size) vector of approximating functions (such as tensor splines based on the splines of each conditioning variable). For any given sample size, results are robust to the number of expanded moments. We derive Bernstein-von Mises theorems for the behavior of the posterior distribution under both correct and incorrect specification of the conditional moments, subject to growth rate conditions (slower under misspecification) on the number of approximating functions. A large-sample theory for comparing different conditional moment models is also developed. The central result is that the marginal likelihood criterion selects the model that is less misspecified. We also introduce sparsity-based model search for high-dimensional conditioning variables, and provide efficient MCMC computations for high-dimensional parameters. Along with clarifying examples, the framework is illustrated with real-data applications to risk-factor determination in finance, and causal inference under conditional ignorability.

math.ST

Revisiting identification concepts in Bayesian analysis

This paper studies the role played by identification in the Bayesian analysis of statistical and econometric models. First, for unidentified models we demonstrate that there are situations where the introduction of a non-degenerate prior distribution can make a parameter that is nonidentified in frequentist theory identified in Bayesian theory. In other situations, it is preferable to work with the unidentified model and construct a Markov Chain Monte Carlo (MCMC) algorithms for it instead of introducing identifying assumptions. Second, for partially identified models we demonstrate how to construct the prior and posterior distributions for the identified set parameter and how to conduct Bayesian analysis. Finally, for models that contain some parameters that are identified and others that are not we show that marginalizing out the identified parameter from the likelihood with respect to its conditional prior, given the nonidentified parameter, allows the data to be informative about the nonidentified and partially identified parameter. The paper provides examples and simulations that illustrate how to implement our techniques.

econ.EM

When are Google data useful to nowcast GDP? An approach via pre-selection and shrinkage

Alternative data sets are widely used for macroeconomic nowcasting together with machine learning--based tools. The latter are often applied without a complete picture of their theoretical nowcasting properties. Against this background, this paper proposes a theoretically grounded nowcasting methodology that allows researchers to incorporate alternative Google Search Data (GSD) among the predictors and that combines targeted preselection, Ridge regularization, and Generalized Cross Validation. Breaking with most existing literature, which focuses on asymptotic in-sample theoretical properties, we establish the theoretical out-of-sample properties of our methodology and support them by Monte-Carlo simulations. We apply our methodology to GSD to nowcast GDP growth rate of several countries during various economic periods. Our empirical findings support the idea that GSD tend to increase nowcasting accuracy, even after controlling for official variables, but that the gain differs between periods of recessions and of macroeconomic stability.

econ.EM

Bayesian MIDAS Penalized Regressions: Estimation, Selection, and Prediction

We propose a new approach to mixed-frequency regressions in a high-dimensional environment that resorts to Group Lasso penalization and Bayesian techniques for estimation and inference. In particular, to improve the prediction properties of the model and its sparse recovery ability, we consider a Group Lasso with a spike-and-slab prior. Penalty hyper-parameters governing the model shrinkage are automatically tuned via an adaptive MCMC algorithm. We establish good frequentist asymptotic properties of the posterior of the in-sample and out-of-sample prediction error, we recover the optimal posterior contraction rate, and we show optimality of the posterior predictive density. Simulations show that the proposed models have good selection and forecasting performance in small samples, even when the design matrix presents cross-correlation. When applied to forecasting U.S. GDP, our penalized regressions can outperform many strong competitors. Results suggest that financial variables may have some, although very limited, short-term predictive content.

econ.EM

Ill-posed Estimation in High-Dimensional Models with Instrumental Variables

This paper is concerned with inference about low-dimensional components of a high-dimensional parameter vector $\beta^0$ which is identified through instrumental variables. We allow for eigenvalues of the expected outer product of included and excluded covariates, denoted by $M$, to shrink to zero as the sample size increases. We propose a novel estimator based on desparsification of an instrumental variable Lasso estimator, which is a regularized version of 2SLS with an additional correction term. This estimator converges to $\beta^0$ at a rate depending on the mapping properties of $M$ captured by a sparse link condition. Linear combinations of our estimator of $\beta^0$ are shown to be asymptotically normally distributed. Based on consistent covariance estimation, our method allows for constructing confidence intervals and statistical tests for single or low-dimensional components of $\beta^0$. In Monte-Carlo simulations we analyze the finite sample behavior of our estimator.

econ.EM

Gaussian processes and Bayesian moment estimation

Given a set of moment restrictions (MRs) that overidentify a parameter $\theta$, we investigate a semiparametric Bayesian approach for inference on $\theta$ that does not restrict the data distribution $F$ apart from the MRs. As main contribution, we construct a degenerate Gaussian process prior that, conditionally on $\theta$, restricts the $F$ generated by this prior to satisfy the MRs with probability one. Our prior works even in the more involved case where the number of MRs is larger than the dimension of $\theta$. We demonstrate that the corresponding posterior for $\theta$ is computationally convenient. Moreover, we show that there exists a link between our procedure, the Generalized Empirical Likelihood with quadratic criterion and the limited information likelihood-based procedures. We provide a frequentist validation of our procedure by showing consistency and asymptotic normality of the posterior distribution of $\theta$. The finite sample properties of our method are illustrated through Monte Carlo experiments and we provide an application to demand estimation in the airline market.

math.ST

Bayesian Estimation and Comparison of Moment Condition Models

In this paper we consider the problem of inference in statistical models characterized by moment restrictions by casting the problem within the Exponentially Tilted Empirical Likelihood (ETEL) framework. Because the ETEL function has a well defined probabilistic interpretation and plays the role of a nonparametric likelihood, a fully Bayesian semiparametric framework can be developed. We establish a number of powerful results surrounding the Bayesian ETEL framework in such models. One major concern driving our work is the possibility of misspecification. To accommodate this possibility, we show how the moment conditions can be reexpressed in terms of additional nuisance parameters and that, even under misspecification, the Bayesian ETEL posterior distribution satisfies a Bernstein-von Mises result. A second key contribution of the paper is the development of a framework based on marginal likelihoods and Bayes factors to compare models defined by different moment conditions. Computation of the marginal likelihoods is by the method of Chib (1995) as extended to Metropolis-Hastings samplers in Chib and Jeliazkov (2001). We establish the model selection consistency of the marginal likelihood and show that the marginal likelihood favors the model with the minimum number of parameters and the maximum number of valid moment restrictions. When the models are misspecified, the marginal likelihood model selection procedure selects the model that is closer to the (unknown) true data generating process in terms of the Kullback-Leibler divergence. The ideas and results in this paper provide a further broadening of the theoretical underpinning and value of the Bayesian ETEL framework with likely far-reaching practical consequences. The discussion is illuminated through several examples.

stat.ME

Adaptive Bayesian estimation in indirect Gaussian sequence space models

In an indirect Gaussian sequence space model lower and upper bounds are derived for the concentration rate of the posterior distribution of the parameter of interest shrinking to the parameter value $\theta^\circ$ that generates the data. While this establishes posterior consistency, however, the concentration rate depends on both $\theta^\circ$ and a tuning parameter which enters the prior distribution. We first provide an oracle optimal choice of the tuning parameter, i.e., optimized for each $\theta^\circ$ separately. The optimal choice of the prior distribution allows us to derive an oracle optimal concentration rate of the associated posterior distribution. Moreover, for a given class of parameters and a suitable choice of the tuning parameter, we show that the resulting uniform concentration rate over the given class is optimal in a minimax sense. Finally, we construct a hierarchical prior that is adaptive. This means that, given a parameter $\theta^\circ$ or a class of parameters, respectively, the posterior distribution contracts at the oracle rate or at the minimax rate over the class. Notably, the hierarchical prior does not depend neither on $\theta^\circ$ nor on the given class. Moreover, convergence of the fully data-driven Bayes estimator at the oracle or at the minimax rate is established.

math.ST

Semi-parametric Bayesian Partially Identified Models based on Support Function

We provide a comprehensive semi-parametric study of Bayesian partially identified econometric models. While the existing literature on Bayesian partial identification has mostly focused on the structural parameter, our primary focus is on Bayesian credible sets (BCS's) of the unknown identified set and the posterior distribution of its support function. We construct a (two-sided) BCS based on the support function of the identified set. We prove the Bernstein-von Mises theorem for the posterior distribution of the support function. This powerful result in turn infers that, while the BCS and the frequentist confidence set for the partially identified parameter are asymptotically different, our constructed BCS for the identified set has an asymptotically correct frequentist coverage probability. Importantly, we illustrate that the constructed BCS for the identified set does not require a prior on the structural parameter. It can be computed efficiently for subset inference, especially when the target of interest is a sub-vector of the partially identified parameter, where projecting to a low-dimensional subset is often required. Hence, the proposed methods are useful in many applications. The Bayesian partial identification literature has been assuming a known parametric likelihood function. However, econometric models usually only identify a set of moment inequalities, and therefore using an incorrect likelihood function may result in misleading inferences. In contrast, with a nonparametric prior on the unknown likelihood function, our proposed Bayesian procedure only requires a set of moment conditions, and can efficiently make inference about both the partially identified parameter and its identified set. This makes it widely applicable in general moment inequality models. Finally, the proposed method is illustrated in a financial asset pricing problem.

stat.ME