Search arXivSearch

arXiv subjects

Eric Slud

Publications and source records attributed to Eric Slud.

5 recordsLinked to original sources

SDR Variance Estimates in Small Domains

Successive Difference Replication (SDR) is a replication based method of variance estimation introduced by Fay and Train (1995) for estimators based on complex multistage surveys, especially those including a final systematic sampling stage. The method has been used for many years as the primary variance-estimation methodology in large national household surveys administered by the Census Bureau, including the American Community Survey and also the Current Population Survey's monthly estimates based on self-representing strata. In settings where it is applied, generally no second method of variance estimation has been available, so the performance of SDR has been studied via simulation by various authors, for variances of survey totals and of nonlinear survey estimators. This paper begins with a thorough exposition of the SDR method and review of previously published results on the small-domain biases of SDR variance estimation. It is shown that the number D of cycles used in implementing SDR should be 3 or larger, in order to control the variability of SDR estimates, but need not be larger than 5. Beyond that, the value of D is virtually irrelevant to the occurrence of small-domain bias in SDR. The SDR method is shown via theoretical formulas and simulation to inflate average estimated variances in small domains by amounts that vary systematically with the patterns of attribute means and variances and survey weights in consecutively enumerated strata. The degree of average variance inflation is generally moderate, no more than 15 percent in domains with sample size 20, but can be larger in special settings. Moreover, SDR estimates are extremely variable in small domains, with standard deviations often far larger than any biases.

stat.ME

Analyzing the Machine Learning Conference Review Process

Mainstream machine learning conferences have seen a dramatic increase in the number of participants, along with a growing range of perspectives, in recent years. Members of the machine learning community are likely to overhear allegations ranging from randomness of acceptance decisions to institutional bias. In this work, we critically analyze the review process through a comprehensive study of papers submitted to ICLR between 2017 and 2020. We quantify reproducibility/randomness in review scores and acceptance decisions, and examine whether scores correlate with paper impact. Our findings suggest strong institutional bias in accept/reject decisions, even after controlling for paper quality. Furthermore, we find evidence for a gender gap, with female authors receiving lower scores, lower acceptance rates, and fewer citations per paper than their male counterparts. We conclude our work with recommendations for future conference organizers.

cs.LG

An Open Review of OpenReview: A Critical Analysis of the Machine Learning Conference Review Process

Mainstream machine learning conferences have seen a dramatic increase in the number of participants, along with a growing range of perspectives, in recent years. Members of the machine learning community are likely to overhear allegations ranging from randomness of acceptance decisions to institutional bias. In this work, we critically analyze the review process through a comprehensive study of papers submitted to ICLR between 2017 and 2020. We quantify reproducibility/randomness in review scores and acceptance decisions, and examine whether scores correlate with paper impact. Our findings suggest strong institutional bias in accept/reject decisions, even after controlling for paper quality. Furthermore, we find evidence for a gender gap, with female authors receiving lower scores, lower acceptance rates, and fewer citations per paper than their male counterparts. We conclude our work with recommendations for future conference organizers.

cs.LG

Introduction

The Statistics Consortium at the University of Maryland, College Park, hosted a two-day workshop on Bayesian Methods that Frequentists Should Know during April 30--May 1, 2008. The event was co-sponsored by the Institute of Mathematical Statistics (IMS), Office of Research and Methodology, National Center for Health Statistics, Survey Research Methods Section (SRMS) of the American Statistical Association, and Washington Statistical Society. The workshop was intended to bring out the positive features of Bayesian statistics in solving real-life problems, including complex problems in sample surveys and production of high-quality official statistics.

stat.ME