Home Business & Economics Improving Statistical Inference with Clustered Data
Article
Licensed
Unlicensed Requires Authentication

Improving Statistical Inference with Clustered Data

  • Jeffrey J. Harden
Published/Copyright: January 4, 2012
Become an author with De Gruyter Brill

Political science data often contain grouped observations, which produces unobserved "cluster effects" in statistical models. Typical solutions include (1) ignoring the impact on coefficients and only adjusting the standard errors of generalized linear models (GLM) or (2) addressing clustering in coefficient estimation while relying on a parametric assumption for the cluster effects and/or a large number of clusters for standard errors. I show that both approaches are problematic for inference. Through simulation I demonstrate that multilevel modeling (MLM) and generalized estimating equations (GEE) produce more efficient coefficients than does GLM. Next, I show that commonly-used MLM and GEE standard error methods can be biased downward, while bootstrapping by resampling clusters (BCSE) performs better, even with a misspecified error distribution and/or few clusters. I recommend the use of MLM or GEE to estimate coefficients and BCSE to estimate uncertainty, and show that this approach can produce divergent conclusions in applied research.

Published Online: 2012-1-4

©2012 Walter de Gruyter GmbH & Co. KG, Berlin/Boston

Downloaded on 29.1.2026 from https://www.degruyterbrill.com/document/doi/10.2202/2151-7509.1026/html
Scroll to top button