IASS Webinar 70: Robust imputation procedures in the presence of influential survey data based on adaptative tuning constants by Sixia Chen

Map Unavailable

Date/Time
Date(s) - 20/10/2026
4:00 pm - 5:30 pm

Category(ies)


Date/time: Tuesday, October 20th, 2026, 4:00pm – 5:30pm New York time

Title: Robust imputation procedures in the presence of influential survey data based on adaptative tuning constants

Speaker: Sixia Chen

Abstract:

Missing data due to nonresponse is a common issue in survey sampling and, without
appropriate treatment, may lead to substantial bias in point estimation. Item nonresponse
is typically handled through imputation, and under a Missing at Random (MAR) mechanism
with a correctly specified imputation model, nonresponse bias can be eliminated.
However, even when this condition holds, the presence of influential units among
respondents may severely affect the efficiency of estimators. Influential units arise
frequently in practice, particularly in business and biomedical surveys, where the
distribution of the study variable may be highly skewed. These units may have a
disproportionate impact on estimates due to high leverage, large residuals, or large
sampling weights. Unlike nonresponse, which primarily induces bias, influential units
mainly inflate the variance of estimators, leading to unstable results and large mean
squared error. A common approach to mitigate the effect of influential observations is to
use robust regression methods at the imputation stage (e.g., M-estimators; see Maronna et
al., 2019; Andersen, 2008). These methods typically rely on a fixed tuning constant to
control the level of robustness. However, while such procedures may perform well for
describing the behavior of the inlier population, they may introduce substantial bias when
estimating finite population quantities. In particular, a fixed tuning constant does not adapt
to the sample size and may lead to an unfavorable bias–variance trade-off, especially in the
presence of asymmetric outliers. This paper addresses the problem of constructing
imputation procedures that properly account for influential units while maintaining good
bias and efficiency properties for finite population estimation.