Bayesian Mixture Models with Focused Clustering for Mixed Ordinal and Nominal Data

被引:11
|
作者
DeYoreo, Maria [1 ,3 ]
Reiter, Jerome P. [1 ,4 ]
Hillygus, D. Sunshine [2 ,5 ]
机构
[1] Duke Univ, Dept Stat Sci, Durham, NC 27708 USA
[2] Duke Univ, Dept Polit Sci, Durham, NC USA
[3] Duke Univ, Durham, NC USA
[4] Duke Univ, Stat Sci, Durham, NC USA
[5] Duke Univ, Polit Sci, Durham, NC USA
来源
BAYESIAN ANALYSIS | 2017年 / 12卷 / 03期
基金
美国国家科学基金会;
关键词
categorical; missing; mixture model; multiple imputation; MULTIPLE IMPUTATION; CATEGORICAL-DATA; BINARY;
D O I
10.1214/16-BA1020
中图分类号
O1 [数学];
学科分类号
0701 ; 070101 ;
摘要
In some contexts, mixture models can fit certain variables well at the expense of others in ways beyond the analyst's control. For example, when the data include some variables with non-trivial amounts of missing values, the mixture model may fit the marginal distributions of the nearly and fully complete variables at the expense of the variables with high fractions of missing data. Motivated by this setting, we present a mixture model for mixed ordinal and nominal data that splits variables into two groups, focus variables and remainder variables. The model allows the analyst to specify a rich sub-model for the focus variables and a simpler sub-model for remainder variables, yet still capture associations among the variables. Using simulations, we illustrate advantages and limitations of focused clustering compared to mixture models that do not distinguish variables. We apply the model to handle missing values in an analysis of the 2012 American National Election Study, estimating relationships among voting behavior, ideology, and political party affiliation.
引用
收藏
页码:679 / 703
页数:25
相关论文
共 50 条
  • [21] Log-multiplicative association models as latent variable models for nominal and/or ordinal data
    Anderson, CJ
    Vermunt, JK
    SOCIOLOGICAL METHODOLOGY 2000, VOL 30, 2000, 30 : 81 - 121
  • [22] A Bayesian method for analyzing combinations of continuous, ordinal, and nominal categorical data with missing values
    Zhang, Xiao
    Boscardin, W. John
    Belin, Thomas R.
    Wan, Xiaohai
    He, YuLei
    Zhang, Kui
    JOURNAL OF MULTIVARIATE ANALYSIS, 2015, 135 : 43 - 58
  • [23] Bayesian Mixture of AR Models for Time Series Clustering
    Venkatararnana, Kini B.
    Sekhar, C. Chandra
    ICAPR 2009: SEVENTH INTERNATIONAL CONFERENCE ON ADVANCES IN PATTERN RECOGNITION, PROCEEDINGS, 2009, : 35 - 38
  • [24] Bayesian mixture of AR models for time series clustering
    B. Venkataramana Kini
    C. Chandra Sekhar
    Pattern Analysis and Applications, 2013, 16 : 179 - 200
  • [25] BClass: A Bayesian approach based on mixture models for clustering and classification of heterogeneous biological data
    Medrano-Soto, A
    Christen, JA
    Collado-Vides, J
    JOURNAL OF STATISTICAL SOFTWARE, 2005, 13 (02): : 1 - 18
  • [26] Bayesian mixture of AR models for time series clustering
    Kini, B. Venkataramana
    Sekhar, C. Chandra
    PATTERN ANALYSIS AND APPLICATIONS, 2013, 16 (02) : 179 - 200
  • [27] Bayesian analysis of multivariate mixed longitudinal ordinal and continuous data
    Zhang, Xiao
    AUSTRALIAN & NEW ZEALAND JOURNAL OF STATISTICS, 2024, 66 (03) : 325 - 346
  • [28] Clustering large mixed-type data with ordinal variables
    Szepannek, Gero
    Aschenbruck, Rabea
    Wilhelm, Adalbert
    ADVANCES IN DATA ANALYSIS AND CLASSIFICATION, 2024,
  • [29] Evaluation of Bayesian models for focused clustering in health data (vol 18, pg 871, 2007)
    Ma, Bo
    Lawson, Andrew B.
    Liu, Yuan
    ENVIRONMETRICS, 2007, 18 (08) : 891 - 891
  • [30] Multipartition clustering of mixed data with Bayesian networks
    Rodriguez-Sanchez, Fernando
    Bielza, Concha
    Larranaga, Pedro
    INTERNATIONAL JOURNAL OF INTELLIGENT SYSTEMS, 2022, 37 (03) : 2188 - 2218