The Latent Maximum Entropy Principle

被引：1

作者：

Wang, Shaojun ^{[1
]}

Schuurmans, Dale ^{[2
]}

Zhao, Yunxin ^{[3
]}

机构：

[1] Wright State Univ, Dept Comp Sci & Engn, Dayton, OH 45435 USA

[2] Univ Alberta, Dept Comp Sci, Edmonton, AB T6G 2E8, Canada

[3] Univ Missouri, Dept Comp Sci, Columbia, MO 65211 USA

来源：

ACM TRANSACTIONS ON KNOWLEDGE DISCOVERY FROM DATA | 2012年 / 6卷 / 02期

关键词：

Maximum entropy; iterative scaling; expectation maximization; latent variable models; information geometry; EM; ALGORITHM; MINIMIZATION; LIKELIHOOD; GEOMETRY; MODELS;

D O I：

暂无

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

We present an extension to Jaynes' maximum entropy principle that incorporates latent variables. The principle of latent maximum entropy we propose is different from both Jaynes' maximum entropy principle and maximum likelihood estimation, but can yield better estimates in the presence of hidden variables and limited training data. We first show that solving for a latent maximum entropy model poses a hard nonlinear constrained optimization problem in general. However, we then show that feasible solutions to this problem can be obtained efficiently for the special case of log-linear models-which forms the basis for an efficient approximation to the latent maximum entropy principle. We derive an algorithm that combines expectation-maximization with iterative scaling to produce feasible log-linear solutions. This algorithm can be interpreted as an alternating minimization algorithm in the information divergence, and reveals an intimate connection between the latent maximum entropy and maximum likelihood principles. To select a final model, we generate a series of feasible candidates, calculate the entropy of each, and choose the model that attains the highest entropy. Our experimental results show that estimation based on the latent maximum entropy principle generally gives better results than maximum likelihood when estimating latent variable models on small observed data samples.

引用

页数：42

共 50 条

[1] The latent maximum entropy principle
Department of Computer Science and Engineering, Wright State University, Dayton, OH 45435, United States
不详
不详
ACM Trans. Knowl. Discov. Data, 2
[2] The latent maximum entropy principle
Wang, SJ
Rosenfeld, R
Zhao, YX
Schuurmans, D
ISIT: 2002 IEEE INTERNATIONAL SYMPOSIUM ON INFORMATION THEORY, PROCEEDINGS, 2002, : 131 - 131
[3] Latent maximum entropy principle for statistical language modeling
Wang, SJ
Rosenfeld, R
Zhao, YX
ASRU 2001: IEEE WORKSHOP ON AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING, CONFERENCE PROCEEDINGS, 2001, : 182 - 185
[4] Learning mixture models with the regularized latent maximum entropy principle
Wang, SJ
Schuurmans, D
Peng, FC
Zhao, YX
IEEE TRANSACTIONS ON NEURAL NETWORKS, 2004, 15 (04): : 903 - 916
[5] Combining statistical language models via the latent maximum entropy principle
Wang, SJ
Schuurmans, D
Peng, FC
Zhao, YX
MACHINE LEARNING, 2005, 60 (1-3) : 229 - 250
[6] Combining Statistical Language Models via the Latent Maximum Entropy Principle
Shaojun Wang
Dale Schuurmans
Fuchun Peng
Yunxin Zhao
Machine Learning, 2005, 60 : 229 - 250
[7] Semantic N-gram language modeling with the latent maximum entropy principle
Wang, SJ
Schuurmans, D
Peng, FC
Zhao, YX
2003 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, VOL I, PROCEEDINGS: SPEECH PROCESSING I, 2003, : 376 - 379
[8] Maximum entropy principle revisited
Dreyer, W
Kunik, M
CONTINUUM MECHANICS AND THERMODYNAMICS, 1998, 10 (06) : 331 - 347
[9] The principle of the maximum entropy method
Sakata, M
Takata, M
HIGH PRESSURE RESEARCH, 1996, 14 (4-6) : 327 - 333
[10] THE MAXIMUM-ENTROPY PRINCIPLE
FELLGETT, PB
KYBERNETES, 1987, 16 (02) : 125 - 125

← 1 2 3 4 5 →