An overview on the distribution of word counts in Markov chains

被引:28
|
作者
Schbath, S [1 ]
机构
[1] INRA, Biometr Unit, F-78352 Jouy En Josas, France
关键词
word count distribution; Markovian random sequence; overlapping occurrences; renewals; clumps;
D O I
10.1089/10665270050081469
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
In this paper, me give an overview about the different results existing on the statistical distribution of word counts in a Markovian sequence of letters. Results concerning the number of overlapping occurrences, the number of renewals and the number of clumps mill be presented, Counts of single words and also multiple words are considered. Most of the results are approximations as the length of the sequence tends to infinity. We will see that Gaussian approximations switch to (compound) Poisson approximations for rare words, Modeling DNA sequences or proteins by stationary Markov chains, these results can be used to study the statistical frequency of motifs in a given sequence.
引用
收藏
页码:193 / 201
页数:9
相关论文
共 50 条