Your paper used CMU-Dict 0.7b released in 2014. But I find previous papers use an old version CMU-Dict and partition, which might result in lower WER.
I downloaded it from https://code.google.com/p/phonetisaurus/ (according to Encoding linear models as weighted finite-state transducers, Ke Wu), and it corresponds to data set description in most previous works I found.
Could retry your model on it?
Your paper used CMU-Dict 0.7b released in 2014. But I find previous papers use an old version CMU-Dict and partition, which might result in lower WER.
I downloaded it from https://code.google.com/p/phonetisaurus/ (according to Encoding linear models as weighted finite-state transducers, Ke Wu), and it corresponds to data set description in most previous works I found.
Could retry your model on it?