Skip to content

Data set issue #1

Description

@OneMoreSecond

Your paper used CMU-Dict 0.7b released in 2014. But I find previous papers use an old version CMU-Dict and partition, which might result in lower WER.
I downloaded it from https://code.google.com/p/phonetisaurus/ (according to Encoding linear models as weighted finite-state transducers, Ke Wu), and it corresponds to data set description in most previous works I found.
Could retry your model on it?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions