Skip to content

GPT-NeoX has only minimal inference support #3293

Description

@cebtenzzre

Steps to reproduce:

  1. Download https://huggingface.co/EleutherAI/gpt-neox-20b
  2. Convert the model and attempt to use it:
$ TMPDIR=/var/tmp ./convert-gptneox-hf-to-gguf.py gpt-neox-20b 1 --outfile gpt-neox-20b.f16.gguf
$ ./main -m gpt-neox-20b.f16.gguf
<snip>
llama_model_loader: - type  f32:  354 tensors
llama_model_loader: - type  f16:  178 tensors
error loading model: cannot find tokenizer scores in model file

llama_load_model_from_file: failed to load model
llama_init_from_gpt_params: error: failed to load model 'gpt-neox-20b.f16.gguf'
main: error: unable to load model

Activity

  1. cebtenzzre commented on Sep 21, 2023

    @cebtenzzre
    CollaboratorAuthor

    Even if you add dummy scores and token types in the conversion script, it fails here:
    https://github.com/ggerganov/llama.cpp/blob/bc9d3e3971e5607a10ff4c24e39568ce1ac87271/llama.cpp#L2288
    Was GPT-NeoX ever even implemented in GGUF?

  2. changed the title [-]GPT-NeoX : cannot find tokenizer scores in model file[/-] [+]GPT-NeoX has a conversion script but cannot be loaded or used for inference[/+] on Sep 21, 2023
  3. jackie1218 commented on Sep 22, 2023

    @jackie1218

    Was GPT-NeoX ever even implemented in GGUF?

    Yes, example inference code exists here: https://github.com/ggerganov/llama.cpp/blob/master/examples/gptneox-wip/gptneox-main.cpp

  4. cebtenzzre commented on Sep 22, 2023

    @cebtenzzre
    CollaboratorAuthor

    Oh, it has a separate implementation. So I can't currently use it with any third-party software that uses the llama.cpp API.

    edit: This file is not listed in either of the build scripts. It doesn't seem to have GPU acceleration. It seems like that could be improved.

  5. changed the title [-]GPT-NeoX has a conversion script but cannot be loaded or used for inference[/-] [+]GPT-NeoX has only minimal inference support[/+] on Sep 22, 2023
  6. ggerganov commented on Sep 28, 2023

    @ggerganov
    Member

    Yeah, there's just a poc implementation. We should add it in llama.cpp eventually

  7. maddes8cht commented on Nov 12, 2023

    @maddes8cht
    Contributor

    The situation is now that we do have code in this repository to "successfully" convert and quantize a gpt-neo-x model, but no way to run these models.
    https://github.com/ggerganov/ggml/tree/master/examples/gpt-neox does have its own convert-script. The conversion here in the convert-hf-to-gguf.py does not seem to have any purpose at all.

  8. maddes8cht commented on Nov 22, 2023

    @maddes8cht
    Contributor

    I still would like to bring this forward again:
    There is code inside the convert-hf-to-gguf.py since its first release in #3838 and before in the seperate convert-gptneox-script to somehow "sucessfully" convert gpt-neo-x models into gguf models. But there is no code whatsoever to run inference on a model that the convert script labels as a gptneox model.

  9. Galunid commented on Nov 22, 2023

    @Galunid
    Contributor

    I believe you can run this using https://github.com/ggerganov/ggml/blob/master/examples/gpt-neox/main.cpp
    it's just not in llama.cpp. It'll be supported eventually.

  10. maddes8cht commented on Nov 22, 2023

    @maddes8cht
    Contributor

    Oh, I alredy compiled that example for testing. It seems to expect the old ggml .bin files, which can be created using the convert program in the same example directory. It doesn't run the gguf files that are built using the convet-hf-to-gguf.pyi script. Right now, there is no code that can run these gguf files converted from gpt-neo-x models.

  11. ggerganov commented on Nov 23, 2023

    @ggerganov
    Member

    It's much easier to add new arches to llama.cpp now (I hope) - PRs welcome

  12. github-actions commented on Apr 3, 2024

    @github-actions
    Contributor

    This issue was closed because it has been inactive for 14 days since being marked as stale.

  13. cebtenzzre commented on Apr 3, 2024

    @cebtenzzre
    CollaboratorAuthor

    I'm still interested in this.

  14. github-actions commented on May 18, 2024

    @github-actions
    Contributor

    This issue was closed because it has been inactive for 14 days since being marked as stale.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions