Repository navigation
GPT-NeoX has only minimal inference support #3293
Description
Activity
Even if you add dummy scores and token types in the conversion script, it fails here:
https://github.com/ggerganov/llama.cpp/blob/bc9d3e3971e5607a10ff4c24e39568ce1ac87271/llama.cpp#L2288
Was GPT-NeoX ever even implemented in GGUF?Reacted by Austin- changed the title
[-]GPT-NeoX : cannot find tokenizer scores in model file[/-][+]GPT-NeoX has a conversion script but cannot be loaded or used for inference[/+]on Sep 21, 2023 Was GPT-NeoX ever even implemented in GGUF?
Yes, example inference code exists here: https://github.com/ggerganov/llama.cpp/blob/master/examples/gptneox-wip/gptneox-main.cpp
Oh, it has a separate implementation. So I can't currently use it with any third-party software that uses the llama.cpp API.
edit: This file is not listed in either of the build scripts. It doesn't seem to have GPU acceleration. It seems like that could be improved.
- changed the title
[-]GPT-NeoX has a conversion script but cannot be loaded or used for inference[/-][+]GPT-NeoX has only minimal inference support[/+]on Sep 22, 2023 Yeah, there's just a poc implementation. We should add it in
llama.cppeventuallyReacted by AustinThe situation is now that we do have code in this repository to "successfully" convert and quantize a gpt-neo-x model, but no way to run these models.
https://github.com/ggerganov/ggml/tree/master/examples/gpt-neox does have its own convert-script. The conversion here in the convert-hf-to-gguf.py does not seem to have any purpose at all.I still would like to bring this forward again:
There is code inside theconvert-hf-to-gguf.pysince its first release in #3838 and before in the seperate convert-gptneox-script to somehow "sucessfully" convert gpt-neo-x models into gguf models. But there is no code whatsoever to run inference on a model that the convert script labels as a gptneox model.I believe you can run this using https://github.com/ggerganov/ggml/blob/master/examples/gpt-neox/main.cpp
it's just not in llama.cpp. It'll be supported eventually.Oh, I alredy compiled that example for testing. It seems to expect the old ggml .bin files, which can be created using the convert program in the same example directory. It doesn't run the gguf files that are built using the
convet-hf-to-gguf.pyi script. Right now, there is no code that can run these gguf files converted from gpt-neo-x models.It's much easier to add new arches to
llama.cppnow (I hope) - PRs welcomeThis issue was closed because it has been inactive for 14 days since being marked as stale.
I'm still interested in this.
github-actions commented
on May 18, 2024 on May 18, 2024 – with GitHub ActionsContributorMore actionsThis issue was closed because it has been inactive for 14 days since being marked as stale.
- linked a pull request that will close this issueAdd missing inference support for GPTNeoXForCausalLM (Pythia and GPT-NeoX base models) #7461
on May 22, 2024
Steps to reproduce: