Skip to content

llama.cpp需求 #28

Description

@hipudding
  • 优化图顺序,多流并行 #15850
  • ROPE算子:MROPE模式支持
  • ROPE算子:VISION模式支持
  • ROPE算子:支持部分headSize旋转(n_dims <= src0->ne0)
  • CONV_TRANSPOSE_1D算子:支持 (op->src[0]->ne[0] - 1) > 255 场景
  • OUT_PROD算子:新算子支持
  • GATED_LINEAR_ATTN算子:新算子支持
  • L2_NORM算子:新算子支持
  • CROSS_ENTROPY_LOSS算子:新算子支持
  • RWKV_WKV6算子:新算子支持
  • RWKV_WKV7算子:新算子支持
  • SSM_CONV算子:新算子支持
  • SSM_SCAN算子:新算子支持
  • 重构:acl graph中,将图命中的校验沉淀至lru cache中
  • 针对某些特殊模型,支持使用aclnnRopeWithSinCosCache融合算子(工作量较大)
  • 优化set_device
  • 新增量化方法支持Q4_1,Q8_1
  • 重构:所有调用aclnn的方法,全部提供静态方法和注释进行封装,并替换之前的使用

后续维护在:noemotiovon/llama.cpp#1

Activity

  1. noemotiovon commented on Sep 10, 2025

    @noemotiovon
    Contributor

    ROPE算子:MROPE模式支持

    验收条件:

    1. 移除下面的判断条件
            case GGML_OP_ROPE: {
                // TODO: with ops-test v == 1
                // TODO: n_dims <= ne0
                if (op->src[0]->ne[0] != op->op_params[1]) {
                    return false;
                }
    
                const int mode = ((const int32_t *) op->op_params)[2];
                if (mode & GGML_ROPE_TYPE_MROPE) {
                    return false;
                }
                if (mode & GGML_ROPE_TYPE_VISION) {
                    return false;
                }
    #ifdef ASCEND_310P
                if(!ggml_is_contiguous(op->src[0])){
                    return false;
                }
    #endif
                return true;
            }
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o ROPE
    
  2. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    ROPE算子:VISION模式支持

    验收条件:

    1. 移除下面的判断条件
            case GGML_OP_ROPE: {
                // TODO: with ops-test v == 1
                // TODO: n_dims <= ne0
                if (op->src[0]->ne[0] != op->op_params[1]) {
                    return false;
                }
    
                const int mode = ((const int32_t *) op->op_params)[2];
                if (mode & GGML_ROPE_TYPE_MROPE) {
                    return false;
                }
                if (mode & GGML_ROPE_TYPE_VISION) {
                    return false;
                }
    #ifdef ASCEND_310P
                if(!ggml_is_contiguous(op->src[0])){
                    return false;
                }
    #endif
                return true;
            }
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o ROPE
    
  3. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    ROPE算子:支持部分headSize旋转(n_dims <= src0->ne0)

    验收条件:

    1. 移除下面的判断条件
            case GGML_OP_ROPE: {
                // TODO: with ops-test v == 1
                // TODO: n_dims <= ne0
                if (op->src[0]->ne[0] != op->op_params[1]) {
                    return false;
                }
    
                const int mode = ((const int32_t *) op->op_params)[2];
                if (mode & GGML_ROPE_TYPE_MROPE) {
                    return false;
                }
                if (mode & GGML_ROPE_TYPE_VISION) {
                    return false;
                }
    #ifdef ASCEND_310P
                if(!ggml_is_contiguous(op->src[0])){
                    return false;
                }
    #endif
                return true;
            }
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o ROPE
    
  4. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    CONV_TRANSPOSE_1D算子:支持 (op->src[0]->ne[0] - 1) > 255 场景

    验收条件:

    1. 移除下面的判断条件
            case GGML_OP_CONV_TRANSPOSE_1D:
                // TODO: ((weightL - 1) * dilationW - padLeft)=1336 should not be larger than 255.
                return (op->src[0]->ne[0] - 1) <= 255;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o CONV_TRANSPOSE_1D
    
  5. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    OUT_PROD算子:新算子支持

    验收条件:

    1. 在ggml_backend_cann_supports_op中注册GGML_OP_OUT_PROD,
            case GGML_OP_OUT_PROD:
                return true;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o OUT_PROD
    
  6. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    GATED_LINEAR_ATTN算子:新算子支持

    验收条件:

    1. 在ggml_backend_cann_supports_op中注册GGML_OP_GATED_LINEAR_ATTN,
            case GGML_OP_GATED_LINEAR_ATTN:
                return true;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o GATED_LINEAR_ATTN
    
  7. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    L2_NORM算子:新算子支持

    验收条件:

    1. 在ggml_backend_cann_supports_op中注册GGML_OP_L2_NORM,
            case GGML_OP_L2_NORM:
                return true;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o L2_NORM
    
  8. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    CROSS_ENTROPY_LOSS算子:新算子支持

    验收条件:

    1. 在ggml_backend_cann_supports_op中注册GGML_OP_CROSS_ENTROPY_LOSS,
            case GGML_OP_CROSS_ENTROPY_LOSS:
                return true;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o CROSS_ENTROPY_LOSS
    
  9. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    RWKV_WKV6算子:新算子支持

    验收条件:

    1. 在ggml_backend_cann_supports_op中注册GGML_OP_RWKV_WKV6,
            case GGML_OP_RWKV_WKV6:
                return true;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o RWKV_WKV6
    
  10. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    RWKV_WKV7算子:新算子支持

    验收条件:

    1. 在ggml_backend_cann_supports_op中注册GGML_OP_RWKV_WKV7,
            case GGML_OP_RWKV_WKV6:
                return true;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o RWKV_WKV7
    
  11. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    SSM_CONV算子:新算子支持

    验收条件:

    1. 在ggml_backend_cann_supports_op中注册GGML_OP_SSM_CONV,
            case GGML_OP_SSM_CONV:
                return true;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o SSM_CONV
    
  12. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    SSM_SCAN算子:新算子支持

    验收条件:

    1. 在ggml_backend_cann_supports_op中注册GGML_OP_SSM_SCAN,
            case GGML_OP_SSM_SCAN:
                return true;
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o SSM_SCAN
    
  13. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    重构:acl graph中,将图命中的校验沉淀至lru cache中

    验收条件:
    在 acl graph 模型推理正常,无精度问题。

    ./bin/llama-cli -m path_to_model -p "Building a website can be done in 10 steps:" -ngl 32
    
  14. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    针对某些特殊模型,支持使用aclnnRopeWithSinCosCache融合算子(工作量较大)

    拆分为两个部分:

    1. 对于GGML_OP_ROPE,使用aclnnRopeWithSinCosCache代替之前的实现。之前的ROPE不要删,这个可以很对Qwen2.5-0.5B来做。
    2. 当前是在q,k上分别调用了GGML_OP_ROPE算子,使用融合算子,来代替之前的两次调用。

    验收条件:

    1. 确保所有的测试用例都通过,且模型推理正常,无精度问题。
    ./bin/test-backend-ops test -b CANN0 -o ROPE
    ./bin/llama-cli -m path_to_model -p "Building a website can be done in 10 steps:" -ngl 32
    
    1. 确保所有的测试用例都通过,且模型推理正常,无精度问题。
    ./bin/test-backend-ops test -b CANN0 -o ROPE
    ./bin/llama-cli -m path_to_model -p "Building a website can be done in 10 steps:" -ngl 32
    

    参考资料:
    算子融合方法:

  15. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    重构:所有调用aclnn的方法,全部提供静态方法和注释进行封装,并替换之前的使用

    验收条件:
    确保所有的运算符测试用例都通过,且模型推理正常,无精度问题。

    ./bin/test-backend-ops test -b CANN0
    ./bin/llama-cli -m path_to_model -p "Building a website can be done in 10 steps:" -ngl 32
    
  16. noemotiovon commented on Sep 11, 2025

    @noemotiovon
    Contributor

    新增量化方法支持Q4_1,Q8_1(矩阵乘支持支持量化格式Q4_1和Q8_1)

    验收条件:

    1. 在ggml_backend_cann_supports_op中的GGML_OP_MUL_MAT,新增case GGML_TYPE_Q8_1和case GGML_TYPE_Q4_1的支持。
            case GGML_OP_MUL_MAT: {
                switch (op->src[0]->type) {
                    case GGML_TYPE_F16:
                    case GGML_TYPE_F32:
                        return true;
                    case GGML_TYPE_Q8_0:
                    case GGML_TYPE_Q4_0:
    #ifdef ASCEND_310P
                        // Q4 && Q8 per group is not support on 310p device
                        return false;
    #endif
                        // only support contiguous for quantized types.
                        return ggml_is_contiguous(op->src[0]) &&
                                ggml_is_contiguous(op->src[1]);
                    default:
                        return false;
                }
            }
    
    1. 确保所有的测试用例都通过
    ./bin/test-backend-ops test -b CANN0 -o MUL_MAT
    
  17. noemotiovon commented on Sep 15, 2025

    @noemotiovon
    Contributor

    后续维护在:
    noemotiovon/llama.cpp#1

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions