Repository navigation
llama.cpp需求 #28
Copy link
Copy link
Closed
Description
Activity
ROPE算子:MROPE模式支持
验收条件:
- 移除下面的判断条件
case GGML_OP_ROPE: { // TODO: with ops-test v == 1 // TODO: n_dims <= ne0 if (op->src[0]->ne[0] != op->op_params[1]) { return false; } const int mode = ((const int32_t *) op->op_params)[2]; if (mode & GGML_ROPE_TYPE_MROPE) { return false; } if (mode & GGML_ROPE_TYPE_VISION) { return false; } #ifdef ASCEND_310P if(!ggml_is_contiguous(op->src[0])){ return false; } #endif return true; }- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o ROPEROPE算子:VISION模式支持
验收条件:
- 移除下面的判断条件
case GGML_OP_ROPE: { // TODO: with ops-test v == 1 // TODO: n_dims <= ne0 if (op->src[0]->ne[0] != op->op_params[1]) { return false; } const int mode = ((const int32_t *) op->op_params)[2]; if (mode & GGML_ROPE_TYPE_MROPE) { return false; } if (mode & GGML_ROPE_TYPE_VISION) { return false; } #ifdef ASCEND_310P if(!ggml_is_contiguous(op->src[0])){ return false; } #endif return true; }- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o ROPEROPE算子:支持部分headSize旋转(n_dims <= src0->ne0)
验收条件:
- 移除下面的判断条件
case GGML_OP_ROPE: { // TODO: with ops-test v == 1 // TODO: n_dims <= ne0 if (op->src[0]->ne[0] != op->op_params[1]) { return false; } const int mode = ((const int32_t *) op->op_params)[2]; if (mode & GGML_ROPE_TYPE_MROPE) { return false; } if (mode & GGML_ROPE_TYPE_VISION) { return false; } #ifdef ASCEND_310P if(!ggml_is_contiguous(op->src[0])){ return false; } #endif return true; }- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o ROPECONV_TRANSPOSE_1D算子:支持 (op->src[0]->ne[0] - 1) > 255 场景
验收条件:
- 移除下面的判断条件
case GGML_OP_CONV_TRANSPOSE_1D: // TODO: ((weightL - 1) * dilationW - padLeft)=1336 should not be larger than 255. return (op->src[0]->ne[0] - 1) <= 255;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o CONV_TRANSPOSE_1DOUT_PROD算子:新算子支持
验收条件:
- 在ggml_backend_cann_supports_op中注册GGML_OP_OUT_PROD,
case GGML_OP_OUT_PROD: return true;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o OUT_PRODGATED_LINEAR_ATTN算子:新算子支持
验收条件:
- 在ggml_backend_cann_supports_op中注册GGML_OP_GATED_LINEAR_ATTN,
case GGML_OP_GATED_LINEAR_ATTN: return true;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o GATED_LINEAR_ATTNL2_NORM算子:新算子支持
验收条件:
- 在ggml_backend_cann_supports_op中注册GGML_OP_L2_NORM,
case GGML_OP_L2_NORM: return true;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o L2_NORMCROSS_ENTROPY_LOSS算子:新算子支持
验收条件:
- 在ggml_backend_cann_supports_op中注册GGML_OP_CROSS_ENTROPY_LOSS,
case GGML_OP_CROSS_ENTROPY_LOSS: return true;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o CROSS_ENTROPY_LOSSRWKV_WKV6算子:新算子支持
验收条件:
- 在ggml_backend_cann_supports_op中注册GGML_OP_RWKV_WKV6,
case GGML_OP_RWKV_WKV6: return true;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o RWKV_WKV6RWKV_WKV7算子:新算子支持
验收条件:
- 在ggml_backend_cann_supports_op中注册GGML_OP_RWKV_WKV7,
case GGML_OP_RWKV_WKV6: return true;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o RWKV_WKV7SSM_CONV算子:新算子支持
验收条件:
- 在ggml_backend_cann_supports_op中注册GGML_OP_SSM_CONV,
case GGML_OP_SSM_CONV: return true;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o SSM_CONVSSM_SCAN算子:新算子支持
验收条件:
- 在ggml_backend_cann_supports_op中注册GGML_OP_SSM_SCAN,
case GGML_OP_SSM_SCAN: return true;- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o SSM_SCAN重构:acl graph中,将图命中的校验沉淀至lru cache中
验收条件:
在 acl graph 模型推理正常,无精度问题。./bin/llama-cli -m path_to_model -p "Building a website can be done in 10 steps:" -ngl 32针对某些特殊模型,支持使用aclnnRopeWithSinCosCache融合算子(工作量较大)
拆分为两个部分:
- 对于GGML_OP_ROPE,使用aclnnRopeWithSinCosCache代替之前的实现。之前的ROPE不要删,这个可以很对Qwen2.5-0.5B来做。
- 当前是在q,k上分别调用了GGML_OP_ROPE算子,使用融合算子,来代替之前的两次调用。
验收条件:
- 确保所有的测试用例都通过,且模型推理正常,无精度问题。
./bin/test-backend-ops test -b CANN0 -o ROPE ./bin/llama-cli -m path_to_model -p "Building a website can be done in 10 steps:" -ngl 32- 确保所有的测试用例都通过,且模型推理正常,无精度问题。
./bin/test-backend-ops test -b CANN0 -o ROPE ./bin/llama-cli -m path_to_model -p "Building a website can be done in 10 steps:" -ngl 32参考资料:
算子融合方法:重构:所有调用aclnn的方法,全部提供静态方法和注释进行封装,并替换之前的使用
验收条件:
确保所有的运算符测试用例都通过,且模型推理正常,无精度问题。./bin/test-backend-ops test -b CANN0 ./bin/llama-cli -m path_to_model -p "Building a website can be done in 10 steps:" -ngl 32新增量化方法支持Q4_1,Q8_1(矩阵乘支持支持量化格式Q4_1和Q8_1)
验收条件:
- 在ggml_backend_cann_supports_op中的GGML_OP_MUL_MAT,新增case GGML_TYPE_Q8_1和case GGML_TYPE_Q4_1的支持。
case GGML_OP_MUL_MAT: { switch (op->src[0]->type) { case GGML_TYPE_F16: case GGML_TYPE_F32: return true; case GGML_TYPE_Q8_0: case GGML_TYPE_Q4_0: #ifdef ASCEND_310P // Q4 && Q8 per group is not support on 310p device return false; #endif // only support contiguous for quantized types. return ggml_is_contiguous(op->src[0]) && ggml_is_contiguous(op->src[1]); default: return false; } }- 确保所有的测试用例都通过
./bin/test-backend-ops test -b CANN0 -o MUL_MAT后续维护在:
noemotiovon/llama.cpp#1
Metadata
Metadata
Assignees
Labels
No labels
后续维护在:noemotiovon/llama.cpp#1