Skip to content

Enable support for Gemma 3n-E4B #138

Description

@shahkarnav115-beep

Name and Version

dev_backend_openvino

Operating systems

Windows

GGML backends

OpenVINO

Hardware

13th Gen Intel(R) Core(TM) i5-13450HX + Nvidia RTX 3050

Models

gemma-3n-E4B-it-Q4_K_M.gguf

Problem description & steps to reproduce

gemma 3n-E4B not running with openvino backend.

GGML_OPENVINO:ON

llama-cli.exe -m "path/to/gemma-3n-E4B-it-Q4_K_M.gguf"

First Bad Commit

No response

Relevant log output

OpenVINO: using device CPU

Loading model... /Doesn't handle node name: node_2 op: REPEAT                                                          -Doesn't handle node name: node_2 op: REPEAT
Doesn't handle node name: node_4 op: SQR
Doesn't handle node name: node_5 op: SUM_ROWS
Doesn't handle node name: node_6 op: SQRT
Doesn't handle node name: node_8 op: SQR
Doesn't handle node name: node_9 op: SUM_ROWS
Doesn't handle node name: node_10 op: SQRT
Doesn't handle node name: node_11 op: DIV
Doesn't handle node name: inp_stacked op: CONCAT
GGML OpenVINO backend ov::Exception: Check 'dynamic_dim_value == node->ne[m_node_dynamic_dims[node]]' failed at C:\Users\karnav\openvino.genai\thirdparty\llama.cpp\ggml\src\ggml-openvino\ggml-decoder.cpp:1142:
Dynamic dim value mismatch for node: inp_stacked (permuted) and its src[0]: inp_stacked

graph_compute: ggml_backend_sched_graph_compute_async failed with error -1
process_ubatch: failed to compute graph, compute status: -1
llama_decode: failed to decode, ret = -3

Activity

  1. shahkarnav115-beep commented on Apr 28, 2026

    @shahkarnav115-beep
    Author

    @ravi9, @wine99 and @cavusmustafa, I was working on these Issue, and I have successfully passed the node not found error. Now, I have encountered to Tensor mismatch error:

    Loading model... OpenVINO: using device CPU
    |GGML OpenVINO backend ov::Exception: Exception from src\inference\src\cpp\infer_request.cpp:157:
    Exception from src\inference\src\cpp\infer_request.cpp:67:
    Check 'shape.compatible(ov::PartialShape(tensor->get_shape())) || tensor->get_size() == 0' failed at src\plugins\intel_cpu\src\infer_request.cpp:442:
    Can't set the output tensor with index: 0, because the model output tensor (shape=[1,1,2,2048]) and the current tensor (shape=(1.3.2.2048)) are incompatible
    
    
    
    graph_compute: ggml_backend_sched_graph_compute_async failed with error -1
    process_ubatch: failed to compute graph, compute status: -1
    llama_decode: failed to decode, ret = -3                                   /GGML OpenVINO backend ov::Exception: Exception from src\inference\src\cpp\infer_request.cpp:157:
    Exception from src\inference\src\cpp\infer_request.cpp:67:
    Check 'shape.compatible(ov::PartialShape(tensor->get_shape())) || tensor->get_size() == 0' failed at src\plugins\intel_cpu\src\infer_request.cpp:442:
    Can't set the output tensor with index: 0, because the model output tensor (shape=[1,1,2,2048]) and the current tensor (shape=(1.3.2.2048)) are incompatible
    
    
    
    graph_compute: ggml_backend_sched_graph_compute_async failed with error -1
    process_ubatch: failed to compute graph, compute status: -1
    llama_decode: failed to decode, ret = -3
    common_speculative_is_compat: llama_decode() failed: -3
    
    
    ▄▄ ▄▄
    ██ ██
    ██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
    ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
    ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                        ██    ██
                                        ▀▀    ▀▀
    
    build      : b8728-36daf2491
    model      : gemma-3n-E4B-it-Q4_K_M.gguf
    modalities : text
    
    available commands:
      /exit or Ctrl+C     stop or exit
      /regen              regenerate the last response
      /clear              clear the chat history
      /read <file>        add a text file
      /glob <pattern>     add text files using globbing pattern
    
    
    > hii
    
    |GGML OpenVINO backend ov::Exception: Exception from src\inference\src\cpp\infer_request.cpp:157:
    Exception from src\inference\src\cpp\infer_request.cpp:67:
    Check 'shape.compatible(ov::PartialShape(tensor->get_shape())) || tensor->get_size() == 0' failed at src\plugins\intel_cpu\src\infer_request.cpp:442:
    Can't set the output tensor with index: 2, because the model output tensor (shape=[1,2,2048,?]) and the current tensor (shape=(1.6.2048.4)) are incompatible
    
    
    
    graph_compute: ggml_backend_sched_graph_compute_async failed with error -1
    process_ubatch: failed to compute graph, compute status: -1
    llama_decode: failed to decode, ret = -3                                   srv  update_slots: Compute error. i = 0, n_batch = 2048, ret = -3
    srv    send_error: task id = 0, error: Compute error.
    Error: Compute error.
    
    
    [ Prompt: 0.0 t/s | Generation: 0.0 t/s ]
    
    >
    

    So, I wanted to ask, should I open a PR first Introducing the node support or open a Final PR solving the whole issue at once. I don't know the current reviewing system.

    Thank you!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions