Skip to content

hrm : fix layer placement of z_l_init weight - #29512

Merged
ggerganov merged 1 commit into
masterfrom
gg/hrm-fix-z-placement
Sep 27, 2026
Merged

ggerganov merged 1 commit into
masterfrom
gg/hrm-fix-z-placement

Conversation

@ggerganov

Copy link
Copy Markdown
Member

Overview

fix #29484
cont #27625

The embd + hrm.z_l_init op was done in the input layer which is inefficient and causes graph reallocations. Move it to the first repeating layer.

Requirements

@ggerganov
ggerganov requested a review from CISC as a code owner September 27, 2026 06:51
@github-actions github-actions Bot added the model Model specific label Sep 27, 2026
@ggerganov
ggerganov merged commit 85ca3b5 into master Sep 27, 2026
13 checks passed
@ggerganov
ggerganov deleted the gg/hrm-fix-z-placement branch September 27, 2026 07:16
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model Model specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval bug: test-save-load-state aborts on hrm_text with GGML_SCHED_NO_REALLOC (input-layer tensor makes the backend assignment depend on batch size)

1 participant