Skip to content

test-save-load-state : print a per-model results table in --models mode - #29316

Merged
ggerganov merged 3 commits into
masterfrom
gg/test-save-load-clean-up-logs
Sep 24, 2026
Merged

ggerganov merged 3 commits into
masterfrom
gg/test-save-load-clean-up-logs

Conversation

@ggerganov

@ggerganov ggerganov commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

Overview

The --models output of test-save-load-state was very heavy: every model printed its token dumps, per-test headers and PASS lines, making it hard to see which models pass.

  • in --models mode, silence all logging except the table itself (common_log_set_verbosity_thold(0) leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with one column per test (colored PASS/FAIL/SKIP cells), row by row
  • the suite now continues past failures: tests 3-5 are SKIPped when the baseline (test 1) fails, model init failure skips all tests; results are stored in a dynamic std::vector<test_status>
  • per-test token dumps, test headers and PASS lines are demoted to LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode
  • add a print_usage callback (passed to common_params_parse) so -h documents the tool-specific --models option with example commands
  • single-model output and exit codes are unchanged

Additional information

sample output from --models over the 10 arch-test models:

Model                  baseline  seq_rm  state_load  cp_h  cp_d  cp_h_s  cp_d_s  rt  
bailingmoe3-moe.gguf   PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
dream-dense.gguf       PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
hrm_text-dense.gguf    PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
kimi-k3-moe.gguf       PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
llama-moe.gguf         PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
minicpm-moe.gguf       PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
mistral3-dense.gguf    PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
mpt-dense.gguf         PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
nemotron_h-dense.gguf  PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
qwen2-dense.gguf       PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS
0.00.998.926 I main: summary: 10 passed, 0 failed (of 10)

Requirements

in --models mode the output was very heavy: every model printed its
token dumps, per-test headers and PASS lines. instead, silence all
logging except the table itself (common_log_set_verbosity_thold(0)
leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with
one column per test, colored PASS/FAIL/SKIP cells, row by row.

- run_save_load_tests_for_model returns a test_suite with a dynamic
  std::vector<test_status> and continues past failures: tests 3-5 are
  SKIPped when the baseline (test 1) fails, model init failure skips all
- per-test token dumps, test headers and PASS lines are demoted to
  LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode
- the table header/rows derive their columns from test_names; the
  model name is printed and flushed before the suite runs so the model
  currently in flight is always visible
- single-model output and exit codes are unchanged

Assisted-by: pi:llama.cpp/Qwen3.8-27B
add a print_usage callback passed to common_params_parse, so -h/--help
also shows example commands for the tool-specific --models option and
the -lv verbosity level

Assisted-by: pi:llama.cpp/Qwen3.8-27B
@github-actions github-actions Bot added the testing Everything test related label Sep 23, 2026
@ggerganov
ggerganov marked this pull request as ready for review September 23, 2026 14:15
Comment thread tests/test-save-load-state.cpp Outdated
Comment thread tests/test-save-load-state.cpp Outdated
Comment thread tests/test-save-load-state.cpp Outdated
Comment thread tests/test-save-load-state.cpp Outdated
ref: #29316

Assisted-by: pi:llama.cpp/Qwen3.8-27B
@ggerganov
ggerganov merged commit 4c5957c into master Sep 24, 2026
12 checks passed
@ggerganov
ggerganov deleted the gg/test-save-load-clean-up-logs branch September 24, 2026 06:17
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…de (ggml-org#29316)

* test-save-load-state : print a per-model results table in --models mode

in --models mode the output was very heavy: every model printed its
token dumps, per-test headers and PASS lines. instead, silence all
logging except the table itself (common_log_set_verbosity_thold(0)
leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with
one column per test, colored PASS/FAIL/SKIP cells, row by row.

- run_save_load_tests_for_model returns a test_suite with a dynamic
  std::vector<test_status> and continues past failures: tests 3-5 are
  SKIPped when the baseline (test 1) fails, model init failure skips all
- per-test token dumps, test headers and PASS lines are demoted to
  LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode
- the table header/rows derive their columns from test_names; the
  model name is printed and flushed before the suite runs so the model
  currently in flight is always visible
- single-model output and exit codes are unchanged

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* test-save-load-state : print example usage on -h

add a print_usage callback passed to common_params_parse, so -h/--help
also shows example commands for the tool-specific --models option and
the -lv verbosity level

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* test-save-load-state : remove comments

ref: ggml-org#29316

Assisted-by: pi:llama.cpp/Qwen3.8-27B
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…de (ggml-org#29316)

* test-save-load-state : print a per-model results table in --models mode

in --models mode the output was very heavy: every model printed its
token dumps, per-test headers and PASS lines. instead, silence all
logging except the table itself (common_log_set_verbosity_thold(0)
leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with
one column per test, colored PASS/FAIL/SKIP cells, row by row.

- run_save_load_tests_for_model returns a test_suite with a dynamic
  std::vector<test_status> and continues past failures: tests 3-5 are
  SKIPped when the baseline (test 1) fails, model init failure skips all
- per-test token dumps, test headers and PASS lines are demoted to
  LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode
- the table header/rows derive their columns from test_names; the
  model name is printed and flushed before the suite runs so the model
  currently in flight is always visible
- single-model output and exit codes are unchanged

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* test-save-load-state : print example usage on -h

add a print_usage callback passed to common_params_parse, so -h/--help
also shows example commands for the tool-specific --models option and
the -lv verbosity level

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* test-save-load-state : remove comments

ref: ggml-org#29316

Assisted-by: pi:llama.cpp/Qwen3.8-27B
(cherry picked from commit 4c5957c)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant