Repository navigation
test-save-load-state : print a per-model results table in --models mode - #29316
Merged
Merged
Conversation
in --models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead, silence all logging except the table itself (common_log_set_verbosity_thold(0) leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with one column per test, colored PASS/FAIL/SKIP cells, row by row. - run_save_load_tests_for_model returns a test_suite with a dynamic std::vector<test_status> and continues past failures: tests 3-5 are SKIPped when the baseline (test 1) fails, model init failure skips all - per-test token dumps, test headers and PASS lines are demoted to LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode - the table header/rows derive their columns from test_names; the model name is printed and flushed before the suite runs so the model currently in flight is always visible - single-model output and exit codes are unchanged Assisted-by: pi:llama.cpp/Qwen3.8-27B
add a print_usage callback passed to common_params_parse, so -h/--help also shows example commands for the tool-specific --models option and the -lv verbosity level Assisted-by: pi:llama.cpp/Qwen3.8-27B
ggerganov
marked this pull request as ready for review
September 23, 2026 14:15
ggerganov
commented
Sep 23, 2026
ref: #29316 Assisted-by: pi:llama.cpp/Qwen3.8-27B
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
…de (ggml-org#29316) * test-save-load-state : print a per-model results table in --models mode in --models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead, silence all logging except the table itself (common_log_set_verbosity_thold(0) leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with one column per test, colored PASS/FAIL/SKIP cells, row by row. - run_save_load_tests_for_model returns a test_suite with a dynamic std::vector<test_status> and continues past failures: tests 3-5 are SKIPped when the baseline (test 1) fails, model init failure skips all - per-test token dumps, test headers and PASS lines are demoted to LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode - the table header/rows derive their columns from test_names; the model name is printed and flushed before the suite runs so the model currently in flight is always visible - single-model output and exit codes are unchanged Assisted-by: pi:llama.cpp/Qwen3.8-27B * test-save-load-state : print example usage on -h add a print_usage callback passed to common_params_parse, so -h/--help also shows example commands for the tool-specific --models option and the -lv verbosity level Assisted-by: pi:llama.cpp/Qwen3.8-27B * test-save-load-state : remove comments ref: ggml-org#29316 Assisted-by: pi:llama.cpp/Qwen3.8-27B
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 7, 2026
…de (ggml-org#29316) * test-save-load-state : print a per-model results table in --models mode in --models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead, silence all logging except the table itself (common_log_set_verbosity_thold(0) leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with one column per test, colored PASS/FAIL/SKIP cells, row by row. - run_save_load_tests_for_model returns a test_suite with a dynamic std::vector<test_status> and continues past failures: tests 3-5 are SKIPped when the baseline (test 1) fails, model init failure skips all - per-test token dumps, test headers and PASS lines are demoted to LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode - the table header/rows derive their columns from test_names; the model name is printed and flushed before the suite runs so the model currently in flight is always visible - single-model output and exit codes are unchanged Assisted-by: pi:llama.cpp/Qwen3.8-27B * test-save-load-state : print example usage on -h add a print_usage callback passed to common_params_parse, so -h/--help also shows example commands for the tool-specific --models option and the -lv verbosity level Assisted-by: pi:llama.cpp/Qwen3.8-27B * test-save-load-state : remove comments ref: ggml-org#29316 Assisted-by: pi:llama.cpp/Qwen3.8-27B (cherry picked from commit 4c5957c)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
The
--modelsoutput oftest-save-load-statewas very heavy: every model printed its token dumps, per-test headers and PASS lines, making it hard to see which models pass.--modelsmode, silence all logging except the table itself (common_log_set_verbosity_thold(0)leaves onlyLOG/LOG_LEVEL_OUTPUT) and print one row per model with one column per test (colored PASS/FAIL/SKIP cells), row by rowstd::vector<test_status>LOGV(LOG_LEVEL_INFO, ...)so they still show in single-model modeprint_usagecallback (passed tocommon_params_parse) so-hdocuments the tool-specific--modelsoption with example commandsAdditional information
sample output from
--modelsover the 10 arch-test models:Requirements