Repository navigation
server: refactor sleep handling, allow access /metrics during sleep - #27376
Conversation
|
I think /slots should keep resetting the idle timer, only /metrics should skip it, otherwise monitoring no longer keeps the server awake but still forces a model reload on every polling round.
I pushed directly: |
313e807 to
0baf73a
Compare
|
for /slots vs /metrics, it's better to separate them to 2 different tasks, which was done in c1fc93f that also resolved the |
|
Nice. Splitting the task type kills the root cause instead of patching the symptom, and keeping the last window until a scrape reads it is more correct than dropping it on sleep |
An upstream refactor can merge and compile clean and still break the memory endpoints. The sleep refactor (ggml-org#27376) did exactly that: it moved /metrics behind a cached snapshot, and the memory gauges went missing from the sleeping path while the build stayed green. Run tools/server/tests after the compile step, the same tests.sh upstream's own server workflow uses. The build takes about 4 minutes and the tests about the same, so the existing 60 minute job timeout still has room.
Upstream restructured the code the memory work sits on. server: refactor sleep handling (ggml-org#27376) split slot reporting out of SERVER_TASK_TYPE_METRICS into a new SERVER_TASK_TYPE_SLOT_GET task, rewrote get_metrics around a cached snapshot served during sleep, and extracted the props body into get_res_props(). memory_data stays on server_task_result_metrics, the memory gauges keep their place on /metrics, and endpoint_memory is back in get_res_props(). /metrics during sleep must equal the last awake scrape. The model is unloaded by then, so update_cached_responses() now renders the memory series while the model is still loaded and use_cached_metrics() appends the cached string. The rendering moved into render_memory_metrics() and the breakdown into server_context_impl::get_memory_data(), so the awake and sleeping paths share one implementation. common: add json.h abstraction (ggml-org#27511) replaced nlohmann::json with common_json, which has no find(). The gauge loop uses contains() and at() instead.
…gml-org#27376) * add cached responses * refactor on_sleeping_state * allow accessing metrics during sleep * metrics task should not reset timer * updated docs * fix * fix get_res_model_info * add test * fix a race condition * split metrics and slots tasks / results * should_reset_buckets
…gml-org#27376) * add cached responses * refactor on_sleeping_state * allow accessing metrics during sleep * metrics task should not reset timer * updated docs * fix * fix get_res_model_info * add test * fix a race condition * split metrics and slots tasks / results * should_reset_buckets
…gml-org#27376) * add cached responses * refactor on_sleeping_state * allow accessing metrics during sleep * metrics task should not reset timer * updated docs * fix * fix get_res_model_info * add test * fix a race condition * split metrics and slots tasks / results * should_reset_buckets
…gml-org#27376) * add cached responses * refactor on_sleeping_state * allow accessing metrics during sleep * metrics task should not reset timer * updated docs * fix * fix get_res_model_info * add test * fix a race condition * split metrics and slots tasks / results * should_reset_buckets
Overview
server_routeswill read cached response during sleepon_sleeping_stateas a side-effect,
/metricscan now be accessed during sleepRequirements