Skip to content

server: refactor batch construction - #24843

Merged
ngxson merged 15 commits into
masterfrom
xsn/server_refactor_batch
Jun 21, 2026
Merged

ngxson merged 15 commits into
masterfrom
xsn/server_refactor_batch

Conversation

@ngxson

@ngxson ngxson commented Jun 20, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

Multiple motivations for this refactoring:

Additional information

The detailed implementation:

  • update_slots() now contains 3 steps: pre-decode / decode / post-decode
  • a new struct server_batch is added to abstract out the construction of llama_batch
  • metric updates are queued inside server_batch and is applied once the batch decoding complete
  • sub-batch (smaller batch_view) will be handled by server_batch

Logic flow:

  1. pre-decode add tokens to the server_batch
  2. decode processes the batch, may require a smaller batch_view if it doesn't fit
  3. post-decode sample the token

TODO:

  • better handling i_batch
  • implement sub-batch
  • queue metrics --> will be a dedicated PR
  • per-slot error handling

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: 99% code is hand-written, AI only used for validation and nit fixes

Comment on lines +2663 to +2669
t_prev = t_start;
SRV_INF("n_pre_decode = %" PRId64 "\n", n_pre_decode);
SRV_INF("avg t_pre_decode = %f ms\n", (double) t_pre_decode / n_pre_decode / 1000.0);
SRV_INF("avg t_decode = %f ms\n", (double) t_decode / n_decode / 1000.0);
SRV_INF("avg t_post_decode = %f ms\n", (double) t_post_decode / n_post_decode / 1000.0);
SRV_INF("avg t_sampl = %f ms\n", (double) t_sampl / n_sampl / 1000.0);
}

@ngxson ngxson Jun 20, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok so I did a quick test on how the number of parallel requests affect the timings; turns out, t_post_decode grows proportionally. so probably worth making sampling multi-thread.

2 parallel requests:

0.22.465.968 I srv  update_slots: n_pre_decode      = 1427
0.22.465.971 I srv  update_slots: avg t_pre_decode  = 0.002619 ms
0.22.465.972 I srv  update_slots: avg t_decode      = 12.810192 ms
0.22.465.972 I srv  update_slots: avg t_post_decode = 0.464135 ms
0.22.465.973 I srv  update_slots: avg t_sampl       = 0.230341 ms

10 parallel requests:

0.38.793.972 I srv  update_slots: n_pre_decode      = 207
0.38.793.986 I srv  update_slots: avg t_pre_decode  = 0.030816 ms
0.38.793.993 I srv  update_slots: avg t_decode      = 161.709222 ms
0.38.793.994 I srv  update_slots: avg t_post_decode = 4.672478 ms
0.38.793.995 I srv  update_slots: avg t_sampl       = 0.466093 ms

@ngxson

ngxson commented Jun 20, 2026 •

Copy link
Copy Markdown
Collaborator Author

@ggml-org/maintainers I'm trying to measure how sampling speed affects serving parallel requests on llama-server, and probably make it multi-threaded if it worth improving. Appreciate if someone can run tests on your hardware.

To test it:

0.27.514.119 I srv  update_slots: n_pre_decode      = 1276
0.27.514.123 I srv  update_slots: avg t_pre_decode  = 0.004343 ms
0.27.514.123 I srv  update_slots: avg t_decode      = 17.561716 ms
0.27.514.124 I srv  update_slots: avg t_post_decode = 0.932403 ms
0.27.514.124 I srv  update_slots: avg t_sampl       = 0.231795 ms

Thanks in advanced!


Tested on RTX 5060 Ti:

Results

2 requests:

0.17.421.212 I srv  update_slots: n_pre_decode      = 1283
0.17.421.216 I srv  update_slots: avg t_pre_decode  = 0.001221 ms
0.17.421.217 I srv  update_slots: avg t_decode      = 10.054578 ms
0.17.421.217 I srv  update_slots: avg t_post_decode = 0.500834 ms
0.17.421.217 I srv  update_slots: avg t_sampl       = 0.250060 ms
0.19.107.251 I slot print_timing: id  2 | task 1 | n_decoded =   1438, tg =  95.64 t/s, tg_3s =  93.17 t/s
0.19.107.517 I slot print_timing: id  3 | task 0 | n_decoded =   1438, tg =  95.64 t/s, tg_3s =  93.17 t/s
0.22.114.424 I slot print_timing: id  2 | task 1 | n_decoded =   1717, tg =  95.16 t/s, tg_3s =  92.78 t/s
0.22.114.715 I slot print_timing: id  3 | task 0 | n_decoded =   1717, tg =  95.16 t/s, tg_3s =  92.78 t/s

10 requests:

0.22.505.793 I srv  update_slots: n_pre_decode      = 1014
0.22.505.797 I srv  update_slots: avg t_pre_decode  = 0.006096 ms
0.22.505.798 I srv  update_slots: avg t_decode      = 14.740408 ms
0.22.505.798 I srv  update_slots: avg t_post_decode = 3.481039 ms
0.22.505.798 I srv  update_slots: avg t_sampl       = 0.347986 ms
0.25.421.670 I slot print_timing: id  9 | task 0 | n_decoded =   1166, tg =  55.33 t/s, tg_3s =  52.81 t/s
0.25.437.551 I slot print_timing: id  0 | task 11 | n_decoded =   1166, tg =  55.37 t/s, tg_3s =  52.81 t/s
0.25.437.957 I slot print_timing: id  1 | task 10 | n_decoded =   1166, tg =  55.37 t/s, tg_3s =  52.81 t/s

10 requests with multi-threaded sampling (PoC):

0.22.523.109 I srv  update_slots: n_pre_decode      = 1008
0.22.523.114 I srv  update_slots: avg t_pre_decode  = 0.008549 ms
0.22.523.115 I srv  update_slots: avg t_decode      = 14.840249 ms
0.22.523.115 I srv  update_slots: avg t_post_decode = 1.690424 ms
0.22.523.116 I srv  update_slots: avg t_sampl       = 1.670142 ms
0.24.254.467 I slot print_timing: id  9 | task 0 | n_decoded =   1107, tg =  61.32 t/s, tg_3s =  58.40 t/s
0.24.288.705 I slot print_timing: id  0 | task 9 | n_decoded =   1108, tg =  61.36 t/s, tg_3s =  58.40 t/s
0.24.288.709 I slot print_timing: id  1 | task 8 | n_decoded =   1108, tg =  61.36 t/s, tg_3s =  58.40 t/s

More:

20 requests:

single-threaded:

0.12.283.381 I srv  update_slots: n_pre_decode      = 130
0.12.283.386 I srv  update_slots: avg t_pre_decode  = 3.050846 ms
0.12.283.387 I srv  update_slots: avg t_decode      = 35.677700 ms
0.12.283.387 I srv  update_slots: avg t_post_decode = 7.097292 ms
0.12.283.388 I srv  update_slots: avg t_sampl       = 0.358927 ms
0.14.175.071 I slot print_timing: id 19 | task 0 | n_decoded =    177, tg =  24.06 t/s, tg_3s =  25.36 t/s
0.14.207.762 I slot print_timing: id  0 | task 20 | n_decoded =    177, tg =  25.45 t/s, tg_3s =  25.36 t/s
0.14.208.145 I slot print_timing: id  1 | task 19 | n_decoded =    177, tg =  25.45 t/s, tg_3s =  25.36 t/s


multi-threaded:

0.12.032.348 I srv  update_slots: n_pre_decode      = 166
0.12.032.353 I srv  update_slots: avg t_pre_decode  = 2.386849 ms
0.12.032.356 I srv  update_slots: avg t_decode      = 35.066090 ms
0.12.032.357 I srv  update_slots: avg t_post_decode = 4.716886 ms
0.12.032.357 I srv  update_slots: avg t_sampl       = 4.702030 ms
0.12.627.081 I slot print_timing: id 19 | task 0 | n_decoded =    181, tg =  25.51 t/s, tg_3s =  26.92 t/s
0.12.664.347 I slot print_timing: id  0 | task 20 | n_decoded =    181, tg =  27.04 t/s, tg_3s =  26.91 t/s
0.12.664.350 I slot print_timing: id  1 | task 19 | n_decoded =    181, tg =  27.04 t/s, tg_3s =  26.91 t/s

---

30 requests:

single-threaded:

0.17.498.687 I srv  update_slots: n_pre_decode      = 183
0.17.498.691 I srv  update_slots: avg t_pre_decode  = 3.243945 ms
0.17.498.692 I srv  update_slots: avg t_decode      = 48.384131 ms
0.17.498.692 I srv  update_slots: avg t_post_decode = 10.854809 ms
0.17.498.693 I srv  update_slots: avg t_sampl       = 0.363318 ms
0.18.947.927 I slot print_timing: id 14 | task 15 | n_decoded =    208, tg =  17.51 t/s, tg_3s =  17.89 t/s
0.18.948.328 I slot print_timing: id 15 | task 14 | n_decoded =    208, tg =  17.51 t/s, tg_3s =  17.89 t/s
0.18.948.712 I slot print_timing: id 16 | task 13 | n_decoded =    208, tg =  17.52 t/s, tg_3s =  17.89 t/s

multi-threaded:

0.17.659.210 I srv  update_slots: n_pre_decode      = 193
0.17.659.213 I srv  update_slots: avg t_pre_decode  = 3.112720 ms
0.17.659.214 I srv  update_slots: avg t_decode      = 48.381777 ms
0.17.659.214 I srv  update_slots: avg t_post_decode = 7.596896 ms
0.17.659.214 I srv  update_slots: avg t_sampl       = 7.580375 ms
0.18.874.459 I slot print_timing: id 29 | task 0 | n_decoded =    215, tg =  18.07 t/s, tg_3s =  18.96 t/s
0.18.927.561 I slot print_timing: id  0 | task 30 | n_decoded =    215, tg =  19.06 t/s, tg_3s =  18.96 t/s
0.18.927.565 I slot print_timing: id  1 | task 29 | n_decoded =    215, tg =  19.06 t/s, tg_3s =  18.96 t/s

@ngxson
ngxson marked this pull request as ready for review June 20, 2026 18:45
@ngxson
ngxson requested a review from a team as a code owner June 20, 2026 18:45
@ngxson
ngxson requested a review from ggerganov June 20, 2026 18:46
@ggerganov ggerganov self-assigned this Jun 20, 2026
@angt

angt commented Jun 21, 2026 •

Copy link
Copy Markdown
Member

here my results on macbook air m3:

$ python run_parallel_cmpl.py 1
==> Sending 1 parallel requests to http://localhost:8080/chat/completions
    Model: llama
    [0] a mermaid cursed to sing only in dad jokes

--- Results (1 ok, 0 failed, wall-clock 22.1s) ---
  [0] 200  22.1s  ~0 words  a mermaid cursed to sing only in dad jokes

$ python run_parallel_cmpl.py 2
==> Sending 2 parallel requests to http://localhost:8080/chat/completions
    Model: llama
    [0] a chef who can taste emotions and cooks meals that change people's moods
    [1] a runaway prince who joins a traveling circus of magical misfits

--- Results (2 ok, 0 failed, wall-clock 64.0s) ---
  [0] 200  64.0s  ~0 words  a chef who can taste emotions and cooks meals that change people's moods
  [1] 200  64.0s  ~0 words  a runaway prince who joins a traveling circus of magical misfits

$ python run_parallel_cmpl.py 5
==> Sending 5 parallel requests to http://localhost:8080/chat/completions
    Model: llama
    [0] a mermaid cursed to sing only in dad jokes
    [1] a robot learning to paint sunsets in a post-human world
    [2] a pirate crew sailing a sea made of clouds, hunting for a floating island of gold
    [3] a time traveler who keeps accidentally adopting stray animals from every era they visit
    [4] a detective solving crimes in a world where everyone's shadow has a mind of its own

--- Results (5 ok, 0 failed, wall-clock 104.7s) ---
  [0] 200  104.7s  ~0 words  a mermaid cursed to sing only in dad jokes
  [1] 200  72.2s  ~0 words  a robot learning to paint sunsets in a post-human world
  [2] 200  72.2s  ~213 words  a pirate crew sailing a sea made of clouds, hunting for a floating island of gold
  [3] 200  72.2s  ~0 words  a time traveler who keeps accidentally adopting stray animals from every era they visit
  [4] 200  72.2s  ~0 words  a detective solving crimes in a world where everyone's shadow has a mind of its own
  
$ python run_parallel_cmpl.py 10
==> Sending 10 parallel requests to http://localhost:8080/chat/completions
    Model: llama
    [0] a detective solving crimes in a world where everyone's shadow has a mind of its own
    [1] a cat who discovers it can teleport but only to places it has already napped
    [2] a pirate crew sailing a sea made of clouds, hunting for a floating island of gold
    [3] a snowman who comes to life every winter and has been alive for 300 years
    [4] a mermaid cursed to sing only in dad jokes
    [5] a snowman who comes to life every winter and has been alive for 300 years
    [6] a dragon who is terrible at hoarding gold but excellent at collecting rare books
    [7] a dragon who is terrible at hoarding gold but excellent at collecting rare books
    [8] a cat who discovers it can teleport but only to places it has already napped
    [9] a runaway prince who joins a traveling circus of magical misfits

--- Results (10 ok, 0 failed, wall-clock 241.8s) ---
  [0] 200  173.0s  ~0 words  a detective solving crimes in a world where everyone's shadow has a mind of its own
  [1] 200  241.8s  ~0 words  a cat who discovers it can teleport but only to places it has already napped
  [2] 200  82.2s  ~779 words  a pirate crew sailing a sea made of clouds, hunting for a floating island of gold
  [3] 200  173.0s  ~0 words  a snowman who comes to life every winter and has been alive for 300 years
  [4] 200  241.8s  ~0 words  a mermaid cursed to sing only in dad jokes
  [5] 200  173.0s  ~0 words  a snowman who comes to life every winter and has been alive for 300 years
  [6] 200  82.1s  ~0 words  a dragon who is terrible at hoarding gold but excellent at collecting rare books
  [7] 200  173.0s  ~0 words  a dragon who is terrible at hoarding gold but excellent at collecting rare books
  [8] 200  82.2s  ~0 words  a cat who discovers it can teleport but only to places it has already napped
  [9] 200  82.1s  ~138 words  a runaway prince who joins a traveling circus of magical misfits

@ngxson

ngxson commented Jun 21, 2026 •

Copy link
Copy Markdown
Collaborator Author

@angt thanks but I forgot to mention that the timing log is on llama-server log (the script simply sends the requests, no timings info)

could you please post the log lines in llama-server? like this:

0.17.421.212 I srv  update_slots: n_pre_decode      = 1283
0.17.421.216 I srv  update_slots: avg t_pre_decode  = 0.001221 ms
0.17.421.217 I srv  update_slots: avg t_decode      = 10.054578 ms
0.17.421.217 I srv  update_slots: avg t_post_decode = 0.500834 ms
0.17.421.217 I srv  update_slots: avg t_sampl       = 0.250060 ms
0.19.107.251 I slot print_timing: id  2 | task 1 | n_decoded =   1438, tg =  95.64 t/s, tg_3s =  93.17 t/s
0.19.107.517 I slot print_timing: id  3 | task 0 | n_decoded =   1438, tg =  95.64 t/s, tg_3s =  93.17 t/s
0.22.114.424 I slot print_timing: id  2 | task 1 | n_decoded =   1717, tg =  95.16 t/s, tg_3s =  92.78 t/s
0.22.114.715 I slot print_timing: id  3 | task 0 | n_decoded =   1717, tg =  95.16 t/s, tg_3s =  92.78 t/s

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Here are timings on RTX 5090 with gpt-oss-20b:

==> Sending 1 parallel requests to http://localhost:8033/chat/completions

14.13.933.949 I srv  update_slots: n_pre_decode      = 1930
14.13.933.952 I srv  update_slots: avg t_pre_decode  = 0.000589 ms
14.13.933.952 I srv  update_slots: avg t_decode      = 2.484524 ms
14.13.933.953 I srv  update_slots: avg t_post_decode = 0.104937 ms
14.13.933.953 I srv  update_slots: avg t_sampl       = 0.104513 ms

14.12.040.453 I slot print_timing: id  9 | task 0 | n_decoded =   1190, tg = 396.34 t/s, tg_3s = 396.34 t/s


==> Sending 2 parallel requests to http://localhost:8033/chat/completions

0.07.027.191 I srv  update_slots: n_pre_decode      = 991
0.07.027.195 I srv  update_slots: avg t_pre_decode  = 0.001558 ms
0.07.027.195 I srv  update_slots: avg t_decode      = 3.664242 ms
0.07.027.196 I srv  update_slots: avg t_post_decode = 0.211232 ms
0.07.027.196 I srv  update_slots: avg t_sampl       = 0.105371 ms

0.09.292.503 I slot print_timing: id  8 | task 1 | n_decoded =   1583, tg = 263.52 t/s, tg_3s = 262.32 t/s
0.09.292.628 I slot print_timing: id  9 | task 0 | n_decoded =   1583, tg = 263.54 t/s, tg_3s = 262.32 t/s


==> Sending 10 parallel requests to http://localhost:8033/chat/completions

0.15.067.358 I slot print_timing: id  0 | task 4 | n_decoded =   1106, tg =  91.98 t/s, tg_3s =  92.40 t/s
0.15.067.466 I slot print_timing: id  1 | task 8 | n_decoded =   1106, tg =  91.98 t/s, tg_3s =  92.40 t/s
0.15.067.572 I slot print_timing: id  2 | task 7 | n_decoded =   1106, tg =  91.98 t/s, tg_3s =  92.40 t/s
0.15.067.679 I slot print_timing: id  3 | task 9 | n_decoded =   1106, tg =  91.99 t/s, tg_3s =  92.40 t/s
0.15.067.785 I slot print_timing: id  4 | task 3 | n_decoded =   1106, tg =  91.99 t/s, tg_3s =  92.40 t/s
0.15.067.892 I slot print_timing: id  5 | task 6 | n_decoded =   1106, tg =  91.99 t/s, tg_3s =  92.40 t/s
0.15.067.998 I slot print_timing: id  6 | task 0 | n_decoded =   1106, tg =  91.99 t/s, tg_3s =  92.40 t/s
0.15.068.103 I slot print_timing: id  7 | task 5 | n_decoded =   1106, tg =  92.00 t/s, tg_3s =  92.40 t/s
0.15.068.208 I slot print_timing: id  8 | task 2 | n_decoded =   1106, tg =  92.00 t/s, tg_3s =  92.40 t/s
0.15.068.319 I slot print_timing: id  9 | task 1 | n_decoded =   1106, tg =  92.00 t/s, tg_3s =  92.40 t/s

0.17.026.005 I srv  update_slots: n_pre_decode      = 1288
0.17.026.012 I srv  update_slots: avg t_pre_decode  = 0.006169 ms
0.17.026.013 I srv  update_slots: avg t_decode      = 9.931544 ms
0.17.026.013 I srv  update_slots: avg t_post_decode = 1.049071 ms
0.17.026.013 I srv  update_slots: avg t_sampl       = 0.104443 ms


struct server_batch {
llama_batch batch;
bool batch_rendered = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

batch_rendered is never used

Comment thread tools/server/server-context.cpp Outdated
batch_rendered = true;
}

llama_batch get_view(int32_t off, int32_t n_tokens) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
llama_batch get_view(int32_t off, int32_t n_tokens) {
llama_batch get_view(int32_t off, int32_t n_tokens) const {

@ngxson

ngxson commented Jun 21, 2026

Copy link
Copy Markdown
Collaborator Author

@ggerganov Thanks for testing. If I calculate correctly, doing multi-threaded sampling with your hardware config (-np 10 case) will improve the tg from 92 -> 98 t/s. If possible could you also try this PoC? https://github.com/ggml-org/llama.cpp/tree/xsn/tmp_smpl_parallel

Then I will let you decide if it worth adding (with a cleaner implementation)

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, there is a ~5% improvement in this case using multi-threaded sampling:

0.15.705.360 I slot print_timing: id  0 | task 9 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.363 I slot print_timing: id  1 | task 4 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.364 I slot print_timing: id  2 | task 5 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.365 I slot print_timing: id  3 | task 7 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.366 I slot print_timing: id  4 | task 8 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.366 I slot print_timing: id  5 | task 6 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.367 I slot print_timing: id  6 | task 2 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.368 I slot print_timing: id  7 | task 3 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.369 I slot print_timing: id  8 | task 1 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.15.705.369 I slot print_timing: id  9 | task 0 | n_decoded =   1156, tg =  96.16 t/s, tg_3s =  94.77 t/s
0.17.119.197 I srv  update_slots: n_pre_decode      = 1293
0.17.119.201 I srv  update_slots: avg t_pre_decode  = 0.007119 ms
0.17.119.202 I srv  update_slots: avg t_decode      = 9.956678 ms
0.17.119.203 I srv  update_slots: avg t_post_decode = 0.553848 ms
0.17.119.203 I srv  update_slots: avg t_sampl       = 0.545325 ms

I think it's worth it.

@ngxson
ngxson merged commit bddfd2b into master Jun 21, 2026
25 checks passed
adrianhoehne pushed a commit to adrianhoehne/llama.cpp that referenced this pull request Jul 5, 2026
* server: refactor batch construction

* wip

* wip 2

* wip 3

* wip 4

* add abort_all_slots

* handle batch full more carefully

* fix assert

* rm debug log

* small nits

* (debug) add timings

* debug: force llama_synchronize for accurate timings

* address comments

* disable DEBUG_TIMINGS
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
* server: refactor batch construction

* wip

* wip 2

* wip 3

* wip 4

* add abort_all_slots

* handle batch full more carefully

* fix assert

* rm debug log

* small nits

* (debug) add timings

* debug: force llama_synchronize for accurate timings

* address comments

* disable DEBUG_TIMINGS
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* server: refactor batch construction

* wip

* wip 2

* wip 3

* wip 4

* add abort_all_slots

* handle batch full more carefully

* fix assert

* rm debug log

* small nits

* (debug) add timings

* debug: force llama_synchronize for accurate timings

* address comments

* disable DEBUG_TIMINGS
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants