Repository navigation
wc: add a UTF-8 character-count benchmark - #15202
Conversation
| } | ||
|
|
||
| #[divan::bench(args = [100_000])] | ||
| fn wc_chars_utf8(bencher: Bencher, num_lines: usize) { |
There was a problem hiding this comment.
could you please set LC_ALL in the bench itself, like join_bench.rs or sort_locale_utf8_bench.rs?
codspeed won't use the env from the doc
| let file_path = create_test_file(data.as_bytes(), temp_dir.path()); | ||
|
|
||
| bencher | ||
| .with_inputs(|| get_bench_args(&[&"-cm", &file_path]).into_iter()) |
There was a problem hiding this comment.
why -cm and not just -m like wc_chars_large_line_count?
Merging this PR will degrade performance by 14.95%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | three_39_bit_primes |
438.2 ms | 761.9 ms | -42.49% |
| ⚡ | Simulation | five_38_bit_primes |
1.9 s | 1.8 s | +3.83% |
| ⚡ | Simulation | thirteen_39_bit_primes |
9.1 s | 8.8 s | +3.04% |
| 🆕 | Memory | wc_chars_utf8[100000] |
N/A | 16.6 KB | N/A |
| 🆕 | Simulation | wc_chars_utf8[100000] |
N/A | 767.1 µs | N/A |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing darkraider01:wc-utf8-benchmark (9d308ce) with main (33a2889)
Footnotes
-
54 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
Thanks for your PR |
The existing character-count benchmark uses ASCII input. Add a UTF-8 case with two-, three-, and four-byte characters, plus instructions for running it in a UTF-8 locale.
Split from #15139 and #15140 as requested in #15140 (comment).