Repository navigation
wc: count characters as bytes in single-byte locales - #15140
darkraider01 wants to merge 8 commits into
Conversation
| writeln!(stdout) | ||
| } | ||
|
|
||
| static IS_C_LOCALE: LazyLock<bool> = LazyLock::new(|| { |
There was a problem hiding this comment.
Byte counting is limited to C/POSIX because other non-UTF-8 locales can still use multibyte encodings. Those locales keep their existing counting paths.
|
GNU testsuite comparison: |
Merging this PR will regress 8 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | three_39_bit_primes |
630.5 ms | 783.8 ms | -19.56% |
| ❌ | Simulation | wc_chars_large_line_count[100000] |
2.3 ms | 2.8 ms | -18.04% |
| ❌ | Simulation | unexpand_large_file[10] |
274.8 ms | 286.3 ms | -4.01% |
| ❌ | Simulation | unexpand_many_lines[100000] |
131.4 ms | 136.9 ms | -4% |
| ❌ | Simulation | sort_mixed_utf8_locale |
80.8 ms | 83.6 ms | -3.35% |
| ❌ | Simulation | sort_reverse_utf8_locale |
80.6 ms | 83.3 ms | -3.32% |
| ❌ | Simulation | sort_unique_utf8_locale |
83.6 ms | 86.3 ms | -3.19% |
| ❌ | Simulation | cksum_crc32b |
33.2 ms | 34.2 ms | -3.07% |
| ⚡ | Simulation | cut_fields_custom_delim |
65.9 ms | 51.9 ms | +27% |
| ⚡ | Simulation | cut_fields_tab |
57.4 ms | 45.5 ms | +26.24% |
| ⚡ | Simulation | cut_bytes |
18.1 ms | 15.5 ms | +16.66% |
| ⚡ | Simulation | cut_characters |
25.4 ms | 22.9 ms | +10.88% |
| ⚡ | Simulation | five_38_bit_primes |
1.8 s | 1.6 s | +8.4% |
| 🆕 | Memory | wc_chars_utf8[100000] |
N/A | 16.6 KB | N/A |
| 🆕 | Simulation | wc_chars_utf8[100000] |
N/A | 1.5 ms | N/A |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing darkraider01:wc-c-locale-character-counts (0a7e171) with main (9019055)2
Footnotes
-
54 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
main(d027897) during the generation of this report, so 9019055 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
6523468 to
0a7e171
Compare
In single-byte locales (where MB_CUR_MAX == 1), characters should be counted as bytes for
wc -m.Previously, this was inferred by checking if the locale was named "C" or "POSIX", which does not match GNU semantics on platforms where the C locale resolves to UTF-8, nor does it handle other single-byte locales (such as ISO-8859-1).
This updates
wcto detect whether the active locale encoding is genuinely single-byte via native libc locale properties (MB_CUR_MAX == 1), counting bytes directly for-macross all dispatch paths while preserving multibyte counting for UTF-8 and other multibyte locales.Once #15139 merges, I'll rebase this follow-up ^^
Closes #14928.