Backend
VL (Velox)
Bug description
When testing, I find a task with following logs:
I20260920 14:24:00.305145 38755 LocalPartitionWriter.cc:475] Merge 183 spills to 4000 partitions.
W20260920 14:24:43.041427 38755 Utils.cc:283] madvise willneed failed: Invalid argument
W20260920 14:24:55.489629 38755 Utils.cc:283] madvise willneed failed: Invalid argument
...
W20260920 14:26:52.805511 38755 Utils.cc:283] madvise willneed failed: Invalid argument
W20260920 14:26:52.808156 38755 Utils.cc:283] madvise willneed failed: Invalid argument
I20260920 14:26:52.812170 38755 LocalPartitionWriter.cc:507] Merge spills done in 172 s, detail: mergeTempFileTime:159 s payloadCacheWriteTime:0 s payloadMergerWriteTime:8 s totalTempFileSize:1433658590 bytes writeSize:1564238492 bytes
After some investigation, I find there is a bug in this line of code. We shouldn't calculate fetchLen using pos_, because pos_ can lag behind posFetch_. If we use pos_ to calculate fetchLen, we will have two problem:
- Unaligned madvise addresses, prefetch silently dropped. When the result is not page-aligned,
posFetch_ lands on a non-page-aligned offset inside the file and stays there. Every subsequent madvise(MADV_WILLNEED) then passes an unaligned address, which the kernel rejects with EINVAL, so all further prefetching is lost.
- Advised range overruns the mapping. When
pos_ lags behind posFetch_, the result length is too large by the gap, so the advised range extends past the end of the mapping. The overrunning part silently operates on whatever mapping happens to follow in the address space (e.g. other mmap'd files or allocator arenas).
Gluten version
No response
Spark version
None
Spark configurations
No response
System information
No response
Relevant logs
Backend
VL (Velox)
Bug description
When testing, I find a task with following logs:
After some investigation, I find there is a bug in this line of code. We shouldn't calculate
fetchLenusingpos_, becausepos_can lag behindposFetch_. If we usepos_to calculatefetchLen, we will have two problem:posFetch_lands on a non-page-aligned offset inside the file and stays there. Every subsequentmadvise(MADV_WILLNEED)then passes an unaligned address, which the kernel rejects withEINVAL, so all further prefetching is lost.pos_lags behindposFetch_, the result length is too large by the gap, so the advised range extends past the end of the mapping. The overrunning part silently operates on whatever mapping happens to follow in the address space (e.g. other mmap'd files or allocator arenas).Gluten version
No response
Spark version
None
Spark configurations
No response
System information
No response
Relevant logs