Repository navigation
[GSD-12696] [Xe2/BMG-G31] multi-rank MPI + Level Zero: resource_info.cpp:15 abort after upgrading CR 26.05 → 26.14 (Ubuntu 26.04, xe driver) #922
Description
Activity
- addedType: BugGeneral bug report, unexpected behavior or crashGeneral bug report, unexpected behavior or crashOS: LinuxIssue specific to Linux distributions (Ubuntu, Fedora, RHEL, etc.)Issue specific to Linux distributions (Ubuntu, Fedora, RHEL, etc.)
on May 4, 2026 Thank you for your contribution.
We tried to reproduce the crash with the following reproducer:
# run.py import dpctl q = dpctl.SyclQueue('level_zero:0') print(q.sycl_device)mpirun -np 2 python run.pyHowever, on our side the problem is not reproducible on BMG-G31 e223 + 7.0.0-14-generic + Ubuntu 26.04 + https://github.com/intel/compute-runtime/releases/tag/26.14.37833.4
We could try another reproducer if you provide instructions on how to set up/build the specific benchmark:
mpirun -np 8 foamRun -parallel -solver incompressibleFluidAdding a reproduction with single-process failure on the same hardware (Arc Pro B70 / BMG G31, PCI
0x8086:0xe223), so this may be a broader code path than just the multi-rank MPI race.Environment
OS NixOS 26.05 Kernel 7.0.3, xe driver GPU Arc Pro B70 (BMG G31, 32 GB ECC) PCI 0000:03:00.08086:e223(subsys8086:1701)Reproducer (any of these, all abort identically — single Python process, no MPI):
# 1. torch 2.10.0+xpu (vendored by github.com/utensils/comfyui-nix) # 2. torch 2.13.0.dev20260515+xpu (pytorch.org nightly XPU index) # 3. torch 2.8.0+xpu + intel-extension-for-pytorch 2.8.10+xpu python -c "import torch; print(torch.xpu.device_count())"
Aborts at:
Created new BO with GEM_USERPTR, handle: BO-0 Abort was called at 15 line in file: ../../neo/shared/source/gmm_helper/resource_info.cppNEO + gmmlib versions I tried (same abort in every cell, including with Intel's own prebuilt
.debbinaries unpacked from your releases):NEO gmmlib Result 26.05.37020.3 22.9.0 (the deb pairs you ship) abort 26.05.37020.3 22.10.0 abort 26.14.37833.4 22.9.0 abort 26.14.37833.4 22.10.0 abort 26.18.38308.1 22.10.0 abort NEO debug knobs that did not change behaviour:
EnableLocalMemory=0,ForceDeviceMemoryAlloc=0,EnableDirectSubmission=0,EnableHostUsmAllocationPool=0,EnableDeviceUsmAllocationPool=0,
SetCommandStreamReceiver=1,EnableXeCorePool=1.gdb backtrace at the abort:
#3 NEO::abortUnrecoverable libze_intel_gpu.so #4 NEO::GmmResourceInfo::create libze_intel_gpu.so #5 NEO::Gmm::Gmm(GmmHelper*, void*, size, align, type, StorageInfo&, GmmRequirements&) libze_intel_gpu.so #6 NEO::DrmMemoryManager::makeGmmIfSingleHandle libze_intel_gpu.so #7 NEO::DrmMemoryManager::allocateGraphicsMemoryInDevicePool libze_intel_gpu.so #8 NEO::MemoryManager::allocateGraphicsMemoryInPreferredPool libze_intel_gpu.so #9 NEO::CommandStreamReceiver::initializeResources libze_intel_gpu.so #10 NEO::Device::createEngines libze_intel_gpu.so #11 NEO::DeviceFactory::{lambda} libze_intel_gpu.so #12 L0::Driver::initialize libze_intel_gpu.so #13 std::call_once<L0::Driver::driverInit()...> libze_intel_gpu.so #16 L0::zeInit libze_intel_gpu.so #20 zeInit libze_loader.so.1 #21 ze_lib::context_t::Init libur_adapter_level_zero.so.0 […] #34 c10::xpu::initGlobalDevicePoolState libc10_xpu.so #35 c10::xpu::device_count libc10_xpu.soSo the failing allocation is the very first one
CommandStreamReceiver::initializeResources()requests during eager device init inzeInit. The kernel-side ioctl succeeds (Created new BO with GEM_USERPTR, handle: BO-0), thenmakeGmmIfSingleHandlebuilds anAllocationData+GmmRequirementntoGmmResourceInfo::create, and gmmlib returns aGMM_RESOURCE_INFO*whose underlyinghandle is zero — so theUNRECOVERABLE_IF(resourceInfo->peekHandle() == 0)onresource_info.cpp:15` fires.What works on the exact same system, with the same NEO/gmmlib/kernel (i.e. not a hardware/firmware issue):
clinfoenumerates the B70 fully, includingcl_intel_subgroup_matrix_multiply_accumulate_tf32andcl_intel_unified_shared_memory.vulkaninforeportsIntel(R) Graphics (BMG G31)with apiVersion 1.4.335, driver 26.0- llama.cpp-SYCL via
libur_adapter_opencl.so(ONEAPI_DEVICE_SELECTOR=opencl:gpu) runs models on the card to completion (~16 t/s on Qwen-32B Q4_K_M). - ONNX Runtime with OpenVINO Execution Provider (
libopenvino_intel_gpu_plugin.so) driveace-rec workloads continuously.
So gmmlib itself is functional on this card via the OpenCL ICD path; it only fails whcts the request during NEO's L0 driver init.
Hypothesis (not a fix, just where to look): on xe-kmd the
GEM_USERPTRioctl still sadata isn't shaped the same as on i915 — andDrmMemoryManager::makeGmmIfSingleHandlesynthesizes theGmmRequirements/StorageInfotriple under i915-era assumptions. The OpenCL path doesn't pass through the samemakeGmmIfSingleHandleearly-CSR-init code, which is why every other consumer of the same gmmlib works.Happy to test patches or add more instrumentation — let me know what would help.
Hi @heikogleu-dev and @pshirshov I was able to reproduce the issue on mentioned driver, but I can see that the crash is no longer reproducible on the latest public release: 26.18.38308.1 could you please verify whether upgrading to 26.18.38308.1 resolves the issue?
Is L0 supposed to work now?
No its still broken, also 26.18 did not fix it with the 7.0.3 kernel. https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8022
- changed the title
[-][Xe2/BMG-G31] multi-rank MPI + Level Zero: resource_info.cpp:15 abort after upgrading CR 26.05 → 26.14 (Ubuntu 26.04, xe driver)[/-][+][GSD-12696] [Xe2/BMG-G31] multi-rank MPI + Level Zero: resource_info.cpp:15 abort after upgrading CR 26.05 → 26.14 (Ubuntu 26.04, xe driver)[/+]on May 27, 2026 I believe I am running into this issue with Blender on Fedora 44 (a custom Bazzite build, specifically). When I select the oneAPI option in Blender, I get a crash:
Abort was called at 15 line in file: /builddir/build/BUILD/intel-compute-runtime-26.18.38308.1-build/compute-runtime-26.18.38308.1/shared/source/gmm_helper/resource_info.cpp fish: Job 1, './blender --factory-startup' terminated by signal SIGABRT (Abort)I have these dependencies installed, which I know to have previously worked in my setup, though I don't know which specific version I last used that worked:
➜ dnf5 list --installed "intel-level-zero*" "intel-compute-runtime*" "oneapi*" Installed packages (available for reinstall, available for upgrade) intel-compute-runtime.x86_64 26.18.38308.1-2.fc44 updates intel-level-zero.x86_64 26.18.38308.1-2.fc44 updates intel-level-zero-gpu-raytracing.x86_64 1.2.3-1.fc44 updates oneapi-level-zero.x86_64 1.28.6-1.fc44 updatesI can reproduce what looks like the same issue on Ubuntu 25.10 (2x ArcPro B60) with Intel GPU compute packages from the kobuk-team Intel graphics PPA.
After upgrading to:
intel-opencl-icd 26.18.38308.1-1~25.10~ppa1 libze-intel-gpu1 26.18.38308.1-1~25.10~ppa1 intel-ocloc 26.18.38308.1-1~25.10~ppa1 libigdgmm12 22.10.0-1~25.10~ppa1 intel-igc-core-2 2.28.4 intel-igc-opencl-2 2.28.4both
clinfoandsycl-lsabort with:Abort was called at 15 line in file: ./shared/source/gmm_helper/resource_info.cpp Aborted (core dumped)This happens before running any higher-level workload such as PyTorch or vLLM, so it appears to be in the OpenCL / Level Zero / GMM stack itself.
/etc/OpenCL/vendors/contains:intel.icd intel64.icdI also tested after initializing oneAPI with:
source /opt/intel/oneapi/setvars.shbut
sycl-lsstill aborts with the sameresource_info.cpp:15message.For now I will probably roll back / hold the affected packages, but I wanted to add another confirmation that the issue is still reproducible with
26.18.38308.1.Hi @encodatamHirmer,
We are currently attempting to reproduce the issue you reported, but so far we have been unable to replicate it - everything is functioning correctly in our environment.
lsb_release -a No LSB modules are available. Distributor ID: Ubuntu Description: Ubuntu 25.10 Release: 25.10uname -a ... 6.17.0-5sudo dpkg --list | grep -iE "igc|gmm|opencl|libze" ii clinfo 3.0.25.02.14-1 amd64 Query OpenCL system information ii intel-opencl-icd 26.18.38308.1-1~25.10~ppa1 amd64 Intel graphics compute runtime for OpenCL ii libigc2 2.34.4-1260~25.10 amd64 Core libraries for Intel(R) Graphics Compiler for OpenCL(TM) ii libigdfcl2 2.34.4-1260~25.10 amd64 OpenCL library for Intel(R) Graphics Compiler for OpenCL(TM) ii libigdgmm12:amd64 22.10.0-1~25.10~ppa1 amd64 Intel Graphics Memory Management Library -- shared library ii libze-intel-gpu1 26.18.38308.1-1~25.10~ppa1 amd64 Intel oneAPI L0 support implementation for Intel GPUs -- shared library ii libze1:amd64 1.28.2-1~25.10~ppa1 amd64 oneAPI Level Zero -- share libraries ii ocl-icd-libopencl1:amd64 2.3.3-1To help us investigate further, could you please:
- Share your system information Provide the output of the following commands:
lsb_release -auname -asudo dpkg --list | grep -iE "igc|gmm|opencl|libze"
- Upgrade your IGC (
2.28.4) (Intel Graphics Compiler) to the recommended version:
- Compute Runtime: https://github.com/intel/compute-runtime/releases/tag/26.18.38308.1
- Intel Graphics Compiler: https://github.com/intel/intel-graphics-compiler/releases/tag/v2.34.4
- Provide the following logs after the upgrade:
dmesgLD_DEBUG=libs clinfo -l
Thank you for your cooperation.
Hi @kgibala,
Thanks for the hint about the IGC version mismatch.
I checked with:
LD_DEBUG=libs clinfo -l
and found that my system was loading old manually installed IGC libraries from
/usr/local/libinstead of the packaged ones.After moving those old libraries away and running
ldconfig,clinfo -lnow loads the correct libraries from/lib/x86_64-linux-gnu/and works again.So this was most likely a mixed local IGC installation on my machine, not a Compute Runtime package issue.
Thanks for pointing me in the right direction and sorry for the noise & the extra work.
Reacted by Krzysztof Gibala, Aleksandra Nizio and Jaroslaw WarchulskiHi @encodatamHirmer,
I'm glad to hear that you managed to resolve your setup issue.
Hi @heikogleu-dev / @pshirshov / @Stoney49th / @amarevite
Could you please review the previous comment and confirm that you are using the correct libraries, loaded from the proper locations? There is a high possibility that some leftover dependencies on your system may be causing conflicts.
If you are still experiencing this issue, please provide the logs as mentioned previously (#922 (comment)).
Thank you for your collaboration.
- addedStatus: Needs FeedbackWaiting for additional information from reporterWaiting for additional information from reporter
on Jun 10, 2026 Unless there was a recent update to the Fedora packages that fixed the issue, I can confirm that running (sudo)
ldconfigfixed things for me. I will add this step to my Bazzite image build process to prevent the issue in the future. Thanks for helping narrow the issue down!Reacted by Krzysztof Gibala- added a commit that references this issue
on Jun 15, 2026 Follow-up for #922 — persists through CR 26.18, kernel-independent, pure-Level-Zero minimal reproducer
Update on GSD-12696. Three new data points that should make this much
easier to bisect on Intel's side.1. Still present in CR 26.18 (latest)
The
resource_info.cpp:15abort is unchanged from 26.14 through
26.18.38308.1 (current latest, June 2026). Three releases, no fix.
The 26.18 driver prints the abort with an empty__FILE__:Abort was called at 15 line in file:but the stack and behaviour are identical to the 26.14
gmm_helper/resource_info.cpp:15abort reported originally.2. A newer kernel does NOT fix it
The original report was on kernel
7.0.0-15-generic. We are now on
7.0.0-22-generic(xe modulesrcversion ACC2C75180429CFD259EDF2)
and the abort is identical. This weakens the "ABI mismatch with the xe
kernel module" hypothesis from the original report — a newer xe module
did not change anything.GMM userspace:
libigdgmm12 22.10.0.3. Pure-Level-Zero minimal reproducer (no SYCL / MPI-runtime-SYCL needed)
The original reproducer was a full SYCL + MPI application. Here is a
~120-line program that reproduces the abort with **onlylibze_loader- MPI** — no SYCL, no oneMKL, no application framework. Each rank just
doeszeInit → zeDriverGet → zeDeviceGet → zeContextCreate → zeMemAllocDevice:
mpirun -np 2 ./diag-mpi-l0 # ABORT at resource_info.cpp:15 mpirun -np 1 ./diag-mpi-l0 # OK
Characterisation matrix (B70 Pro, BMG-G31, CR 26.18, kernel 7.0.0-22):
Config Result np=1 OK np=2 abort np=8 abort np=8, ranks staggered 200 ms apart abort (timing does not help) np=8, LD_LIBRARY_PATH→ CR 26.05libze_intel_gpu.soall ranks OK Staggering not helping indicates this is deterministic, not a
start-up timing race: the second process to touch the device aborts in
GMM resource-info init regardless of when it arrives.4. Env-var workarounds that do NOT work
Tried the IPC/GMM-related NEO debug keys as environment variables
(plain names) withdiag-mpi-l0 -np 8on 26.18 — all still abort:EnablePidFdOrSocketsForIpc=1/=0ForceIpcSocketFallback=0EnableIpcSocketFallback=0
(Presumably these are not honoured as plain env vars in a release
build.) If there is a supported runtime switch to make GMM
resource-info init multi-process-safe, that would be a great interim
workaround — otherwise the only fix remains pinninglibze-intel-gpu1
to 26.05.Why this matters
Multi-process single-GPU is the normal execution mode for MPI HPC
(CFD, etc.), not an edge case. Any Level-Zero MPI workload on
Battlemage is currently blocked on CR ≥ 26.14 unless the user knows to
downgrade. The pure-L0 reproducer above should let you confirm and
bisect without any of our CFD stack.Happy to run debug builds or
NEO_*-prefixed debug keys if you can
point at the right enable mechanism.- MPI** — no SYCL, no oneMKL, no application framework. Each rank just
Could you please follow the instructions provided by @kgibala in the previous comment?
In particular, we would appreciate it if you could share your system information by providing the output of the following commands:
- lsb_release -a
- uname -a
- sudo dpkg --list | grep -iE "igc|gmm|opencl|libze"
Additionally, please upgrade your Intel Graphics Compiler (IGC) to the recommended version:
- Compute Runtime: https://github.com/intel/compute-runtime/releases/tag/26.22.38646.4
- Intel Graphics Compiler: https://github.com/intel/intel-graphics-compiler/releases/tag/v2.36.3
After the upgrade, please provide the following logs:
- dmesg
- LD_DEBUG=libs clinfo -l
These details will greatly help us better understand the issue and continue the investigation.
Hi @heikogleu-dev,
We’d like to know if this issue is still affecting you. If so, please provide an update or any additional information. Otherwise, we’ll close this issue after 30 days of inactivity. Your feedback is appreciated.
- addedStatus: Marked as StaleNo activity for 30+ days, trending toward closureNo activity for 30+ days, trending toward closure
on Jul 20, 2026 Hi @heikogleu-dev,
We're closing this issue due to inactivity. We haven't received the requested information or updates for an extended period.
If you encounter this or a similar problem in the future, please feel free to create a new issue. When submitting, please follow our Issue Submission Guide to help us understand and address your concern efficiently.
Reacted by Paul S.Great joke
Hi @pshirshov,
Could you please let us know if you are currently encountering this issue? If so, could you check if the steps provided earlier by @kgibala resolve it for you?
If the problem persists, we will gladly reopen this issue. Thank you for your cooperation.
Pre-submission Checklist
GPU Hardware
Intel Arc Pro B70 Pro (BMG-G31, Battlemage Xe2, 32 GB GDDR6 ECC) PCI Device ID: 0xe223
DRI Devices Information
intel_gpu_top -L:
card0 Intel Battlemage (Gen20) pci:vendor=8086,device=E223,card=0
└─renderD129
sycl-ls output:
[level_zero:gpu][level_zero:0] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Graphics [0xe223] 20.2.0 [1.14.37020]
GPU Detailed Information (lspci output)
04:00.0 Display controller: Intel Corporation Battlemage [Arc Pro B70] (rev 08)
Subsystem: Intel Corporation Device 0000
Kernel driver in use: xe
Kernel modules: xe
Driver Version
26.14.37833.4 (failing) 26.05.37020.3-1 (working — rolled back)
Installed GPU Driver Packages
intel-igc-core-2 2.32.7+21184
intel-igc-opencl-2 2.32.7+21184
intel-opencl-icd 26.14.37833.4-0 ← failing version
intel-ocloc 26.14.37833.4-0
libze-intel-gpu1 26.14.37833.4-0
libigdgmm12 22.9.0+ds1-1 (Ubuntu)
Working (rolled back) versions:
intel-opencl-icd 26.05.37020.3-1
libze-intel-gpu1 26.05.37020.3-1
intel-ocloc 26.05.37020.3-1
Driver Installation Details
Installed via Intel release .deb packages downloaded from:
https://github.com/intel/compute-runtime/releases/tag/26.14.37833.4
Command:
sudo dpkg -i intel-opencl-icd_26.14.37833.4-0_amd64.deb
libze-intel-gpu1_26.14.37833.4-0_amd64.deb
intel-ocloc_26.14.37833.4-0_amd64.deb
intel-igc-core-2_2.32.7+21184_amd64.deb
intel-igc-opencl-2_2.32.7+21184_amd64.deb
libigdgmm12_22.9.0_amd64.deb
Previous working version (26.05.37020.3-1) was installed from
Ubuntu 26.04 Universe repository via apt.
Linux Distribution
Other (please specify below)
Other Linux Distribution
Ubuntu 26.04 LTS (Noble Numbat)
Kernel Version & Boot Parameters
uname -r: 7.0.0-15-generic
Loaded GPU modules:
xe 4415488 16
i915 5095424 3
lsmod | grep xe:
xe 4415488 16
Actual Behavior
After upgrading Compute Runtime from 26.05.37020.3-1 to 26.14.37833.4,
all multi-rank MPI SYCL workloads crash. Single-process SYCL works normally.
Two crash patterns observed:
Pre-reboot (immediate after install):
All 8 MPI ranks abort in zeInit — race condition during concurrent
Level Zero initialization:
Abort was called at 15 line in file:
../../neo/shared/source/gmm_helper/resource_info.cpp
[Signal: Aborted (6)] × 8 ranks
[22] libze_loader.so.1(zeInit+0x71)
Post-reboot (after system restart):
Crash shifts to first GPU memory allocation in multi-rank context:
Abort was called at 15 line in file:
../../neo/shared/source/gmm_helper/resource_info.cpp
Stack trace:
[0] gko::experimental::mpi::communicator::all_to_all_v(
shared_ptr, void const*, int const*, int const*,
ompi_datatype_t*, void*, int const*, int const*, ompi_datatype_t*)
[1] communicate_values(...device-pointer...)
[2] update_impl(...)
[3] incompressibleFluid::correctPressure → GKOCG::solve
Note: tested with BOTH a newly rebuilt application library AND
the original pre-update backup — identical crash in both cases.
This confirms the bug is in CR 26.14 / libze, not in the application.
Expected Behavior
Multi-rank MPI SYCL workloads run normally as they did with CR 26.05.
Single-process SYCL already works correctly on CR 26.14 — the bug
is specific to concurrent multi-process GPU initialization.
Reproduction Rate
Always reproduces - 100%
Steps to Reproduce
sudo dpkg -i intel-opencl-icd_26.14.37833.4-0_amd64.deb
libze-intel-gpu1_26.14.37833.4-0_amd64.deb
intel-ocloc_26.14.37833.4-0_amd64.deb
mpirun -np 8 <sycl_application>
Minimum reproducer (if available):
mpirun -np 2 python3 -c "
import dpctl
q = dpctl.SyclQueue('level_zero:0')
print(q.sycl_device)
"
(may reproduce with simpler approach — OpenFOAM+OGL confirmed)
Is this a regression?
Last Known Working Driver Version
26.05.37020.3-1
First Known Failing Driver Version
26.14.37833.4
API Call Logs
Level Zero trace (SYCL_PI_TRACE=1 pre-reboot, abbreviated):
[pi_level_zero] ZE ---> zeInit
[pi_level_zero] ZE ---> zeDriverGet
→ Abort in resource_info.cpp before zeInit completes in all ranks
strace Logs
N/A — crash occurs before strace-traceable syscalls
System Logs / dmesg Output
[ xxx.xxxxxx] xe 0000:04:00.0: [drm] VM worker error: -12
[ xxx.xxxxxx] xe 0000:04:00.0: [drm] Aborting with incomplete GM operations
Backtrace (if crash or hang occurred)
Full backtrace from crash output (8 ranks, showing rank 0):
[Tavea-Station:XXXXX] Signal: Aborted (6)
[0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45cb0)
[1] /usr/lib/x86_64-linux-gnu/libc.so.6(abort+0xd3)
[11] /lib/x86_64-linux-gnu/libze_intel_gpu.so.1(+0x6a83d6)
[12] /lib/x86_64-linux-gnu/libze_intel_gpu.so.1(+0x6bd5ea)
[13] /lib/x86_64-linux-gnu/libze_intel_gpu.so.1(+0x7f4414)
[18] /lib/x86_64-linux-gnu/libze_intel_gpu.so.1(+0x1c427b)
[19] /lib/x86_64-linux-gnu/libze_loader.so.1(+0x4c9f7)
[22] /lib/x86_64-linux-gnu/libze_loader.so.1(zeInit+0x71)
[23] /opt/intel/oneapi/compiler/2026.0/lib/libur_adapter_level_zero.so.0(+0x1b4803)
[29] Abort was called at 15 line in file:
../../neo/shared/source/gmm_helper/resource_info.cpp
Source Code / Reproducer
OpenFOAM 13 + OGL (OpenFOAM Ginkgo Layer) + Ginkgo 1.10 SYCL backend.
Source: https://github.com/hpsim/OGL
Full reproduction environment documented at:
https://github.com/heikogleu-dev/Openfoam13---GPU-Offloading-Intel-B70-Pro/blob/main/findings/13_stack_update_zeinit_race.md
Command Line / Application Details
mpirun -np 8 foamRun -parallel -solver incompressibleFluid
(OpenFOAM CFD solver, 34M cell mesh, 8 MPI ranks on single GPU)
Environment:
export ONEAPI_DEVICE_SELECTOR=level_zero:0
source /opt/intel/oneapi/setvars.sh
oneAPI Version (if applicable)
Intel oneAPI 2026.0.0
icpx --version: Intel(R) oneAPI DPC++/C++ Compiler 2026.0.0
Screenshots / Video
No response
Additional Notes
Key observations:
Single-process SYCL works fine on CR 26.14:
IGC 2.32.7 is NOT the cause:
libigdgmm12 is NOT the cause:
Suspected root cause:
ABI mismatch between libze_intel_gpu.so.1 from CR 26.14 and the
xe kernel module version shipped with Ubuntu 26.04 kernel 7.0.0-15.
The concurrent multi-process zeInit exposes a race condition or
API incompatibility not triggered by single-process initialization.
Workaround:
apt-mark hold intel-opencl-icd libze-intel-gpu1 intel-ocloc
(pin at CR 26.05.37020.3-1 from Ubuntu Universe)