Skip to content

[GSD-12696] [Xe2/BMG-G31] multi-rank MPI + Level Zero: resource_info.cpp:15 abort after upgrading CR 26.05 → 26.14 (Ubuntu 26.04, xe driver) #922

Description

@heikogleu-dev

Pre-submission Checklist

  • I am using the latest GPU driver version (releases)
  • I have searched for similar issues and found none

GPU Hardware

Intel Arc Pro B70 Pro (BMG-G31, Battlemage Xe2, 32 GB GDDR6 ECC) PCI Device ID: 0xe223

DRI Devices Information

intel_gpu_top -L:
card0 Intel Battlemage (Gen20) pci:vendor=8086,device=E223,card=0
└─renderD129

sycl-ls output:
[level_zero:gpu][level_zero:0] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Graphics [0xe223] 20.2.0 [1.14.37020]

GPU Detailed Information (lspci output)

04:00.0 Display controller: Intel Corporation Battlemage [Arc Pro B70] (rev 08)
Subsystem: Intel Corporation Device 0000
Kernel driver in use: xe
Kernel modules: xe

Driver Version

26.14.37833.4 (failing) 26.05.37020.3-1 (working — rolled back)

Installed GPU Driver Packages

intel-igc-core-2 2.32.7+21184
intel-igc-opencl-2 2.32.7+21184
intel-opencl-icd 26.14.37833.4-0 ← failing version
intel-ocloc 26.14.37833.4-0
libze-intel-gpu1 26.14.37833.4-0
libigdgmm12 22.9.0+ds1-1 (Ubuntu)

Working (rolled back) versions:

intel-opencl-icd 26.05.37020.3-1
libze-intel-gpu1 26.05.37020.3-1
intel-ocloc 26.05.37020.3-1

Driver Installation Details

Installed via Intel release .deb packages downloaded from:
https://github.com/intel/compute-runtime/releases/tag/26.14.37833.4

Command:
sudo dpkg -i intel-opencl-icd_26.14.37833.4-0_amd64.deb
libze-intel-gpu1_26.14.37833.4-0_amd64.deb
intel-ocloc_26.14.37833.4-0_amd64.deb
intel-igc-core-2_2.32.7+21184_amd64.deb
intel-igc-opencl-2_2.32.7+21184_amd64.deb
libigdgmm12_22.9.0_amd64.deb

Previous working version (26.05.37020.3-1) was installed from
Ubuntu 26.04 Universe repository via apt.

Linux Distribution

Other (please specify below)

Other Linux Distribution

Ubuntu 26.04 LTS (Noble Numbat)

Kernel Version & Boot Parameters

uname -r: 7.0.0-15-generic

Loaded GPU modules:
xe 4415488 16
i915 5095424 3

lsmod | grep xe:
xe 4415488 16

Actual Behavior

After upgrading Compute Runtime from 26.05.37020.3-1 to 26.14.37833.4,
all multi-rank MPI SYCL workloads crash. Single-process SYCL works normally.

Two crash patterns observed:

  1. Pre-reboot (immediate after install):
    All 8 MPI ranks abort in zeInit — race condition during concurrent
    Level Zero initialization:

    Abort was called at 15 line in file:
    ../../neo/shared/source/gmm_helper/resource_info.cpp
    [Signal: Aborted (6)] × 8 ranks
    [22] libze_loader.so.1(zeInit+0x71)

  2. Post-reboot (after system restart):
    Crash shifts to first GPU memory allocation in multi-rank context:

    Abort was called at 15 line in file:
    ../../neo/shared/source/gmm_helper/resource_info.cpp

    Stack trace:
    [0] gko::experimental::mpi::communicator::all_to_all_v(
    shared_ptr, void const*, int const*, int const*,
    ompi_datatype_t*, void*, int const*, int const*, ompi_datatype_t*)
    [1] communicate_values(...device-pointer...)
    [2] update_impl(...)
    [3] incompressibleFluid::correctPressure → GKOCG::solve

Note: tested with BOTH a newly rebuilt application library AND
the original pre-update backup — identical crash in both cases.
This confirms the bug is in CR 26.14 / libze, not in the application.

Expected Behavior

Multi-rank MPI SYCL workloads run normally as they did with CR 26.05.
Single-process SYCL already works correctly on CR 26.14 — the bug
is specific to concurrent multi-process GPU initialization.

Reproduction Rate

Always reproduces - 100%

Steps to Reproduce

  1. Start with working CR 26.05.37020.3-1 on Ubuntu 26.04, xe driver
  2. Install CR 26.14.37833.4 via Intel release .deb:
    sudo dpkg -i intel-opencl-icd_26.14.37833.4-0_amd64.deb
    libze-intel-gpu1_26.14.37833.4-0_amd64.deb
    intel-ocloc_26.14.37833.4-0_amd64.deb
  3. Run any multi-rank SYCL workload:
    mpirun -np 8 <sycl_application>
  4. All ranks abort in resource_info.cpp:15

Minimum reproducer (if available):
mpirun -np 2 python3 -c "
import dpctl
q = dpctl.SyclQueue('level_zero:0')
print(q.sycl_device)
"
(may reproduce with simpler approach — OpenFOAM+OGL confirmed)

Is this a regression?

  • Yes, this is a regression - functionality that previously worked is now broken

Last Known Working Driver Version

26.05.37020.3-1

First Known Failing Driver Version

26.14.37833.4

API Call Logs

Level Zero trace (SYCL_PI_TRACE=1 pre-reboot, abbreviated):
[pi_level_zero] ZE ---> zeInit
[pi_level_zero] ZE ---> zeDriverGet
→ Abort in resource_info.cpp before zeInit completes in all ranks

strace Logs

N/A — crash occurs before strace-traceable syscalls

System Logs / dmesg Output

[ xxx.xxxxxx] xe 0000:04:00.0: [drm] VM worker error: -12
[ xxx.xxxxxx] xe 0000:04:00.0: [drm] Aborting with incomplete GM operations

Backtrace (if crash or hang occurred)

Full backtrace from crash output (8 ranks, showing rank 0):

[Tavea-Station:XXXXX] Signal: Aborted (6)
[0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45cb0)
[1] /usr/lib/x86_64-linux-gnu/libc.so.6(abort+0xd3)
[11] /lib/x86_64-linux-gnu/libze_intel_gpu.so.1(+0x6a83d6)
[12] /lib/x86_64-linux-gnu/libze_intel_gpu.so.1(+0x6bd5ea)
[13] /lib/x86_64-linux-gnu/libze_intel_gpu.so.1(+0x7f4414)
[18] /lib/x86_64-linux-gnu/libze_intel_gpu.so.1(+0x1c427b)
[19] /lib/x86_64-linux-gnu/libze_loader.so.1(+0x4c9f7)
[22] /lib/x86_64-linux-gnu/libze_loader.so.1(zeInit+0x71)
[23] /opt/intel/oneapi/compiler/2026.0/lib/libur_adapter_level_zero.so.0(+0x1b4803)
[29] Abort was called at 15 line in file:
../../neo/shared/source/gmm_helper/resource_info.cpp

Source Code / Reproducer

OpenFOAM 13 + OGL (OpenFOAM Ginkgo Layer) + Ginkgo 1.10 SYCL backend.
Source: https://github.com/hpsim/OGL

Full reproduction environment documented at:
https://github.com/heikogleu-dev/Openfoam13---GPU-Offloading-Intel-B70-Pro/blob/main/findings/13_stack_update_zeinit_race.md

Command Line / Application Details

mpirun -np 8 foamRun -parallel -solver incompressibleFluid
(OpenFOAM CFD solver, 34M cell mesh, 8 MPI ranks on single GPU)

Environment:
export ONEAPI_DEVICE_SELECTOR=level_zero:0
source /opt/intel/oneapi/setvars.sh

oneAPI Version (if applicable)

Intel oneAPI 2026.0.0
icpx --version: Intel(R) oneAPI DPC++/C++ Compiler 2026.0.0

Screenshots / Video

No response

Additional Notes

Key observations:

  1. Single-process SYCL works fine on CR 26.14:

    • sycl-ls detects B70 Pro normally
    • FP64 benchmark: 1346 GFLOPS (94% of spec) ✅
    • Regression is STRICTLY multi-process
  2. IGC 2.32.7 is NOT the cause:

    • Tested CR 26.05 + IGC 2.32.7 → works correctly
    • Only CR 26.14 + libze 26.14 causes the crash
  3. libigdgmm12 is NOT the cause:

    • Tested both Intel's 22.9.0 and Ubuntu's 22.9.0+ds1-1 → same crash
  4. Suspected root cause:
    ABI mismatch between libze_intel_gpu.so.1 from CR 26.14 and the
    xe kernel module version shipped with Ubuntu 26.04 kernel 7.0.0-15.
    The concurrent multi-process zeInit exposes a race condition or
    API incompatibility not triggered by single-process initialization.

  5. Workaround:
    apt-mark hold intel-opencl-icd libze-intel-gpu1 intel-ocloc
    (pin at CR 26.05.37020.3-1 from Ubuntu Universe)

Activity

  1. added
    Type: BugGeneral bug report, unexpected behavior or crash
    OS: LinuxIssue specific to Linux distributions (Ubuntu, Fedora, RHEL, etc.)
    on May 4, 2026
  2. jchabere commented on May 13, 2026

    @jchabere
    Contributor

    Hi @heikogleu-dev

    Thank you for your contribution.

    We tried to reproduce the crash with the following reproducer:

    # run.py
    import dpctl
    q = dpctl.SyclQueue('level_zero:0')
    print(q.sycl_device) 
    
    mpirun -np 2 python run.py
    

    However, on our side the problem is not reproducible on BMG-G31 e223 + 7.0.0-14-generic + Ubuntu 26.04 + https://github.com/intel/compute-runtime/releases/tag/26.14.37833.4

    We could try another reproducer if you provide instructions on how to set up/build the specific benchmark:

    mpirun -np 8 foamRun -parallel -solver incompressibleFluid
    
  3. pshirshov commented on May 15, 2026

    @pshirshov

    Adding a reproduction with single-process failure on the same hardware (Arc Pro B70 / BMG G31, PCI 0x8086:0xe223), so this may be a broader code path than just the multi-rank MPI race.

    Environment

    OS NixOS 26.05
    Kernel 7.0.3, xe driver
    GPU Arc Pro B70 (BMG G31, 32 GB ECC)
    PCI 0000:03:00.0 8086:e223 (subsys 8086:1701)

    Reproducer (any of these, all abort identically — single Python process, no MPI):

    # 1. torch 2.10.0+xpu (vendored by github.com/utensils/comfyui-nix)
    # 2. torch 2.13.0.dev20260515+xpu (pytorch.org nightly XPU index)
    # 3. torch 2.8.0+xpu + intel-extension-for-pytorch 2.8.10+xpu
    
    python -c "import torch; print(torch.xpu.device_count())"

    Aborts at:

    Created new BO with GEM_USERPTR, handle: BO-0
    Abort was called at 15 line in file:
    ../../neo/shared/source/gmm_helper/resource_info.cpp
    

    NEO + gmmlib versions I tried (same abort in every cell, including with Intel's own prebuilt .deb binaries unpacked from your releases):

    NEO gmmlib Result
    26.05.37020.3 22.9.0 (the deb pairs you ship) abort
    26.05.37020.3 22.10.0 abort
    26.14.37833.4 22.9.0 abort
    26.14.37833.4 22.10.0 abort
    26.18.38308.1 22.10.0 abort

    NEO debug knobs that did not change behaviour: EnableLocalMemory=0, ForceDeviceMemoryAlloc=0, EnableDirectSubmission=0, EnableHostUsmAllocationPool=0, EnableDeviceUsmAllocationPool=0,
    SetCommandStreamReceiver=1, EnableXeCorePool=1.

    gdb backtrace at the abort:

    #3  NEO::abortUnrecoverable                            libze_intel_gpu.so
    #4  NEO::GmmResourceInfo::create                       libze_intel_gpu.so
    #5  NEO::Gmm::Gmm(GmmHelper*, void*, size, align,
                       type, StorageInfo&, GmmRequirements&) libze_intel_gpu.so
    #6  NEO::DrmMemoryManager::makeGmmIfSingleHandle       libze_intel_gpu.so
    #7  NEO::DrmMemoryManager::allocateGraphicsMemoryInDevicePool
                                                           libze_intel_gpu.so
    #8  NEO::MemoryManager::allocateGraphicsMemoryInPreferredPool
                                                           libze_intel_gpu.so
    #9  NEO::CommandStreamReceiver::initializeResources    libze_intel_gpu.so
    #10 NEO::Device::createEngines                         libze_intel_gpu.so
    #11 NEO::DeviceFactory::{lambda}                       libze_intel_gpu.so
    #12 L0::Driver::initialize                             libze_intel_gpu.so
    #13 std::call_once<L0::Driver::driverInit()...>        libze_intel_gpu.so
    #16 L0::zeInit                                         libze_intel_gpu.so
    #20 zeInit                                             libze_loader.so.1
    #21 ze_lib::context_t::Init                            libur_adapter_level_zero.so.0
    […]
    #34 c10::xpu::initGlobalDevicePoolState                libc10_xpu.so
    #35 c10::xpu::device_count                             libc10_xpu.so
    

    So the failing allocation is the very first one CommandStreamReceiver::initializeResources() requests during eager device init in zeInit. The kernel-side ioctl succeeds (Created new BO with GEM_USERPTR, handle: BO-0), then makeGmmIfSingleHandle builds an AllocationData + GmmRequirementnto GmmResourceInfo::create, and gmmlib returns a GMM_RESOURCE_INFO*whose underlyinghandle is zero — so theUNRECOVERABLE_IF(resourceInfo->peekHandle() == 0)onresource_info.cpp:15` fires.

    What works on the exact same system, with the same NEO/gmmlib/kernel (i.e. not a hardware/firmware issue):

    • clinfo enumerates the B70 fully, including cl_intel_subgroup_matrix_multiply_accumulate_tf32 and cl_intel_unified_shared_memory.
    • vulkaninfo reports Intel(R) Graphics (BMG G31) with apiVersion 1.4.335, driver 26.0
    • llama.cpp-SYCL via libur_adapter_opencl.so (ONEAPI_DEVICE_SELECTOR=opencl:gpu) runs models on the card to completion (~16 t/s on Qwen-32B Q4_K_M).
    • ONNX Runtime with OpenVINO Execution Provider (libopenvino_intel_gpu_plugin.so) driveace-rec workloads continuously.

    So gmmlib itself is functional on this card via the OpenCL ICD path; it only fails whcts the request during NEO's L0 driver init.

    Hypothesis (not a fix, just where to look): on xe-kmd the GEM_USERPTR ioctl still sadata isn't shaped the same as on i915 — and DrmMemoryManager::makeGmmIfSingleHandlesynthesizes the GmmRequirements/StorageInfo triple under i915-era assumptions. The OpenCL path doesn't pass through the same makeGmmIfSingleHandle early-CSR-init code, which is why every other consumer of the same gmmlib works.

    Happy to test patches or add more instrumentation — let me know what would help.

  4. ola2308 commented on May 18, 2026

    @ola2308
    Contributor

    Hi @heikogleu-dev and @pshirshov I was able to reproduce the issue on mentioned driver, but I can see that the crash is no longer reproducible on the latest public release: 26.18.38308.1 could you please verify whether upgrading to 26.18.38308.1 resolves the issue?

  5. pshirshov commented on May 18, 2026

    @pshirshov

    Is L0 supposed to work now?

  6. Stoney49th commented on May 18, 2026

    @Stoney49th

    No its still broken, also 26.18 did not fix it with the 7.0.3 kernel. https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8022

  7. changed the title [-][Xe2/BMG-G31] multi-rank MPI + Level Zero: resource_info.cpp:15 abort after upgrading CR 26.05 → 26.14 (Ubuntu 26.04, xe driver)[/-] [+][GSD-12696] [Xe2/BMG-G31] multi-rank MPI + Level Zero: resource_info.cpp:15 abort after upgrading CR 26.05 → 26.14 (Ubuntu 26.04, xe driver)[/+] on May 27, 2026
  8. amarevite commented on Jun 2, 2026

    @amarevite

    I believe I am running into this issue with Blender on Fedora 44 (a custom Bazzite build, specifically). When I select the oneAPI option in Blender, I get a crash:

    Abort was called at 15 line in file:
    /builddir/build/BUILD/intel-compute-runtime-26.18.38308.1-build/compute-runtime-26.18.38308.1/shared/source/gmm_helper/resource_info.cpp
    fish: Job 1, './blender --factory-startup' terminated by signal SIGABRT (Abort)
    

    I have these dependencies installed, which I know to have previously worked in my setup, though I don't know which specific version I last used that worked:

    ➜ dnf5 list --installed "intel-level-zero*" "intel-compute-runtime*" "oneapi*"
    Installed packages (available for reinstall, available for upgrade)
    intel-compute-runtime.x86_64           26.18.38308.1-2.fc44 updates
    intel-level-zero.x86_64                26.18.38308.1-2.fc44 updates
    intel-level-zero-gpu-raytracing.x86_64 1.2.3-1.fc44         updates
    oneapi-level-zero.x86_64               1.28.6-1.fc44        updates
    
  9. encodatamHirmer commented on Jun 2, 2026

    @encodatamHirmer

    I can reproduce what looks like the same issue on Ubuntu 25.10 (2x ArcPro B60) with Intel GPU compute packages from the kobuk-team Intel graphics PPA.

    After upgrading to:

    intel-opencl-icd  26.18.38308.1-1~25.10~ppa1
    libze-intel-gpu1  26.18.38308.1-1~25.10~ppa1
    intel-ocloc       26.18.38308.1-1~25.10~ppa1
    libigdgmm12       22.10.0-1~25.10~ppa1
    intel-igc-core-2  2.28.4
    intel-igc-opencl-2 2.28.4
    

    both clinfo and sycl-ls abort with:

    Abort was called at 15 line in file:
    ./shared/source/gmm_helper/resource_info.cpp
    Aborted (core dumped)
    

    This happens before running any higher-level workload such as PyTorch or vLLM, so it appears to be in the OpenCL / Level Zero / GMM stack itself.

    /etc/OpenCL/vendors/ contains:

    intel.icd
    intel64.icd
    

    I also tested after initializing oneAPI with:

    source /opt/intel/oneapi/setvars.sh

    but sycl-ls still aborts with the same resource_info.cpp:15 message.

    For now I will probably roll back / hold the affected packages, but I wanted to add another confirmation that the issue is still reproducible with 26.18.38308.1.

  10. kgibala commented on Jun 10, 2026

    @kgibala
    Contributor

    Hi @encodatamHirmer,

    We are currently attempting to reproduce the issue you reported, but so far we have been unable to replicate it - everything is functioning correctly in our environment.

    lsb_release -a
    No LSB modules are available.
    Distributor ID: Ubuntu
    Description:    Ubuntu 25.10
    Release:        25.10
    
    uname -a
    ... 6.17.0-5
    
    sudo dpkg --list | grep -iE "igc|gmm|opencl|libze"
    ii  clinfo                                                      3.0.25.02.14-1                             amd64        Query OpenCL system information
    ii  intel-opencl-icd                                            26.18.38308.1-1~25.10~ppa1                 amd64        Intel graphics compute runtime for OpenCL
    ii  libigc2                                                     2.34.4-1260~25.10                          amd64        Core libraries for Intel(R) Graphics Compiler for OpenCL(TM)
    ii  libigdfcl2                                                  2.34.4-1260~25.10                          amd64        OpenCL library for Intel(R) Graphics Compiler for OpenCL(TM)
    ii  libigdgmm12:amd64                                           22.10.0-1~25.10~ppa1                       amd64        Intel Graphics Memory Management Library -- shared library
    ii  libze-intel-gpu1                                            26.18.38308.1-1~25.10~ppa1                 amd64        Intel oneAPI L0 support implementation for Intel GPUs -- shared library
    ii  libze1:amd64                                                1.28.2-1~25.10~ppa1                        amd64        oneAPI Level Zero -- share libraries
    ii  ocl-icd-libopencl1:amd64                                    2.3.3-1    
    

    To help us investigate further, could you please:

    1. Share your system information Provide the output of the following commands:
    • lsb_release -a
    • uname -a
    • sudo dpkg --list | grep -iE "igc|gmm|opencl|libze"
    1. Upgrade your IGC (2.28.4) (Intel Graphics Compiler) to the recommended version:
    1. Provide the following logs after the upgrade:
    • dmesg
    • LD_DEBUG=libs clinfo -l

    Thank you for your cooperation.

  11. encodatamHirmer commented on Jun 10, 2026

    @encodatamHirmer

    Hi @kgibala,

    Thanks for the hint about the IGC version mismatch.

    I checked with:

    LD_DEBUG=libs clinfo -l

    and found that my system was loading old manually installed IGC libraries from /usr/local/lib instead of the packaged ones.

    After moving those old libraries away and running ldconfig, clinfo -l now loads the correct libraries from /lib/x86_64-linux-gnu/ and works again.

    So this was most likely a mixed local IGC installation on my machine, not a Compute Runtime package issue.

    Thanks for pointing me in the right direction and sorry for the noise & the extra work.

  12. kgibala commented on Jun 10, 2026

    @kgibala
    Contributor

    Hi @encodatamHirmer,

    I'm glad to hear that you managed to resolve your setup issue.


    Hi @heikogleu-dev / @pshirshov / @Stoney49th / @amarevite

    Could you please review the previous comment and confirm that you are using the correct libraries, loaded from the proper locations? There is a high possibility that some leftover dependencies on your system may be causing conflicts.

    If you are still experiencing this issue, please provide the logs as mentioned previously (#922 (comment)).

    Thank you for your collaboration.

  13. amarevite commented on Jun 10, 2026

    @amarevite

    Unless there was a recent update to the Fedora packages that fixed the issue, I can confirm that running (sudo) ldconfig fixed things for me. I will add this step to my Bazzite image build process to prevent the issue in the future. Thanks for helping narrow the issue down!

  14. heikogleu-dev commented on Jun 17, 2026

    @heikogleu-dev
    Author

    Follow-up for #922 — persists through CR 26.18, kernel-independent, pure-Level-Zero minimal reproducer

    Update on GSD-12696. Three new data points that should make this much
    easier to bisect on Intel's side.

    1. Still present in CR 26.18 (latest)

    The resource_info.cpp:15 abort is unchanged from 26.14 through
    26.18.38308.1 (current latest, June 2026). Three releases, no fix.
    The 26.18 driver prints the abort with an empty __FILE__:

    Abort was called at 15 line in file:
    

    but the stack and behaviour are identical to the 26.14
    gmm_helper/resource_info.cpp:15 abort reported originally.

    2. A newer kernel does NOT fix it

    The original report was on kernel 7.0.0-15-generic. We are now on
    7.0.0-22-generic (xe module srcversion ACC2C75180429CFD259EDF2)
    and the abort is identical. This weakens the "ABI mismatch with the xe
    kernel module" hypothesis from the original report — a newer xe module
    did not change anything.

    GMM userspace: libigdgmm12 22.10.0.

    3. Pure-Level-Zero minimal reproducer (no SYCL / MPI-runtime-SYCL needed)

    The original reproducer was a full SYCL + MPI application. Here is a
    ~120-line program that reproduces the abort with **only libze_loader

    • MPI** — no SYCL, no oneMKL, no application framework. Each rank just
      does zeInit → zeDriverGet → zeDeviceGet → zeContextCreate → zeMemAllocDevice:
    mpirun -np 2 ./diag-mpi-l0     # ABORT at resource_info.cpp:15
    mpirun -np 1 ./diag-mpi-l0     # OK

    Full source: https://github.com/heikogleu-dev/Openfoam13---GPU-Offloading-Intel-B70-Pro/blob/main/findings/code/gpu-diag/diag-mpi-l0.cpp

    Characterisation matrix (B70 Pro, BMG-G31, CR 26.18, kernel 7.0.0-22):

    Config Result
    np=1 OK
    np=2 abort
    np=8 abort
    np=8, ranks staggered 200 ms apart abort (timing does not help)
    np=8, LD_LIBRARY_PATH → CR 26.05 libze_intel_gpu.so all ranks OK

    Staggering not helping indicates this is deterministic, not a
    start-up timing race: the second process to touch the device aborts in
    GMM resource-info init regardless of when it arrives.

    4. Env-var workarounds that do NOT work

    Tried the IPC/GMM-related NEO debug keys as environment variables
    (plain names) with diag-mpi-l0 -np 8 on 26.18 — all still abort:

    • EnablePidFdOrSocketsForIpc=1 / =0
    • ForceIpcSocketFallback=0
    • EnableIpcSocketFallback=0

    (Presumably these are not honoured as plain env vars in a release
    build.) If there is a supported runtime switch to make GMM
    resource-info init multi-process-safe, that would be a great interim
    workaround — otherwise the only fix remains pinning libze-intel-gpu1
    to 26.05.

    Why this matters

    Multi-process single-GPU is the normal execution mode for MPI HPC
    (CFD, etc.), not an edge case. Any Level-Zero MPI workload on
    Battlemage is currently blocked on CR ≥ 26.14 unless the user knows to
    downgrade. The pure-L0 reproducer above should let you confirm and
    bisect without any of our CFD stack.

    Happy to run debug builds or NEO_*-prefixed debug keys if you can
    point at the right enable mechanism.

  15. ola2308 commented on Jun 18, 2026

    @ola2308
    Contributor

    Hi @heikogleu-dev

    Could you please follow the instructions provided by @kgibala in the previous comment?

    In particular, we would appreciate it if you could share your system information by providing the output of the following commands:

    • lsb_release -a
    • uname -a
    • sudo dpkg --list | grep -iE "igc|gmm|opencl|libze"

    Additionally, please upgrade your Intel Graphics Compiler (IGC) to the recommended version:

    After the upgrade, please provide the following logs:

    • dmesg
    • LD_DEBUG=libs clinfo -l

    These details will greatly help us better understand the issue and continue the investigation.

  16. jwarchul commented on Jul 20, 2026

    @jwarchul
    Contributor

    Hi @heikogleu-dev,

    We’d like to know if this issue is still affecting you. If so, please provide an update or any additional information. Otherwise, we’ll close this issue after 30 days of inactivity. Your feedback is appreciated.

  17. jwarchul commented on Aug 26, 2026

    @jwarchul
    Contributor

    Hi @heikogleu-dev,

    We're closing this issue due to inactivity. We haven't received the requested information or updates for an extended period.

    If you encounter this or a similar problem in the future, please feel free to create a new issue. When submitting, please follow our Issue Submission Guide to help us understand and address your concern efficiently.

  18. pshirshov commented on Aug 26, 2026

    @pshirshov

    Great joke

  19. jwarchul commented on Aug 27, 2026

    @jwarchul
    Contributor

    Hi @pshirshov,

    Could you please let us know if you are currently encountering this issue? If so, could you check if the steps provided earlier by @kgibala resolve it for you?

    If the problem persists, we will gladly reopen this issue. Thank you for your cooperation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    OS: LinuxIssue specific to Linux distributions (Ubuntu, Fedora, RHEL, etc.)Status: Marked as StaleNo activity for 30+ days, trending toward closureStatus: Needs FeedbackWaiting for additional information from reporterType: BugGeneral bug report, unexpected behavior or crash

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions