Skip to content

[.NET main] ubuntu-latest build job killed without log output on every push since 2026-06-18 #1430

Description

@tamirdresher

Summary

The .NET main workflow has been failing on every push to main since 2026-06-18 (8+ consecutive failures). Always the same shape: build (windows-latest) passes, build (ubuntu-latest) fails in ~3–4 minutes with no Build step log file persisted — classic signature of the runner being killed before it can flush.

Evidence

Run Date Commit windows ubuntu Build log captured?
28141532703 2026-06-25 d8c5674 (Squad #1394) ✅ 13m11s ❌ 3m41s No — only Set-up/Checkout/Restore steps in the archive
28005035836 2026-06-23 typescript-apphost-backfill merge ✅ ❌ No
28004682918 2026-06-23 SQLite audit suppress ✅ ❌ No
28003556555 2026-06-23 Bitwarden SM #1329 ✅ ❌ No
27738932971 2026-06-18 XML comments #1422 ✅ ❌ No
27731598150 2026-06-18 Harden tests #1415 ✅ ❌ No
27729982356 2026-06-18 SeaweedFS #1349 ✅ ❌ No
27727583582 2026-06-18 SurrealDB v3 #1294 ✅ ✅ — last green run
27726664815 2026-06-17 dbx #1391 ✅ 8m30s ❌ 4m7s No — first failure

Likely cause

The breakpoint (#1391 dbx) lands right when the solution crossed some size threshold. Several new hosting integrations have landed since (dbx, SeaweedFS, Bitwarden, Squad), each adding 3–6 new csproj.

  • GitHub-hosted ubuntu-latest runners have ~14 GB free disk; windows-latest has ~80 GB.
  • A full dotnet build -c Release of the whole toolkit now produces a LOT of bin/ + obj/ output across 100+ projects.
  • The "Build step has no log file at all" signature is classic runner-process killed without flushing — strongly suggests disk exhaustion (or OOM).
  • The matching test job Hosting.Java.Tests-ubuntu-latest also timed out at exactly 1h on the first failure (run 27726664815) — additional symptom of the runner being unhealthy.

What dotnet-ci.yml (per-PR) does differently

PRs pass CI because dotnet-ci.yml only builds affected projects (using the path-based filter you set up), not the whole solution. dotnet-main.yml does a full dotnet restore + dotnet build on every push, which is what's blowing past the disk/memory budget.

Suggested fixes (any one should be enough)

  1. Add disk cleanup step — easiest:
    - uses: jlumbroso/free-disk-space@main
    This reclaims 30+ GB on ubuntu-latest by removing Android SDK, GHC, .NET preinstalls we don't need, etc.

  2. Shard the build — split dotnet build across N matrix shards by Directory.Solution.props slice.

  3. Use --no-incremental + clean between project builds — slower but bounded memory.

  4. Switch build (ubuntu-latest) to ubuntu-latest-large or ubuntu-22.04-large — paid runner with more disk.

Happy to send a PR for option 1 if you want — it's the lowest-risk and matches what other large .NET repos (including dotnet/aspire itself) do.

cc @aaronpowell — flagging because main is currently red and v13.5.0 can't be cut until this is fixed (release workflow runs run-tests → package → sign → publish-nuget).

Activity

  1. jmezach commented on Jul 6, 2026

    @jmezach
    Contributor

    I am wondering if it is going to be sustainable to have a single huge pipeline that builds, tests and publishes all the integrations all at the same time. Wouldn't it make sense to have separate pipelines for each integration? I realise this will take some effort, but I feel like it is worth the effort.

  2. Odonno commented on Jul 6, 2026

    @Odonno
    Contributor

    I am wondering if it is going to be sustainable to have a single huge pipeline that builds, tests and publishes all the integrations all at the same time. Wouldn't it make sense to have separate pipelines for each integration? I realise this will take some effort, but I feel like it is worth the effort.

    Yep, probably just publish 1 artifact (nupack) per job, and then have a single job dependent of all the other jobs to publish the packages.

  3. aaronpowell commented on Jul 9, 2026

    @aaronpowell
    Member

    There's a collection of unique problems that we have with the monorepo design that we have.

    At the time of writing, we have 260 projects in the solution, so that's not even covering the TypeScript app hosts and the non-.NET sample applications. This is, simply put, a huge project - even though there really isn't that much code in there.

    Because of this, it's entirely impractical to run the tests in a single agent, that why there's the matrix design for running tests (if anyone wants to know more on it check out the blog I wrote) - this radically reduces the resources required for doing test runs, which is what is most resource intensive.

    We also have to do some tricks with Docker and caching - we hit rate limits with Docker a while back which we have worked around using ACR for our CI builds (#601) that caches Docker images for us so we don't hit the Docker registry directly.

    There's minimal caching setup on things like NPM, NuGet, etc. that should improve the performance of setting up actions, that happens in https://github.com/CommunityToolkit/Aspire/blob/main/.github/actions/setup-runtimes-caching/action.yml and I'm pretty sure it's not very ruggedised so doesn't cache as aggressively as it should.

    I made the decision when I was first setting up the CI that "push to main" (and release) would bump the version of every NuGet package, not just the one(s) that had changes in the diff since last time. This does mean that we push a lot of NuGet packages each time, often times with packages that have no changes since their last release other than dependencies. This reduces the overhead on the maintainers - don't have to be pickup on what is to be published, but at the code of a really aggressive publishing approach.

    Recently we had a publishing failure because the NuGet API key was expired. I've changed the design to use the new Trusted Publishing on NuGet (you can read how I did that here), which means that we no longer use a long-lived API token, instead CI requests one on demand. This is a lot more secure and means that we don't have to worry about expiring tokens.

    Lastly, I added some auto retry logic into the pipeline to handle failing tests. Because of the volume of runners we spin up to do the tests (it's like, 120+ now) I think we might hit some caps that cause tests to just randomly fail at an infra level, Docker not starting is a common one. The retry logic is pretty solid but if we're hitting cases where it should auto-rerun, let's tweak the rerun workflow to cover more cases.

    Hopefully that gives some overall context of the problem we face with this project, and the things that have been added along the way to attempt to address the issues. The most common failures we see on the main branch now seem to be from a handful of integrations that "randomly fail" and if re-run, the tests pass. The problem is, they appear as tests themselves that are failing, not infrastructure around the tests, which makes it hard to identify the failures separately from legitimate test failures. Maybe we should have just a default "retry N times on failure" rather than attempting to do some dynamic detection of failure conditions?

    I also know that most of these solutions have just come from myself doing things. Generally that's just been because I'll see a problem and be like "well, I'll throw Copilot at it" (or, back in the day, just code up a solution 🫨), but if people have ideas, raise them for discussion or prototype them out, and we can look at their viability. Things like larger runners (paid) runners are a harder thing to tackle because someone has to pay for them, but if it's the only possible solution, then it's something I'll have to look into.

  4. jmezach commented on Jul 9, 2026

    @jmezach
    Contributor

    @aaronpowell Thanks for the extensive write-up and explanation of the reasoning behind most of these choices.

    One thing that stands out to me is the decision that "push to main" (and release) bumps the version of every NuGet package. While I realise that makes things easier from the perspective of not having to figure out what has actually changed it does mean we push a lot of packages with no actual changes which could be confusing from the users perspective.

    I wonder if we could take some inspiration from the Backstage Community Plugins repository (https://github.com/backstage/community-plugins/). They've created workspaces, which are folders with a bunch of files that are logically related. Each workspace has its own isolated release, so only those with updates are pushed.

    I realise this is an entirely different ecosystem (C# and NuGet vs. TypeScript and npm), but I do think it works fairly well there. It avoids the problem of pushing too many packages, while also keeping CI relatively small.

  5. aaronpowell commented on Jul 10, 2026

    @aaronpowell
    Member

    I wonder if we could take some inspiration from the Backstage Community Plugins repository (https://github.com/backstage/community-plugins/). They've created workspaces, which are folders with a bunch of files that are logically related. Each workspace has its own isolated release, so only those with updates are pushed.

    I'm not familiar with that project (prior to now) but having a look around the repo, it would appear that they are using yarn workspaces. This is pretty analogous to what we are able to do with a csproj + slnx + Directory.Build.props + Directory.Packages.props - workspaces are the project (csproj), common dependency versions are managed at a root level (package.json -> Directory.Package.props), and common metadata is shared (package.json -> Directory.Build.props).

    We could probably leverage the logic that we have to detect changes in a PR and only run a subset of the tests to do targeted releases, but I'd question if that really would reduce overhead that much. Originally, the idea I had for versioning was major -> Aspire major, minor -> release new integration(s), patch -> bug fixes. This would mean that we could rev the minor.patch part of our version scheme separate from Aspire and that a hypothetical Toolkit 13.6 != Aspire 13.6. Now, new integrations would still ship early, but they would be in the .beta tag for a patch release, not becoming "stable" until the minor bump happened.

    In reality though, this hasn't been the case for a while, mostly because we haven't released all that much, and because Aspire had been shipping more frequently so we could hold integrations. But with more contributors, we might be able to push releases more often which would disconnect us from the Aspire minor version release cycle (we wouldn't rev major without an Aspire major release).

    I do wonder though, if we do start holding releases for unchanged packages, say we cut 13.5 to release stable Squad, SeaweedFS, and a few others, does it become confusing when you have some integrations at 13.5 while others are 13.4?

  6. Odonno commented on Jul 11, 2026

    @Odonno
    Contributor

    I am a big fan of turborepo when it comes to managing a monorepo. It is primarily made for a JS ecosystem but it can work with other programming languages. I have some repository context where I setup turborepo against a multi-language repository by having Rust, C# and TypeScript apps/packages. One extra features (that I do not use often but are interesting) are Remote cache and affected packages (a git diff in the CI).

    For remote caching, you can use GitHub Actions cache or the default cloud provider (Vercel Remote Cache). GitHub Actions cache is particularly slow and with some limitations. I heard good feedback about Depot recently. Never used it but I am curious as to how the two combined can behave.

    When speaking of monorepo, I am more of the team "update all versions to latest" so that every package shares the same version at any time, even if no previous change made. (It is easier to read in my package manager file too). This is more of a "big bang" release but that happen not so often.

  7. github-actions commented on Aug 2, 2026

    @github-actions
    Contributor

    We have noticed this issue has not been updated within 21 days. If there is no action on this issue in the next 14 days, we will automatically close it. You can use /stale-extend to extend the window.

  8. github-actions commented on Aug 17, 2026

    @github-actions
    Contributor

    This issue has been stale for 5 weeks and has been automatically closed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions