Repository navigation
[.NET main] ubuntu-latest build job killed without log output on every push since 2026-06-18 #1430
Description
Activity
I am wondering if it is going to be sustainable to have a single huge pipeline that builds, tests and publishes all the integrations all at the same time. Wouldn't it make sense to have separate pipelines for each integration? I realise this will take some effort, but I feel like it is worth the effort.
Reacted by David BottiauI am wondering if it is going to be sustainable to have a single huge pipeline that builds, tests and publishes all the integrations all at the same time. Wouldn't it make sense to have separate pipelines for each integration? I realise this will take some effort, but I feel like it is worth the effort.
Yep, probably just publish 1 artifact (nupack) per job, and then have a single job dependent of all the other jobs to publish the packages.
There's a collection of unique problems that we have with the monorepo design that we have.
At the time of writing, we have 260 projects in the solution, so that's not even covering the TypeScript app hosts and the non-.NET sample applications. This is, simply put, a huge project - even though there really isn't that much code in there.
Because of this, it's entirely impractical to run the tests in a single agent, that why there's the matrix design for running tests (if anyone wants to know more on it check out the blog I wrote) - this radically reduces the resources required for doing test runs, which is what is most resource intensive.
We also have to do some tricks with Docker and caching - we hit rate limits with Docker a while back which we have worked around using ACR for our CI builds (#601) that caches Docker images for us so we don't hit the Docker registry directly.
There's minimal caching setup on things like NPM, NuGet, etc. that should improve the performance of setting up actions, that happens in https://github.com/CommunityToolkit/Aspire/blob/main/.github/actions/setup-runtimes-caching/action.yml and I'm pretty sure it's not very ruggedised so doesn't cache as aggressively as it should.
I made the decision when I was first setting up the CI that "push to main" (and release) would bump the version of every NuGet package, not just the one(s) that had changes in the diff since last time. This does mean that we push a lot of NuGet packages each time, often times with packages that have no changes since their last release other than dependencies. This reduces the overhead on the maintainers - don't have to be pickup on what is to be published, but at the code of a really aggressive publishing approach.
Recently we had a publishing failure because the NuGet API key was expired. I've changed the design to use the new Trusted Publishing on NuGet (you can read how I did that here), which means that we no longer use a long-lived API token, instead CI requests one on demand. This is a lot more secure and means that we don't have to worry about expiring tokens.
Lastly, I added some auto retry logic into the pipeline to handle failing tests. Because of the volume of runners we spin up to do the tests (it's like, 120+ now) I think we might hit some caps that cause tests to just randomly fail at an infra level, Docker not starting is a common one. The retry logic is pretty solid but if we're hitting cases where it should auto-rerun, let's tweak the rerun workflow to cover more cases.
Hopefully that gives some overall context of the problem we face with this project, and the things that have been added along the way to attempt to address the issues. The most common failures we see on the
mainbranch now seem to be from a handful of integrations that "randomly fail" and if re-run, the tests pass. The problem is, they appear as tests themselves that are failing, not infrastructure around the tests, which makes it hard to identify the failures separately from legitimate test failures. Maybe we should have just a default "retry N times on failure" rather than attempting to do some dynamic detection of failure conditions?I also know that most of these solutions have just come from myself doing things. Generally that's just been because I'll see a problem and be like "well, I'll throw Copilot at it" (or, back in the day, just code up a solution 🫨), but if people have ideas, raise them for discussion or prototype them out, and we can look at their viability. Things like larger runners (paid) runners are a harder thing to tackle because someone has to pay for them, but if it's the only possible solution, then it's something I'll have to look into.
Reacted by David Bottiau and Jonathan Mezach@aaronpowell Thanks for the extensive write-up and explanation of the reasoning behind most of these choices.
One thing that stands out to me is the decision that "push to main" (and release) bumps the version of every NuGet package. While I realise that makes things easier from the perspective of not having to figure out what has actually changed it does mean we push a lot of packages with no actual changes which could be confusing from the users perspective.
I wonder if we could take some inspiration from the Backstage Community Plugins repository (https://github.com/backstage/community-plugins/). They've created workspaces, which are folders with a bunch of files that are logically related. Each workspace has its own isolated release, so only those with updates are pushed.
I realise this is an entirely different ecosystem (C# and NuGet vs. TypeScript and npm), but I do think it works fairly well there. It avoids the problem of pushing too many packages, while also keeping CI relatively small.
I wonder if we could take some inspiration from the Backstage Community Plugins repository (https://github.com/backstage/community-plugins/). They've created workspaces, which are folders with a bunch of files that are logically related. Each workspace has its own isolated release, so only those with updates are pushed.
I'm not familiar with that project (prior to now) but having a look around the repo, it would appear that they are using yarn workspaces. This is pretty analogous to what we are able to do with a csproj + slnx +
Directory.Build.props+Directory.Packages.props- workspaces are the project (csproj), common dependency versions are managed at a root level (package.json->Directory.Package.props), and common metadata is shared (package.json->Directory.Build.props).We could probably leverage the logic that we have to detect changes in a PR and only run a subset of the tests to do targeted releases, but I'd question if that really would reduce overhead that much. Originally, the idea I had for versioning was major -> Aspire major, minor -> release new integration(s), patch -> bug fixes. This would mean that we could rev the minor.patch part of our version scheme separate from Aspire and that a hypothetical Toolkit 13.6 != Aspire 13.6. Now, new integrations would still ship early, but they would be in the
.betatag for a patch release, not becoming "stable" until the minor bump happened.In reality though, this hasn't been the case for a while, mostly because we haven't released all that much, and because Aspire had been shipping more frequently so we could hold integrations. But with more contributors, we might be able to push releases more often which would disconnect us from the Aspire minor version release cycle (we wouldn't rev major without an Aspire major release).
I do wonder though, if we do start holding releases for unchanged packages, say we cut 13.5 to release stable Squad, SeaweedFS, and a few others, does it become confusing when you have some integrations at 13.5 while others are 13.4?
I am a big fan of turborepo when it comes to managing a monorepo. It is primarily made for a JS ecosystem but it can work with other programming languages. I have some repository context where I setup turborepo against a multi-language repository by having Rust, C# and TypeScript apps/packages. One extra features (that I do not use often but are interesting) are Remote cache and affected packages (a git diff in the CI).
For remote caching, you can use GitHub Actions cache or the default cloud provider (Vercel Remote Cache). GitHub Actions cache is particularly slow and with some limitations. I heard good feedback about Depot recently. Never used it but I am curious as to how the two combined can behave.
When speaking of monorepo, I am more of the team "update all versions to latest" so that every package shares the same version at any time, even if no previous change made. (It is easier to read in my package manager file too). This is more of a "big bang" release but that happen not so often.
We have noticed this issue has not been updated within 21 days. If there is no action on this issue in the next 14 days, we will automatically close it. You can use
/stale-extendto extend the window.github-actions commented
on Aug 17, 2026 on Aug 17, 2026 – with GitHub ActionsContributorMore actionsThis issue has been stale for 5 weeks and has been automatically closed.
Summary
The
.NET mainworkflow has been failing on every push tomainsince 2026-06-18 (8+ consecutive failures). Always the same shape:build (windows-latest)passes,build (ubuntu-latest)fails in ~3–4 minutes with no Build step log file persisted — classic signature of the runner being killed before it can flush.Evidence
d8c5674(Squad #1394)Likely cause
The breakpoint (#1391 dbx) lands right when the solution crossed some size threshold. Several new hosting integrations have landed since (dbx, SeaweedFS, Bitwarden, Squad), each adding 3–6 new csproj.
ubuntu-latestrunners have ~14 GB free disk;windows-latesthas ~80 GB.dotnet build -c Releaseof the whole toolkit now produces a LOT ofbin/+obj/output across 100+ projects.Hosting.Java.Tests-ubuntu-latestalso timed out at exactly 1h on the first failure (run 27726664815) — additional symptom of the runner being unhealthy.What
dotnet-ci.yml(per-PR) does differentlyPRs pass CI because
dotnet-ci.ymlonly builds affected projects (using the path-based filter you set up), not the whole solution.dotnet-main.ymldoes a fulldotnet restore+dotnet buildon every push, which is what's blowing past the disk/memory budget.Suggested fixes (any one should be enough)
Add disk cleanup step — easiest:
- uses: jlumbroso/free-disk-space@mainThis reclaims 30+ GB on ubuntu-latest by removing Android SDK, GHC, .NET preinstalls we don't need, etc.
Shard the build — split
dotnet buildacross N matrix shards byDirectory.Solution.propsslice.Use
--no-incremental+ clean between project builds — slower but bounded memory.Switch
build (ubuntu-latest)toubuntu-latest-largeorubuntu-22.04-large— paid runner with more disk.Happy to send a PR for option 1 if you want — it's the lowest-risk and matches what other large .NET repos (including
dotnet/aspireitself) do.cc @aaronpowell — flagging because main is currently red and
v13.5.0can't be cut until this is fixed (release workflow runsrun-tests→package→sign→publish-nuget).