Repository navigation
Services in global mode do not reschedule when stopped #2705
Description
Activity
Yeah, this is a P0, actually, this is legit broken in a bad way.
Verified on version 18.03 at commit ID 0b1ad7c
Anything in changes that are staged for 18.03? 0b1ad7c...bump_v18.03
I have verified that this issue is in fact present on master. I vendored swarmkit master into my local moby/moby tree, did
make shell, started the daemon, and created a global service with a constraint.Doing an informal bisect. Swarmkit commit 49a9d7f is unaffected.
I ran a git bisect on the docker tree,
github.com/moby/moby. I managed to narrow down the commit that introduced this issue:333b2f28fef4ba857905e7263e7b9bbbf7c522fc is the first bad commit commit 333b2f28fef4ba857905e7263e7b9bbbf7c522fc Author: Sebastiaan van Stijn <github@gone.nl> Date: Tue Apr 17 13:35:35 2018 -0700 Bump SwarmKit to 9c2aa152c3054371b833483a7ddad8d15052ec4f Relevant changes: - docker/swarmkit#2551 RoleManager will remove deleted nodes from the cluster membership - docker/swarmkit#2574 Scheduler/TaskReaper: handle unassigned tasks marked for shutdown - docker/swarmkit#2561 Avoid predefined error log - docker/swarmkit#2557 Task reaper should delete tasks with removed slots that were not yet assigned - docker/swarmkit#2587 [fips] Agent reports FIPS status - docker/swarmkit#2603 Fix manager/state/store.timedMutex Signed-off-by: Sebastiaan van Stijn <github@gone.nl>The previous commit of swarmkit in the docker tree was this one, which is unaffected:
commit 27749659d5a30999691e401a351221780a483099 Author: Sebastiaan van Stijn <github@gone.nl> Date: Mon Mar 26 21:29:15 2018 +0200 Bump SwarmKit to 831df679a0b8a21b4dccd5791667d030642de7ff Changes included: - Ingress network should not be attachable - [manager/state] Add fernet as an option for raft encryption - Log GRPC server errors - Log leadership changes at manager level - [state/raft] Increase raft ElectionTick to 10xHeartbeatTick - Remove the containerd executor - agent: backoff session when no remotes are available - [ca/manager] Remove root CA key encryption support entirely - Fix agent logging race (fixes https://github.com/docker/swarmkit/issues/2576) - Adding logic to restore networks in order Also adds github.com/fernet/fernet-go as a new dependency Signed-off-by: Sebastiaan van Stijn <github@gone.nl>Here is a link to the diff between the last known good commit and the first known bad one:
Specifically,
git bisectdropped me on moby/moby commit moby/moby@333b2f2 (moby/moby#36880)I just tested on 17.06 and this behavior is not present. What's more another odd behavior I noticed which is the task that just completed was being removed (once retention limit is reached), instead of the oldest task.
I think this is a bug in the task reaper where the most recent task which is in state completed desired state running, gets removed, instead of the oldest task. As a result, the restart fails.
Also, note that this is not related to constraints as we thought was the case originally.
What's more another odd behavior I noticed which is the task that just completed was being removed (once retention limit is reached), instead of the oldest task.
Could that be related to this change? #2265 "Make the task termination order deterministic"
Here's one of the fixes: #2712
- changed the title
[-]Services in global mode with constraints do not reschedule when stopped[/-][+]Services in global mode do not reschedule when stopped[/+]on Jul 26, 2018 second fix: #2677
Here's the breakdown of the root cause:
-
There was a bug in the task sorter. This leads to the task that just reached a terminal state to be examined for cleanup first, instead of the oldest one in the history. Fix for this is here [orchestrator] Fix task sorting #2712.
-
There as a second bug introduced in 88064b5, which causes a task to be prematurely removed. Because of the incorrect sorting order, the task that is still running is inspected for cleanup, and because of an incorrect criteria, is actually deleted. This causes the service to no be rescheduled. The fix is here [manager/orchestrator/reaper] Fix the condition used for skipping over running tasks #2677
Reacted by Olli Janatuinen-
Yeah, this looks like a solid fix.
Verified on version 18.03 at commit ID 0b1ad7c
Problem
Tasks belonging to a service in global mode that have constraints may not be rescheduled when they fall into a terminal state.
Minimal repro
Input the following command and wait a bit, it doesn't happen on the first try:
Investigation
Got hands on a cluster with 1 linux node. There was a global service created whose constraints should have matched the node, but the task had dropped into Desired State
SHUTDOWN, StateCOMPLETE, and nothing was rescheduling.First, I tried updating the service.
As expected for this service, the task went into state
COMPLETE. However, instead of rescheduling, the task was again stuck.The next thing I tried was to watch
docker service ps broken-serviceto see the task state progression. However, this time, when force-updating, the service did reschedule.I thought that the issue might be with the
COMPLETEstate, so I created a new serviceI then watched the task progression. The task stopped in
Complete, was rescheduled, and restarted several times with no trouble.I next decided to try adding a constraint
This time, the task began failing to reschedule.
The next step was to see if this was an interaction between the
COMPLETEstate and constraints, or merely an issue with constraints. To do this, I created a new service:I then watched and did
docker killon the tasks' containers as they came up. After a few attempts, the tasks stopped rescheduling, from theFAILEDstate this time instead of theCOMPLETEstate.The final step was to verify that the issue only affected global-mode services. To do this, I created a replicated service:
The replicated service continued to reschedule correctly.
The logs yielded no useful information.