Skip to content

Services in global mode do not reschedule when stopped #2705

Description

@dperny

Verified on version 18.03 at commit ID 0b1ad7c

Problem

Tasks belonging to a service in global mode that have constraints may not be rescheduled when they fall into a terminal state.

Minimal repro

Input the following command and wait a bit, it doesn't happen on the first try:

$ docker service create --name constraint-test --constraint node.platform.os==linux --mode global busybox sleep 10

Investigation

Got hands on a cluster with 1 linux node. There was a global service created whose constraints should have matched the node, but the task had dropped into Desired State SHUTDOWN, State COMPLETE, and nothing was rescheduling.

First, I tried updating the service.

$ docker service update --force broken-service

As expected for this service, the task went into state COMPLETE. However, instead of rescheduling, the task was again stuck.

The next thing I tried was to watch docker service ps broken-service to see the task state progression. However, this time, when force-updating, the service did reschedule.

I thought that the issue might be with the COMPLETE state, so I created a new service

$ docker service create --name complete-test --mode global busybox sleep 20

I then watched the task progression. The task stopped in Complete, was rescheduled, and restarted several times with no trouble.

I next decided to try adding a constraint

$ docker service update --constraint-add node.platform.os==linux complete-test

This time, the task began failing to reschedule.

The next step was to see if this was an interaction between the COMPLETE state and constraints, or merely an issue with constraints. To do this, I created a new service:

$ docker service create --name constraint-test --constraint node.platform.os==linux --mode global nginx

I then watched and did docker kill on the tasks' containers as they came up. After a few attempts, the tasks stopped rescheduling, from the FAILED state this time instead of the COMPLETE state.

The final step was to verify that the issue only affected global-mode services. To do this, I created a replicated service:

$ docker service create --name replicated-test --constraint node.platform.os==linux --mode replicated busybox sleep 10

The replicated service continued to reschedule correctly.

The logs yielded no useful information.

Activity

  1. dperny commented on Jul 11, 2018

    @dperny
    CollaboratorAuthor

    Yeah, this is a P0, actually, this is legit broken in a bad way.

  2. thaJeztah commented on Jul 12, 2018

    @thaJeztah
    Member

    Verified on version 18.03 at commit ID 0b1ad7c

    Anything in changes that are staged for 18.03? 0b1ad7c...bump_v18.03

  3. dperny commented on Jul 12, 2018

    @dperny
    CollaboratorAuthor

    I have verified that this issue is in fact present on master. I vendored swarmkit master into my local moby/moby tree, did make shell, started the daemon, and created a global service with a constraint.

  4. dperny commented on Jul 12, 2018

    @dperny
    CollaboratorAuthor

    Doing an informal bisect. Swarmkit commit 49a9d7f is unaffected.

  5. dperny commented on Jul 12, 2018

    @dperny
    CollaboratorAuthor

    I ran a git bisect on the docker tree, github.com/moby/moby. I managed to narrow down the commit that introduced this issue:

    9c2aa15

    333b2f28fef4ba857905e7263e7b9bbbf7c522fc is the first bad commit
    commit 333b2f28fef4ba857905e7263e7b9bbbf7c522fc
    Author: Sebastiaan van Stijn <github@gone.nl>
    Date:   Tue Apr 17 13:35:35 2018 -0700
    
        Bump SwarmKit to 9c2aa152c3054371b833483a7ddad8d15052ec4f
    
        Relevant changes:
    
        - docker/swarmkit#2551 RoleManager will remove deleted nodes from the cluster membership
        - docker/swarmkit#2574 Scheduler/TaskReaper: handle unassigned tasks marked for shutdown
        - docker/swarmkit#2561 Avoid predefined error log
        - docker/swarmkit#2557 Task reaper should delete tasks with removed slots that were not yet assigned
        - docker/swarmkit#2587 [fips] Agent reports FIPS status
        - docker/swarmkit#2603 Fix manager/state/store.timedMutex
    
        Signed-off-by: Sebastiaan van Stijn <github@gone.nl>
    

    The previous commit of swarmkit in the docker tree was this one, which is unaffected:

    831df67

    commit 27749659d5a30999691e401a351221780a483099
    Author: Sebastiaan van Stijn <github@gone.nl>
    Date:   Mon Mar 26 21:29:15 2018 +0200
    
        Bump SwarmKit to 831df679a0b8a21b4dccd5791667d030642de7ff
    
        Changes included:
    
        - Ingress network should not be attachable
        - [manager/state] Add fernet as an option for raft encryption
        - Log GRPC server errors
        - Log leadership changes at manager level
        - [state/raft] Increase raft ElectionTick to 10xHeartbeatTick
        - Remove the containerd executor
        - agent: backoff session when no remotes are available
        - [ca/manager] Remove root CA key encryption support entirely
        - Fix agent logging race (fixes https://github.com/docker/swarmkit/issues/2576)
        - Adding logic to restore networks in order
    
        Also adds github.com/fernet/fernet-go as a new dependency
    
        Signed-off-by: Sebastiaan van Stijn <github@gone.nl>
    
  6. dperny commented on Jul 12, 2018

    @dperny
    CollaboratorAuthor

    Here is a link to the diff between the last known good commit and the first known bad one:

    831df67...9c2aa15

  7. dperny commented on Jul 13, 2018

    @dperny
    CollaboratorAuthor

    Specifically, git bisect dropped me on moby/moby commit moby/moby@333b2f2 (moby/moby#36880)

  8. anshulpundir commented on Jul 17, 2018

    @anshulpundir
    Contributor

    It's likely this commit: e3fcf6d (#2574)

    mainly because this is the only commit in the suspect list that touches the scheduler. If so, the problem may not necessarily be related to constraints

  9. anshulpundir commented on Jul 20, 2018

    @anshulpundir
    Contributor

    I just tested on 17.06 and this behavior is not present. What's more another odd behavior I noticed which is the task that just completed was being removed (once retention limit is reached), instead of the oldest task.

    I think this is a bug in the task reaper where the most recent task which is in state completed desired state running, gets removed, instead of the oldest task. As a result, the restart fails.

    Also, note that this is not related to constraints as we thought was the case originally.

  10. thaJeztah commented on Jul 23, 2018

    @thaJeztah
    Member

    What's more another odd behavior I noticed which is the task that just completed was being removed (once retention limit is reached), instead of the oldest task.

    Could that be related to this change? #2265 "Make the task termination order deterministic"

  11. anshulpundir commented on Jul 23, 2018

    @anshulpundir
    Contributor

    Here's one of the fixes: #2712

  12. changed the title [-]Services in global mode with constraints do not reschedule when stopped[/-] [+]Services in global mode do not reschedule when stopped[/+] on Jul 26, 2018
  13. self-assigned this
    on Jul 26, 2018
  14. anshulpundir commented on Jul 26, 2018

    @anshulpundir
    Contributor

    second fix: #2677

  15. anshulpundir commented on Jul 31, 2018

    @anshulpundir
    Contributor

    Here's the breakdown of the root cause:

    1. There was a bug in the task sorter. This leads to the task that just reached a terminal state to be examined for cleanup first, instead of the oldest one in the history. Fix for this is here [orchestrator] Fix task sorting #2712.

    2. There as a second bug introduced in 88064b5, which causes a task to be prematurely removed. Because of the incorrect sorting order, the task that is still running is inspected for cleanup, and because of an incorrect criteria, is actually deleted. This causes the service to no be rescheduled. The fix is here [manager/orchestrator/reaper] Fix the condition used for skipping over running tasks #2677

  16. dperny commented on Aug 3, 2018

    @dperny
    CollaboratorAuthor

    Yeah, this looks like a solid fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions