Skip to content

Task reaper does not consider some tasks for removal #2555

Description

@nishanttotla

#2461 looks at tasks with desired state REMOVE and only considers them for deletion if they have moved past COMPLETED:

		for _, t := range removeTasks {
			if t.Status.State >= api.TaskStateCompleted {
				tr.cleanup = append(tr.cleanup, t.ID)
			}
		}

This misses the case for when the service was removed and perhaps had some tasks that never got to run.

Activity

  1. self-assigned this
    on Mar 13, 2018
  2. eferley commented on Mar 13, 2018

    @eferley

    Hello @nishanttotla if you just read the symptoms described in moby/moby#36527, would you say that falls within the case you describe here ?

  3. prianna commented on Mar 13, 2018

    @prianna

    Originally posted this in the wrong issue. Apologies.

    I'm seeing something that might be related to this. On an environment running Docker 17.12.1, the following was observed: A service was removed, but the task was orphaned, and is lingering in the task list. It seems that an assumption made in task_reaper.go does not always hold. Based on the comment in TaskReaper.run, I would guess that this not expected behavior, but I'm not sure.

    For example:

    [ec2-user@manager1 ~]$ sudo docker node ps uyqwvctka5ow3c3ay1boztslh | grep w8cqmb3ng
    w8cqmb3ngslj        qphwmhh7brgph9msda2r3c7ev.1                                                             manager1:9874/anaconda-0201-1517522214-runscript:latest                      worker1.us-west-2.compute.internal   Running             Failed 3 hours ago           "task: non-zero exit (1)"
    [ec2-user@manager1 ~]$ sudo docker service ps qphwmhh7brgph9msda2r3c7ev
    no such service: qphwmhh7brgph9msda2r3c7ev
    

    We first observed this state after removing the service associated with this task.
    The state persisted despite removing the container on the node that this task was running on.

    We've seen this occur on 17.12.0 and 17.12.1, but not prior to that.

  4. nishanttotla commented on Mar 13, 2018

    @nishanttotla
    ContributorAuthor

    @prianna thanks for reporting. Are you able to inspect these orphaned tasks? What state do they report?

  5. prianna commented on Mar 14, 2018

    @prianna

    @nishanttotla Yes, we can inspect them. Here's the Status from an inspect on taskId w8cqmb3ngslj (referenced in my previous comment):

       "Status":{
          "Timestamp":"2018-03-13T17:34:21.376037195Z",
          "State":"failed",
          "Message":"started",
          "Err":"task: non-zero exit (1)",
          "ContainerStatus":{
             "ContainerID":"3a59862d101bfe8d00574fb4d5ba874ec986c542622a14c40c692737d100b505",
             "ExitCode":1
          },
          "PortStatus":{
          }
       },
       "DesiredState":"running",
    
    
  6. prianna commented on Mar 14, 2018

    @prianna

    @nishanttotla I got another one for you, this one has a different state:

    Salient bits from sudo docker inspect xooskd3382kqz555igrw4h7gj:

    "Status": {
                "Timestamp": "2018-03-14T17:23:05.876919939Z",
                "State": "complete",
                "Message": "finished",
                "ContainerStatus": {
                    "ContainerID": "ca1deaf833240271281d9b08b7183ca5aea5d18b8617281fe7909917a52c7f13"
                },
                "PortStatus": {}
            },
            "DesiredState": "running",
    

    Same situation upon inspect of the service, though:

    sudo docker service ps h8cznnlumfb1kvw6gas73ivqu
    no such service: h8cznnlumfb1kvw6gas73ivqu
    
    sudo docker service inspect h8cznnlumfb1kvw6gas73ivqu
    []
    Status: Error: no such service: h8cznnlumfb1kvw6gas73ivqu, Code: 1
    

    And the container is no longer running on the node associated with that task.

  7. nishanttotla commented on Mar 19, 2018

    @nishanttotla
    ContributorAuthor

    @prianna it seems like the task is in failed state with desired state running. This is particularly tricky, because it doesn't seem like they're being marked as orphaned, so the task reaper will be unable to detect this and remove them. I believe that there is an issue in task reaper, which we're working on.

  8. anshulpundir commented on Mar 19, 2018

    @anshulpundir
    Contributor

    This latest instance looks different from the original problem in the issue. You're prob already thinking of this (in which case ignore my comment), but I'd suggest listing down all the different state combinations that are possible and see how the reaper handles it @nishanttotla

  9. conceptdeluxe commented on May 27, 2018

    @conceptdeluxe

    Here is another one with a different error status:

            "Status": {
                "Timestamp": "2018-05-27T09:10:51.2263348Z",
                "State": "pending",
                "Message": "pending task scheduling",
                "Err": "no suitable node (scheduling constraints not satisfied on 1 node)",
                "PortStatus": {}
            },
    

    It is reproduceable by deploying a stack with [node.role == worker] when only a manager is available.

  10. nishanttotla commented on Jun 5, 2018

    @nishanttotla
    ContributorAuthor

    @conceptdeluxe which version is this latest error status for?

  11. conceptdeluxe commented on Jun 7, 2018

    @conceptdeluxe

    @nishanttotla For me it is reproducable with 18.03.1-ce on both MacOS and Ubuntu

  12. kedarkekan commented on Jul 10, 2018

    @kedarkekan

    @nishanttotla Not sure, why a state: new with Version: 18.03.1-ce, gets into DesiredState: remove, but not removed

            "Status": {
                "Timestamp": "2018-07-10T14:20:21.062436275Z",
                "State": "new",
                "Message": "created",
                "PortStatus": {}
            },
            "DesiredState": "remove"
    
  13. rauschbit commented on Oct 19, 2018

    @rauschbit

    are there any updates on this issue - we have the same problem here...

    Docker Version: 18.03.0~ce on Ubuntu 16.04 (-> 4.4.0-104-generic #127-Ubuntu SMP)
    Swarm with 3 Manager nodes and 19 worker nodes.

    we had the problem of a not starting service because the ip address range of the docker overlay network went full - so no available ip for the service task. we don't know why this happened because the (only) service in this stack had only 2 tasks (replicas) and the standard docker overlay network ip range is /24 - so 255 ip addresses - that should be enough. but that's another problem...

    we wanted to clean up this stack by deleting the stack with "docker stack rm" and the coresponding network but some service tasks (with the above problem) could not be deleted and now have the "CURRENT STATE" of "New ..." and the "DESIRED STATE" of "Remove".
    See the following output: (docker image names are changed)

    ID                  NAME                          IMAGE                 NODE                DESIRED STATE       CURRENT STATE        ERROR               PORTS
    zpb8192dtwwu        i20ry1skr7rfpq0evuqw7ebyz.1   ss1:944d78c9                              Remove              New 2 hours ago                          
    g777ml8mib6t        i20ry1skr7rfpq0evuqw7ebyz.2   ss1:944d78c9                              Remove              New 2 hours ago
    

    as a workaround we started the services in a other stack and network - everything fine with the services now.

    does anyone know when the task reaper get's updated on this?

  14. unglaublicherdude commented on May 2, 2019

    @unglaublicherdude

    We have the same issue with docker version

    Docker version 18.09.5, build e8ff056

    The biggest problem for us is that the ports in the ingress are still blocked by these services.

    ID                  NAME                          IMAGE                        NODE                                          DESIRED STATE       CURRENT STATE             ERROR               PORTS
    g025ydj0v7zf        yhflvl8e9pt9kv0bibaudw06n.1   anoxis/registry-cli:latest   processing-worker-orange-101.X   Remove              Assigned 9 minutes ago
    a7glafei93mf        ufnpxrns2spno2kuypiho4mlh.1   anoxis/registry-cli:latest   processing-worker-orange-101.X   Remove              Assigned 12 minutes ago
    v9lpzbzgiuvj        l6fuvcvn4kvfp2opg2ov6udba.1   registry:2                   processing-worker-orange-101.X   Remove              Ready 46 hours ago
    

    We had to restart the docker daemon on this node, where the tasks were scheduled on. To get the stack to work again.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions