Version
$ openshift-install version
v0.8.0
Platform (aws|libvirt|openstack):
AWS
What happened?
Background information
After performing an install which failed and rebooting, i lost the cluster assets (they were stored in /tmp while working through troubleshooting with @abhinavdahiya). As such, I had to manually reap all resources related to the cluster. In the process of performing this manual cleanup, I missed the following three resources:
- rb-master-role
- rb-bootstrap-role
- rb-worker-role
Installation failure
As the roles already existed in IAM when attempting to perform an installation it failed with the following error:
[bharrington@leviathan OPENSHIFT.RwRC]$ ./openshift-install-linux-amd64 create cluster --dir=.
INFO Consuming "Install Config" from target directory
INFO Creating cluster...
ERROR
ERROR Error: Error applying plan:
ERROR
ERROR 3 errors occurred:
ERROR * module.masters.aws_iam_role.master_role: 1 error occurred:
ERROR * aws_iam_role.master_role: Error creating IAM Role rb-master-role: EntityAlreadyExists: Role with name rb-master-role already exists.
ERROR status code: 409, request id: 3da0f0e0-107b-11e9-815b-3377a72ac1b6
ERROR
ERROR
ERROR * module.bootstrap.aws_iam_role.bootstrap: 1 error occurred:
ERROR * aws_iam_role.bootstrap: Error creating IAM Role rb-bootstrap-role: EntityAlreadyExists: Role with name rb-bootstrap-role already exists.
ERROR status code: 409, request id: 3d9e7fd8-107b-11e9-815b-3377a72ac1b6
ERROR
ERROR
ERROR * module.iam.aws_iam_role.worker_role: 1 error occurred:
ERROR * aws_iam_role.worker_role: Error creating IAM Role rb-worker-role: EntityAlreadyExists: Role with name rb-worker-role already exists.
ERROR status code: 409, request id: 3d9f6a21-107b-11e9-a107-2b53e6606676
ERROR
ERROR
ERROR
ERROR
ERROR
ERROR Terraform does not automatically rollback in the face of errors.
ERROR Instead, your Terraform state file has been partially updated with
ERROR any resources that successfully completed. Please address the error
ERROR above and apply again to incrementally change your infrastructure.
ERROR
ERROR
FATAL failed to fetch Cluster: failed to generate asset "Cluster": failed to create cluster: failed to apply using Terraform
Noting the following message:
"Terraform does not automatically rollback in the face of errors. Instead, your Terraform state file has been partially updated with any resources that successfully completed.
I then used the destroy cluster mechanism of the installer as follows:
1 [bharrington@leviathan OPENSHIFT.RwRC]$ ./openshift-install-linux-amd64 destroy cluster --dir=.
2 INFO Deleted NAT Gateway id=nat-05927ef6d073b9121
3 INFO Deleted NAT Gateway id=nat-05198d50a7ad0ea0d
4 INFO Deleted subnet id=subnet-0c28912fe5073cfb1
5 INFO Deleted NAT Gateway id=nat-0e2f41d8d9b4e3453
6 INFO Deleted load balancer name=rb-ext
7 INFO Removed role rb-master-role from instance profile rb-master-profile
8 INFO Deleted load balancer name=rb-int
9 INFO Deleted subnet id=subnet-0101aba6dbd7f1803
10 INFO deleted profile rb-master-profile
11 INFO Deleted target group name=rb-api-ext
12 INFO Deleted security group id=sg-021ce2083c435246a
13 INFO Deleted target group name=rb-api-int
14 INFO Deleted subnet id=subnet-0f993073b5407c002
15 INFO deleted role rb-master-role
16 INFO Deleted target group name=rb-services
17 INFO Removed role rb-worker-role from instance profile rb-worker-profile
18 INFO deleted profile rb-worker-profile
19 INFO Deleted record rb-api.os4.rvu.io. from r53 zone /hostedzone/Z1LPW7VY401D8I
20 INFO deleted role rb-worker-role
21 INFO Deleted record rb-api.os4.rvu.io. from r53 zone /hostedzone/Z2MCRFQ54FOX8O
22 INFO Deleted route53 zone id=/hostedzone/Z2MCRFQ54FOX8O
23 INFO Removed role rb-bootstrap-role from instance profile rb-bootstrap-profile
24 INFO deleted profile rb-bootstrap-profile
25 INFO deleted role rb-bootstrap-role
26 INFO Emptied bucket name=terraform-20190104234830758200000001
27 INFO Deleted bucket name=terraform-20190104234830758200000001
28 INFO Deleted security group id=sg-0facb240e6af5009f
29 INFO Deleted VPC endpoint id=vpce-0d692d145faedf7ce
30 INFO Deleted route table id=rtb-026f3a544a749f08c
31 INFO Deleted route table id=rtb-073895ff9d3db914a
32 INFO Deleted route table id=rtb-02eaed5eba17b53c2
33 INFO Deleted route table id=rtb-0844bf4d66d7085e4
34 INFO Disassociated route table association id=rtbassoc-07f68e182c77a22ff
35 INFO Disassociated route table association id=rtbassoc-0f99d1bf0d204160a
36 INFO Disassociated route table association id=rtbassoc-0a554c1525b4c8d49
37 INFO Deleted security group id=sg-038242c8afcb4a9c3
38 INFO Deleted security group id=sg-06a6940fa83e766e2
39 INFO Deleted security group id=sg-0c5ce8b88e7940f5c
40 INFO Detached Internet GW igw-0611cdbc2bfa65e38 from VPC vpc-0b68ab55648319ce2
41 INFO Deleted internet gateway id=igw-0611cdbc2bfa65e38
42 INFO Deleted subnet id=subnet-02ceff55dec67667b
43 INFO Deleted subnet id=subnet-03f851d63ac7b7b4f
44 INFO Deleted subnet id=subnet-0d1fb0531e0fabc31
45 INFO Deleted Elastic IP ip=35.161.25.141
46 INFO Deleted VPC id=vpc-0b68ab55648319ce2
47 INFO Deleted Elastic IP ip=52.33.9.142
48 INFO Deleted Elastic IP ip=52.40.180.51
As noted on lines 7, 17, & 18 the installer deleted those roles despite the fact that it failed due to their existence.
What you expected to happen?
I would expect that the installer would only delete "resources that successfully completed", as per it's error message. As it was not able to successfully create those resources it should not have removed them when the cleanup was performed.
How to reproduce it (as minimally and precisely as possible)?
Create conflicting roles, perform an install, destroy the cluster
Version
Platform (aws|libvirt|openstack):
AWS
What happened?
Background information
After performing an install which failed and rebooting, i lost the cluster assets (they were stored in /tmp while working through troubleshooting with @abhinavdahiya). As such, I had to manually reap all resources related to the cluster. In the process of performing this manual cleanup, I missed the following three resources:
Installation failure
As the roles already existed in IAM when attempting to perform an installation it failed with the following error:
Noting the following message:
I then used the
destroy clustermechanism of the installer as follows:As noted on lines 7, 17, & 18 the installer deleted those roles despite the fact that it failed due to their existence.
What you expected to happen?
I would expect that the installer would only delete "resources that successfully completed", as per it's error message. As it was not able to successfully create those resources it should not have removed them when the cleanup was performed.
How to reproduce it (as minimally and precisely as possible)?
Create conflicting roles, perform an install, destroy the cluster