Describe the bug
When fs.s3a.aws.credentials.provider names an AWS SDK profile provider, software.amazon.awssdk.auth.credentials.ProfileCredentialsProvider or com.amazonaws.auth.profile.ProfileCredentialsProvider, native scans resolve the profile with aws-config's ProfileFileCredentialsProvider. With Hadoop 3.4 and later (Spark 4.x), Hadoop maps both names to the Java SDK v2 ProfileCredentialsProvider, and the two SDKs resolve profiles differently. Some of these differences change which identity signs requests or which STS endpoint is called:
- a source profile with both static keys and
credential_process: Java uses credential_process, aws-config the keys;
- assume-role STS region: Java uses each role's own
region, else the default region chain, else the global sts.amazonaws.com; the native side uses one region, from aws-config's default region chain, for every role in the chain;
- a role whose
source_profile names itself: Java reports a cycle, aws-config assumes the role with the profile's own keys;
- SSO keys together with
role_arn: Java uses SSO, aws-config assumes the role.
This is the current behavior on main and branch-1.1. Hadoop's own org.apache.hadoop.fs.s3a.auth.ProfileAWSCredentialsProvider is handled in #5872 and is not part of this issue.
Steps to reproduce
On Spark 4.x, set fs.s3a.aws.credentials.provider=software.amazon.awssdk.auth.credentials.ProfileCredentialsProvider and select an assume-role profile whose source profile has both aws_access_key_id and credential_process. A native read signs AssumeRole with the static key, while Spark with Comet disabled signs with the process credentials.
Expected behavior
Native scans pick the same credentials and STS endpoints as Hadoop for the SDK profile provider names. One option is to route them through HadoopS3ACredentialProviderAdapter, which builds Hadoop's own provider list on the executor. That would add the adapter's classpath requirement and a JNI call per request for static keys to configurations that work today.
Additional context
Found while reviewing #5872.
Describe the bug
When
fs.s3a.aws.credentials.providernames an AWS SDK profile provider,software.amazon.awssdk.auth.credentials.ProfileCredentialsProviderorcom.amazonaws.auth.profile.ProfileCredentialsProvider, native scans resolve the profile with aws-config'sProfileFileCredentialsProvider. With Hadoop 3.4 and later (Spark 4.x), Hadoop maps both names to the Java SDK v2ProfileCredentialsProvider, and the two SDKs resolve profiles differently. Some of these differences change which identity signs requests or which STS endpoint is called:credential_process: Java usescredential_process, aws-config the keys;region, else the default region chain, else the globalsts.amazonaws.com; the native side uses one region, from aws-config's default region chain, for every role in the chain;source_profilenames itself: Java reports a cycle, aws-config assumes the role with the profile's own keys;role_arn: Java uses SSO, aws-config assumes the role.This is the current behavior on
mainandbranch-1.1. Hadoop's ownorg.apache.hadoop.fs.s3a.auth.ProfileAWSCredentialsProvideris handled in #5872 and is not part of this issue.Steps to reproduce
On Spark 4.x, set
fs.s3a.aws.credentials.provider=software.amazon.awssdk.auth.credentials.ProfileCredentialsProviderand select an assume-role profile whose source profile has bothaws_access_key_idandcredential_process. A native read signsAssumeRolewith the static key, while Spark with Comet disabled signs with the process credentials.Expected behavior
Native scans pick the same credentials and STS endpoints as Hadoop for the SDK profile provider names. One option is to route them through
HadoopS3ACredentialProviderAdapter, which builds Hadoop's own provider list on the executor. That would add the adapter's classpath requirement and a JNI call per request for static keys to configurations that work today.Additional context
Found while reviewing #5872.