Skip to content

Allow user to use glob/wildcard in file path #2393

Description

@timvw

Currently it is only possible to use a fully qualified path to a file or a folder, eg:

  • /Users/timvw/src/github/arrow-datafusion/testing/data/wildcard/green_01.csv
  • /Users/timvw/src/github/arrow-datafusion/testing/data/wildcard

In situations when a folder contains multiple "sorts" of files it would be handy to use a glob/wildcard to list relevant files, eg:

  • /Users/timvw/src/github/arrow-datafusion/testing/data/wildcard/green_*.csv
    (This would allow one to exclude files such as red_01.csv in that same folder)

Currently this does not seem to work due to some "incorrect/flawed" logic in local.rs / list_all / tokio::fs::metadata(&prefix).await?;

Activity

  1. tustvold commented on May 1, 2022

    @tustvold
    Contributor

    I'm not sure how I feel about having the local object store treat paths differently from the other implementations. Perhaps we should consistently support glob expressions as part of the ObjectStore trait, or not? Having a mix seems unfortunate...

    FYI @matthewmturner @alamb

  2. timvw commented on May 1, 2022

    @timvw
    ContributorAuthor

    The thing is, my believe is that most objectstores already support globbing (Otherwise, yes, making it explicit on the trait would have been valuable).

    eg: in datafusion-objectstore-s3 when I modify a test in s3.rs to use globbing instead of filename it keeps working:

        let mut files = s3_file_system
            .list_file("data/alltypes_plain.sn*py.parquet") //l379
            .await?;
    

    same for datafusion-objectstore-azure

        let mut files = azure_file_system
            .list_file("parquet-testing-data/alltypes_plai*.snappy.parquet") //283
            .await?;
    

    Todos:

    • Consider to make globbing explicit on ObjectStore trait?
    • Add tests in all objectstores to verify that the globbing is working as expect
    • Make globbing also work on ListingTable

    What do you think?

  3. tustvold commented on May 2, 2022

    @tustvold
    Contributor

    Does this work correctly with multiple files, or wildcards in the path and not the file? I'm very surprised to see this working.

    I ask as it is just calling out to https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjectsV2.html which doesn't support wildcards, let alone glob expressions. S3 doesn't even have a defined concept of a directory, so globbing has to be implemented client side.

    Edit: in fact the S3 tests are written in such a way as they don't fail if the list returns no results. Is this what is going on here? Perhaps we should fix them to assert the correct results are returned?

  4. timvw commented on May 2, 2022

    @timvw
    ContributorAuthor

    @tustvold : Good remarks/questions!

    This is what I can do:

    • have a look at the tests (s3 and azure) and verify that they fail when they should..
    • verify how globbing could work on local, s3, azure and hdfs
  5. timvw commented on May 2, 2022

    @timvw
    ContributorAuthor

    Can confirm that globbing indeed does not work OOTB on azure and s3.

    Added additional assertions to the respective projects to ensure that at least one or more file(s) are processed

  6. alamb commented on May 2, 2022

    @alamb
    Contributor

    timvm -- I wonder if you could do the "globbing" at a higher level (aka what interface would the user decide to type in /Users/timvw/src/github/arrow-datafusion/testing/data/wildcard/green_*.csv)?

    So rather than passing /Users/timvw/src/github/arrow-datafusion/testing/data/wildcard/green_*.csv to datafusion directly, you can implement something that uses the object store's list_all feature to implement whatever globbing semantics you wanted, and then pass the fully resolved list of files to datafusion?

    I am not sure if this would work for your usecase

  7. timvw commented on May 2, 2022

    @timvw
    ContributorAuthor

    @alamb My preference is to not push this to the user, but have it available in datafusion itself (expectation is to have behaviour very similar to apache spark).

    Will have a look at hadoop fs to find some inspiration on how they implemented that (across filesystem implementations) and potentially come back with a proposal ;)

  8. timvw commented on May 2, 2022

    @timvw
    ContributorAuthor

    Currently spark does the following:
    Datasource (paths) -> (when globPaths option is true) -> checkAndGlobPathIfNecessary
    when no glob pattern in path -> fs.listfiles(path)
    when glob pattern in path -> globber.glob(fs.listfiles(path))

    As @tustvold already suggested, adding a glob_files method to ObjectStore seems the appropriate way to implement this feature. This method should then:
    when no glob pattern in path -> simply list_files
    when glob pattern in path -> glob (list_files) (can implement this with file_stream.filter, similar to existing list_file_with_suffix)

    Reworked my code in #2394 to conform with the above.

  9. timvw commented on May 3, 2022

    @timvw
    ContributorAuthor

    With the current code my use-case works:

        let ctx = SessionContext::new();
        let nycdata = "/Users/timvw/nyc/trip data";
        let yellow = format!("{}/yellow_tripdata_2018-0[345].csv", nycdata);
        let options = CsvReadOptions::new();
        let df = ctx.read_csv(yellow, options).await?;
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions