Skip to content

Support Apache Ozone as Deep Storage #20224

Description

@maron546

Description

We would like to request support for Apache Ozone as a deep storage option in Apache Druid.
Apache Ozone provides an S3-compatible object store and is designed to integrate with the Hadoop ecosystem. Since the Hadoop version used by the Druid druid-hdfs-storage extension is compatible with the Hadoop libraries required by Apache Ozone, we believe Ozone could potentially be supported without introducing a completely separate storage implementation.

We would therefore like to ask whether support for Apache Ozone could be added either:

  1. To the existing druid-hdfs-storage extension, leveraging the Hadoop/Ozone filesystem integration; or
  2. Through a dedicated Druid Ozone storage extension, if integrating it into druid-hdfs-storage is not considered appropriate.

Motivation

We are using Apache Ozone as the object storage layer of our Big Data platform and would like to use it as Druid's deep storage as well.
Apache Ozone provides a Hadoop-compatible filesystem interface, allowing applications to access Ozone through the Hadoop FileSystem API. This could potentially allow Druid to access Ozone directly through its existing Hadoop-based deep storage implementation.

Although Apache Ozone provides an S3-compatible API, relying on its S3 Gateway is not ideal for our use case, as it introduces an additional access layer and associated request overhead. Direct access through Ozone's native Hadoop FileSystem integration would avoid this extra layer and provide lower-latency access to deep storage.

Activity

  1. FrankChen021 commented on Sep 2, 2026

    @FrankChen021
    Member

    If we take it as a S3-compatible object storage and use the S3 interface to access OZone, I think the implementation should be much simpler as we have standard S3 and Aliabab Cloud OSS supported already, we can copy code and small dependencies changes( I guess maybe current s3 object storage can also be used to access OZone via S3)

    But, using S3 interface to access OZone might bring performance problems if the QPS at object storage is high, and the s3g might be a bottleneck. We have seen in this many OZone application cases. So from the performance perspective, I would prefer its hadoop client interfaces, no S3 involved. I didn't explore the native API of OZone, not sure if they share SAME APIs with HDFS, but I think if possible, the OZone support can be also provided in a dedicated extension

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions