Repository navigation
docs: 38.0.0 release notes - #19945
Conversation
techdocsmith
left a comment
There was a problem hiding this comment.
Great work. Left some comments
| Kafka ingestion can now publish segments that the Broker prunes at query time without waiting for compaction. Set `tuningConfig.streamingPartitionsSpec.partitionDimensions` to a list of low-to-medium cardinality dimensions; each task records | ||
| the distinct values it observes per dimension and stamps them onto a new `dim_value_set` shard spec. Queries that filter on a declared dimension then skip segments whose values can't match. |
There was a problem hiding this comment.
| Kafka ingestion can now publish segments that the Broker prunes at query time without waiting for compaction. Set `tuningConfig.streamingPartitionsSpec.partitionDimensions` to a list of low-to-medium cardinality dimensions; each task records | |
| the distinct values it observes per dimension and stamps them onto a new `dim_value_set` shard spec. Queries that filter on a declared dimension then skip segments whose values can't match. | |
| Kafka ingestion can now publish segments that the Broker prunes at query time without waiting for compaction. Set `tuningConfig.streamingPartitionsSpec.partitionDimensions` to a list of low-to-medium cardinality dimensions; each task records the distinct values it observes per dimension and stamps them onto a new `dim_value_set` shard spec. Queries that filter on a declared dimension then skip segments whose values don't match. |
odd line break. "values can't match" reads awkwardly
|
|
||
| #### Storage metrics | ||
|
|
||
| The `storage/load/bytes` and `storage/virtual/load/bytes` metrics now measure once the load is complete. Previously, they measured when the load starts. |
There was a problem hiding this comment.
"they measured when the load starts" mix of past/present reads awkwardly. Consider revising
Co-authored-by: Charles Smith <techdocsmith@gmail.com>
|
|
||
| [#19559](https://github.com/apache/druid/pull/19559) | ||
|
|
||
| #### Partial segment loading |
There was a problem hiding this comment.
Shouldn't this section have a description?
There was a problem hiding this comment.
| #### Partial segment loading | |
| #### Partial segment loading (experimental) | |
| Historical tiers now support loading partial data from a segment. The exact data to be downloaded is dictated by partial load rules which may list one or more cluster groups or projections as eligible for download. |
There was a problem hiding this comment.
partial load rules are actually separate from the linked PRs (there are a handful of those PRs too). 19535 is about allowing historicals running in virtual storage mode, with v10 segments which are stored in deep storage without zip, to load on demand at the column level at query time (so dictated by what the query needs to read rather than load rules), when using Dart or MSQ (legacy engine still requires full download up front in virtual mode). The storage accounting of the historical cache is done at the 'bundle' level (base table, projections, cluster groups), where the cache has an entry for each bundle and we just lazy load columns contained in that bundle into it as the query needs.
Partial load rules work with this partial load functionality by allowing virtual mode historicals to eagerly load parts of segments at the 'bundle' level (base table, projections, cluster groups) and make those bundles sticky so that they cannot be evicted/are always 'hot' even if the rest of the segment can be loaded and evicted on demand.
All of this stuff is quite experimental, and will probably be in a lot nicer state next release.
There was a problem hiding this comment.
Hmm, in that case, I feel we should just remove them from the release notes here, as you had suggested earlier. Otherwise, this will just create confusion.
| [#19620](https://github.com/apache/druid/pull/19620) | ||
| [#19535](https://github.com/apache/druid/pull/19535) | ||
|
|
||
| #### Clustered segments |
There was a problem hiding this comment.
| #### Clustered segments | |
| #### Clustered segments (experimental) | |
| Clustered segments contain one or more cluster groups that allow quick pruning of data while querying by scanning only the groups which match a given query. |
There was a problem hiding this comment.
its probably worth mentioning that it only supports v10 segments, which are not really documented either. I think the tricky part with announcing some of this stuff is that probably the only way to use them is to look at PRs or look at the code since none of it is documented yet.
There was a problem hiding this comment.
So, do you advise just removing these items from the release notes for now?
|
as discussed in #20236 (comment) See also #20110 |
| [#19620](https://github.com/apache/druid/pull/19620) | ||
| [#19535](https://github.com/apache/druid/pull/19535) | ||
|
|
||
| #### Clustered segments |
There was a problem hiding this comment.
| #### Clustered segments | |
| #### Clustered segments (experimental) | |
| Clustered segments contain one or more cluster groups that allow quick pruning of data while querying by scanning only the groups which match a given query. |
|
|
||
| [#19559](https://github.com/apache/druid/pull/19559) | ||
|
|
||
| #### Partial segment loading |
There was a problem hiding this comment.
| #### Partial segment loading | |
| #### Partial segment loading (experimental) | |
| Historical tiers now support loading partial data from a segment. The exact data to be downloaded is dictated by partial load rules which may list one or more cluster groups or projections as eligible for download. |
|
|
||
| #### Segment prefetching for Dart | ||
|
|
||
| Dart now supports the runtime property `druid.msq.dart.worker.segmentLoadAheadCount`, which controls the number of segments that Dart prefetches. If set greater than 0 for a worker, this setting becomes the default `segmentLoadAheadCount` value for the worker. If a query includes the `segmentLoadAheadCount` query context parameter, the query context takes precedence. |
There was a problem hiding this comment.
this is only really applicable to virtual storage mode historicals i think, though i suppose the docs for it don't mention that either
There was a problem hiding this comment.
Thanks for calling it out, updating it.
|
|
||
| #### New load rule types | ||
|
|
||
| Adds a new family of retention rules, `loadPartialByPeriod`, `loadPartialByInterval`, `loadPartialForever`, laying the groundwork for partial loading of version 10 segment projections on Historicals. |
There was a problem hiding this comment.
partial load rules allow virtual storage mode historicals (also not documented i think, druid.segmentCache.virtualStorage and druid.segmentCache.virtualStoragePartialDownloadsEnabled both true) to eagerly load "parts" of v10 segments that are stored in deep storage unzipped (druid.storage.zip=false but only s3 supports that). There are matchers that must be set on the partial rule that can select the parts of the v10 segments (v10 segments are internally organized into 'bundles' which are sort of logical containers that hold all of the columns of a projection, where projection could be the base table (or individual cluster groups for clustered segments which when concatenated together form the base table projection), or individual aggregate projections).
I'm hesitant to call this stuff out since like the other partial load stuff it its going to be pretty tough to use without docs.
I wonder if we should consolidate all of this stuff in an 'experimental virtual storage mode improvements' section (partial load on demand and the partial load rules built on top of it) or just remove it as suggested in other thread
There was a problem hiding this comment.
I think removing it for now would make more sense.
Release notes for 38.0.0. This should be caught up to August 9th
Fixes #XXXX.
Description
Fixed the bug ...
Renamed the class ...
Added a forbidden-apis entry ...
Release note
Key changed/added classes in this PR
MyFooOurBarTheirBazThis PR has: