fix(documents): keep both items of a combined "Items 1 and 2" heading - #1385
Merged
Merged
Conversation
When a 10-K's TOC link text is a combined heading ("Items 1 and 2. Business
and Properties", "Items 1. and 2.", "Items 1. and 2.", "Items 7. and 7A."),
every TOC parser (Workiva, Donnelley, generic) dropped the row:
_item_label_from_text matches `item\s+(\d+)`, and the plural's "s" fails it.
The keyword fallback matched nothing, and since 5.34.0's per-form schema the
generic path returns "" for unmatched text instead of the raw title that used
to become a part_i_items_1_and_2_... key. Part I then started at Item 1A and
tenk.business was None on Viper (broken since 5.29.0, the Workiva parser),
Devon (5.34.0), Freeport and Cheniere. The #710 test injected ready-made keys,
so it never exercised detection.
- toc_analyzer: a combined label reads as its first item
(TOCAnalyzer.combined_item_numbers).
- Section.covered_items (new): both items, read from the section's own heading
within its first 5,000 chars. Viper's TOC link lands on a 2,681-char preamble
ahead of the heading, and Talos's TOC splits the label and link across cells.
- TenK.__getitem__ / .items resolve the second item to that section.
- The pattern-augmentation gate counts covered items as found, so a bare
"Properties" sub-heading inside Talos's Item 1 no longer becomes a
1,311-char Item 2 (#1383).
Verified offline on four real filings (new fixtures vnom, lng, fcx, talo) and
live on Devon. Fast lane: 7,799 passed; the one failure
(test_dt1f1_wrapped_item_headers 20-F lengths) fails identically on main.
All 60 section network regression files pass, one file per call, and the
parity and corpus ratchets pass (35) with the new fixtures in the corpus.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Thanks for the quick turnaround. I can confirm the fix on merge commit
I also compared every section against 5.59.1 on these five filings and two Apple 10-Ks. Nothing else changed: the remaining sections are byte-identical in markdown, confidence and detection method. Looking forward to seeing this in the next release 👍🏻 |
|
@dgunning Do you expect it to ship in a 5.x release soon, or will it wait for 6.0? |
Owner
Author
|
This should be in a 5.x release by this weekend |
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #1382
Fixes #1383
Problem
When a 10-K's TOC link text is a combined heading, every TOC parser dropped the row. The Workiva, Donnelley and generic parsers all did this, for headings like:
Items 1. and 2._item_label_from_textmatchesitem\s+(\d+), and the plural's "s" fails that match. The keyword fallback found nothing. Since 5.34.0's per-form schema, the generic path returns""for unmatched text, instead of the raw title that used to become apart_i_items_1_and_2_…key.As a result, Part I started at Item 1A and
tenk.businesswasNoneon Viper, Devon, Freeport and Cheniere. The #710 test injected ready-made keys, so it never exercised detection.On Talos, the TOC puts the label and the link in separate cells, so Item 1 survived. But Item 2 looked missing, and the pattern pass promoted a bare "Properties" sub-heading from inside Item 1 to a 1,311-character Item 2 (#1383).
Change
toc_analyzer: a combined label now reads as its first item, viaTOCAnalyzer.combined_item_numbers.Section.covered_items(new, defaults to()): lists both items. It is read from the section's own heading within its first 5,000 characters, not from the TOC row. Viper's TOC link lands on a 2,681-character preamble before the heading, and Talos's TOC row never names both items.TenK.__getitem__andTenK.items: the second item resolves to the combined section. The oldpart_*_items_N_and_Mkey lookup is kept.covered_items.Verification
0002074176-26-0000100001193125-26-0564850000831259-25-0000060000003570-26-0000050001193125-26-067807propertiesas Item 2tests/issues/regression/test_issue_1382_combined_item_headings.py, 19 tests, fully offline. It parses four new real-filing fixtures end to end (vnom,lng,fcx,talo). On main, 18 of them fail.test_dt1f1_wrapped_item_headers::test_2010_20f_…(20-F item lengths), fails identically on main locally, so it is unrelated.🤖 Generated with Claude Code