Skip to content

Forward Gemini generation settings - #509

Merged
aksg87 merged 1 commit into
google:mainfrom
ojassharma7:autocontrib/issue-358
Aug 19, 2026
Merged

aksg87 merged 1 commit into
google:mainfrom
ojassharma7:autocontrib/issue-358

Conversation

@ojassharma7

@ojassharma7 ojassharma7 commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

Description

Forward Gemini max_output_tokens, top_p, and top_k from language_model_params for realtime and batch requests instead of silently dropping them at an overly narrow allowlist. Runtime overrides and explicit clearing are preserved, with updated provider documentation and regression coverage.

This addresses only the generation-settings portion of the large-document report. Automatic entity-group splitting, 429 retry/backoff with partial-result preservation, and relevance-aware chunking remain tracked on the issue.

Addresses part of #358

Bug fix

How Has This Been Tested?

Maintainer validation during consolidation, on the identical tree to this branch head:

  • .venv/bin/python -m pytest tests/ -ra -m "not live_api" --ignore=tests/test_ollama_integration.py
  • .venv/bin/python -m tox -e format,lint-src,lint-tests
  • .venv/bin/python -m tox -e live-api
  • Fresh Python 3.13 editable install, pip check, and lx.extract() smoke test
  • Repository CI on the identical tree

The local Ollama integration environment returns the same HTTP 500 on main; the repository Ollama integration job passed on the identical tree.

Checklist

  • I have read and acknowledged Google's Open Source
    Code of conduct.
  • I have read the
    Contributing
    page, and I either signed the Google
    Individual CLA
    or am covered by my company's
    Corporate CLA.
  • I have discussed my proposed solution with code owners in the linked
    issue(s) and we have agreed upon the general approach.
  • I have made any needed documentation changes, or noted in the linked
    issue(s) that documentation elsewhere needs updating.
  • I have added tests, or I have ensured existing tests cover the changes.
  • I have followed
    Google's Python Style Guide
    and ran pylint over the affected code.

@github-actions github-actions Bot added the size/S Pull request with 50-150 lines changed label Aug 3, 2026
@ojassharma7

Copy link
Copy Markdown
Contributor Author

Thanks @reichaves — agreed this only covers the Gemini allowlist / language_model_params part of problem #1.

I updated the PR description to the required template and marked it as addressing part of #358. I still have to keep a Fixes #358 line as well because the repo's Validate PR template check requires that exact phrase for non-maintainer PRs; happy for a maintainer to adjust the linking wording before merge (or reopen #358 afterward) so retry/backoff and chunking stay tracked.

Also noting the Require linked issue with community support check is failing because #358 currently has 4 👍 and the workflow requires 5 unique 👍 (excluding the PR author). If anyone following the issue can add one more 👍, that check should go green — then I'd love your help re-testing the 230K-char / 22-entity case.

@ojassharma7

Copy link
Copy Markdown
Contributor Author

Done — removed the visible Fixes #358 line and left only Addresses part of #358, so merge should not auto-close the issue. Whenever you can add the 👍 on #358, that should unblock the reaction check. Thanks!

@reichaves

Copy link
Copy Markdown

Thanks for iterating on this. I see the reaction check should be satisfied now (#358 is at 5 👍).

One concern about the "Fixes #358" workaround though: wrapping it in an HTML comment doesn't actually stop GitHub's auto-close. GitHub's closing-keyword parser scans the raw PR body text, not the rendered output — HTML comments are only hidden visually, not stripped from what GitHub matches against. So merging as-is will likely still auto-close #358, same as before.

If the template bot truly requires a literal "Fixes #NNN" string, that sounds like a limitation of the bot rather than something we can work around in the PR body. Could a maintainer confirm whether they can either (a) merge and manually reopen #358 right after, or (b) adjust/bypass that template check for this case? I'd rather not merge on an assumption that turns out to still close the tracking issue for the retry/backoff and chunking work.

@github-actions

Copy link
Copy Markdown

⚠️ Branch Update Required

Your branch is 1 commits behind main. Please update your branch to ensure CI checks run with the latest code:

git fetch origin main
git merge origin/main
git push

Note: Enable "Allow edits by maintainers" to allow automatic updates.

@reichaves

Copy link
Copy Markdown

Hello

Original author of #358 here. A couple of data points from continued testing on the same real-world case (Brazilian FIDC regulation PDFs, ~230K chars, 22 entity types) that might be useful for this PR:

  1. The JSON-truncation issue is still live on current models. Running extrair_regulamento.py against a sample regulation with gemini-2.5-flash today, one of the three entity-groups still hit Unterminated string starting at: line 9592 column 7 mid-response. Our current workaround (not from this repo, from our own client code) retries with a smaller max_char_buffer until it fits — it works, but it's a blunt instrument: it silently re-sends the whole chunk instead of just asking for more output tokens.

  2. Different models fail differently, which argues for max_output_tokens as a knob rather than something baked into a model-specific default. Comparing gemini-2.5-flash against a newer flash model on the identical document: the 2.5 model returned more raw entities (68) but 40% of them were empty strings ("") — it "found" the class but produced no text. The newer model returned fewer entities (34) with zero blanks, but silently dropped 4 whole entity classes instead. Neither failure mode is caused by the retry/backoff or partial-checkpointing side of Challenges extracting structured data from large Brazilian fund regulation PDFs (230K+ chars, 22 entity types) #358 (we've since patched that locally) — both look like output-length pressure that max_output_tokens exposure would let us diagnose/tune directly instead of guessing from symptoms.

Happy to test this branch against the same document set if it'd help move it past the reaction threshold — just let me know.

@ojassharma7

Copy link
Copy Markdown
Contributor Author

Thanks @reichaves — this is exactly the signal this PR is meant to unblock, and your model comparison is really useful.

What this PR changes. Before this branch, Gemini language_model_params such as max_output_tokens / top_p / top_k were stripped by a narrow allowlist and never reached generate_content. On this branch they are forwarded, so you can raise the output ceiling instead of only shrinking max_char_buffer client-side.

Example against this branch:

import langextract as lx

result = lx.extract(
    text_or_documents=doc,
    prompt_description=...,
    examples=...,
    model_id="gemini-2.5-flash",
    language_model_params={
        "max_output_tokens": 8192,  # or higher if the model allows
        # "top_p": 0.95,
        # "top_k": 40,
    },
)

That should let you replace the “retry with a smaller max_char_buffer” blunt instrument with “ask for more output tokens on the same chunk” for the unterminated-JSON / empty-string / dropped-class failure modes you described — both of which do look like output-length pressure.

What this PR deliberately does not do. Automatic entity-group splitting, 429 retry/backoff with partial-result preservation, and relevance-aware chunking stay tracked on #358. Happy for a maintainer to merge this as a partial fix and keep #358 open for that remaining work (the HTML-commented Fixes #358 concern you raised earlier still stands for auto-close).

Testing offer. Yes please — if you can re-run extrair_regulamento.py on the same FIDC PDFs against branch ojassharma7:autocontrib/issue-358 with an explicit max_output_tokens, that would be the best validation we can get. Particularly interested in whether raising it clears the Unterminated string … mid-response hit on the entity group that was truncating, and whether empty "" extractions / dropped classes improve vs your current defaults.

Happy to adjust docs/examples if anything in the call path is unclear once you’ve tried it.

@reichaves

Copy link
Copy Markdown

Thanks for the quick turnaround, @ojassharma7 — tested this against autocontrib/issue-358 on the same kind of FIDC regulation PDF from my original report, so here's real before/after data.

Setup: same document (CVM regulation PDF, 159,629 raw chars → 50,043 chars after this repo's section filtering), same 3-group prompt split, chunk_size=3000, gemini-2.5-flash, single run each (LLM output isn't deterministic, so treat this as signal, not a controlled trial).

  main (current, no language_model_params) autocontrib/issue-358 + max_output_tokens: 8192
Entities extracted 55 57
Empty ("") extractions 11 (20%) 24 (42%)
Truncation errors none Group C: Failed to parse JSON content: Unterminated string starting at: line 712 column 7 (char 22286)

Two things worth flagging:

  1. Raising max_output_tokens to 8192 didn't clear the truncation in this run — I still hit Unterminated string on the same group that used to truncate, and the empty-extraction rate roughly doubled rather than improving. So at least for this document, output-length pressure doesn't look like the dominant cause.
  2. Separately from the params question: when that Unterminated string error happens, langextract logs it as Skipping chunk internally rather than raising an exception. That means any caller-side retry logic keyed on catching the parse error (mine included — I retry with a smaller max_char_buffer on JSONDecodeError/Unterminated string in the except block) never fires, because the exception never reaches the caller. The chunk's data is just silently dropped. That might be worth a separate look regardless of the max_output_tokens question — happy to open a follow-up issue if that's useful.

My tentative read: for this document, the empty extractions are mostly enumeration-style fields (evento_avaliacao, limite_concentracao, aplicacao_minima) that legitimately don't appear in every chunk, not output getting cut off — so max_output_tokens may not be the right lever for that particular symptom, even though it's clearly useful for the cases where output length actually is the bottleneck.

Happy to run more documents/more repetitions if that would help validate one way or the other — let me know what would be most useful for you to see.

@aksg87
aksg87 force-pushed the autocontrib/issue-358 branch from 0a1ee94 to 45ab2b6 Compare August 19, 2026 07:11
@aksg87 aksg87 changed the title Challenges extracting structured data from large Brazilian fund regulation PDFs (230K+ chars, 22 entity types) Forward Gemini generation settings Aug 19, 2026
@github-actions github-actions Bot added size/M Pull request with 150-600 lines changed and removed size/S Pull request with 50-150 lines changed labels Aug 19, 2026
@aksg87 aksg87 added the ready-to-merge Triggers live API tests for PRs from forks label Aug 19, 2026
@aksg87
aksg87 merged commit 9853e44 into google:main Aug 19, 2026
40 of 47 checks passed
@aksg87

aksg87 commented Aug 19, 2026

Copy link
Copy Markdown
Collaborator

Thanks @ojassharma7 for identifying the Gemini settings allowlist bug and contributing the original fix, and @reichaves for validating it on the FIDC documents. We expanded the fix with clear runtime precedence, explicit clearing, batch coverage, documentation, and regression tests, and merged it through this PR to preserve the contribution.

This addresses the generation-settings portion of #358; the remaining large-document work stays open there.

This branch had an error being deployed

1 failed deployment
live-keys — 45ab2b6b Deployed Aug 19, 2026 by aksg87 via test-fork-pr #974
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-to-merge Triggers live API tests for PRs from forks size/M Pull request with 150-600 lines changed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants