Forward Gemini generation settings - #509
Conversation
|
Thanks @reichaves — agreed this only covers the Gemini allowlist / I updated the PR description to the required template and marked it as addressing part of #358. I still have to keep a Also noting the Require linked issue with community support check is failing because #358 currently has 4 👍 and the workflow requires 5 unique 👍 (excluding the PR author). If anyone following the issue can add one more 👍, that check should go green — then I'd love your help re-testing the 230K-char / 22-entity case. |
|
Done — removed the visible |
|
Thanks for iterating on this. I see the reaction check should be satisfied now (#358 is at 5 👍). One concern about the "Fixes #358" workaround though: wrapping it in an HTML comment doesn't actually stop GitHub's auto-close. GitHub's closing-keyword parser scans the raw PR body text, not the rendered output — HTML comments are only hidden visually, not stripped from what GitHub matches against. So merging as-is will likely still auto-close #358, same as before. If the template bot truly requires a literal "Fixes #NNN" string, that sounds like a limitation of the bot rather than something we can work around in the PR body. Could a maintainer confirm whether they can either (a) merge and manually reopen #358 right after, or (b) adjust/bypass that template check for this case? I'd rather not merge on an assumption that turns out to still close the tracking issue for the retry/backoff and chunking work. |
|
Your branch is 1 commits behind git fetch origin main
git merge origin/main
git pushNote: Enable "Allow edits by maintainers" to allow automatic updates. |
|
Hello Original author of #358 here. A couple of data points from continued testing on the same real-world case (Brazilian FIDC regulation PDFs, ~230K chars, 22 entity types) that might be useful for this PR:
Happy to test this branch against the same document set if it'd help move it past the reaction threshold — just let me know. |
|
Thanks @reichaves — this is exactly the signal this PR is meant to unblock, and your model comparison is really useful. What this PR changes. Before this branch, Gemini Example against this branch: import langextract as lx
result = lx.extract(
text_or_documents=doc,
prompt_description=...,
examples=...,
model_id="gemini-2.5-flash",
language_model_params={
"max_output_tokens": 8192, # or higher if the model allows
# "top_p": 0.95,
# "top_k": 40,
},
)That should let you replace the “retry with a smaller What this PR deliberately does not do. Automatic entity-group splitting, 429 retry/backoff with partial-result preservation, and relevance-aware chunking stay tracked on #358. Happy for a maintainer to merge this as a partial fix and keep #358 open for that remaining work (the HTML-commented Testing offer. Yes please — if you can re-run Happy to adjust docs/examples if anything in the call path is unclear once you’ve tried it. |
|
Thanks for the quick turnaround, @ojassharma7 — tested this against Setup: same document (CVM regulation PDF, 159,629 raw chars → 50,043 chars after this repo's section filtering), same 3-group prompt split,
Two things worth flagging:
My tentative read: for this document, the empty extractions are mostly enumeration-style fields ( Happy to run more documents/more repetitions if that would help validate one way or the other — let me know what would be most useful for you to see. |
0a1ee94 to
45ab2b6
Compare
|
Thanks @ojassharma7 for identifying the Gemini settings allowlist bug and contributing the original fix, and @reichaves for validating it on the FIDC documents. We expanded the fix with clear runtime precedence, explicit clearing, batch coverage, documentation, and regression tests, and merged it through this PR to preserve the contribution. This addresses the generation-settings portion of #358; the remaining large-document work stays open there. |
Description
Forward Gemini
max_output_tokens,top_p, andtop_kfromlanguage_model_paramsfor realtime and batch requests instead of silently dropping them at an overly narrow allowlist. Runtime overrides and explicit clearing are preserved, with updated provider documentation and regression coverage.This addresses only the generation-settings portion of the large-document report. Automatic entity-group splitting, 429 retry/backoff with partial-result preservation, and relevance-aware chunking remain tracked on the issue.
Addresses part of #358
Bug fix
How Has This Been Tested?
Maintainer validation during consolidation, on the identical tree to this branch head:
.venv/bin/python -m pytest tests/ -ra -m "not live_api" --ignore=tests/test_ollama_integration.py.venv/bin/python -m tox -e format,lint-src,lint-tests.venv/bin/python -m tox -e live-apipip check, andlx.extract()smoke testThe local Ollama integration environment returns the same HTTP 500 on
main; the repository Ollama integration job passed on the identical tree.Checklist
Code of conduct.
Contributing
page, and I either signed the Google
Individual CLA
or am covered by my company's
Corporate CLA.
issue(s) and we have agreed upon the general approach.
issue(s) that documentation elsewhere needs updating.
Google's Python Style Guide
and ran
pylintover the affected code.