Repository navigation
Conversation
There was a problem hiding this comment.
🟡 Changes recommended
The implementation and tests currently don’t cover the documented metadata["relevance_score"] fallback, so behavior diverges from the PR description and can drop valid scores.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR enhances the LangChain GenAI instrumentation to include per-document relevance scores in retrieval span content, aligning emitted gen_ai.retrieval.documents data more closely with the GenAI semantic conventions.
Changes:
- Add polymorphic score extraction and include
scorein retrieval document serialization for spans when content capture is enabled. - Add/extend unit and integration test coverage to validate score extraction behavior and JSON safety (non-finite floats).
- Add a changelog fragment documenting the new behavior.
File summaries
| File | Description |
|---|---|
| instrumentation/opentelemetry-instrumentation-genai-langchain/src/opentelemetry/instrumentation/genai/langchain/callback_handler.py | Adds _extract_document_score / _document_to_dict and uses them to populate retrieval documents (including score) on retrieval spans. |
| instrumentation/opentelemetry-instrumentation-genai-langchain/tests/test_callback_handler.py | Adds unit tests around score extraction and document-to-dict conversion. |
| instrumentation/opentelemetry-instrumentation-genai-langchain/tests/test_retriever.py | Adds end-to-end/integration assertions that retrieval spans include score where applicable and exclude invalid/non-finite scores. |
| instrumentation/opentelemetry-instrumentation-genai-langchain/.changelog/670.added | Changelog entry for capturing document relevance scores on retrieval spans. |
Review details
- Files reviewed: 4/4 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Pull request dashboard statusMerged · refreshed 2026-09-17 05:30 UTC Status above doesn't look right?
|
|
Added explicit /dashboard route:reviewers |
…sts and documentation
Description
Fixes #584
Per the OpenTelemetry Semantic Conventions for Generative AI systems (
gen_ai.retrieval.documents), each retrieved document item should capture its relevance score under thescorekey when available from the retriever or vector search engine.This PR adds document relevance score extraction to
OpenTelemetryCallbackHandler._get_retrieval_documentsinopentelemetry-instrumentation-genai-langchain.Key Changes:
scoreby checking document attributes first (getattr(doc, "score", None)or dictionarydoc.get("score")), then falling back to document metadata (metadata.get("score")/metadata.get("relevance_score")).intorfloat.boolinstances (sinceisinstance(True, int)is True in Python).0and0.0scores (usingis not Noneand type checks instead of truthiness).NaN,+Inf,-Inf) viamath.isfinite.langchain_core.documents.Documentobjects and arbitrary duck-typed document objects/dictionaries withpage_contentorcontentandmetadata.scorekey alongsidecontentandidingen_ai.retrieval.documentsJSON serialization.Type of change
How has this been tested?
test_callback_handler.pycovering score extraction, attribute precedence, metadata fallbacks,0/0.0preservation, boolean exclusion,NaN/Infexclusion, invalid types, and duck-typed objects.test_retriever.pycovering synchronous and asynchronous retrievers (invoke,ainvoke,get_relevant_documents,aget_relevant_documents) with scores.instrumentation/opentelemetry-instrumentation-genai-langchain.pyright --level error(0 errors) andmypystrict type check clean.ruff check(0 errors) andruff format --checkclean across all files.Checklist