Skip to content

Fix PESQ aborting a whole batch on one unscorable sample - #3455

Merged
Borda merged 6 commits into
Lightning-AI:masterfrom
Kayvan-Zahiri:fix/pesq-degenerate-sample-3304
Sep 9, 2026
Merged

Borda merged 6 commits into
Lightning-AI:masterfrom
Kayvan-Zahiri:fix/pesq-degenerate-sample-3304

Conversation

@Kayvan-Zahiri

Copy link
Copy Markdown
Contributor

Fixes #3304.

One unscorable sample takes down the whole batch. pesq and pesq_batch both
default to on_error=PesqError.RAISE_EXCEPTION and torchmetrics never passes
on_error, so what happens depends on n_processes:

  • n_processes=1 (the default): pesq() raises NoUtterancesError and it
    propagates out of update(). The _filter_error_msg guard added in Ignore the NoUtterancesError when calculating pesq for a batch #2753 is
    unreachable on this path.
  • n_processes != 1: pesq_batch() catches worker exceptions and returns them
    in the result list, so the guard is reached, but it drops the failed
    entry. A 3-sample batch silently returns 2 values.

Measured on a 3-sample batch with target[1] silent:

functional,   n_processes=1  -> raises NoUtterancesError
functional,   n_processes=2  -> shape (2,)  values [2.5199, 2.1842]   <- one sample vanished
class metric, n_processes=2  -> compute()=2.3521  total=2             <- 3 in, 2 counted

After:

functional,   n_processes=1  -> shape (3,)  values [2.5199, nan, 2.1842]
functional,   n_processes=2  -> shape (3,)  values [2.5199, nan, 2.1842]
class metric, either         -> compute()=2.3521  total=2

Change

Pass on_error=PesqError.RETURN_VALUES at all three call sites and map every
failure report onto nan, keeping one entry per input sample. Failures arrive
two ways and both are handled: negative error codes (-1..-7, which cannot
collide with a real score since valid PESQ is >= -0.5) and the exception
objects pesq_batch collects from its workers.

The class metric excludes nan samples from both the running sum and total,
so a degenerate sample does not poison an epoch's average. That matches what the
multiprocessing path already did by dropping them, so it is not a new policy.

Scores for scorable samples are unchanged: RETURN_VALUES and
RAISE_EXCEPTION return the same value on success (verified,
2.158228635787964 both ways).

Also in this PR

#2753 replaced pesq_val.reshape(preds.shape[:-1]) with
reshape(len(pesq_val)), so a (2, 3, 2100) input returned (6,) instead of
(2, 3), contradicting the documented (...,) shape. Restored, which is safe
now that no entries are dropped. This is a behaviour change for ndim > 2
inputs
relative to 1.4.x-1.9.0. Happy to split it into its own PR if you would
rather keep this one minimal.

Tests

Four cases in tests/unittests/audio/test_pesq.py, all failing on main:

test_degenerate_sample_does_not_abort_batch[1]  NoUtterancesError
test_degenerate_sample_does_not_abort_batch[2]  assert torch.Size([2]) == (3,)
test_degenerate_sample_returns_nan_for_single_sample  NoUtterancesError
test_multidim_input_keeps_shape                 assert torch.Size([6]) == (2, 3)

The two different failure modes for [1] and [2] are the two paths above.

tests/unittests/audio/test_pesq.py: 20 -> 24 passed, same 6 skipped and 3
xfailed. Doctests pass. Wider audio dir: 88 passed. ruff, format and mypy clean.

The pesq backend defaults to on_error=PesqError.RAISE_EXCEPTION, so a
single degenerate sample (e.g. a silent reference, or a prediction whose
amplitude dwarfs the reference after the backend's shared normalisation)
raised NoUtterancesError out of the whole update.

Pass on_error=PesqError.RETURN_VALUES at every call site and map the
returned error codes, as well as the exception objects pesq_batch
collects from its workers, onto nan. The result now keeps one entry per
input sample, so the documented shape contract holds again for inputs
with more than one batch dimension. The class metric leaves nan samples
out of its average instead of propagating them.

Fixes Lightning-AI#3304
@Kayvan-Zahiri
Kayvan-Zahiri force-pushed the fix/pesq-degenerate-sample-3304 branch from e87534b to 8426d03 Compare August 24, 2026 18:35
@mergify mergify Bot removed the has conflicts label Aug 24, 2026
@Borda
Borda requested a lite review from Copilot September 8, 2026 06:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The changes align with the PR description, add targeted regression tests for the reported failure modes, and the only feedback is a minor docstring wording nit.

Pull request overview

This PR updates TorchMetrics’ PESQ functional and class-based metric behavior so that an unscorable sample (e.g., silent reference triggering NoUtterancesError) does not abort or shrink a batch, instead producing nan for that sample while keeping output shape consistent and excluding nan samples from the class metric average.

Changes:

  • Pass on_error=PesqError.RETURN_VALUES to PESQ backend calls and map backend failures to nan while preserving one output per input sample.
  • Restore multidimensional output shaping to preds.shape[:-1] for the functional interface.
  • Add regression tests covering degenerate samples (single + batched, single + multi-process) and multidimensional shape behavior; update changelog entry.
File summaries
File Description
tests/unittests/audio/test_pesq.py Adds tests ensuring degenerate samples yield nan without aborting batches and that multidim inputs preserve shape.
src/torchmetrics/functional/audio/pesq.py Ensures backend errors become nan (not exceptions/dropped entries) and restores documented output shaping.
src/torchmetrics/audio/pesq.py Updates class metric aggregation to ignore nan samples in sum/total.
CHANGELOG.md Documents the PESQ degenerate-sample behavior fix under Unreleased.
Review details
  • Files reviewed: 4/4 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/torchmetrics/functional/audio/pesq.py Outdated
@Borda Borda added the bug / fix Something isn't working label Sep 8, 2026
@codecov-commenter

codecov-commenter commented Sep 8, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 0% with 16 lines in your changes missing coverage. Please review.
✅ Project coverage is 31%. Comparing base (d184220) to head (bdd017d).
⚠️ Report is 20 commits behind head on master.
❗ Your organization needs to install the Codecov GitHub app to enable full functionality.

❗ There is a different number of reports uploaded between BASE (d184220) and HEAD (bdd017d). Click for more details.

HEAD has 828 uploads less than BASE
Flag BASE (d184220) HEAD (bdd017d)
torch2.0.1+cpu 20 0
python3.10 129 5
Windows 30 3
cpu 218 11
torch2.9.1 9 0
macOS 39 0
python3.12 89 6
torch2.9.1+cpu 30 3
torch2.8.0+cpu 30 2
torch2.8.0 10 0
Linux 149 8
torch2.10.0 10 0
torch2.7.1+cpu 20 2
torch2.10.0+cpu 20 2
torch2.4.1+cpu 10 0
torch2.2.2+cpu 10 0
torch2.1.2+cpu 10 0
torch2.5.1+cpu 10 0
torch2.0.1 10 0
torch2.3.1+cpu 9 0
torch2.6.0+cpu 10 2
Additional details and impacted files
@@           Coverage Diff            @@
##           master   #3455     +/-   ##
========================================
- Coverage      37%     31%     -6%     
========================================
  Files         349     349             
  Lines       19901   19899      -2     
========================================
- Hits         7264    6157   -1107     
- Misses      12637   13742   +1105     
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@mergify mergify Bot added the ready label Sep 8, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The change is narrowly scoped, aligns behavior across single/multiprocess paths, preserves documented shapes, and is backed by targeted regression tests for the previously failing scenarios.

Review details
  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

Comment thread src/torchmetrics/functional/audio/pesq.py Outdated
Borda and others added 2 commits September 8, 2026 18:56
Addresses review feedback on Lightning-AI#3455. All three backend calls were made
multi-line by this PR, so name the arguments: the upstream signatures are
pesq(fs, ref, deg, mode, on_error) and pesq_batch(fs, ref, deg, mode,
n_processor, on_error), which makes the target -> ref and preds -> deg
mapping explicit at the call site.

Also hyphenates "class-based metric" in the note added by this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014puKzWNVkoZhuj4LeGZwjD
@mergify

mergify Bot commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Tick the box to add this pull request to the merge queue (same as @mergifyio queue).

  • Queue this pull request

@Borda
Borda enabled auto-merge (squash) September 9, 2026 12:21
@mergify

mergify Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

queue

⚠️ Configuration not compatible with a branch protection setting

Details

The branch protection setting Require branches to be up to date before merging is not compatible with draft PR checks. To keep this branch protection enabled, update your Mergify configuration to enable in-place checks: set merge_queue.max_parallel_checks: 1, set every queue rule batch_size: 1, and avoid two-step CI (make merge_conditions identical to queue_conditions). Otherwise, disable this branch protection.

@Borda
Borda merged commit 093de1e into Lightning-AI:master Sep 9, 2026
54 of 55 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug / fix Something isn't working ready

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PerceptualEvaluationSpeechQuality - NoUtterancesError: b'No utterances detected'

5 participants