🔉 fix: Normalize audio MIME types in STT format validation - #12674
Conversation
Use getFileExtensionFromMime() to normalize non-standard MIME types (e.g. audio/x-m4a, audio/x-wav, audio/x-flac) before checking against the accepted formats list in azureOpenAIProvider. This is the same class of bug as #12608 (text/x-markdown), but for STT audio validation. Only audio/ and video/ MIME prefixes are normalized to prevent non-audio types from matching via the webm default fallback. Export getFileExtensionFromMime for testability. Fixes #12632
Use MIME_TO_EXTENSION_MAP for normalization instead of getFileExtensionFromMime() which falls back to 'webm' for unrecognized types. Gate raw subtype matching on audio/video prefix to prevent non-audio types (e.g. text/webm) from passing validation. Resolves Codex review comment about unknown subtypes silently passing.
There was a problem hiding this comment.
Pull request overview
Normalizes audio MIME types during Azure OpenAI STT format validation so common browser-reported vendor MIME types (e.g. audio/x-m4a) are accepted when they map to an allowed audio format, while unknown/unsafe MIME types remain rejected.
Changes:
- Update Azure OpenAI STT provider validation to normalize via
MIME_TO_EXTENSION_MAPand restrict raw-subtype fallback toaudio/andvideo/MIME prefixes. - Export
MIME_TO_EXTENSION_MAP(andgetFileExtensionFromMime) for unit testing. - Add Jest tests covering MIME normalization and negative cases (unknown audio subtypes, non-audio prefix bypass).
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.
| File | Description |
|---|---|
| api/server/services/Files/Audio/STTService.js | Normalize MIME types before validating accepted audio formats; export MIME map/util for tests. |
| api/server/services/Files/Audio/STTService.spec.js | New unit tests for MIME normalization and format acceptance/rejection scenarios. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| function isFormatAccepted(mimetype) { | ||
| const [mimePrefix, rawFormat = ''] = mimetype.split('/'); | ||
| const isAudioMime = mimePrefix === 'audio' || mimePrefix === 'video'; | ||
| const isKnownMime = mimetype in MIME_TO_EXTENSION_MAP; | ||
| const normalizedFormat = isKnownMime ? MIME_TO_EXTENSION_MAP[mimetype] : null; | ||
| return ( | ||
| acceptedFormats.includes(normalizedFormat) || | ||
| (isAudioMime && acceptedFormats.includes(rawFormat)) | ||
| ); | ||
| } |
There was a problem hiding this comment.
The new tests validate a local isFormatAccepted() helper that re-implements the provider’s MIME/format check, but they never call STTService.azureOpenAIProvider() (the code changed in this PR). This can let the test suite pass even if the real provider logic diverges (e.g., acceptedFormats list or normalization rules change). Consider exercising azureOpenAIProvider() directly with mocked sttSchema/genAzureEndpoint and asserting it throws/doesn’t throw for the same mimetypes instead of duplicating the logic in the test.
| } | ||
|
|
||
| module.exports = { STTService, speechToText }; | ||
| module.exports = { STTService, speechToText, getFileExtensionFromMime, MIME_TO_EXTENSION_MAP }; |
There was a problem hiding this comment.
MIME_TO_EXTENSION_MAP is now exported, which makes it part of the module’s public surface and allows external code/tests to mutate it at runtime. To prevent hard-to-debug side effects, consider freezing the map (e.g., Object.freeze) or exporting a defensive copy so consumers can’t modify the internal lookup table.
| !acceptedFormats.includes(normalizedFormat) && | ||
| !(isAudioMime && acceptedFormats.includes(rawFormat)) | ||
| ) { | ||
| throw new Error(`The audio file format ${rawFormat} is not accepted`); |
There was a problem hiding this comment.
The thrown error now always interpolates rawFormat. For some invalid inputs rawFormat can be empty or misleading (e.g., malformed MIME strings without a subtype), which makes debugging harder. Consider including the full audioFile.mimetype and/or the computed normalizedFormat in the error message so logs clearly show what was rejected.
| throw new Error(`The audio file format ${rawFormat} is not accepted`); | |
| throw new Error( | |
| `The audio file format is not accepted (mimetype: "${audioFile.mimetype}", raw format: "${rawFormat}", normalized format: "${normalizedFormat}")`, | |
| ); |
GitNexus: ❌ deploy failedThe deploy failed — the previous index (if any) continues to be served. |
|
@codex review |
|
Codex Review: Didn't find any major issues. Bravo. ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
…-AI#12674) * fix: normalize audio MIME types in STT format validation Use getFileExtensionFromMime() to normalize non-standard MIME types (e.g. audio/x-m4a, audio/x-wav, audio/x-flac) before checking against the accepted formats list in azureOpenAIProvider. This is the same class of bug as LibreChat-AI#12608 (text/x-markdown), but for STT audio validation. Only audio/ and video/ MIME prefixes are normalized to prevent non-audio types from matching via the webm default fallback. Export getFileExtensionFromMime for testability. Fixes LibreChat-AI#12632 * fix: reject unknown audio subtypes in STT format validation Use MIME_TO_EXTENSION_MAP for normalization instead of getFileExtensionFromMime() which falls back to 'webm' for unrecognized types. Gate raw subtype matching on audio/video prefix to prevent non-audio types (e.g. text/webm) from passing validation. Resolves Codex review comment about unknown subtypes silently passing. --------- Co-authored-by: Tobias Jonas <t.jonas@innfactory.de>
…-AI#12674) * fix: normalize audio MIME types in STT format validation Use getFileExtensionFromMime() to normalize non-standard MIME types (e.g. audio/x-m4a, audio/x-wav, audio/x-flac) before checking against the accepted formats list in azureOpenAIProvider. This is the same class of bug as LibreChat-AI#12608 (text/x-markdown), but for STT audio validation. Only audio/ and video/ MIME prefixes are normalized to prevent non-audio types from matching via the webm default fallback. Export getFileExtensionFromMime for testability. Fixes LibreChat-AI#12632 * fix: reject unknown audio subtypes in STT format validation Use MIME_TO_EXTENSION_MAP for normalization instead of getFileExtensionFromMime() which falls back to 'webm' for unrecognized types. Gate raw subtype matching on audio/video prefix to prevent non-audio types (e.g. text/webm) from passing validation. Resolves Codex review comment about unknown subtypes silently passing. --------- Co-authored-by: Tobias Jonas <t.jonas@innfactory.de>
…-AI#12674) * fix: normalize audio MIME types in STT format validation Use getFileExtensionFromMime() to normalize non-standard MIME types (e.g. audio/x-m4a, audio/x-wav, audio/x-flac) before checking against the accepted formats list in azureOpenAIProvider. This is the same class of bug as LibreChat-AI#12608 (text/x-markdown), but for STT audio validation. Only audio/ and video/ MIME prefixes are normalized to prevent non-audio types from matching via the webm default fallback. Export getFileExtensionFromMime for testability. Fixes LibreChat-AI#12632 * fix: reject unknown audio subtypes in STT format validation Use MIME_TO_EXTENSION_MAP for normalization instead of getFileExtensionFromMime() which falls back to 'webm' for unrecognized types. Gate raw subtype matching on audio/video prefix to prevent non-audio types (e.g. text/webm) from passing validation. Resolves Codex review comment about unknown subtypes silently passing. --------- Co-authored-by: Tobias Jonas <t.jonas@innfactory.de>
…-AI#12674) * fix: normalize audio MIME types in STT format validation Use getFileExtensionFromMime() to normalize non-standard MIME types (e.g. audio/x-m4a, audio/x-wav, audio/x-flac) before checking against the accepted formats list in azureOpenAIProvider. This is the same class of bug as LibreChat-AI#12608 (text/x-markdown), but for STT audio validation. Only audio/ and video/ MIME prefixes are normalized to prevent non-audio types from matching via the webm default fallback. Export getFileExtensionFromMime for testability. Fixes LibreChat-AI#12632 * fix: reject unknown audio subtypes in STT format validation Use MIME_TO_EXTENSION_MAP for normalization instead of getFileExtensionFromMime() which falls back to 'webm' for unrecognized types. Gate raw subtype matching on audio/video prefix to prevent non-audio types (e.g. text/webm) from passing validation. Resolves Codex review comment about unknown subtypes silently passing. --------- Co-authored-by: Tobias Jonas <t.jonas@innfactory.de>
Summary
Browsers commonly report audio files with non-standard MIME types (e.g.
audio/x-m4afor.m4afiles,audio/x-wav,audio/x-flac). The STTazureOpenAIProviderrejects these because the format validation only checks the raw MIME subtype against the accepted formats list.This is the same class of bug as #12608 (
text/x-markdownrejected in file uploads). The fix reuses the existingMIME_TO_EXTENSION_MAP— which already has the correct MIME-to-extension mapping — to normalize the format before validation. Unknown MIME types are correctly rejected instead of silently falling through to awebmdefault. Raw subtype fallback is gated onaudio/videoprefix to prevent non-audio types from bypassing validation.Fixes #12632
Continues #12633
Credit to @jona7o for the original PR.
Changes
MIME_TO_EXTENSION_MAPdirectly for normalization instead ofgetFileExtensionFromMime()(which has awebmdefault fallback that would accept unknown audio subtypes)audio/videoprefix to reject types liketext/webmMIME_TO_EXTENSION_MAPfor test useTest plan
audio/x-m4a,audio/x-wav,audio/x-flac) acceptedaudio/mpeg,audio/wav, etc.) acceptedaudio/aac,audio/somethingelse) rejectedtext/webm) rejectedapplication/ogg(valid Ogg container in the MIME map) accepted