Skip to content

Stop misdetecting Markdown that mentions HTML - #36

Closed
trsdn wants to merge 1 commit into
masterfrom
fix/14-markdown-html-detection
Closed

trsdn wants to merge 1 commit into
masterfrom
fix/14-markdown-html-detection

Conversation

@trsdn

@trsdn trsdn commented Aug 19, 2026

Copy link
Copy Markdown
Owner

Fixes #14.

Independent PR — branches from master, touches only ClipboardContentDetector.cs.

Problem

GetMarkdownScore vetoed the entire score on any HTML-looking tag:

if (ContainsHTML(trimmed))
    return 0;

This only bites when the clipboard carries both text and HTML — which is exactly what editors and browsers publish. So Markdown that merely mentions a tag was classified as rich text, and Ctrl+Enter ran HTML → Markdown on a document that was already Markdown.

Change

  • Markup inside fenced blocks and inline code is ignored, since angle brackets there are content, not markup.
  • Remaining raw HTML is a penalty (-6) weighed against the other signals, not an absolute veto. Markdown that embeds a little raw HTML still wins on structure; text that genuinely is HTML still scores low and stays rich text.

Verification

All cases below have the same syntax-highlighted editor HTML on the clipboard alongside the text, which is what triggers the bug.

Clipboard text Before After
# Title + list + **bold** Markdown Markdown
# Layout + list + `<div class="card">` RichText ❌ Markdown ✅
# Title + list + fenced ```html block RichText ❌ Markdown ✅
<h1>Hi</h1><p>there</p> (HTML source as text) RichText RichText
Word paste: plain prose text + semantic HTML RichText RichText
HTML only, no text RichText RichText
empty Unknown Unknown

No regressions: the two cases that must stay RichText still do, including a paste whose plain-text side has no Markdown syntax at all.

dotnet build -c Release -r win-x64 succeeds with 0 warnings.

GetMarkdownScore returned 0 as soon as the text contained anything looking
like an HTML tag. Markdown may legitimately embed raw HTML, and technical
writing routinely mentions tags inside inline code or fenced blocks, so a
single occurrence vetoed every other signal.

That only mattered when the clipboard carried both text and HTML, which is
exactly what editors and browsers publish - so copying Markdown that
mentions a tag was classified as rich text, and Ctrl+Enter converted a
document that was already Markdown.

Ignore markup inside code when checking, and weigh the remainder against
the other signals instead of vetoing them. Text that really is HTML still
scores low enough to be treated as rich text.

Fixes #14

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
@trsdn

trsdn commented Aug 28, 2026

Copy link
Copy Markdown
Owner Author

Superseded by #53. These commits are already contained in that branch, so this PR would merge nothing on its own. Collapsing the stacked converter PRs into one avoids the base-rewriting problem that made the stack unmergeable.

@trsdn trsdn closed this Aug 28, 2026
@trsdn
trsdn deleted the fix/14-markdown-html-detection branch August 28, 2026 21:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Markdown containing HTML is misdetected as rich text

1 participant