Skip to content

Fix the Markdown/HTML conversion bugs and add a test project - #53

Merged
trsdn merged 11 commits into
masterfrom
fix/32-blank-line-after-list
Aug 28, 2026
Merged

trsdn merged 11 commits into
masterfrom
fix/32-blank-line-after-list

Conversation

@trsdn

@trsdn trsdn commented Aug 28, 2026 •

Copy link
Copy Markdown
Owner

This collapses the whole Markdown/HTML conversion stack into a single pull request. It was previously spread across seven stacked PRs (#29, #30, #31, #35, #36, #41, #51, #52, #53), which turned out to be unreviewable and unmergeable in practice: every merge rewrote the base of the next PR, and squash merges broke the ancestry so diffs kept re-expanding.

All of that work is already in this branch as ten commits, so merging this one PR lands the entire set.

Closes #15
Closes #19
Closes #16
Closes #21
Closes #14
Closes #23
Closes #34
Closes #33
Closes #32

What is in here

Conversion fixes (HTML to Markdown)

  • Nested list items were duplicated, because HtmlAgilityPack's InnerText returns every descendant.
  • Inline formatting was lost inside list items and table cells.
  • <pre> blocks dropped their content when there was no single direct <code> child.
  • Converted text was not escaped, so literal * or 1. produced unintended formatting.
  • Whitespace between block elements leaked a stray leading space onto the next block. The fix has to keep whitespace between inline elements, otherwise <em>a</em> <em>b</em> collapses to *a**b*, so it keys off whether the output buffer already ends with a newline.
  • A list was not separated from the block after it, so a following paragraph became a lazy continuation of the last list item and got swallowed into the list.

Conversion fixes (Markdown to HTML)

  • Markdown that merely mentions HTML was misdetected as rich text.
  • Task list checkboxes are rendered through a real Markdig renderer instead of brittle string replacement.
  • Stripping CSS classes also stripped language-*, losing the code language on every fenced block. Now stripped with a negative lookahead so the language survives.

Test project

md2loop.Core is split out of the WinUI app so the conversion logic can be tested without UI or clipboard dependencies, plus tests/md2loop.Tests covering it.

Verification

87 passed, 0 failed, 0 skipped. The suite originally carried several [Fact(Skip = ...)] tests documenting known-broken behaviour; all of them are now un-skipped and passing, so no bug the suite records is still open.

Worth flagging for review: fixing the code-language and list bugs broke three tests that had asserted the old buggy output. Those were rewritten rather than forced to pass. For the "non-language classes are still stripped" assertion I checked what Markdig actually emits with a throwaway probe rather than guessing, so the test cannot pass vacuously.

master is merged in and the one conflict, a README section added on both sides, was resolved additively. dotnet build -c Release -r win-x64 succeeds with 0 warnings and 0 errors.

trsdn and others added 11 commits August 19, 2026 08:50
Both list items and table cells pulled their content from InnerText, which
flattens markup to plain text. The converter already handles strong, em,
code, links and strikethrough correctly, but that pipeline was bypassed in
the two places formatting appears most often, so bold text and links were
silently dropped from every bullet and every cell.

Run both through ConvertChildren instead. List items convert a copy with
the nested lists detached, so nested content is still emitted separately as
indented items. Cell and item text is collapsed onto a single line, and
pipes inside cells are escaped so the table stays well formed.

Before:
    - See docs for details
    | Bold | see link |

After:
    - See [docs](https://example.com) for **details**
    | **Bold** | see [link](https://x.test) |

Fixes #15

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
The pre branch looked up a single direct <code> child and converted only
that node, so anything else inside the block was discarded: a second
<code> element, a <code> wrapped in a div (which several editors emit),
and text sitting alongside a <code> sibling. The user saw a successful
"Markdown copied" message with code missing from the result.

Take the language hint from the first descendant <code>, but always
convert the whole <pre> subtree so no children are lost. Also size the
fence to be longer than the longest backtick run in the content, which
previously produced a broken nested fence whenever a code block contained
one.

Fixes #19

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
Text nodes were written straight into the output, so prose containing
Markdown punctuation was re-interpreted as formatting the next time it was
rendered.

Escape the inline characters that can open a construct, escape a leading
character that would turn a paragraph or list item into a heading, list,
blockquote or thematic break, and size inline code delimiters to be longer
than any backtick run they contain.

Underscores are only escaped at word boundaries so identifiers such as
file_name_here stay readable, and an ordered marker is neutralised on its
delimiter ("1\.") because a backslash before a digit is not an escape
sequence and would render literally.

Escaping is skipped inside pre and inline code, where the content is
already literal.

Fixes #16

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
Task lists were broken. The Unicode checkbox substitution matched literal
strings that Markdig 0.38 does not emit - it renders
<input disabled="disabled" type="checkbox" checked="checked" />, while the
replacement looked for <input checked="" disabled="" type="checkbox" />.
Neither replacement ever fired, so raw disabled checkbox elements were
pasted into Loop instead of the checkboxes it can display.

Replace the TaskList renderer instead, so the output does not depend on
Markdig's attribute order or spacing.

TransformTaskLists is removed. It scanned for a LiteralInline starting
with "[x] ", which the TaskList extension enabled by UseAdvancedExtensions
never produces, so it had no effect.

Before: <li><input disabled="disabled" type="checkbox" checked="checked" /> done</li>
After:  <li>&#9745; done</li>

Fixes #21

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
GetMarkdownScore returned 0 as soon as the text contained anything looking
like an HTML tag. Markdown may legitimately embed raw HTML, and technical
writing routinely mentions tags inside inline code or fenced blocks, so a
single occurrence vetoed every other signal.

That only mattered when the clipboard carried both text and HTML, which is
exactly what editors and browsers publish - so copying Markdown that
mentions a tag was classified as rich text, and Ctrl+Enter converted a
document that was already Markdown.

Ignore markup inside code when checking, and weigh the remainder against
the other signals instead of vetoing them. Text that really is HTML still
scores low enough to be treated as rich text.

Fixes #14

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
…nd 'origin/fix/14-markdown-html-detection' into test/23-converter-tests
The conversion code had no tests at all, which is why several of the bugs
fixed in this batch shipped: task list checkboxes were emitted as raw
<input> markup, nested lists were duplicated, and inline formatting was
dropped from list items and table cells. All of those are cheap to catch
with a unit test and impossible to catch by eye.

Nothing in the converters depends on WinUI or the clipboard, but they
lived in the app project, which targets net8.0-windows and cannot be
referenced by an ordinary test project. They move to a new md2loop.Core
library so they can be tested directly; the app now references it and is
otherwise unchanged.

The suite has 84 tests across the detector, both converters, and a
round-trip class that pushes Markdown through Loop HTML and back, which
is what a user actually does when they paste into Loop and copy the
result out later.

Three tests are marked Skip rather than deleted, one per known open bug
(#32, #33, #34). They record the expected behaviour and will start
passing when those bugs are fixed.

Also adds bin/ and obj/ to .gitignore. They were never listed and the
repo only stayed clean because files were staged individually, which is
no longer safe with three projects in the tree.

Fixes #23

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
Markdig records a fenced code block's language as class="language-xxx" on the
<code> element. The cleanup regex stripped every class attribute, so the only
record of the language was destroyed and a round-trip turned ```csharp back
into a bare ```.

The HTML to Markdown direction already reads the language-* class correctly, so
the fix is limited to not throwing it away: the regex now skips class attributes
whose value starts with "language-" and continues to strip everything else,
which is what Loop needs.

Un-skips the round-trip test that documented the bug. Two tests that asserted
the old behaviour are updated: one claimed no class attribute survives at all,
and the other expected the language to be lost. Non-language stripping is now
covered by a custom container, which renders <div class="warning">.

Fixes #34

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
Pretty-printed HTML puts a newline between block elements, and that newline is
its own text node. It was collapsed to a single space and appended, so
"<p>one</p>\n<p>two</p>" converted to "one\n\n two" with a stray leading space
on the second block.

The same whitespace-only node between inline elements is a real word separator,
so it cannot simply be discarded. A block always ends by appending a newline,
which makes the buffer a reliable signal: if the output so far ends with a
newline the whitespace is layout and is dropped, otherwise it is kept as a
single space.

Fixes #33

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
List items were emitted with a trailing newline but no blank line, so
"<ul>...</ul><p>After the list.</p>" produced "- one\n- two\nAfter the list.".
Re-parsing that Markdown makes the paragraph a lazy continuation of the last
list item, so pasting a list followed by any other block silently merged them.

A nested list has to stay attached, since it is rendered as further indented
items of the same list, so the blank line is skipped when the list has a list
item ancestor.

Fixes #32

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b721ad5f-9add-4f55-b9c8-7e6d7b0fb3e7
# Conflicts:
#	README.md
@trsdn trsdn changed the title Separate a list from the block that follows it Fix the Markdown/HTML conversion bugs and add a test project Aug 28, 2026
@trsdn
trsdn changed the base branch from fix/33-whitespace-between-blocks to master August 28, 2026 21:04
@trsdn
trsdn merged commit 7a82447 into master Aug 28, 2026
1 check passed
@trsdn
trsdn deleted the fix/32-blank-line-after-list branch August 28, 2026 21:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment