mdurl.encode() treats '[' and ']' as unsafe and percent-encodes them,
so IPv6 address literals such as http://[2001:db8::1]/ were emitted as
http://%5B2001:db8::1%5D/. Encode URL components individually and let
mdurl.format() add the IPv6 brackets back as URL syntax delimiters.
A backslash followed by a space inside a link destination is a literal
backslash (a space is not ASCII punctuation, so it is not escaped), and
the space ends the destination. markdown-it instead broke at the
backslash, dropping it from the destination and failing to parse the
link.
Repro:
md.render('[a](/url\\ )')
Expected (matches the CommonMark 0.31.2 reference renderer):
<p><a href="/url%5C">a</a></p>
Actual on master:
<p>[a](/url\ )</p>
In lib/helpers/parse_link_destination.mjs the unenclosed-destination
loop, on seeing `\` followed by a space, called `break`, which excludes
the backslash from the destination and leaves `\ )` unconsumed so the
link fails. Advance past the backslash instead (`pos++; continue`) so it
is kept literal and the following space triggers the normal terminator.
The existing commonmark/CommonMark#493 case `[a](a\ b)` is unaffected:
content after the space still prevents the link from parsing, matching
the reference.
A backslash followed by a space was baking that space into the escape
token, so the newline rule saw only one trailing space and emitted a
soft break instead of the mandated hard line break.
Per CommonMark 0.31.2 section 6.1, a backslash before a space is a
literal backslash and the space is not consumed. Leaving the space in
the text stream lets section 6.7 two-space hard-break detection fire.
Input `a\ \nb` (backslash + two trailing spaces) now renders
`<p>a\<br>\nb</p>` instead of `<p>a\ \nb</p>`.
HTML blocks of types 1-5 (e.g. `<!--` comments) must continue until
their closing sequence (`-->` for comments), per CommonMark. Only
types 6 and 7 end on a blank line.
The block scanner broke out of the continuation loop whenever a line
was outdented below the current block indent. Inside a container such
as a list item, a blank line is outdented (its sCount is 0), so a
comment block was terminated early instead of running to `-->`.
Skip that break for outdented blank lines unless the block type
actually ends on a blank line, so comment blocks no longer depend on
indentation. Adds regression fixtures for issue #1144.
Co-authored-by: ATOM00blue <219721791+ATOM00blue@users.noreply.github.com>
The spec update changes these things:
* It simplifies the HTML regex so that `<!-- a -- b -->` is an HTML
comment. HTML5 reports this as an error, but still parses it.
* It changes the set of known HTML block elements to match HTML5, adding
`search` and removing `source`.
* It adds Unicode Symbols to the set of punctuation characters that are
used to evaluate flankingness.
This commit also changes the declaration HTML regex to match lowercase,
even though that change was technically made in spec version 0.30.
Spec is not clear on how to handle this. Three variations exist:
```
$ echo '' | /home/user/commonmark.js/bin/commonmark
<p><img src="image.png" alt="text <textarea> text" /></p>
$ echo '' | /home/user/cmark/build/src/cmark
<p><img src="image.png" alt="text <textarea> text" /></p>
$ echo '' | /home/user/.local/bin/commonmark
<p><img src="image.png" alt="text text" /></p>
```
Prior to this commit:
- when HTML tags are enabled, tags were removed (as in Haskell version)
- when HTML tags are disabled, tags were escaped (as in C version)
After this commit:
- tags will be escaped (as in C version) regardless of HTML flag
+ render hardbreaks as newlines, same as cmark