Changelog
Add $html_parser (#10569): tolerant HTML parser reusing the $xml_node DOM + CSS selectors
2026-05-30 18:47 UTC · claude (#13505)
New $html_parser:parse(text) returns the #document $xml_node for tag-soup HTML, reusing the $xml_node DOM so :select/:child/:text/CSS selectors work for free.
Single pcre_match tokenizer pass + stack tree builder. Tolerant: void elements stay childless (br/img/meta/...), script/style kept as raw text, <li>/<p>/table-cell auto-close, unquoted and boolean attrs, case-folded tag/attr names, comments/doctype/PI skipped, multi-root fragments preserved.
Rule tables on #10569: .void_tags, .raw_tags, .closes_p, .autoclose. Helpers :_parse_tag (tolerant attrs), :_find_tag_end.
Known limits: simple <tag[^>]*> open pattern mis-splits a > inside a quoted attr value (fancy pattern blew PCRE JIT stack); only the 5 core entities decoded; no adoption-agency for mis-nested formatting tags.
Verified on live google.com (extracted <title> "Google", zero void mis-nesting vs 15 in the XML parser). Tests test_tag_soup and test_html_attrs_and_roots pass.