core.html ========= .. py:module:: core.html Attributes ---------- .. autoapisummary:: core.html.SANE_HTML_TAGS core.html.SANE_HTML_ATTRS core.html.VALID_PLAINTEXT_CHARACTERS core.html.EMPTY_LINK core.html.sanitize_policy Functions --------- .. autoapisummary:: core.html.sanitize_html core.html.sanitize_svg core.html.html_to_text Module Contents --------------- .. py:data:: SANE_HTML_TAGS .. py:data:: SANE_HTML_ATTRS .. py:data:: VALID_PLAINTEXT_CHARACTERS .. py:data:: EMPTY_LINK .. py:data:: sanitize_policy .. py:function:: sanitize_html(html: str | None) -> markupsafe.Markup Takes the given html and strips all but a whitelisted number of tags from it. .. py:function:: sanitize_svg[T: str](svg: T) -> T I couldn't find a good svg sanitiser function yet, so for now this function will be a no-op, though it will try to detect svg files which are harmful. I tried to go with bleach/html5lib, but the lack of xml namespace support makes those options a no go. In the future we want a proper SVG sanitiser here! .. py:function:: html_to_text(html: str, *, body_width: int = 0, ignore_images: bool = True, single_line_break: bool = True, ignore_emphasis: bool = False, ul_item_mark: str = '*', strong_mark: str = '**', emphasis_mark: str = '_') -> str Takes the given HTML text and extracts the text from it. The result is markdown. The driver behind it is turbohtml.