core.html
=========
.. py:module:: core.html
Attributes
----------
.. autoapisummary::
core.html.SANE_HTML_TAGS
core.html.SANE_HTML_ATTRS
core.html.VALID_PLAINTEXT_CHARACTERS
core.html.EMPTY_LINK
core.html.sanitize_policy
Functions
---------
.. autoapisummary::
core.html.sanitize_html
core.html.sanitize_svg
core.html.html_to_text
Module Contents
---------------
.. py:data:: SANE_HTML_TAGS
.. py:data:: SANE_HTML_ATTRS
.. py:data:: VALID_PLAINTEXT_CHARACTERS
.. py:data:: EMPTY_LINK
.. py:data:: sanitize_policy
.. py:function:: sanitize_html(html: str | None) -> markupsafe.Markup
Takes the given html and strips all but a whitelisted number of tags
from it.
.. py:function:: sanitize_svg[T: str](svg: T) -> T
I couldn't find a good svg sanitiser function yet, so for now
this function will be a no-op, though it will try to detect
svg files which are harmful.
I tried to go with bleach/html5lib, but the lack of xml namespace support
makes those options a no go.
In the future we want a proper SVG sanitiser here!
.. py:function:: html_to_text(html: str, *, body_width: int = 0, ignore_images: bool = True, single_line_break: bool = True, ignore_emphasis: bool = False, ul_item_mark: str = '*', strong_mark: str = '**', emphasis_mark: str = '_') -> str
Takes the given HTML text and extracts the text from it.
The result is markdown. The driver behind it is turbohtml.