HTML to Text — Strip Tags and Extract Content
Extract readable text from HTML with paragraph and link options
Runs in your browser · nothing is uploaded
Extracts text without executing or displaying the HTML. Scripts, styles, hidden and aria-hidden content are omitted. CSS visibility is not computed, and whitespace in pre blocks is normalized.
How to use
- Paste HTML from a page source, editor or email.
- Choose whether to keep paragraph breaks.
- Enable link destinations if you want each anchor followed by its address.
- Copy the plain-text result into notes or documents.
Extraction rules
The tool uses the browser HTML parser rather than removing tags with a regular expression. HTML entities are decoded once, and Korean, Chinese and emoji remain intact. Paragraphs, divisions, headings, list items and table rows introduce line boundaries. Table cells are separated with whitespace. Single-line mode collapses all whitespace to individual spaces.
Scripts, styles, document titles, templates, noscript blocks and embedded content are omitted. Elements with hidden or aria-hidden=“true” are skipped. Image alternative text and form field values are not extracted separately. The tool does not reproduce CSS visibility or visual layout. When enabled, link destinations are appended as ordinary text and cannot execute.
Examples
<p>한글 & 中文</p><p>😀</p> becomes two text paragraphs when paragraph mode is enabled. <a href="/guide">Guide</a> becomes Guide (/guide) when destinations are included.
Limits and privacy
Input is limited to one million characters and DOM traversal to 128 levels. Malformed HTML follows the browser parser recovery rules. Original whitespace inside pre elements is not preserved. All processing is local; input and results are not uploaded.
FAQ
Will scripts run or external images load?
No. Input is parsed in an inert template that is never attached to the page. The result is displayed only as text, and script and style content is omitted.
Does the result exactly match visible page text?
No. Elements marked hidden or aria-hidden are omitted, but CSS styles, external stylesheets and layout are not evaluated.
Is original whitespace preserved?
No. Paragraph mode retains block and br boundaries, while repeated spaces and whitespace inside pre elements are normalized.