Markdown and HTML to Unicode: convert a whole document

A document can be perfectly formatted and still arrive as an unstyled wall of text the moment you paste it into a chat, a bio or a comment box. Paste the source HTML or Markdown into this converter instead, and every heading, bold run, list, link and table comes back out as styled plain text that reads the same everywhere.

Why convert HTML or Markdown instead of just pasting it

Formatting on the web is not stored in the words. It is stored alongside them, in tags and attributes that a receiving application is free to ignore. Select ship it by Friday on a page and the clipboard takes two things: the letters, and an instruction. The instruction is the part that vanishes.

That instruction has a different name depending on where the document came from. On a web page it is markup such as <strong>. In a README it is **ship it**. In a Word file it is a run-property flag that never leaves the document at all, and in Google Docs it is a span of styled HTML. None of these are text. A chat app has no stylesheet, no table layout and no reason to run markup from a stranger, so it keeps the characters it recognises and discards the rest.

For one sentence the fix is trivial: retype it in the platform's own syntax. For a document it is not. A pasted README loses its heading levels and its code fences. A pasted release note loses its comparison table and collapses the rows into a single paragraph. A Word paste arrives with line breaks that no target will reproduce and indentation made of tabs that render as random gaps. Nothing is strictly lost — the structure is simply no longer anywhere you can see it.

Conversion rewrites the document into characters instead of instructions. A heading becomes a styled heading line. Bold becomes the bold version of each letter. A table becomes box-drawing characters whose columns line up. A link becomes a visible address. Nothing in the result depends on the receiving app cooperating, which is exactly why the same text survives a Discord channel, a 160-character SMS field, a profile bio, and someone retyping it from a screenshot.

How to convert a document in four steps

  1. Copy the formatted content. Select the finished text wherever it lives. Pressing Ctrl+A then Ctrl+C inside a Word document, a Google Doc, a Notion page or a rendered GitHub README puts HTML on the clipboard, which is exactly what the converter wants. In a browser, the page source or the Inspect panel gives you the markup directly if you would rather paste the tags than the rendered result.
  2. Open the tool. Go to the Unicode Formatter and put the cursor in the editor. There is no import step, no file picker and no upload — paste and continue.
  3. Paste into the editor and adjust. Anything that looks like markup is detected on the spot. HTML and rich text from Word, Google Docs, Notion or any browser editor is parsed, translated and inserted, and a toast confirms what was found. Markdown is handled identically, with no mode to pick first. After conversion you can still change how headings are rendered, how links are written and how colour highlights are drawn.
  4. Pick the target and copy. Choose the output profile that matches where the text is going, hit Copy, and paste it in. The character counter next to the output already knows the budget for that destination, so it will warn you before the message is rejected.

Everything it translates

The same block model drives both input formats, so a heading written as # Heading and a heading written as <h1>Heading</h1> end up in exactly the same place in the output.

Each document structure, how you write it, and what the converter produces.
InputHow you write itWhat you get out
Heading# Heading or <h1> to <h3>A weighted heading line at the level you chose, or plain text if you set headings to plain
Bold**bold** or <strong>Bold Unicode letters, or the destination's own bold syntax
Italic*italic* or <em>Italic Unicode letters, readable inside a paragraph of normal text
Strikethrough~~gone~~ or <s>Struck-through Unicode, degraded to a marker on targets with no strike
Inline code`code` or <code>Monospace Unicode that stays legible mid-sentence
Blockquote> quoted or <blockquote>An indented quote with a visible marker on every line
Bulleted list- item or <ul><li>One bullet per item, with indentation preserved for nested levels
Numbered list1. item or <ol><li>Items renumbered in the target's own convention so they stay in order
Link[label](url) or <a href>Address shown, address in brackets, or link text only — all three are selectable
TableA pipe table, or any <table> from HTMLA box-drawn grid with columns aligned by display width

Underline, colour and table borders are options, not surprises. Underline is translated too, drawn with combining characters on the Unicode profile and with the platform's own syntax elsewhere. Colour highlights can be rendered as a colour block or written as [! markers. Table borders are yours to pick from several styles, grid lines can be switched off for a lighter result, and the header row can be repeated or dropped. Every setting can be changed after the paste, so you can convert first and tune afterwards.

Where the output goes: choosing a target

Seven output profiles, each written for a different destination. A style a platform cannot express degrades predictably instead of disappearing — underline on Telegram becomes bold, for instance — and the note on each pill tells you what will happen before you copy.

Output profiles, what each is for, and how it behaves.
ProfileBest forNotes
Unicode (Universal)Platforms with no formatting support at allReal Unicode glyphs rather than markup, so they render in a bio box, an SMS field or a plain comment thread
DiscordServer messages, forum posts and descriptionsDiscord markdown: bold, underline, italic, strike, spoiler, code and quotes
TelegramChannel posts and group messagesTelegram Markdown: bold, italic and code. Underline and strike fall back to bold
SlackChannel messages someone will read as mrkdwnSlack syntax: *bold*, _italic_, ~strike~, backtick code
WhatsAppStatus updates, group announcements, order messages*bold*, _italic_, ~strike~, triple-backtick monospace
MarkdownREADMEs, issues, release notes, documentationCommonMark / GitHub flavoured Markdown, with the blank line between blocks the spec requires
Plain TextBios, alt text, meta descriptions, SEO snippetsAll formatting stripped. The point of this one is that nothing has to render at all

Beyond conversion: tables, banners and ASCII art

The converter sits on the same page as the rest of the toolkit, so the pieces that normally need a second tool are one click away. There is a Unicode table generator that accepts CSV, TSV, a Markdown pipe table or an HTML table and aligns every column by display width rather than character count — which is the only way a column of Japanese text or a row of emoji can line up beside a column of Latin letters. There is a block-letter banner generator with a 5×5 font and 1×, 2× or 3× scale for large headings. ASCII art converts into shaded Unicode with the shade ramp intact, dividers come in a range of weights, and the reverse decoder turns styled Unicode back into readable plain text.

┌──────┬──────┐Table top rule
╔══════╦══════╗Heavy box table
├──────┼──────┤Grid junction
█▀▄█5×5 block font
░▒▓█Shade ramp
░░██░░Shaded ASCII art
▄▀▄▀Banner scale
─────Dividers

Character limits by destination

Styled Unicode is not free. Most bold and italic letters live above U+FFFF and are stored as a surrogate pair, so a 15-character word in bold can cost close to 30 against a platform's budget. The tool offers a budget per target, so you can watch the real number rather than the visible one.

Per-message and per-field character budgets for the platforms this converter targets.
PlatformLimitNote
Discord2,000 — 4,000 with NitroStyling costs the most here, so prefer single-character styles for long messages
Telegram4,096Measured after entities are parsed, so visible markup can hide extra characters
X / Twitter280Counted in weighted length; every link costs the same 23 characters however long it is
Instagram2,200 caption · 150 bioThe bio is the binding constraint, and styled characters eat it fastest
LinkedIn3,000The first 140 characters show before the “see more” cut-off, so the opening line carries the weight
WhatsApp65,536Effectively unlimited for a single message
Reddit40,000 · title 300The title is the tight field; the body is not the problem
TikTok2,200Caption length. The bio is a separate and much smaller field
Threads500The smallest posting budget here, and easy to overrun with a table
Bluesky300Short by design; plan the structure, not the styling
Mastodon500Default limit; a self-hosted instance can raise it
SMS160The GSM-7 budget. Anything outside that alphabet is split into 70-character segments

Your document never leaves your browser

There is no server call at all. Parsing, conversion and rendering all happen in the page you are looking at. Nothing is uploaded, nothing is stored, and no part of your document is sent anywhere.

That guarantee is stronger than a promise, because it is architectural. The conversion engine ships with its own HTML tokenizer, its own block model and its own Markdown parser instead of calling the browser's DOMParser. That choice buys three things at once.

First, pasted markup cannot inject anything into the output. Because the parser is code we control, an unexpected tag in a pasted document is treated as data to be interpreted or ignored; it can never become an element that escapes into what you copy. Second, the same code path runs in the browser and under Node, which is what makes the engine's assertion suite possible at all — the conversion logic can be tested without a DOM. Third, the result is deterministic: the same document produces the same output everywhere instead of depending on which browser engine happens to be parsing it.

The practical upshot for documents is simple. You can paste an unreleased changelog, a draft press release, a table of unreleased pricing or a client's private notes without thinking about it, because at no point is there a network request carrying any of it.

Frequently asked questions

Does this converter upload my document anywhere?

No. The converter runs entirely in the page: there is no server call at all, no upload endpoint and nothing measuring your text. Pasting markup, parsing it and rendering the styled output all happen in your browser's own memory.

An unpublished release note, a draft announcement or a private table stays on the machine you pasted it on.

Can I paste raw HTML straight from a web page?

Yes, and that is the normal way to use it. Select a paragraph or a whole article in any browser and paste it into the editor, or open the page source and copy the tags directly.

Rich text pasted from Word, Google Docs, Notion or any browser editor is detected the same way and converted automatically, with a toast confirming it.

Does it read Markdown as well as HTML?

Yes, in the CommonMark and GitHub style, including headings, bold and italic runs, fenced and inline code, blockquotes, both list types, links and pipe tables.

There is no mode to choose before you paste — the parser looks at what arrives and decides, so the same editor handles a rendered README and a page of raw tags.

What happens to a table I paste from Word or Google Docs?

It comes out as a real box-drawn grid rather than a run of tabs and spaces, so the columns stay aligned when the text is pasted somewhere that uses a monospace font.

Border style is yours to pick, grid lines can be turned off, and the header row can be repeated or dropped. Columns are aligned by display width, so CJK characters and emoji line up instead of pushing the table out of shape.

Why do the links in my output look different from the originals?

Because a clickable link cannot survive as a clickable link in plain text. The destination only receives characters.

There are three output styles to choose from: show the full address, put it in brackets after the label, or keep only the link text. A visible URL matters in a bio or a printed table but clutters a chat message.

Can I convert styled Unicode text back into plain text?

Yes. Paste the styled output into the Decode tool on the same page and it returns readable plain text with the styling characters removed.

The reverse path is what you want for alt text, meta descriptions, bios and search snippets. The Plain Text output profile does the same job in one step, straight from the editor.

Part of the free Unicode Formatter on aczh.dev — no sign-up, no ads, and no text ever leaves your browser. Converter behaviour verified against the shipped engine in October 2026.