Why your formatting disappears, and how to make it permanent

You select a bold sentence in a PDF, a Word document or a web page, paste it into a message box, and the bold is gone. Nothing crashed. Bold was never in the text — it was attached to the text, and the attachment was thrown away at the boundary. This page explains exactly where formatting lives, why every paste loses it, and the two techniques that make it survive.

The one sentence that explains everything

Formatting is not in the text. It is an instruction stored next to the text, and it only survives if the place you paste into understands the same instruction.

Take the phrase ship it. A web page stores it as <b>ship it</b>. Microsoft Word stores seven characters with a bold font attribute attached to the run. Discord stores it as **ship it**. Those are three different files, but in all of them the text is the same seven characters — nothing that you or I could read tells one apart from the other.

Now copy from the browser and paste into Discord. The letters arrive perfectly. The bold does not, because Discord received an HTML tag and looked for asterisks. The tag was not rejected and not stripped by a filter. It was simply never read.

Everything else here follows from that asymmetry: the characters are guaranteed to copy; the styling is not. Any method that keeps the styling must stop sending an instruction the recipient may not be able to read.

What actually happens when you press copy

  1. The source serialises the selection. It converts what it has into a portable form: a browser usually offers two flavours: text/plain, the characters alone, and text/html, the characters plus their tags. A word processor may offer text/plain and an RTF flavour, and some applications offer plain text only.
  2. The clipboard holds all of them. There is no negotiation and no handshaking. The clipboard is a slot holding several representations of the same selection, and the destination picks one it can parse.
  3. The destination parses only what it knows. A rich-text field that understands HTML keeps the <b>. Discord scans for **. Telegram runs MarkdownV2, which wants its own arrangement of asterisks and underscores. A plain-text box takes text/plain and stops. Anything the parser has no rule for is not refused — it is never examined.
  4. The styling is discarded silently. No error, no warning, no fallback. The letters come through untouched and the instruction does not, which is why it always looks as though the source lost the formatting when the destination threw it away.

This is also why styled text does not round-trip. Copy bold text back out of a Discord message and you get the asterisks again, because there the bold is a property of the message, not of the characters you selected.

The four ways formatting is stored

Everything above reduces to four representations. Three of them carry the styling separately from the text; only the fourth bakes it into the characters themselves.

How each representation is stored, what understands it, and how far it travels.
How it is storedWhere it worksHow far it travels
HTML tags — <b>, <strong>, <i>, <u>Browsers, web mail, CMS editors, some rich-text fieldsPartly. HTML-aware editors and rich-text fields that parse the same flavour. Dead in plain-text boxes, in apps with their own dialect, and in every code editor.
CSS — font-weight: 700The page onlyNever. It belongs to the stylesheet, not to the text, and the clipboard has nothing to carry.
Markdown and app dialects — **bold**, __underline__Only inside the application that implements that exact dialectOne app, one version. Pasting into a different app shows you the raw markers, or nothing at all.
Unicode styled letters — 𝐚, 𝗮Anywhere text is rendered, in any language, in any appEverywhere. The styling is the character, so there is nothing left to interpret or drop.

The dialects disagree in ways that produce silent failures. In Discord, __bold__ is an underline. In Telegram, one asterisk is bold but two asterisks under MarkdownV2 render literally or fail outright. WhatsApp will not bold mid*word*bold. Every one of these claims to be Markdown; no two of them parse the same string.

That is the general problem. Emphasis delimiters are boundary-sensitive in nearly every dialect, so pre**fix** is quietly ignored exactly where a reader expects it to work. Syntax that looks universal is in fact specific, and it fails silently.

Technique 1: change the characters

The first technique removes the instruction instead of trying to transmit it. You replace each letter with the styled letterform of that letter. The plain a at U+0061 becomes 𝐚 at U+1D41A MATHEMATICAL BOLD SMALL A. The glyph is different; the letter is the same to a reader. There is no tag, no parser and no dialect, so there is nothing for a destination application to discard.

𝐁𝐨𝐥𝐝Bold
𝐵𝑜𝑙𝑑Italic
𝑩𝒐𝒍𝒅Bold italic
ℬℴ𝓁𝒹Script
𝓑𝓸𝓵𝓭Bold script
𝔅𝔬𝔩𝔡Fraktur
𝔹𝕠𝕝𝕕Double struck
𝗕𝗼𝗹𝗱Sans bold
𝙱𝚘𝚕𝚍Monospace
BoldFullwidth

That is the whole mechanism, and the reason these styles survive. Copy one out of a browser, paste it into a message box, forward it, screenshot it, re-type it on a phone, upload it to a form with no rich-text mode: it arrives as written, because there was never a second thing that could be lost.

The letterforms come from five areas of Unicode. Mathematical Alphanumeric Symbols at U+1D400 supplies bold, italic, bold italic, script, bold script, fraktur, double-struck, sans, sans italic and monospace. Letterlike Symbols around U+2100 supplies the script and fraktur capitals the mathematical block has no room for, including the italic h. Fullwidth Forms at U+FF01–U+FF5E widen every ASCII letter. Enclosed Alphanumerics at U+2460 give circled, squared and parenthesised letters. Modifier letters supply small capitals, superscripts and subscripts.

The one real cost: these are lookalikes, not the original letters. A styled 𝐬 is a different code point from s, so search, spell-check, find-in-page, OCR and text-to-speech all treat it as an unrelated symbol. Use it for display — names, headings, bios, quotes, announcements. Do not use it where a machine has to read it back: alt text, code, file paths, URLs, or any username you want people to be able to search for.

Technique 2: combining marks, for what substitution cannot do

There is no underline character in ASCII and no underline alphabet in Unicode. That is not an oversight: an underline is not a letter shape, it is a line drawn over one. Unicode models it the way it models every other non-letter decoration — with a mark that attaches to the character before it.

The mark is U+0332 COMBINING LOW LINE, appended after every character. One mark at the end of a run draws a line under the final glyph only. Repeating it after each character paints a continuous line. The same mechanism gives you U+0336 COMBINING LONG STROKE OVERLAY for a strikethrough, U+0330 COMBINING TILDE BELOW for a wavy line, and U+0333 COMBINING DOUBLE LOW LINE for a double underline.

P̲e̲r̲s̲i̲s̲t̲e̲n̲c̲e̲U+0332 underline
d̶i̶s̶c̶a̶r̶d̶e̶d̶U+0336 strike
m̰a̰y̰ ̰n̰o̰t̰ ̰r̰ḛn̰d̰ḛr̰U+0330 wavy
S̲h̲i̲p̲ ̲i̲t̲U+0332 on Unicode
Persistence̲One mark only
𝐁𝐨𝐥𝐝Bold needs no mark

Combining marks are the least portable output in the whole system. A combining character has no width of its own; the text engine must position it relative to the previous glyph, and plenty of them do not. Older Android WebViews drop them. Game clients filter them out. Some mobile keyboards strip them as you type. A renderer that ignores zero-advance marks leaves the text looking entirely unstyled, with no sign that anything was lost. For an unknown audience, prefer a standalone letterform from U+1D400 over anything from the U+0300–U+036F block.

Why some characters break, and how generators avoid it

Substitution looks mechanical until you look at the target block. Mathematical Alphanumeric Symbols is not a clean alphabet repeated a dozen times. It is full of holes, and the holes are not random — Unicode left them unassigned on purpose, because those characters already existed elsewhere under different names.

A generator that computes base + offset will happily emit one of these holes. Here are the ones that matter most:

Unassigned code points in the mathematical block, and where those characters actually live.
Code pointWhat the offset arithmetic gives youWhere that character really is
U+1D455italic small hUnassigned. The italic h used in mathematics is U+210E, a Letterlike Symbols character originally named PLANCK CONSTANT.
U+1D506fraktur capital CUnassigned. The fraktur capital C lives at U+212D as BLACK-LETTER CAPITAL C.
U+1D50Bfraktur capital HUnassigned. It lives at U+210C.
U+1D50Cfraktur capital IUnassigned. It lives at U+2111.
U+1D515fraktur capital RUnassigned. It lives at U+211C.
U+1D51Dfraktur capital ZUnassigned. It lives at U+2128.

The gaps are almost all capitals. In most of these styles the lowercase half is solid, so a fraktur small h is always safe and a fraktur capital H never is. The exception is italic, whose only hole is a lowercase h. Bold italic is the cleanest case of all: U+1D482 MATHEMATICAL BOLD ITALIC SMALL A through U+1D49B MATHEMATICAL BOLD ITALIC SMALL Z is contiguous, h included.

Holes per style, counted against the Unicode Character Database.
StyleLowercase rangeUnassigned in that range
BoldU+1D41A–U+1D433None
ItalicU+1D44E–U+1D467One: h at U+1D455
Bold italicU+1D482–U+1D49BNone
ScriptU+1D4B6–U+1D4CFThree: e, g, o — all three live in Letterlike Symbols
Bold scriptU+1D4F0–U+1D509One: w at U+1D506, with no Letterlike equivalent
FrakturU+1D51E–U+1D537None
Double struckU+1D552–U+1D56BNone
Sans, sans bold, sans italic, sans bold italic, monospaceU+1D5BA–U+1D6A3None in any of the five

This matters because an unassigned code point does not fail consistently. A device with no glyph for it shows a hollow box. A device that substitutes silently shows a different letter. A device that drops it shows plain text. Three devices, three results from the same string — worse than a consistent failure, because you cannot predict it.

The fix is unglamorous and not optional: read the character name straight out of the Unicode Character Database and refuse to emit any code point whose name lookup fails, substituting the Letterlike Symbols character where one exists. An offset table with no name check will eventually hand somebody U+1D455, and that bug surfaces on their device, not yours.

The cost of persistent formatting

Persistent formatting is not free, and the bill arrives at the worst possible moment: when you hit a character limit.

Every code point above U+FFFF is stored as a surrogate pair: one visible glyph, two code units, two counted characters. Every style from the Mathematical Alphanumeric Symbols block sits far above that line, so all of them double. Persistence is eleven letters; in mathematical bold it costs twenty-two. That is why a forty-character bio gets rejected for being over an eighty-character limit even though it looks like forty on screen.

The counter above the message box shows the real number, because the limit is enforced on what you send rather than on what you see. So the practical rule is: when a limit is close, choose a style that lives below U+FFFF and costs one character per letter.

The single-character styles. Fullwidth forms (U+FF01–U+FF5E), circled and squared letters (U+2460–U+24FF), small capitals, superscripts and subscripts, and the script and fraktur characters borrowed from Letterlike Symbols all sit below U+2140, so each costs one character. Only the lookalike letters from the mathematical block double up. If a limit matters, use a shape-based style rather than a face-based one.

The portable checklist

U+0000–U+007F is universal; nothing above it is guaranteed by the standard, only by statistics. Latin-1 Supplement and Latin Extended-A are effectively universal in practice. Everything above U+2000 is where coverage stops being a given and becomes a per-device question.

Frequently asked questions

Why does my formatting disappear when I copy and paste it?

Because formatting was never part of the text. It is an instruction stored alongside the characters — usually HTML tags, a font attribute, or an application-specific markup dialect. When you copy, the clipboard carries the characters plus whatever flavour of markup the source offers.

The destination reads only the flavours it understands and discards the rest without warning, so the styling is dropped even though the text is identical. This is not a bug in the source application: the instruction simply had nowhere to live on the other side.

How do I make text formatting permanent?

Stop sending markup and send characters instead. The main technique is character substitution: replace each letter with its styled Unicode lookalike, so a becomes the mathematical bold small a at U+1D41A. There is no instruction to lose, so the styling survives every paste, forward and re-upload.

For effects substitution cannot express — underline, strikethrough, a wavy line — append a combining mark after each character instead. U+0332 COMBINING LOW LINE draws an underline. Combining marks are less portable, and are the first thing to break on older devices.

How do I make text bold without Markdown?

Replace every letter with its bold Unicode letterform. Mathematical Bold runs from U+1D400 for the capitals and U+1D41A for the lowercase letters, with no unassigned gaps in either half. U+1D41A is MATHEMATICAL BOLD SMALL A.

If you are counting characters, remember that every one of these letters sits above U+FFFF and is stored as a surrogate pair, so each one costs two characters against a length limit. Sans bold, fullwidth and circled letters do not have that problem.

How do I underline text on a website that has no formatting options?

There is no underline character in ASCII and no underline alphabet in Unicode, because an underline is a line drawn over a letter rather than a letter shape. It is drawn instead with combining marks: U+0332 COMBINING LOW LINE appended after every character.

One mark at the end of a run underlines only the final glyph, so the mark has to be repeated after each character to produce a continuous line. Combining marks are the least portable output in the system, so test on the devices you actually care about.

What is a combining character, and why does Unicode underline break on some devices?

A combining character is a zero-width mark that modifies the glyph before it instead of standing on its own. Rendering one requires the font and the text engine to position it relative to the previous character, and not every renderer does that.

Older Android WebViews, many game clients and some mobile keyboards strip combining marks or drop them entirely. A renderer that ignores zero-advance marks leaves the text looking completely plain. That is why standalone Unicode bold, which is a letterform in its own right, is safe everywhere and combining-mark underline is not.

Why does Unicode text count as more characters than I typed?

Because everything above U+FFFF is stored as a surrogate pair: one visible glyph, two code units, two counted characters. All the Mathematical Alphanumeric Symbols styles are astral, so an eleven-letter word in mathematical bold costs twenty-two characters.

This is why a forty-character bio gets rejected for exceeding an eighty-character limit. Styles that live below U+FFFF avoid the doubling entirely: fullwidth forms, circled and squared letters, small caps, superscripts, subscripts, and the script and fraktur capitals in Letterlike Symbols.

Part of the free Unicode Formatter on aczh.dev — no sign-up, no ads, and no text ever leaves your browser. Every code point on this page was checked against the Unicode Character Database, version 15.0, in October 2026.