Base64, URL Encoding, and HTML Entities: What the "Encode" Tools Do
"Encode" tools are scattered across every developer toolkit, and to a non-developer they look interchangeable. They are not — each one solves a specific, different problem. This article explains the three you will actually run into — Base64, URL encoding, and HTML entities — in plain language, so you know which button to press and why.
Base64: turning binary into safe text
Some systems can only handle plain text, but the data you want to send — an image, a file, a chunk of binary — is not text. Base64 solves this by representing any binary data as a string of letters, numbers, and a couple of symbols. It is not encryption and it is not compression; it is just a translation into an alphabet that can survive being pasted into a text field or an email.
You see Base64 constantly: images embedded directly in emails or web pages, files attached to API payloads, and data tucked into JSON Web Tokens. The giveaway is a string that ends with = or == (padding). Decoding it turns the text back into the original bytes.
The thing to remember: Base64 makes data about a third larger, because it maps three bytes onto four characters. It is a compatibility tool, not a space-saver.
URL encoding: making text safe for a web address
URLs can only contain a limited set of characters. Spaces, &, ?, #, and non-ASCII characters are all either illegal or meaningful in a URL, so they cannot appear in a value raw. URL encoding (also called percent-encoding) replaces each problem character with a % followed by its numeric code. A space becomes %20, an ampersand becomes %26, and a Chinese character becomes a sequence like %E4%BD%A0.
You have seen this in every browser address bar without noticing: search for "café menu" and the URL becomes ?q=caf%C3%A9%20menu. Decoding reverses it, turning the percent codes back into the human-readable text.
HTML entities: escaping text so it renders correctly
HTML uses characters like <, >, and & as syntax. If you want to actually display those characters on a page — say, to show a snippet of code — you cannot type them raw, because the browser would interpret them as markup. HTML entities solve this by giving every special character a named or numeric code. < displays as a literal less-than sign, & as an ampersand.
This matters for two reasons beyond cosmetics. First, if you paste raw user text into HTML without escaping it, you create the conditions for a cross-site-scripting (XSS) attack. Second, escaping is why code samples on a page show you the actual <div> instead of rendering it as an invisible element.
Which one do I need?
The choice is almost always obvious once you know the destination:
- Sending a file or binary through a text-only channel → Base64.
- Putting a value with spaces or symbols into a URL or query string → URL encoding.
- Displaying
<,>, or&on a web page → HTML entities.
Each of these is a one-way-ish, reversible, deterministic transform. None of them hide data (that is encryption's job), and none of them shrink it (that is compression's job). Once you know what problem each one solves, the "encode/decode" section of a toolkit stops looking like magic and starts looking like three clearly labeled tools.