How binary works — the language computers speak natively#
A binary number uses only two digits — 0 and 1 — to represent every possible value. Where decimal (base‑10) has ten digits per place and each place is ten times the previous one (1s, 10s, 100s…), binary has two digits per place and each place is double the previous one (1s, 2s, 4s, 8s, 16s…). The decimal number 13, for example, is written 1101 in binary: (1 × 8) + (1 × 4) + (0 × 2) + (1 × 1).
A single binary digit is called a bit. Eight bits make one byte, and a byte is the fundamental unit computers use to represent a character of text. The byte 01100001 is decimal 97 — and in ASCII that is the letter a. Each character you type is stored, transmitted and processed as one or more bytes of binary, whether you see it on screen or not.
ASCII vs UTF-8 — where they overlap and where they diverge#
ASCII (American Standard Code for Information Interchange) was designed in the 1960s and defines 128 characters — the English alphabet (upper and lower case), digits 0‑9, punctuation, and a set of control codes. Every ASCII character fits in 7 bits, so it occupies exactly one byte with the highest bit always 0. This covers English text completely but nothing else — no accented letters, no Greek, no Cyrillic, no CJK, no emoji.
UTF-8 is the dominant encoding on the modern web (over 98 % of all web pages use it). It is backward-compatible with ASCII: the first 128 code points are identical to ASCII and still use a single byte. Characters beyond that range use two, three, or four bytes, with a clever structure that makes it self-synchronising — even if you jump into the middle of a UTF-8 stream, you can never mistake a continuation byte for a new character start. This compact variable-length design is what lets UTF-8 represent every character in the Unicode standard while remaining a drop-in replacement for ASCII for English text.
Why computers use binary instead of decimal#
At the hardware level, a computer has no concept of the digit 2, 3, 4, or 9. Every transistor, logic gate, and memory cell works in two states: on (1) and off (0). Voltage is either above a threshold or below it — there is no reliable ten-way voltage split that would be fast, cheap, and power-efficient. Binary electronics are simpler, more noise-tolerant, and consume far less energy than a decimal machine would.
That binary constraint cascades upward. The CPU operates on bits, the ALU adds binary numbers, registers hold binary words, memory addresses are binary values, and every instruction a processor executes arrives as a binary opcode. Higher-level representations — decimal, hexadecimal, octal — are human conveniences layered on top of the machine's native binary. A computer never sees the digit A; it sees 01000001.
How emoji survive in binary#
An emoji like 😊 (U+1F60A, SMILING FACE WITH SMILING EYES) lives at code point 128,522 in the Unicode standard — far beyond the 128 values ASCII can represent and beyond the 256 values a single byte can hold. UTF-8 encodes this as a 4-byte sequence: 11110000 10011111 10011000 10001010. A decoder that recognises the leading byte pattern (11110xxx) knows four bytes follow, reassembles them into code point 0x1F60A, and renders the smiley face.
This is also why mixing encodings breaks text. If a system expects ASCII and encounters a UTF-8 emoji byte sequence, it sees four unrelated bytes rather than a single character — the familiar "mojibake" of box symbols, question marks, or garbled glyphs. Modern systems agree on UTF-8 exactly so that text survives the round trip no matter which emoji or script it contains.
Reading binary output from this tool#
When you encode text to binary, each character is displayed as an 8‑bit chunk separated by spaces. A single byte per character means ASCII, and multiple bytes per character means the input needed more than one byte (accented letters, CJK, emoji). For example, encoding A gives 01000001, while encoding ñ gives 11000011 10110001 — two bytes, because the tilde-n falls outside the ASCII range.
Watching the binary grow provides an intuitive feel for storage cost. A short sentence that takes 50 bytes in ASCII might take 60 or more bytes if it contains accented characters, and an emoji-heavy message inflates bytes faster than character count suggests because each emoji takes 4 bytes. This real-time feedback is useful for understanding how encoding choices affect file size, API payloads, and database storage.