Skip to main content

WhitePages picoCTF 2019 Solution

Uncover a hidden message encoded using invisible whitespace characters in a text file.

Published: April 2, 2026Updated: September 20, 2026

Description

I stopped using color in my terminal. Decode the binary in the whitepages file.

Download the file.
bash
wget <url>/whitepages.txt

Solution

Want to try it yourself first?

The guided walkthrough reveals hints one step at a time.

Walk me through it
  1. Step 1Examine the file with xxd
    Observation
    The file is called whitepages and holds nothing but whitespace. So the data is hidden in the byte values of whitespace characters that look identical but are distinct, not in any visible text.
    The file appears to contain only whitespace. Run xxd to see the actual hex values of each byte. You will find two different whitespace characters being used - for example, regular space (0x20) and a special Unicode whitespace character.
    bash
    xxd whitepages.txt | head -20

    Expected output

    00000000: e280 8320 2020 e280 83e2 8083 e280 83e2  ...   ..........
    00000010: 8083 e280 8320 20e2 8083 20e2 8083 e280  .....  ... .....
    00000020: 8320 e280 8320 20e2 8083 e280 83e2 8083  . ...  .........
    ...
    What didn't work first

    Tried: Open whitepages.txt in a text editor or cat it to the terminal to look for hidden content.

    Both approaches render whitespace as blank space, so the file looks completely empty. Terminals and editors collapse the distinction between U+0020 and U+2003 entirely. Only a hex dump tool like xxd shows the actual byte values, which is what makes the two-character alphabet visible.

    Tried: Run strings whitepages.txt hoping to extract embedded printable text directly.

    strings filters for runs of at least 4 consecutive printable ASCII bytes. Regular space (0x20) is printable, but the em-space bytes (E2 80 83) between them are not, so the only thing strings can ever emit here is a handful of short runs of blank spaces, never readable text. The message is a binary sequence encoded in the choice of whitespace character, not raw ASCII.

    Learn more

    Unicode contains many whitespace characters beyond the regular space (U+0020). Common ones used in steganography challenges include: em space (U+2003, UTF-8: E2 80 83), en space (U+2002), thin space (U+2009), and others. All look identical in most text editors.

  2. Step 2Map whitespace characters to binary bits
    Observation
    The xxd output shows exactly two distinct byte sequences: 0x20, and the three-byte UTF-8 sequence E2 80 83 for em-space. Treat them as a two-symbol binary alphabet and group every 8 symbols into a byte.
    Identify the two different whitespace byte sequences. Assign one to bit 0 and the other to bit 1. Group every 8 bits into a byte and convert to ASCII.
    python
    python3 << 'EOF'
    with open('whitepages.txt', 'rb') as f:
        data = f.read()
    
    # Identify the two whitespace types from xxd output
    # e.g., 0x20 = space = 1, 0xe2 0x80 0x83 = em-space = 0
    bits = ''
    i = 0
    while i < len(data):
        if data[i:i+3] == b'\xe2\x80\x83':
            bits += '0'
            i += 3
        elif data[i] == 0x20:
            bits += '1'
            i += 1
        else:
            i += 1
    
    # Convert bits to ASCII
    result = ''
    for j in range(0, len(bits) - 7, 8):
        byte = int(bits[j:j+8], 2)
        result += chr(byte)
    print(result)
    EOF
    What didn't work first

    Tried: Swap the bit assignments so that em-space maps to 1 and regular space maps to 0, then decode the result.

    Reversing the bit mapping gives a garbled sequence of non-printable bytes rather than readable ASCII. Work out the right assignment by checking which byte sequence appears more often, since the dominant character is usually 0, or just try both and see which produces valid ASCII. Without confirming against real output, the decoded bytes will be wrong.

    Tried: Use the SNOW steganography tool (stegsnow) to decode the file, since SNOW is a known whitespace steganography tool.

    SNOW encodes data using only trailing tabs and spaces at the ends of lines, and it needs a file format with newlines. This challenge encodes with Unicode em-space versus regular space throughout the body of the file, an entirely different scheme. Run stegsnow on it and you get nothing or an error, because the format is not what SNOW expects.

    Learn more

    This technique is called whitespace steganography. The SNOW tool and the Whitespace programming language both use similar concepts of encoding information in invisible characters. It is effective against casual inspection but immediately visible under hex analysis.

    The key insight is that while the text looks blank, it encodes a binary string where each character of the message is represented by 8 bits of whitespace characters.

Interactive tools
  • StegallDrop any file and Stegall runs every applicable steg technique in parallel: LSB sweeps, bit planes, spectrograms, polyglot carving, metadata, whitespace decode, and a 6-layer base/ROT/XOR/zlib cascade. Recursively unpacks results and surfaces flag matches.
  • Hex ViewerView text or raw hex bytes as a xxd-style hex dump with byte offset, hex columns, and ASCII sidebar. Highlights printable characters and null bytes.
  • Strings ExtractorPull printable text from any binary, library, or image. ASCII and UTF-16 detection, configurable minimum length, flag-like highlight, no command line needed.

Flag

Reveal flag

picoCTF{not_all_spaces_are_created_equal_...}

Per-instance flag. The prefix picoCTF{not_all_spaces_are_created_equal_ is stable across instances; the hex suffix after it is generated per instance, so decode the whitespace on your own copy to get the full value.

Key takeaway

Unicode defines hundreds of whitespace code points that render identically to an ordinary space in any editor or terminal but differ at the byte level. A file containing just two distinct invisible characters is a binary channel: assign one symbol to 0 and the other to 1, group into 8-bit bytes, and the hidden message falls out. The same trick appears in covert communication research, e-book watermarking, and source-code exfiltration where tabs and spaces are swapped interchangeably.

Related reading

Useful tools for Forensics

Where to go next