SadServers
  • Scenarios
  • Labs
    All Labs Linux & Bash Web Servers Databases Data Processing Docker Kubernetes CI/CD Infrastructure As Code Observability Tooling / Applications
  • Dashboard
  • Solutions
    For Individuals For Businesses
  • Ranking
  • Newsletter
  • FAQ
  • Documentation
    Support Pro Accounts Pro+ Accounts Business Accounts Gift API CLI/TUI Privacy Troubleshooting Interviews
  • Blog
  • Pricing
  • Gift
    Gift Purchase Gift Redeem
  • About
Log In - Sign Up

Check out our new interface!

(Beta for registered users; some features still in progress)

SadServers Linux & DevOps Troubleshooting Scenarios

Linux & Bash

  • - Linux commands, Bash scripting
  • - Systemd
  • - Networking, DNS
  • - Storage
  • - SSH, Firewall
  • - Libraries
  • - Cron and more...

Web Servers

  • - Nginx
  • - Apache
  • - HAProxy
  • - Caddy
  • - Gunicorn
  • - uWSGI
  • - HTTPS/TLS

Databases

  • - PostgreSQL
  • - MySQL
  • - SQLite
  • - Redis
  • - ClickHouse
  • - MongoDB
  • - etcd

Data Processing

  • - CSV
  • - JSON
  • - SQL queries

Docker

  • - Building images
  • - Multi-stage builds
  • - Volumes
  • - Networks
  • - Docker Compose
  • - Podman

Kubernetes

  • - kubectl
  • - Helm
  • - K8S Roles & Permissions
  • - Services
  • - Namespaces
  • - Deployments, StatefulSets
  • - ConfigMaps, Secrets

Infrastructure As Code

  • - Ansible
  • - Terraform

Observability

  • - ELK
  • - Prometheus

Tooling / Applications

  • - Git
  • - Rabbitmq
  • - Envoy
  • - Vault
  • - Harbor
  • - Jenkins

Hacking

  • - Capture the Flag (CTF) Challenges
  • - Code Vulnerabilities
  • - Privilege Escalation

Languages

  • - Python
  • - Golang
  • - PHP
  • - Java
  • - Node.js
  • - C
Previous Next
Packages advent2025 ai ansible apache aws cli bash c caddy clickhouse cron csv data processing disk volumes dns docker elk envoy etcd financial finops ftp git gitea golang gunicorn hack haproxy harbor hashicorp vault helm java jenkins json kubernetes linux-other mongodb mysql nginx node.js php podman postgres prometheus python rabbitmq redis sql sqlite ssh ssl supervisord systemd terraform traefik
realistic / interviews new pro business

Easy

# Name Time Type
1 "Istanbul": Delete unattached AWS EBS volumes 15 m Do No Registration New
"Istanbul": Delete unattached AWS EBS volumes

Scenario: "Istanbul": Delete unattached AWS EBS volumes

Level: Easy

Type: Do

Access: Public

Description: Finance found orphaned EBS volumes racking up cost in this account. Your job is to find every volume that is not attached to an EC2 instance and delete it.

Do not delete volumes that are attached (in-use). Do not stop or terminate instances.

The AWS CLI is already configured against the local endpoint. Start with:
aws ec2 describe-volumes

NOTE: this scenario was done using Floci, a local AWS emulator; there are no real AWS resources.

Test: Delete every EBS volume that is not attached to an EC2 instance. Attached (in-use) volumes and running instances must remain.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can see and execute.

Time to Solve: 15 minutes.

Medium

# Name Time Type
1 "Spa": The Docker Compose healthcheck that does not heal 30 m Fix Pro New
"Spa": The Docker Compose healthcheck that does not heal

Scenario: "Spa": The Docker Compose healthcheck that does not heal

Level: Medium

Type: Fix

Access: Paid

Description: The thermal baths booking API should answer on port :5000. After last week's outage the previous engineer added a Docker HEALTHCHECK and restart: unless-stopped so the stack would recycle itself if the probe failed. The front desk is still down.

The Compose project is in /home/admin/spa. There is a short handover note in that directory.

The booking API application itself is fine when it is running: if you restart the container, GET /health comes back. You do not need to rewrite or "fix" the API service code. The real gap is recovery — when the API process inside the container dies again, nothing brings the service back on its own. Make the stack recover automatically from that failure so /health returns {"status":"ok"} without a manual restart.

The process will not usually die on its own during the exercise. To replay the incident, stop the API worker inside the container (leave the container running): docker exec spa-api sh -c 'kill "$(cat /tmp/spa.pid)"' — before a fix, /health stays down. After you add automatic recovery, the same kill should bring /health back without a manual docker restart. The recovery must still work after a host reboot (do not leave the API dead on boot).

Test: curl -s http://127.0.0.1:5000/health returns {"status":"ok"} and the spa-api container is healthy.
A one-off docker restart is not enough: after the API worker dies again inside the container, the service must recover on its own.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can see and execute.

Time to Solve: 30 minutes.

2 "Kyoto": The Gion ticket API will not start 20 m Fix New
"Kyoto": The Gion ticket API will not start

Scenario: "Kyoto": The Gion ticket API will not start

Level: Medium

Type: Fix

Access: Email

Description: The Gion ticket office API should answer on port :80. After last night's change the office never came back.

Manifests are in /home/admin/app. There is a short handover note in /home/admin/HANDOVER.txt. Wait until the Kubernetes node is Ready after boot.

GET /health must return {"status":"ok","office":"gion-tickets"}. Do not move the API off port 80.

TIP: You can use k as an alias for kubectl, and it has autocomplete enabled.

Test: curl -s http://127.0.0.1/health returns {"status":"ok","office":"gion-tickets"} and the gion-api pod in the gion namespace is Ready (1/1).

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can see and execute.

Time to Solve: 20 minutes.

3 "Dakar": The half-installed package 10 m Fix New
"Dakar": The half-installed package

Scenario: "Dakar": The half-installed package

Level: Medium

Type: Fix

Access: Email

Description: The checkout platform's local health agent stopped working after a package upgrade was interrupted. The checkout-agent binary may already be on disk, but its systemd service is not healthy.

Restore the package and service so the agent runs again and still comes back after a reboot. Do not wipe the package database under /var/lib/dpkg, and do not replace the service with a hand-written unit of your own.

Test: checkout-agent is install ok installed; checkout-agent.service is active and enabled; its unit under /lib/systemd/system/ is owned by the package.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can see and execute.

Time to Solve: 10 minutes.

4 "Salzburg": The TLS key that still matches 20 m Fix New
"Salzburg": The TLS key that still matches

Scenario: "Salzburg": The TLS key that still matches

Level: Medium

Type: Fix

Access: Email

Description: The Salzburg Festival box-office site should be served over HTTPS on port :443. The certificate expired, and nginx will not start.

We still have the public key at /home/admin/festival.pub. A key-archive cleanup left many leftover private keys under /opt/festival/keys. Find the private key that matches that public key and copy it to /home/admin/festival.key. Issue a new certificate for the same host using that same key (do not generate a new key pair) and get the site serving HTTPS again.

There is a short handover note in /home/admin/HANDOVER.txt.

Test: /home/admin/festival.key and /home/admin/festival.pub are a matching key pair.

curl -k https://127.0.0.1/ returns the box-office page.

The certificate presented on :443 is not expired, its subject is the original festival host, and it was issued with the same key as /home/admin/festival.pub.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can see and execute.

Time to Solve: 20 minutes.

5 "Auckland": S3 delete & archive old backups 20 m Do New
"Auckland": S3 delete & archive old backups

Scenario: "Auckland": S3 delete & archive old backups

Level: Medium

Type: Do

Access: Email

Description: Ops left years of database dumps in s3://auckland-backups. Keys use an absolute date prefix:
backups/YYYY-MM-DD/...

Apply this retention policy (use the date in the key, ignore LastModified):

  • Before 2023-01-01 — delete those objects
  • From 2023-01-01 through 2024-12-31 — keep them but set storage class to GLACIER
  • On or after 2025-01-01 — leave as Standard
  • Keys without that date prefix — leave untouched
The AWS CLI is ready (no configuration or authentication needed). You can start with:
aws s3 ls s3://auckland-backups --recursive

NOTE: this scenario uses Floci, a local AWS emulator; there are no real AWS resources.

Test: In s3://auckland-backups: objects with key date < 2023-01-01 are gone; dates 2023-01-01.. 2024-12-31 are GLACIER (or deep archive); dates ≥ 2025-01-01 and non-dated keys remain Standard.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can see and execute.

Time to Solve: 20 minutes.

6 "Trade": FIX drop reconciliation 20 m Fix New
"Trade": FIX drop reconciliation

Scenario: "Trade": FIX drop reconciliation

Level: Medium

Type: Fix

Access: Email

Description: FIX (Financial Information eXchange) is a common electronic protocol for trading. Messages are tag=value fields separated by SOH (ASCII 0x01; logs here use |). A message starts with 8= (BeginString), 9= (BodyLength), 35= (MsgType) and ends with 10= (CheckSum). See FIX TagValue encoding.

Our order router sends FIX orders to the exchange gateway and gets fills back. Yesterday the session dropped for a few minutes: the gateway kept executing, the router kept sending, but some fills never made it back.

Both sides log raw FIX 4.2 under /home/admin/fix/: router.log (what the router saw) and gateway.log (authoritative). Useful tags:

35  MsgType: D = NewOrderSingle (order), 8 = ExecutionReport (fill),
             5 = Logout, A = Logon, 4 = SequenceReset (session markers)
11  ClOrdID: client order ID on orders
17  ExecID: fill ID, unique per fill (e.g. EX000431)
Reconcile the two logs, find every fill the gateway executed but the router missed, and recover them by appending each missing 17=ExecID (one per line) to /home/admin/fix/fills.dat, which already holds the fills the router saw. Do not add duplicates or fills the router already has.

Credit: Alex Elliot

Test: fills.dat holds exactly the gateway's fill set: every 17=ExecID from gateway.log is present once, with no extras or duplicates.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can read and execute.

Time to Solve: 20 minutes.

7 "Flip": algo symbol flips 25 m Fix New
"Flip": algo symbol flips

Scenario: "Flip": algo symbol flips

Level: Medium

Type: Fix

Access: Email

Description: FIX (Financial Information eXchange) is an electronic protocol for trading. Messages are tag=value fields separated by SOH (ASCII 0x01; logs here use |). A message starts with 8= (BeginString), 9= (BodyLength), 35= (MsgType) and ends with 10= (CheckSum). See FIX encoding.

Our execution algo has a bug: on some fills it stamps the wrong symbol, e.g. an AAPL execution labeled DELL, GOOG labeled INTC.

You get three files: reference market data at /home/admin/market/BANDS.csv (symbol,min_px,max_px bands that never overlap), plus /home/admin/fix/orders.log (35=D orders with 11=ClOrdID and true 55=Symbol) and /home/admin/fix/executions.log (35=8 fills with 17=ExecID, 11=ClOrdID, stamped 55=Symbol, 31=LastPx).

FIX tag dictionary (lines look like 8=FIX.4.2|9=...|...|10=...|, | in place of the SOH delimiter):

35  MsgType: D = NewOrderSingle (order), 8 = ExecutionReport (fill)
11  ClOrdID: client order ID (links a fill back to its order)
17  ExecID: fill ID, unique per fill (e.g. EX000123)
55  Symbol: true symbol on orders, as-stamped (possibly wrong) on fills
31  LastPx: price of this fill
There are two independent ways to catch a flip: show the execution price is impossible for the stamped symbol against the market data bands (e.g. a $75 fill labeled GOOG whose band is $20-$30), or join 11=ClOrdID back to the order's true 55=. Create /home/admin/fix/corrections.csv with one ExecID,correct_symbol per line listing only the flipped fills with their true symbols.

Credit Alex Elliot

Test: corrections.csv lists exactly the flipped fills (~25 of ~960): every ExecID paired with its true symbol, no missing, extras, or duplicates.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can read and execute.

Time to Solve: 25 minutes.

Hard

# Name Time Type
1 "Cadiz": Cut the live wires 40 m Fix Pro New
"Cadiz": Cut the live wires

Scenario: "Cadiz": Cut the live wires

Level: Hard

Type: Fix

Access: Paid

Description: Your host is wired into a demolition circuit. cable processes are the wires. They talk only to a root supervisor, fuse, to keep it convinced the circuit is intact. Some cables are live; the others are decoys. deton processes (also root) are the charges. They talk only to fuse, never to the cables. That supervisor is what fires them. You cannot kill deton or fuse. You do not have general root access. You can only cut cables.

To cut a cable, kill -9 it. Cut all the live cables (and only those), and cut them together. Any other cut resets the circuit: fuse fires the deton charges and rolls a new live set.

Test: The charges are gone (deton has exited) and the root-owned stamp /var/lib/fuse/disarmed exists. That stamp is created only by the supervisor when the live wires drop together. Creating a file in your home directory does not count. The stamp remains across reboot.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can see and execute.

Time to Solve: 40 minutes.

2 "Pos": position manager down + reconcile 30 m Fix New
"Pos": position manager down + reconcile

Scenario: "Pos": position manager down + reconcile

Level: Hard

Type: Fix

Access: Email

Description: FIX (Financial Information eXchange) is a common electronic protocol for trading. Messages are tag=value fields separated by SOH (ASCII 0x01; logs here use |). A message starts with 8= (BeginString), 9= (BodyLength), 35= (MsgType) and ends with 10= (CheckSum). See FIX TagValue encoding.

Dev just called: a deployment of the position manager service (posmgr, a systemd service) went wrong in prod. They've asked you to take a look — posmgr show just errors. Get the service back to active first.

Once it is up, query the service for its CURRENT view (posmgr show). That view is stale: the authoritative trade history is /home/admin/fix/executions.log, raw FIX 4.2 35=8 fills carrying 55=Symbol, 54=Side (1=Buy, 2=Sell) and 32=LastQty.

FIX tag dictionary (lines look like 8=FIX.4.2|9=...|35=8|...|10=...|, | in place of the SOH delimiter):

35  MsgType: D = NewOrderSingle (order), 8 = ExecutionReport (fill)
11  ClOrdID: client order ID (links a fill back to its order)
17  ExecID: fill ID, unique per fill (e.g. EX000151)
55  Symbol: e.g. AAPL, MSFT, TSLA, NVDA, AMD
54  Side: 1 = Buy, 2 = Sell
32  LastQty: shares in this fill
31  LastPx: price of this fill
Reconcile the two and write the TRUE positions to /home/admin/fix/positions.csv as SYMBOL,QTY per line (net = buys minus sells). The file does not exist yet; you create it.

Credit Alex Elliot

Test: posmgr.service is active and positions.csv holds exactly the nets from executions.log: every symbol with its true net qty, no missing, extras or dups.

The "Check My Solution" button runs the script /home/admin/agent/check.sh, which you can read and execute.

Time to Solve: 30 minutes.

Kubernetes Playgrounds

# Name Time Type
1 K8s Playground - Free 20 m Playground
K8s Playground - Free

Playground: K8s Playground - Free

Level: Easy

Type: Playground

Access: Email

Description: This is a Kubernetes sandbox for you to play with and experiment.

It comes with an nginx.yaml playbook. You can try for example k apply -f nginx.yaml (you can use "k" as an alias for "kubectl".)

The Helm binary is also installed.

Free account:
The free account sandbox runs on a 1 GB of RAM VM. As usual, there is no Internet access.

Paid accounts (Pro/Pro+/Business:
The playground in this case runs on a 2 GB of RAM VM. It has Internet access (to pull your own images for example) and twice the time.

Time to Play: 20 minutes.

2 K8s Playground - Pro 60 m Playground Pro
K8s Playground - Pro

Playground: K8s Playground - Pro

Level: Easy

Type: Playground

Access: Paid

Description: This is a Kubernetes sandbox for you to play with and experiment.

It comes with an nginx.yaml playbook. You can try for example k apply -f nginx.yaml (you can use "k" as an alias for "kubectl".)

The Helm binary is also installed.

Free account:
The free account sandbox runs on a 1 GB of RAM VM. As usual, there is no Internet access.

Paid accounts (Pro/Pro+/Business:
The playground in this case runs on a 2 GB of RAM VM. It has Internet access (to pull your own images for example) and twice the time.

Time to Play: 60 minutes.

Send Us Feedback
Get Notified
For announcements like new scenarios. We'll never share your email with anyone else.
SadServersSadServers

Real-world Linux and DevOps scenarios for hands-on learning and technical assessment.

Uptime Robot ratio (30 days)
Product
  • Scenarios
  • For Individuals
  • For Businesses
  • Pricing
Resources
  • FAQ
  • Blog
  • Newsletter
Company
  • About Us
  • Support
  • Privacy Policy
  • Terms of Service
  • Contact
Connect With Us
info@sadservers.com

Made in Canada 🇨🇦
Updated: 2026-09-27 14:36 UTC – c78fe89