Skip to content

Cracking password-protected Office and PDF files with hashcat

Extract the hash from an encrypted Word, Excel or PDF file with office2john / pdf2john, then pick the right hashcat mode — 9400 to 10700 — to recover the password.

Published on 4 min read

A .docx locked with "Restrict Editing" and a PDF that demands a password before Acrobat will even open it look like different problems. Underneath, they are the same problem: an encryption key derived from a password, and a key derivation function you have to run once per guess. Extract the right hash, pick the right hashcat mode, and both fall to the same wordlist-and-rules workflow you already use everywhere else.

Step one: get the hash out of the file

Neither hashcat nor John the Ripper reads .docx, .xlsx or .pdf directly. You extract the encryption parameters first with the matching *2john script from the John the Ripper jumbo distribution (in its run/ directory):

python3 office2john.py protected.docx > office.hash
python3 pdf2john.py protected.pdf > pdf.hash

Both print a name:hash line. Strip the filename prefix before feeding it to hashcat — John accepts the line as-is:

cut -d: -f2- office.hash > office_hashcat.txt

Picking the Office mode

Office hashes are self-describing: the string starts with $office$*2007*, $office$*2010* or $office$*2013*, so you do not have to guess:

Office versionHash prefixhashcat modeKDF
2007$office$*2007*-m 9400SHA-1, 50,000 iterations
2010$office$*2010*-m 9500SHA-1, 100,000 iterations
2013 / 2016 / 2019 / 365$office$*2013*-m 9600SHA-512, 100,000 iterations
hashcat -m 9600 office_hashcat.txt rockyou.txt -r /usr/share/hashcat/rules/best64.rule

John autodetects all three under one format:

john --format=office office.hash --wordlist=rockyou.txt

2013 is the one to budget time for. Office jumped from SHA-1 to SHA-512 and kept the 100,000-iteration count, which lands somewhere in the low hundreds of thousands of guesses per second on a strong GPU — similar territory to bcrypt, not the billions per second you get on raw NTLM.

Picking the PDF mode

PDF encryption versioned differently across Acrobat releases, and hashcat splits it into several modes. The two ends of the range are on this site:

PDF revisionhashcat modeNotes
1.1–1.3 (Acrobat 2–4)-m 10400RC4 40 or 128-bit, no per-guess KDF — fast
1.4–1.6 (Acrobat 5–8)-m 10500RC4 or AES-128, single MD5 round
1.7 Level 3 (Acrobat 9)-m 10600AES-128, SHA-256
1.7 Level 8 (Acrobat 10–11)-m 10700AES-256, SHA-256 x 64 rounds — the slow one

pdf2john.py prints the revision inside the $pdf$ string, so you read the mode off the hash rather than guessing from the file:

hashcat -m 10700 pdf_hashcat.txt rockyou.txt -r /usr/share/hashcat/rules/best64.rule

John, again, covers the whole family under one format:

john --format=pdf pdf.hash --wordlist=rockyou.txt

An older PDF (10400) cracks at close to raw-hash speed. A modern AES-256 PDF (10700) is deliberately slow, the same design trade-off as sha512crypt or DCC2: the iteration count exists specifically to make brute force expensive, so a targeted wordlist earns its keep far more here than a blind mask attack does. If you have not read choosing a wordlist and rules, start there before you burn hours on the wrong list.

The PDF gotcha: two passwords, one file

A PDF can carry a user password (required to open the document at all) and an owner password (required to change permissions like printing or copying), independently. Plenty of "protected" PDFs set only an owner password and leave the user password empty, meaning the file opens with no password prompt even though pdf2john.py will happily extract a hash and hashcat will happily crack it. Before spending GPU time, just try opening the file with an empty password. If it opens, the "crack" you actually need is prying the owner-password restrictions off, which most PDF tooling does the moment you supply the empty user password — no cracking required.

When it is not worth the GPU time

If identifying the hash tells you it is Office 2013 or PDF 10700 protecting a document you have no other lead on, treat it the way you would treat a strong sha512crypt hash: a short, targeted list beats a long generic one, and an exhaustive mask attack past six or seven characters is not realistic on consumer hardware. Get the provenance right first — see finding the right hashcat mode if the extracted string does not obviously match one of the tables above — then spend your time on the wordlist, not the keyspace.

Related articles

A practical comparison of hashcat and John the Ripper — GPU vs CPU strengths, autodetection, -m modes, jumbo formats, wordlists and rules — with example commands.
Found a mystery hash? Learn the signals that reveal its type — length, character set and prefixes like $2y$ or $6$ — and how to identify it privately in your browser.
MD5 and SHA-1 fall to a GPU in seconds because they are fast and often unsalted. Learn why slow KDFs like bcrypt and Argon2 resist — and what defenders should do.