Cracking password-protected Office and PDF files with hashcat
Extract the hash from an encrypted Word, Excel or PDF file with office2john / pdf2john, then pick the right hashcat mode — 9400 to 10700 — to recover the password.
A .docx locked with "Restrict Editing" and a PDF that demands a password before Acrobat will even open it look like different problems. Underneath, they are the same problem: an encryption key derived from a password, and a key derivation function you have to run once per guess. Extract the right hash, pick the right hashcat mode, and both fall to the same wordlist-and-rules workflow you already use everywhere else.
Step one: get the hash out of the file
Neither hashcat nor John the Ripper reads .docx, .xlsx or .pdf directly. You extract the encryption parameters first with the matching *2john script from the John the Ripper jumbo distribution (in its run/ directory):
python3 office2john.py protected.docx > office.hash
python3 pdf2john.py protected.pdf > pdf.hash
Both print a name:hash line. Strip the filename prefix before feeding it to hashcat — John accepts the line as-is:
cut -d: -f2- office.hash > office_hashcat.txt
Picking the Office mode
Office hashes are self-describing: the string starts with $office$*2007*, $office$*2010* or $office$*2013*, so you do not have to guess:
| Office version | Hash prefix | hashcat mode | KDF |
|---|---|---|---|
| 2007 | $office$*2007* | -m 9400 | SHA-1, 50,000 iterations |
| 2010 | $office$*2010* | -m 9500 | SHA-1, 100,000 iterations |
| 2013 / 2016 / 2019 / 365 | $office$*2013* | -m 9600 | SHA-512, 100,000 iterations |
hashcat -m 9600 office_hashcat.txt rockyou.txt -r /usr/share/hashcat/rules/best64.rule
John autodetects all three under one format:
john --format=office office.hash --wordlist=rockyou.txt
2013 is the one to budget time for. Office jumped from SHA-1 to SHA-512 and kept the 100,000-iteration count, which lands somewhere in the low hundreds of thousands of guesses per second on a strong GPU — similar territory to bcrypt, not the billions per second you get on raw NTLM.
Picking the PDF mode
PDF encryption versioned differently across Acrobat releases, and hashcat splits it into several modes. The two ends of the range are on this site:
| PDF revision | hashcat mode | Notes |
|---|---|---|
| 1.1–1.3 (Acrobat 2–4) | -m 10400 | RC4 40 or 128-bit, no per-guess KDF — fast |
| 1.4–1.6 (Acrobat 5–8) | -m 10500 | RC4 or AES-128, single MD5 round |
| 1.7 Level 3 (Acrobat 9) | -m 10600 | AES-128, SHA-256 |
| 1.7 Level 8 (Acrobat 10–11) | -m 10700 | AES-256, SHA-256 x 64 rounds — the slow one |
pdf2john.py prints the revision inside the $pdf$ string, so you read the mode off the hash rather than guessing from the file:
hashcat -m 10700 pdf_hashcat.txt rockyou.txt -r /usr/share/hashcat/rules/best64.rule
John, again, covers the whole family under one format:
john --format=pdf pdf.hash --wordlist=rockyou.txt
An older PDF (10400) cracks at close to raw-hash speed. A modern AES-256 PDF (10700) is deliberately slow, the same design trade-off as sha512crypt or DCC2: the iteration count exists specifically to make brute force expensive, so a targeted wordlist earns its keep far more here than a blind mask attack does. If you have not read choosing a wordlist and rules, start there before you burn hours on the wrong list.
The PDF gotcha: two passwords, one file
A PDF can carry a user password (required to open the document at all) and an owner password (required to change permissions like printing or copying), independently. Plenty of "protected" PDFs set only an owner password and leave the user password empty, meaning the file opens with no password prompt even though pdf2john.py will happily extract a hash and hashcat will happily crack it. Before spending GPU time, just try opening the file with an empty password. If it opens, the "crack" you actually need is prying the owner-password restrictions off, which most PDF tooling does the moment you supply the empty user password — no cracking required.
When it is not worth the GPU time
If identifying the hash tells you it is Office 2013 or PDF 10700 protecting a document you have no other lead on, treat it the way you would treat a strong sha512crypt hash: a short, targeted list beats a long generic one, and an exhaustive mask attack past six or seven characters is not realistic on consumer hardware. Get the provenance right first — see finding the right hashcat mode if the extracted string does not obviously match one of the tables above — then spend your time on the wordlist, not the keyspace.