.::metadata::.

updated 2026-07 · 8 min read · index

a file is its content plus everything stuck to it: gps coordinates baked into phone photos, author names and revision history in office documents, software chains and hidden layers in pdfs, and the filename itself. you always share more than you see. the fix is a two-step habit, look then clean, and a single principle that trips up even careful people. redaction is removal, not covering up.

what actually hides in a file

the leak is different per format, which is why one habit beats one tool.

look first

# everything the file is willing to admit
$ exiftool report.pdf

# just the location tags, across a whole folder
$ exiftool -gps:all -r ./photos

run it once on a photo straight from your phone and enjoy the shiver: exact model, precise timestamp, and usually the coordinates where you were standing. inspecting before sending, for anything that matters, should be a reflex.

clean

# mat2 writes a cleaned copy next to the original, never touching it
$ mat2 report.pdf

# see what mat2 can find before stripping
$ mat2 --show report.pdf

# images, surgical: strip every tag in place
$ exiftool -all= photo.jpg

mat2 handles images, pdf, office and audio formats and never modifies the original. for office documents, also run the built-in inspector (word's "inspect document", libreoffice's remove-personal-information option), because comments and tracked changes need explicit removal. when in doubt, export to pdf and then run mat2 on the pdf.

black boxes are not redaction

the most dangerous mistake, and a repeat headline. drawing a black rectangle over text in a pdf hides it on screen while the words stay in the file, one select-all-and-copy away. the same is true of "highlighting" in black or leaving an un-applied redaction annotation. some of the biggest leaks of the past decade were exactly this. real redaction deletes the underlying content, strips the metadata, and bakes the result into a new flattened file. the nsa published step-by-step guidance on doing it safely after several government incidents, and the reliable low-tech method is to flatten the page to an image (print-to-image or screenshot) so no text layer survives. always verify by trying to copy from the redacted area, and by reopening in a different reader.

two crude tools that work

a screenshot of a document carries none of the original's hidden layers, only what is visibly on screen (mind that the screenshot file itself may still get exif from your device). and paper is not safe either. many color laser printers stamp near-invisible yellow tracking dots that encode the printer's serial number and a timestamp, documented for years by the eff. the analog world has metadata too.

the full send routine

clean, encrypt, upload, then send the link and the passphrase through different channels. the complete recipe, with commands, is already written at 0807.st/privacy. same kitchen, go read it.

sources

[ home ]

.::  eof  ::.