Advokat Frida

Toolkit

Redactorium: Share the File, Not the PI

Redactorium flags common forms of personal information, lets you choose what to remove or replace, and returns an updated copy with a record of the changes.

· 5 min read
An anonymous Victorian shopkeeper welcomes visitors to REDACTORIUM, surrounded by ornate full-face masks, half-masks, and feathered eye masks.

Typical scenario: you need to share a file, but the recipient doesn’t need the personal information in it. Removing it manually is a pain in the butt because that means hunting through columns and notes, and then you're livin' on a prayer. Yup, only halfway there. Whoa.

Redactorium—our brand-spanking-new tool in the brand-spanking-new AF Toolkit—finds common patterns of personal information for you. It runs locally in your browser, so your file isn’t uploaded to us. Drop in a spreadsheet, a Word file, or even a PDF. The tool lists the personal information detected and includes a justification for each flag. From there, you choose what happens to the data and get back an updated document plus a receipt of what changed.

By hand

Squint real hard, delete, send, pray

You hunt column by column, miss the email address in the notes field, forget to keep a record of what you took out, and hope for the best.

With Redactorium

Find, decide, keep a receipt

Every finding is listed with the reason it was flagged, including that email address in the notes you missed before. You choose what level of de-identification to apply. The tool produces a receipt.

A masked Victorian tailor measures a customer's Guy Fawkes mask, with full-face, half-face, and eye masks displayed on the counter.

What can it do with the data I upload?

Five options, because "nothing" is also an option (yes, how I spend my weekends is a lifestyle option).

Redact

Replaces the value with [REDACTED]. It's gone from this copy.

Use it when nobody who gets the file needs the value. For example, when sending a file export to a vendor who needs the data to fix a bug or test a feature but not the customers' names and email addresses.

Replace with a code (keyed hashing)

Values get 16-character codes within a run or batch. By default, each run uses a fresh, unique key. Set your own key to match codes across runs.

Use it when someone needs to count or match records, like an analyst or an auditor, without knowing who's who.

Make less exact (generalize)

Keeps a rougher version: a birth date becomes the year, and a ZIP code keeps its first three digits. A phone number keeps its prefix and masks the last seven digits.

Use it when the recipient needs a rough value, like the region or the birth year, but not the exact one.

Swap for fakes (synthetic data)

Replaces the value with test data, including email addresses on example domains and North American phone numbers in the reserved 555-01xx range.

Use it when the file has to look and work like the real thing. For example, a customer demo.

A masked waiter unveils a tiny eye mask for two Victorian diners, one wearing a full-face mask and the other a half-mask.

What does the tool output?

Two files.

  • The clean file, in the same format you put in, with your choices applied. Anything the tool missed, or you chose to keep, is still in it.
  • The receipt is a JSON file you can open in any text editor. It lists the patterns checked, the findings and counts, your treatments, and SHA-256 fingerprints of the input and output. It leaves out matched values and scrubs recognized personal information from the filename. Check the filename yourself too: a name the tool misses can remain in the file and its receipt.
A masked inspector examines a peacock wearing a tiny eye mask while its enormous fan of tail feathers remains exposed.

How do I use it?

  1. 1

    Drop in your file

    A spreadsheet (CSV or Excel), a Word file, a PDF, or a text or log file. Large files work too; processing time depends on the file and your device. No file handy? Try the built-in sample. Several files? Use the Batch tab.

  2. 2

    Check what it found

    Findings are listed per row. The tool checks the column's name and values. In a document or a notes column, each kind of data gets a line, like all the email addresses, with a count.

  3. 3

    Decide, apply, download

    Each line comes with a suggested treatment. Change any you disagree with, press Apply treatments, then download the clean file and the receipt.

Redactorium preview

Redactorium's findings for its sample file. At the top, a key titled "Four ways to anonymize a value" shows what each treatment does to the phone number +1 415 555 0134: Redact gives [REDACTED], Replace with a code gives a 16-character code, Make less exact gives +1 415 *** ****, and Swap for fakes gives +1 212 555 0187. Below it, eleven columns from full_name to job_title, each with the kind of data detected, a confidence score, a citation for why it was flagged, and a treatment menu set to a suggestion.
The built-in sample: a key to the four treatments, then each flagged column, why it was flagged, and a suggested treatment.

But how accurate is it?

Redactorium matches patterns. It's good at data with a fixed format, like email addresses and card numbers, and weak at anything you'd have to read the sentence to spot.

Recognizes common formats

  • Email addresses
  • Web addresses starting with http://, https://, or www., including destinations behind Word link labels. Redact replaces a visible address and removes a Word link’s destination while keeping its label. Keep leaves the address and link in place. Review the label too: it can contain identifying text that does not match another pattern.
  • Phone numbers and US Social Security numbers in their usual formats
  • Payment card numbers, IBANs (bank account numbers), and UK NHS numbers, checked with their built-in check digits
  • IP addresses and MAC addresses (network hardware IDs)

Depends on labels or context

  • Names in a name column or after a “Name:” label. In documents, a short opening line can also be flagged as a name, and recognized full names are matched again elsewhere in that document. Check these guesses.
  • US city-and-state patterns, such as “Austin, TX.” The tool checks the state name or abbreviation; it does not confirm that the city exists or that the text describes a location. Make less exact keeps only the state.
  • Street addresses
  • Dates of birth (any date can look like one)
  • Company names and job titles in spreadsheet columns
  • Passport and driver's license numbers, when a column name or a label says so

Needs manual review

  • Names the tool has not recognized elsewhere in the document, including a first name such as “Ada” in a sentence.
  • Identifying context in a resume: employers, schools, job history, and regions such as “Greater Los Angeles Area.” Removing contact details does not remove everything that can identify the applicant. Read the exported copy, check the filename, and inspect any links you kept before sharing it.
  • Details written in sentences, like a health condition in a case note
  • Text inside pictures, like screenshots and scans
  • Passwords, API keys, and user or session IDs
  • Your company's own IDs, like employee, customer, or ticket numbers, until you add a custom rule

These patterns can miss personal information or flag something harmless. Human review is still crucial.

Got identifiers only your company uses, like employee or ticket numbers? Open Custom rules under the drop zone. Give the rule a name and a pattern (a regular expression, like EMP-\d{6} for numbers such as EMP-004211), and try it on an example. Redactorium then treats those numbers like everything else.

Locally processed in your browser

Your IT or security team can check all of this. The repository's guide, "Don't trust us, check," shows how: watch the network while it works, run it offline, compare the live tool to the published code, or build their own copy.

Five Victorians pose in identical Guy Fawkes masks. The founder in the framed portrait behind them wears the same mask.

Catch something wrong, or just want to argue with me? hello@advokatfrida.com