← Writing
Build Tooling

Stripping the client out of case notes without destroying them

Blanking every hostname makes a set of notes unreadable, which is why nobody does it and why nobody shares their notes. Stable pseudonyms keep the timeline followable.

If you have ever wanted to send someone a piece of a case, to get a second opinion or to report a bug in a tool, you already know the problem. The notes are your client's, they're covered by an agreement, and sanitizing them by hand takes long enough that you give up and don't send them at all.

I hit the same wall from the other side. I have been building tooling for incident report writing and I cannot get real case notes, and nobody should send me any. A developer asking a stranger to email him a live case file is asking for something he has no business receiving. So I wrote the tool I wanted to exist, and this is what I learned building it. It is free and it is at benchnotes.io/tools/sanitize.

Redaction destroys the document

Replace every identifier with a marker and you get something like this:

02:14 UTC  RDP to [REDACTED] from [REDACTED], account [REDACTED]
02:31 UTC  Dropper written to [REDACTED]
05:48 UTC  SMB session from [REDACTED] to [REDACTED] using [REDACTED]

Every fact survives and the document is worthless. You cannot tell whether the machine on line 1 is the machine on line 3. Lateral movement is a story about which host talked to which other host, and that story is exactly what you just deleted. Whoever you sent it to, for a second opinion or to reproduce a bug, has nothing left to reason about.

So the requirement is not "remove the identifiers". It is "remove the identity and keep the structure", and those are different jobs.

Stable pseudonyms, numbered by first appearance

The same source value gets the same token every time, and tokens are numbered in the order they first appear. The same three lines come out like this:

02:14 UTC  RDP to HOST-01 from IP-01, account ORG-01\NAME-01
02:31 UTC  Dropper written to C:\Users\NAME-01\AppData\Local\Temp\inv.exe
05:48 UTC  SMB session from HOST-01 to HOST-02 using NAME-01

Now the document still works. HOST-01 is the same machine on line 1 and line 3, so the movement to HOST-02 is legible. The account that logged in is the same account whose profile the dropper was written into, which is the sort of thing you notice only when the names line up.

None of this is a novel idea and I am not claiming it is. It's worth writing down because the obvious implementation, a regular expression and a replacement string, gives you the first output rather than the second, and the difference between them is whether anyone can use what you sent.

Three things it refuses to touch

Timestamps

They're the timeline. A sanitizer that redacts times has removed the one axis the document is organised on. That sounds obvious until you write a date pattern that's slightly too greedy and eat 2026-07-14 out of the middle of a share path.

Hashes

An MD5 or a SHA256 identifies a file. It doesn't identify your client. Keeping them means the person you sent the notes to can look the sample up, which is usually the entire reason you are sending the notes.

Public network indicators, by default

This is the one I expected to argue with people about. In an intrusion, the attacker's addresses and domains are the indicators worth keeping, and it is the internal estate that identifies the client. So the default keeps a public IP address and strips an RFC1918 one:

02:50 UTC  Beacon from IP-01 to 185.220.101.44:443 every 60s
           SHA256 9f2a1b3c4d5e6f70...  Mask 255.255.255.0

Amber is replaced, green is kept. The subnet mask stays too, because a dotted quad of contiguous ones is describing a network, and turning it into IP-04 invents a machine that doesn't exist.

There is a toggle if your threat intelligence is itself sensitive. I recommend leaving it off, on the grounds that an indicator you cannot pivot on is not much of an indicator, but you know your engagement and I don't.

The part patterns cannot do, and the way around it

A pattern finds things shaped like a hostname or an address. It does not find Aldridge Manufacturing Ltd sitting in the middle of one of your sentences, and that is the leak that actually matters. My first version replaced ALDRIDGE\jmoreno and jane@aldridge-mfg.com correctly, and then left "Aldridge SOC lead" standing three lines later. If I had shipped that, you would have sent your client's name to a stranger while looking at a page that told you it was clean.

The fix wasn't another pattern. It was noticing that the tool already knew. By the time the scan finishes, the client's identity is sitting in the values it just matched. So it walks those matches a second time, pulls the identifying label out of each one, and offers them back as suggestions:

Found in your notes, not yet a term:   + aldridge 3   + aldridge-mfg 2

Accepting one closes the prose leak. It also closes a second leak I had not planned for: the client's own public web address. https://portal.aldridge-mfg.com/hr/onboarding.doc survives the keep-public-indicators default, because on shape alone it is indistinguishable from an attacker's domain. A term is what tells the tool which side of that line a domain sits on, and the tool proposes the term itself.

Public mail providers and SaaS vendors are filtered out of the suggestions, so it never proposes gmail or service-now as your client's name.

One person, one token

Names have to be typed in, because no pattern finds them. A box that did nothing more than find-and-replace would not be worth the typing, so one entry expands into the forms a name takes in real notes. I prefer this over trying to detect names automatically, because a detector that guesses wrong on a name is a detector that quietly mangles your notes, and you would have no way to tell which names it missed. Typing Jane Moreno also catches Moreno, MORENO, J. Moreno, Moreno, Jane, jmoreno and jane.moreno.

All of them become the same NAME-01, including the account name in a DOMAIN\user string and the username buried in C:\Users\jmoreno\AppData\Local\Temp. Without that, one human comes back as NAME-01 in the prose, USER-01 in the account and USER-02 in the path, and you've reintroduced exactly the problem that made blanket redaction useless.

A surname that is also an ordinary word is left alone on its own. Adding Mark Brown does not redact every "mark" and "brown" in the notes, because that destroys the document in a different way. The full name still matches, and putting the bare word on its own line is treated as an instruction to match it anyway.

A leak its own tests caught

This is the part worth the post. Protected spans, the things that must never be replaced, were originally applied per candidate: if any part of a candidate overlapped a protected span, the whole candidate was dropped. Then this appeared in a test fixture:

\\FILESRV02\2026-07-14\dump

The date is protected, correctly. The date sits inside the share path, so the whole candidate was dropped, so FILESRV02 stayed in the clear. The output looked sanitized. It had a server name in it.

That is the failure mode that makes this class of tool dangerous, and it is silent. A pattern stops matching, the document still looks clean, and something identifying rides along. There's no error message, and on a document of any real length you won't catch it by eye. That's why every replacement is listed in a review table rather than just applied, and why you want to read that table before you export anything.

The fix was to apply protection per part rather than per candidate, and to exempt the anchored detectors: a match that only fires on ://, \\, X:\ or @ can't be a timestamp by accident. The wider lesson is that this is a tool that needs tests more than most, precisely because I have no real case data to catch a regression by hand. There are 99 of them and they run in about a second.

What it does not do

Stated plainly, because a tool like this is worse than useless if you trust it further than it goes. READ THE OUTPUT BEFORE YOU SEND IT ANYWHERE. If it misses something, the disclosure is yours and it is not recoverable:

  • It's a first pass, not a guarantee. Read the output before you send it anywhere. Every replacement is listed in a review table so you can check and reject.
  • Person names are only caught if you type them in.
  • Phone numbers, postal addresses and GUIDs are not detected at all. Those patterns are noisy internationally and I would rather leave them out than pretend.
  • It cannot find a place or a project named only in prose unless you add it as a term.

How it is built

One HTML file, no dependencies, no build step, no framework. It makes zero third-party requests: no webfont, no analytics, no script from a CDN. Once the page has loaded you can disconnect from the internet and it still works. Before you paste anything real into it, I recommend doing exactly that: load the page, turn your wifi off, and watch it keep working. That is a thirty second check and it is worth more than my saying so.

That is not a performance decision. The proposition is that nothing leaves your browser, and that claim is only worth making to people who'll open the network tab and check it. It is also why the page looks plainer than the rest of this site.

I am a developer, not an incident responder. The shapes this thing knows are the shapes that turn up in published incident reports, because that's where I learned them. If you want to see the kind of writing I was reading, The DFIR Report publishes full intrusion reports for free and they are the best public examples of the deliverable I know of. The formats they print are the formats these tools are built to survive.

The tool is at benchnotes.io/tools/sanitize, free and with nothing uploaded. The honest test of it is one shift of your own notes: if the pseudonymised version is still something you could reason from, it works, and if it is not, it does not. Either way you will know in about two minutes.

It will meet a set of notes shaped in a way I never saw, and when it does I would rather hear about it: hello@benchnotes.io.