If you have ever wanted to send someone a piece of a case, to get a second opinion or to report a bug in a tool, you already know the problem. The notes are your client's, they're covered by an agreement, and sanitizing them by hand takes long enough that you give up and don't send them at all.
I hit the same wall from the other side. I have been building tooling for incident report writing and I cannot get real case notes, and nobody should send me any. A developer asking a stranger to email him a live case file is asking for something he has no business receiving. So I wrote the tool I wanted to exist, and this is what I learned building it. It is free and it is at benchnotes.io/tools/sanitize.
Redaction destroys the document
Replace every identifier with a marker and you get something like this:
02:14 UTC RDP to[REDACTED]from[REDACTED], account[REDACTED]02:31 UTC Dropper written to[REDACTED]05:48 UTC SMB session from[REDACTED]to[REDACTED]using[REDACTED]
Every fact survives and the document is worthless. You cannot tell whether the machine on line 1 is the machine on line 3. Lateral movement is a story about which host talked to which other host, and that story is exactly what you just deleted. Whoever you sent it to, for a second opinion or to reproduce a bug, has nothing left to reason about.
So the requirement is not "remove the identifiers". It is "remove the identity and keep the structure", and those are different jobs.
Stable pseudonyms, numbered by first appearance
The same source value gets the same token every time, and tokens are numbered in the order they first appear. The same three lines come out like this:
02:14 UTC RDP to HOST-01 from IP-01, account ORG-01\NAME-01 02:31 UTC Dropper written to C:\Users\NAME-01\AppData\Local\Temp\inv.exe 05:48 UTC SMB session from HOST-01 to HOST-02 using NAME-01
Now the document still works. HOST-01 is the same machine on line 1 and line 3,
so the movement to HOST-02 is legible. The account that logged in is the same
account whose profile the dropper was written into, which is the sort of thing you notice
only when the names line up.
None of this is a novel idea and I am not claiming it is. It's worth writing down because the obvious implementation, a regular expression and a replacement string, gives you the first output rather than the second, and the difference between them is whether anyone can use what you sent.
Three things it refuses to touch
Timestamps
They're the timeline. A sanitizer that redacts times has removed the one axis the document
is organised on. That sounds obvious until you write a date pattern that's slightly too
greedy and eat 2026-07-14 out of the middle of a share path.
Hashes
An MD5 or a SHA256 identifies a file. It doesn't identify your client. Keeping them means the person you sent the notes to can look the sample up, which is usually the entire reason you are sending the notes.
Public network indicators, by default
This is the one I expected to argue with people about. In an intrusion, the attacker's addresses and domains are the indicators worth keeping, and it is the internal estate that identifies the client. So the default keeps a public IP address and strips an RFC1918 one:
02:50 UTC Beacon from IP-01 to 185.220.101.44:443 every 60s
SHA256 9f2a1b3c4d5e6f70... Mask 255.255.255.0
Amber is replaced, green is kept. The subnet mask stays too, because a dotted quad of
contiguous ones is describing a network, and turning it into IP-04 invents a
machine that doesn't exist.
There is a toggle if your threat intelligence is itself sensitive. I recommend leaving it off, on the grounds that an indicator you cannot pivot on is not much of an indicator, but you know your engagement and I don't.
The part patterns cannot do, and the way around it
A pattern finds things shaped like a hostname or an address. It does not find
Aldridge Manufacturing Ltd sitting in the middle of one of your sentences, and
that is the leak that actually matters. My first version replaced
ALDRIDGE\jmoreno and jane@aldridge-mfg.com correctly, and then
left "Aldridge SOC lead" standing three lines later. If I had shipped that, you would have
sent your client's name to a stranger while looking at a page that told you it was clean.
The fix wasn't another pattern. It was noticing that the tool already knew. By the time the scan finishes, the client's identity is sitting in the values it just matched. So it walks those matches a second time, pulls the identifying label out of each one, and offers them back as suggestions:
Found in your notes, not yet a term: + aldridge 3 + aldridge-mfg 2
Accepting one closes the prose leak. It also closes a second leak I had not planned for:
the client's own public web address. https://portal.aldridge-mfg.com/hr/onboarding.doc
survives the keep-public-indicators default, because on shape alone it is indistinguishable
from an attacker's domain. A term is what tells the tool which side of that line a domain
sits on, and the tool proposes the term itself.
Public mail providers and SaaS vendors are filtered out of the suggestions, so it never
proposes gmail or service-now as your client's name.
One person, one token
Names have to be typed in, because no pattern finds them. A box that did nothing more than
find-and-replace would not be worth the typing, so one entry expands into the forms a name
takes in real notes. I prefer this over trying to detect names automatically,
because a detector that guesses wrong on a name is a detector that quietly mangles your
notes, and you would have no way to tell which names it missed. Typing Jane Moreno also catches
Moreno, MORENO, J. Moreno,
Moreno, Jane, jmoreno and jane.moreno.
All of them become the same NAME-01, including the account name in a
DOMAIN\user string and the username buried in
C:\Users\jmoreno\AppData\Local\Temp. Without that, one human comes back as
NAME-01 in the prose, USER-01 in the account and
USER-02 in the path, and you've reintroduced exactly the problem that made
blanket redaction useless.
A surname that is also an ordinary word is left alone on its own. Adding
Mark Brown does not redact every "mark" and "brown" in the notes, because that
destroys the document in a different way. The full name still matches, and putting the bare
word on its own line is treated as an instruction to match it anyway.
A leak its own tests caught
This is the part worth the post. Protected spans, the things that must never be replaced, were originally applied per candidate: if any part of a candidate overlapped a protected span, the whole candidate was dropped. Then this appeared in a test fixture:
\\FILESRV02\2026-07-14\dump
The date is protected, correctly. The date sits inside the share path, so the whole
candidate was dropped, so FILESRV02 stayed in the clear. The
output looked sanitized. It had a server name in it.
That is the failure mode that makes this class of tool dangerous, and it is silent. A pattern stops matching, the document still looks clean, and something identifying rides along. There's no error message, and on a document of any real length you won't catch it by eye. That's why every replacement is listed in a review table rather than just applied, and why you want to read that table before you export anything.
The fix was to apply protection per part rather than per candidate, and to exempt the
anchored detectors: a match that only fires on ://, \\,
X:\ or @ can't be a timestamp by accident. The wider lesson is
that this is a tool that needs tests more than most, precisely because I have no real case
data to catch a regression by hand. There are 99 of them and they run in about a second.
What it does not do
Stated plainly, because a tool like this is worse than useless if you trust it further than it goes. READ THE OUTPUT BEFORE YOU SEND IT ANYWHERE. If it misses something, the disclosure is yours and it is not recoverable:
- It's a first pass, not a guarantee. Read the output before you send it anywhere. Every replacement is listed in a review table so you can check and reject.
- Person names are only caught if you type them in.
- Phone numbers, postal addresses and GUIDs are not detected at all. Those patterns are noisy internationally and I would rather leave them out than pretend.
- It cannot find a place or a project named only in prose unless you add it as a term.
How it is built
One HTML file, no dependencies, no build step, no framework. It makes zero third-party requests: no webfont, no analytics, no script from a CDN. Once the page has loaded you can disconnect from the internet and it still works. Before you paste anything real into it, I recommend doing exactly that: load the page, turn your wifi off, and watch it keep working. That is a thirty second check and it is worth more than my saying so.
That is not a performance decision. The proposition is that nothing leaves your browser, and that claim is only worth making to people who'll open the network tab and check it. It is also why the page looks plainer than the rest of this site.
I am a developer, not an incident responder. The shapes this thing knows are the shapes that turn up in published incident reports, because that's where I learned them. If you want to see the kind of writing I was reading, The DFIR Report publishes full intrusion reports for free and they are the best public examples of the deliverable I know of. The formats they print are the formats these tools are built to survive.
The tool is at benchnotes.io/tools/sanitize, free and with nothing uploaded. The honest test of it is one shift of your own notes: if the pseudonymised version is still something you could reason from, it works, and if it is not, it does not. Either way you will know in about two minutes.
It will meet a set of notes shaped in a way I never saw, and when it does I would rather hear about it: hello@benchnotes.io.