Case study · Public sector

Anonymise 5,000 PDFs a year without opening a single one

RelevX Redactor automatically detects and removes personal data from official documents. We built it for Santa Coloma de Gramenet City Council and it gave them back over 1,600 hours a year.

+1.600 h
Saved per year
20 min
Per document, before
3 s
Per document, now
Before and after

This is what it does in 3 seconds

On the left, a document as it comes out. On the right, the same document ready to publish.

Resolución 2024/0417
At the request of NAME with ID DNI
residing at ADDRESS
phone PHONE and email EMAIL
and payment into account IBAN

Original · 6 personal data items

Resolución 2024/0417
At the request of with ID
residing at
phone and email
and payment into account

Anonymised · ready to publish

The problem: 20 minutes per document, 5,000 documents a year

Every public body that publishes documents ( judgments, contracts, minutes, resolutions ) is required to strip out the personal data they contain first. Both the GDPR and Spanish data protection law demand it. It is not optional and there are no shortcuts.

The problem is that doing it properly is slow. You have to read the whole document, find every name, every ID number, every address, every account number, and cover them one by one. In Santa Coloma that meant about 20 minutes per document. Multiplied by 5,000 documents a year, that is over 1,600 hours: the equivalent of one full-time person doing nothing but redacting data.

And the worst part is not the time. It is that a manual, repetitive process at that scale fails. One missed surname on page eleven of a thirty-page document is enough for that PDF to be published with personal data inside. In a public body that is not an oversight: it is a data breach that must be reported.

How it works

Drag, review, export

Three steps. The second is optional and is exactly what makes the result trustworthy.

01

You drag in the documents

One or two hundred. The application opens them, extracts the text and locates the personal data: names and surnames, national ID numbers, postal addresses, phone numbers, emails, IBANs and bank cards, number plates and case references.

02

You review what is uncertain

The application flags what it has found and separates what it is sure about from what it is not. An IBAN has an unmistakable format; a proper name that is also a place name does not. It shows you that second list so you decide, and you can manually add anything you want covered even if it was not detected.

Every correction is remembered

03

You export the clean PDF

The document comes out with the data genuinely removed, not hidden under a black box that can be lifted by copying the text. It also generates a log of what was anonymised in each file.

RelevX Redactor procesando un lote de seis PDF con el panel de validación de entidades
A batch of six documents in progress. Left, the original with entities highlighted by colour; centre, the result; right, the 35 detected entities to validate one by one.

Why running on your computer matters

Most anonymisation tools upload your documents to a server. That changes the legal problem entirely.

On-device, like Redactor
  • The document never leaves the computer or the organisation network
  • There is no data transfer to document or justify
  • No data processor agreement with a third party is needed
  • It works without an internet connection
  • No external provider can suffer a breach with your documents inside
In the cloud
  • Every uploaded document is a data transfer to a third party
  • It requires a data processor agreement and a risk assessment
  • If the server is outside the EU, the international transfer must be justified
  • It depends on the connection and on the service continuing to exist
  • A breach at the provider is your breach, and the notification is yours

What it recognises

And in any of the ways it may be written, not only the one you expect.

PersonNames and surnamesMaría Pérez García
IDs / DocsNational ID and tax numbers12345678Z
LocationAddressesC/ Mayor 14, 3º B
OtherPhone numbers600 123 456
OtherEmail addressesnombre@dominio.es
IDs / DocsIBANES91 2100 0418 45…
IDs / DocsCards4532 •••• •••• 9021
OtherNumber plates1234 BCD
IDs / DocsCase numbersEXP-2024/0417
IDs / DocsSocial security08 12345678 90
OrganisationCompanies and public bodiesAsteria Sistemas, S.L.
+ whatever your organisation needs
custom fields can be added
Person Location Organisation IDs and documents Other

This is the baseline, not the limit. The set expands with whatever fields each organisation needs: a court, a council and a clinic do not handle the same data nor call it the same, and some information identifies someone in one context and not in another.

What makes it different

It learns from your documents

Every validation you make teaches it something. The system has memory and sharpens itself with use.

When you confirm a fragment was a name, or manually flag something it had missed, that correction is not lost when you close the document. The system keeps it and applies it to the next ones.

The practical effect shows quickly: the first documents from an organisation need more review, because the system does not yet know how you write, your case-file formats or how you name the parties. As more are processed, the list of doubts shortens and reviewing goes from being a task to being a glance.

That is the difference between a tool that always does the same thing and one that adapts to how you work. And it is also why the validation step is not a nuisance: every minute you spend correcting today is time you will not spend tomorrow.

Single documents or whole batches

You can work document by document or load a whole batch and process it at once. Each file in the batch keeps its own validation panel, and once all are reviewed they download together.

For a monthly publication of resolutions this changes the scale of the work: it is not anonymising one document fifty times, it is launching the batch, reviewing what the system flags as uncertain and downloading.

If you want to see it in action, at redactor.relevx.com you will find the product site, with more technical detail and the option to book a demo. It is available in Spanish, Catalan and English.

Una sentencia civil antes y después de pasar por RelevX Redactor, con los datos personales eliminados del PDF final
The original document with the detected entities and, above it, the already anonymised PDF. The data is not covered with a box: it is out of the file.

Why find and replace is not enough

It is the first question everyone asks, and it is a fair one. The answer is that personal data is not always written the same way.

The same person can appear in very different ways within the same document. An ID number may or may not carry a hyphen. An address may be split across three lines.

How it appears in the document Find and replace RelevX Redactor
María Pérez García yes yes
Pérez García, María no yes
Dña. María Pérez no yes
la demandante no yes
Data found 1 of 4 4 of 4

That is why the system does not look for text strings: it interprets the document. It understands that a fragment is acting as a person's name from the role it plays in the sentence, even if it has never seen it before. And that is exactly why the review step exists: when the interpretation is not certain, it asks rather than deciding on its own.

Anonymising is not the same as pseudonymising

The distinction sounds like a nuance and it is not: it decides whether the document is still subject to the GDPR or no longer is.

Pseudonymise

The data is replaced, but it can be undone

María PérezParty 1lookup table

It is still personal data. Re-identification is possible, so the document remains within the scope of the GDPR.

Anonymise

The data disappears from the file

María Pérezno data

It is no longer personal data. With no way to recover it, the document falls outside the scope of the GDPR. That is what Redactor does.

Redactor does the second: the data disappears from the file. And this connects to the most repeated mistake in manual anonymisation, which is drawing a black box over the text in a PDF viewer. Visually it looks covered, but the text is still in the file: it can be recovered by selecting and copying it, or by opening the PDF with any extractor. Documents published by public bodies have leaked data exactly this way.

By hand, with find and replace, or with Redactor

 By handFind and replaceRelevX Redactor
Time per document~20 min~8 min2-3 s
Detects variants of the same nameYesNoYes
Finds data you did not know was thereSometimesNoYes
Removes the data from the fileDepends how it is doneYesYes
Leaves a log of what was anonymisedNoNoYes
Consistent at page 300NoYesYes
Processes batches of documentsNoNoYes

The by-hand and find-and-replace times are those measured in the Santa Coloma de Gramenet City Council project on their own documents.

FAQ

What people ask about Redactor

Yes. The problem is the same in law firms anonymising judgments to publish or share them, in clinics releasing records for studies, and in any organisation that must hand documents to a third party without the personal data inside. What changes between sectors are the fields to detect and the document format, and that is configurable.
No, and be wary of anyone who tells you otherwise. No automatic system is, just as no person reviewing thirty pages at seven in the evening is either. That is why the design includes the review step: the application separates what it detected with certainty from what it did not, and a person decides on the uncertain. That cuts the time drastically without giving up control.
It improves. The system has memory: when you validate an entity or add something it had missed, that correction is kept and applied to subsequent documents. The first files from an organisation need more review, because it does not yet know your formats or how you write; with use, the list of doubts shortens by itself.
No. It is a common mistake in manual anonymisation: a black box is drawn over the text, but the text is still underneath and can be recovered by copying it or opening the file with another tool. Here the data is removed from the document, not covered.
The core is already built. What takes time is adapting it to your documents: which fields to detect, how they are written and what format they come in. With a sample of your real documents we can give you a specific timeframe in the free diagnosis.

See it working on your documents

The fastest way to know whether it fits is to see it with one of your documents. Book a demo and we will try it on your case.

Redactor has its own site at redactor.relevx.com, also available in Catalan and English.

Message us on WhatsApp