Key takeaway: Removing a client’s name is not enough. Before using a document with an AI tool, remove direct identifiers, look for combinations of details that can reveal the person or company indirectly, check hidden information such as comments and metadata, and consider whether the document remains identifiable even after cleaning.
You have a long client document and want AI to summarize it, restructure it, extract action items, or help improve the writing.
Sending the original file may be unnecessary.
A cleaner approach is to create a working copy and remove information the AI does not need.
But there is an important distinction first.
“Anonymized” and “pseudonymized” are not always the same thing
In everyday conversation, people often say they have “anonymized” a document after replacing names with labels such as `[CLIENT]` or `[PERSON_1]`.
From a data-protection perspective, that may be more accurately described as pseudonymization if you still keep information that allows you to reconnect those placeholders to the real people.
For example:
Maria Jensen → [PERSON_1]
If you maintain a private note saying `[PERSON_1] = Maria Jensen`, the link has been separated, not necessarily eliminated.
True anonymization is a much higher bar: the information should no longer be reasonably linkable to an identifiable person.
For normal client workflows, your practical goal is often therefore:
remove unnecessary identifying information and make re-identification much harder before the data reaches the AI system.
That reduces risk, but it does not automatically remove every privacy or contractual obligation.
What should you look for?
Start with direct identifiers.
These include:
- full names;
- email addresses;
- phone numbers;
- home addresses;
- usernames;
- employee IDs;
- account numbers;
- customer numbers;
- passport or identification numbers;
- URLs containing identifying information.
Then look beyond the obvious.
A document can identify someone indirectly through combinations such as:
“the only Finnish sales manager in the Dubai office”
or:
“the employee who joined the Rovaniemi branch on 14 February and manages the Norwegian account.”
No name appears, but the description may still point to one person.
Other details to consider include:
- exact job titles;
- unusual professions;
- exact dates;
- locations;
- ages;
- project names;
- department names;
- specific customer relationships;
- rare events;
- unusual financial figures;
- unique complaints or incidents.
Company names also deserve attention. A company name is not automatically personal data, but it may reveal the client relationship or commercially confidential information that the AI does not need.
Use a simple placeholder system
Do not replace every name with “John Doe.”
Use consistent neutral labels instead.
For example:
| Original information | Replacement |
|---|---|
| Acme Consulting Oy | `[CLIENT]` |
| Maria Jensen | `[PERSON_1]` |
| Daniel Smith | `[PERSON_2]` |
| Helsinki | `[CITY]` |
| Project Aurora | `[PROJECT]` |
| €83,750 | `[BUDGET]` |
| October 14, 2026 | `[DEADLINE]` |
Consistency matters.
If `[PERSON_1]` appears ten times in the original document, use `[PERSON_1]` all ten times. The AI can then follow the relationships without knowing the real name.
Keep the mapping key outside the AI conversation.
For example:
`[PERSON_1] = Maria Jensen` `[CLIENT] = Acme Consulting Oy`
Do not upload that key with the cleaned document. Doing so would defeat the purpose of replacing the identifiers.
Step 1: Make a copy
Never begin by editing the only version of a client document.
Create a separate working copy.
Give it a neutral filename such as:
`AI_WORKING_COPY.docx`
rather than:
`Maria_Jensen_Northbridge_Employee_Complaint_Sept2026.docx`
Keep the untouched original in its normal approved location.
The sanitized copy should be disposable.
Step 2: Replace obvious identifiers
Use find-and-replace for information that appears repeatedly.
Search for:
- people’s first and last names;
- company names;
- email addresses;
- telephone numbers;
- project names;
- customer IDs;
- addresses.
Replace each with a consistent placeholder.
Do not forget shortened forms.
If the document uses:
Maria Jensen Maria Ms Jensen M. Jensen
search for each version.
Automated replacement helps, but it should not be your final check.
Step 3: Scan for indirect identifiers
This is where simple find-and-replace stops being enough.
Read through the document and ask:
Could someone work out who this person, client or project is from the remaining details?
Consider this sentence:
The 43-year-old finance director from our Oulu office who transferred from Stockholm in March raised the complaint.
Removing the person’s name may do very little. The combination of age, role, location and employment history might identify the person.
A safer version might be:
`[EMPLOYEE]` raised the complaint.
How far you generalize depends on what the AI actually needs.
If age is relevant to the task, perhaps use an age band.
If the exact date is irrelevant, use:
`[DATE]`
or:
“in early 2026.”
If the exact city does not matter, use:
`[LOCATION]`.
The goal is not to preserve every fact. It is to preserve the facts required for the task.
Step 4: Check numbers, dates and unusual events
People tend to look for names and forget everything else.
Numbers can also identify a client or project.
Review:
- contract values;
- invoice totals;
- salaries;
- exact revenue;
- account balances;
- order numbers;
- dates;
- case numbers;
- transaction references.
Consider whether the AI needs the exact value.
Instead of:
€247,392
you may be able to use:
`[CONTRACT_VALUE]`
or:
“approximately €250,000”
depending on the task.
Be especially careful with combinations.
A particular value + city + date + project type can be surprisingly identifying even when every name has been removed.
Step 5: Check hidden information
Cleaning the visible paragraphs is not necessarily enough if you plan to upload the actual document.
Files may contain information outside the main text.
Check for:
File names
A perfectly sanitized document called:
`Termination_Maria_Jensen_ClientX.docx`
still reveals information.
Rename the working copy.
Comments
Review comments left by:
- you;
- the client;
- colleagues;
- editors.
Comments may contain real names or information removed from the main text.
Tracked changes
A document can visually show a clean final version while retaining deleted text or previous edits through revision history or tracked changes.
Accept or remove tracked changes from the working copy when appropriate.
Document properties and metadata
Files can contain properties such as:
- author;
- organization;
- title;
- subject;
- editing information.
Check the application’s document-inspection or metadata controls before uploading the file.
Headers, footers and cover pages
These commonly contain:
- company names;
- document numbers;
- confidentiality notices;
- addresses;
- client names.
Images and screenshots
Removing a person’s name from the text does not remove it from a screenshot embedded on page 12.
Review images, charts and scanned pages as well.
Step 6: Do a final read-through
Now stop searching mechanically and read the cleaned document like a stranger.
Ask:
Who is this about?
What organization is involved?
What location is involved?
Could I find the client by searching a distinctive phrase from this document?
Are there enough unique clues left to identify the project?
Is any remaining detail unnecessary for the AI task?
This final human review often catches things automated replacement misses.
Step 7: Use only the cleaned version
Upload or paste the cleaned version—not the original.
Then keep the prompt narrow.
Instead of:
Analyze everything in this client file and tell me what you find.
prefer:
Review this sanitized project brief. Identify unclear deliverables, missing deadlines and questions that should be clarified before work begins. Do not infer missing facts.
A narrow task reduces both unnecessary disclosure and unnecessary AI output.
Step 8: Reinsert real information afterward
If you use AI to produce a draft, keep the generated version generic while you review it.
For example:
Dear `[PERSON_1]`, Thank you for confirming the revised deadline for `[PROJECT]`.
After you have reviewed the content, replace the placeholders locally:
Dear Maria, Thank you for confirming the revised deadline for Project Aurora.
This keeps real identifiers out of the AI interaction when they are not needed to generate the text.
When removing identifiers is not enough
Some documents are identifiable because the substance itself is unique.
Imagine a document describing:
- a highly publicized dispute;
- a rare medical condition;
- a specific court case;
- an unpublished acquisition;
- a unique product failure;
- a small company’s only employee in a particular country;
- a confidential incident known to only a few people.
Removing the names may not make those documents meaningfully anonymous.
Someone familiar with the situation could immediately recognize who or what the document concerns.
This is why anonymization cannot be reduced to a checklist of fields to delete.
You must consider the whole information set.
When to skip the AI tool
Do not force anonymization into a workflow where the sensible answer is simply not to send the document.
That may be appropriate when:
- your contract prohibits the use;
- the document contains credentials or security secrets;
- you cannot remove sensitive information without destroying the usefulness of the document;
- the remaining context still obviously identifies the person or client;
- the information belongs in an approved internal system only;
- you cannot verify how the chosen AI account handles the data.
A local workflow, approved enterprise environment, synthetic example, or ordinary manual review may be more appropriate.
Client Document AI Prep Checklist
Use this before pasting or uploading a client document.
| Check | Done |
|---|---|
| I created a separate working copy. | ☐ |
| I confirmed that using an AI tool is permitted for this task. | ☐ |
| Names have been replaced. | ☐ |
| Email addresses and phone numbers have been removed. | ☐ |
| Addresses and account/customer numbers have been removed. | ☐ |
| Client and company names have been removed when unnecessary. | ☐ |
| Project names and unique identifiers have been replaced. | ☐ |
| Exact dates and locations have been generalized where possible. | ☐ |
| Sensitive or unnecessary financial figures have been removed. | ☐ |
| I checked combinations of details that could identify someone indirectly. | ☐ |
| The file name does not reveal the client or person. | ☐ |
| Comments have been reviewed. | ☐ |
| Tracked changes and deleted text have been reviewed. | ☐ |
| Document metadata has been checked. | ☐ |
| Headers, footers, images and screenshots have been reviewed. | ☐ |
| The placeholder key is stored separately and will not be uploaded. | ☐ |
| I completed a final human read-through. | ☐ |
| The AI receives only the information necessary for the task. | ☐ |
| Real details will be reinserted locally after reviewing the AI output. | ☐ |
Suggested download: Turn this table into a one-page Client-Safe AI: Pre-Upload Document Checklist PDF and link it from this article.
The simplest rule
Removing a person’s name is only the beginning.
A useful client-safe workflow asks three questions:
Can I remove this information?
Can someone still identify the client or person from what remains?
Does the AI need this detail at all?
If the AI can perform the task without a piece of client information, the safest place for that information is usually outside the prompt.
For the broader questions to check before allowing any AI service to process client information, see AI Privacy Checklist: 15 Questions Before Uploading Client Data.