

By: Ralf Ellspermann
25-Year, Multi-Awarded BPO Veteran
Published: 28 September 2026

Reviewed By: John Maczynski
Former EVP, World's Largest Contact Center
Updated: 28 September 2026
Through layered pipelines combining pattern matching, named entity recognition and human adjudication of flagged spans. Most of what those pipelines produce is pseudonymised rather than anonymised, which means it remains personal data and the obligations continue — a distinction worth settling before the contract is signed.
Key Takeaways
- Tokenization is reversible, and that is what it is for. The vault that preserves relational integrity is the same vault that permits re-identification. Calling it irreversible substitution describes two incompatible things.
- Pseudonymised data stays fully in scope. Under GDPR Recital 26, data attributable to a person using additional information is still information on an identifiable person. Retention limits, subject rights and transfer rules all survive the scrubbing pipeline.
- Accuracy is the wrong metric; recall and precision are. Four operating points scoring between 99.51% and 99.97% on a single accuracy figure differ fortyfold in identifiers missed and twenty-four-fold in context destroyed.
- A 99.95% threshold is a residual, not a completion criterion. Across 500,000 chat logs at three identifiers each, it leaves roughly 750 live identifiers in the delivered corpus.
- Differential privacy belongs at training, not at scrubbing. It is applied to gradients with a stated privacy budget. Noise added to a conversational transcript does not produce a usable transcript.
- Named entity recognition degrades on exactly this material. It performs well on edited prose and less well on conversational text, which is what customer interaction logs are.
What Techniques Are Used, and What Does Each Produce?
Pattern matching for structured identifiers, named entity recognition for names and places, tokenization or hashing for account values, and generalisation for quasi-identifiers. They differ in whether the result can be reversed, and that difference determines the legal status of the output.
The pipeline described across this market is broadly the right one. Automated pattern matching catches the structured identifiers first, contextual recognition handles proper nouns, and human review adjudicates the ambiguous remainder. Where the accounts go wrong is in what they claim each technique delivers.

Figure 1. Four techniques, and what each leaves you holding.
Tokenization is the case worth being precise about. It is regularly described as irreversible substitution that preserves relational data integrity, and those two properties are incompatible. The relational integrity comes from the token vault — the same token always maps to the same original value, which is what lets you count distinct customers after scrubbing. That vault is also, by construction, the means of reversal. A tokenization scheme with no vault and no key is a hash, and a hash does not preserve relational integrity across systems in the way tokenization is bought to do.
None of this makes tokenization a poor choice. It is often exactly the right one, because reversibility is operationally necessary and the vault can be held under far tighter control than the corpus. It simply means the output is pseudonymised, which is a different thing from anonymised.
Does Scrubbing Take the Data Out of Scope?
Rarely. Anonymised data falls outside the regulation entirely, but only where re-identification is not reasonably possible by any means likely to be used. Pseudonymised data — anything with a vault, key or mapping still in existence — remains personal data and carries the full set of obligations.
This is the distinction that determines what a buyer has to do after the pipeline runs, and the word chosen in the contract frequently does not match the technique deployed.

Figure 2. Two outputs with very different consequences.
GDPR Recital 26 sets the test explicitly: data is anonymous only where the person is no longer identifiable, assessed against all means reasonably likely to be used, taking into account cost, time and available technology. Data that could be attributed to a person through additional information is expressly stated to be information on an identifiable natural person. The Philippine Data Privacy Act follows the same logic.
For conversational training corpora the bar is high. Even with every explicit identifier removed, a customer service transcript may contain a combination of date, product, location and circumstance that identifies one person among a client’s records. That is why the FAQ answer this market gives — that the hardest PII is implicit, appearing as references to unique life events or unusual circumstances — is correct, and why it sits awkwardly beside a claim of absolute compliance.
The practical consequence is to plan for the data remaining in scope. Retention schedules, deletion obligations, subject access, and cross-border transfer instruments all continue to apply to a scrubbed corpus, and a programme that budgeted for them ending at the pipeline has a gap.
What Quality Metric Should Govern a Scrubbing Pipeline?
Two, stated separately: a recall floor governing how many identifiers escape, and a precision floor governing how much context is destroyed. A single accuracy figure conflates errors with entirely different consequences and specifies neither.
The threshold in common use — 99.95% accuracy across anonymised batches — sounds demanding and cannot do the job asked of it.

Figure 3. Four operating points that look identical on one number.
The two failures are not symmetric. A missed identifier is a privacy failure, potentially notifiable, and its cost does not scale with how many other spans were handled correctly. An over-redacted span is a utility failure: the model loses the context it needed, which is the whole difficulty the case study in this market identifies — stripping financial detail without stripping semantic meaning. Pushing recall toward 100% drives precision down sharply, because the only way to catch every identifier is to redact aggressively.
Figure 3 shows four settings of the same scrubber. All four score between 99.51% and 99.97% on a single accuracy measure. They differ by a factor of forty in identifiers missed and by more than twenty in spans destroyed. A contract naming only accuracy has left the provider free to choose any of them, and the provider’s rational choice will not be the buyer’s.
The specification that works names a recall floor, a precision floor, and the seeded set both are measured against — identifiers planted in live batches so that recall is measured rather than asserted.
What Does the Threshold Leave Behind?
More than the percentage suggests. At 99.95% recall across 500,000 chat logs containing three identifiers each, roughly 750 live identifiers remain in the delivered corpus. Under either the Data Privacy Act or the GDPR, one identifiable person is a notifiable event.
A residual expressed as a percentage becomes a count once the corpus is large, and counts are what a regulator sees.

Figure 4. Identifiers remaining after a 99.95% pass.
The threshold is not the problem. 99.95% recall on unstructured conversational text is a genuinely demanding standard and most pipelines will not reach it. The problem is treating it as a completion criterion — as the point at which the corpus becomes clean — rather than as a description of a residual that has to be managed.
Managing it means three things in the programme design. A downstream control on the corpus itself, so that a surviving identifier is contained by access restrictions rather than by the assumption that none exists. A discovery process, so that identifiers found later have somewhere to go. And an agreed response, drafted in advance, covering who assesses notifiability and on what timescale. None of that is expensive; all of it is considerably cheaper than improvising after an identifier surfaces in a model output.
Where Does Differential Privacy Actually Apply?
At model training, not at data scrubbing. In machine learning it is applied by clipping and adding noise to gradients, and the resulting guarantee attaches to the trained model rather than to the corpus. It is not a text sanitisation technique and does not belong beside pattern matching in a pipeline table.
Differential privacy appears in this market’s descriptions as a scrubbing method that injects controlled statistical noise into datasets while preserving aggregate trends. That describes differential privacy applied to aggregate statistics — counts, means, query responses — which is a real and valuable application and a different one.

Figure 5. Where each control sits in the pipeline.
You cannot add noise to a conversational transcript and still have a conversational transcript. What is done instead, where the guarantee is wanted, is differentially private training: the gradients computed during optimisation are clipped and perturbed, producing a model about which a formal statement can be made — that its output would be nearly the same had any single training record been absent.
Two consequences follow for a buyer. The first is scope: this is work the party training the model performs, not work an annotation provider delivers, so it does not belong in a scrubbing statement of work. The second is that a differential privacy claim without a stated privacy budget is not a claim at all. The epsilon is the entire content of the guarantee, and a proposal offering differential privacy without naming one is offering a word.
How Is Quality Controlled at Scale?
Through role-based access that limits annotators to segmented snippets rather than complete customer histories, dual-key review where two independent specialists must approve a masking decision on flagged high-risk entries, and continuous error tracking against a stated threshold.
The compartmentalisation is sound and is the right control for insider risk during manual review. Two refinements make the rest of it measurable.
- Seed known identifiers into live batches. Recall asserted from an error log measures what the pipeline noticed. Recall against planted identifiers measures what it catches, which is the number the threshold is supposed to govern.
- Report recall and precision separately, with the seeded set size. One figure for privacy, one for utility, and the sample behind both, so a governance meeting can tell a real movement from ordinary variation.
- Track the two reviewers’ disagreement rate, not just their verdict. Dual-key review resolves ambiguous spans; how often they disagree is the best available signal that the masking taxonomy needs revising.
- Measure entity recognition performance on your own material. Recognition quality is considerably lower on conversational text than on edited prose, and vendor benchmarks are usually drawn from the latter.
- Define the masking taxonomy before the pilot, not during it. The case study’s own lesson is that early alignment on masking taxonomies reduces downstream rework, and that is the cheapest intervention available.
What Do Industry Leaders Say About Data Preparation?
That the difficulty of cleaning unstructured conversational data is routinely underestimated, and that the Philippine advantage is institutional maturity and legal alignment rather than labour cost alone.
The emphasis on unstructured conversational data is the right one, because it is precisely the material on which automated techniques perform worst.
Enterprise leaders often underestimate the sheer complexity of cleaning unstructured conversational data for advanced language models. The competitive advantage in the Philippine market is not merely low-cost labor; it is the institutional maturity, robust legal alignment, and rigorous multi-layered data governance that tier-one providers have cultivated over decades of servicing global Fortune 500 enterprises.
— John Maczynski, CEO, Cynergy BPO
Pattern matching handles structured identifiers well because they have structure. Entity recognition handles names and places reasonably because they are lexically distinctive. Neither catches a customer describing a circumstance that, combined with a date and a branch, identifies exactly one person in the client’s records. That residual is human work, and it is the part that justifies the governance the quotation describes — not because people are more accurate in general, but because they are the only layer capable of noticing an identifier that does not look like one.
How Did One Fintech Sanitise Its Chat Logs?
A North American financial technology firm needed 500,000 customer service chat logs cleansed of embedded account details without destroying the semantics a fraud-detection model required. Fourteen providers were assessed; the selected Manila partner ran automated entity recognition with a secure human adjudication layer, finishing two weeks early with no leakage incidents recorded.
The architecture was right for the problem. Automated recognition handles volume, human adjudication handles the flagged remainder, and the difficulty the client identified — stripping financial detail without stripping meaning — is exactly the precision-recall trade-off Figure 3 describes. Setting the masking taxonomy early, which the engagement records as its main lesson, is what makes that trade-off negotiable rather than accidental.

Figure 6. Reported outcomes from a financial chat log sanitisation.
Two figures need labelling before they travel. The 94% reduction is in manual data preparation cost — automation displacing human effort — which is a different quantity from the 40% to 60% programme savings quoted across this category. Read side by side without that label, it invites a comparison that does not hold.
And no leakage incidents recorded is a measure of what the monitoring detected as much as of what occurred. For a scrubbing engagement specifically, the complementary evidence is straightforward to produce: recall measured against a seeded set of known identifiers planted in live batches. That converts an absence of reports into a demonstrated detection rate, and it is the number a risk committee should be asking for.
Why Do Organizations Work with Cynergy BPO on Data Sanitisation Sourcing?
Cynergy BPO is an independent, vendor-neutral outsourcing advisory firm headquartered in Manila, representing a vetted network of more than 100 Philippine providers. It maps requirements against performance data to produce a shortlist within days and manages competitive negotiation on the buyer’s behalf.
Who Is Cynergy BPO?
Cynergy BPO is an independent outsourcing advisory and consultancy firm headquartered in Manila, founded by industry veterans with more than 65 years of combined operational experience governing major global accounts. It specialises in connecting mid-market and enterprise organisations with vetted Philippine BPO providers across voice, back-office and AI data operations.
How Does Cynergy BPO Differ from Traditional Outsourcing Brokers?
Traditional brokers are transactional and are compensated by the providers they place, which shapes which provider is recommended. Cynergy BPO applies an advisory-led methodology, mapping exact technical, regulatory and commercial requirements against performance data rather than against availability. On sanitisation work, where every provider will claim the same pipeline, an adviser with no placement incentive can establish which claims rest on measured recall.
How Does Cynergy BPO’s Network of 100+ Vetted Philippine BPO Providers Benefit Organizations?
The network establishes what recall is actually achievable on conversational material of a given kind before a threshold is written into a contract, and which providers can evidence it against a seeded set rather than assert it from an error log.
How Does Cynergy BPO’s Advisory-Led Vendor Matching Process Work?
Requirements are mapped against operational, regulatory and commercial criteria, a tailored shortlist of vetted providers is delivered within a few working days, and the firm then manages competitive proposal and negotiation processes on the buyer’s behalf. Masking taxonomy, recall and precision floors and residual handling are settled as part of that process rather than after selection.
Why Do Organizations Use Cynergy BPO?
Because the pipelines all look alike in a proposal and differ in the one place that matters: whether the provider can demonstrate what it catches rather than describe what it does. That distinction is invisible from outside the market and expensive to discover once a corpus is built.
Frequently Asked Questions
Which PII is hardest to remove from conversational data?
Implicit identifiers: references to unusual life events, distinctive circumstances, localised conditions or internal project names that identify a person indirectly. These have no lexical pattern, so neither regular expressions nor entity recognition reaches them, and they are the residual human review exists for.
Does tokenization anonymise data?
No. It pseudonymises it. The vault that makes tokens consistent across records is also the means of reversal, which is why the output remains personal data. That is often the right design choice; it simply does not end the obligations.
When is data genuinely anonymised?
When re-identification is not reasonably possible by any means likely to be used, judged against cost, time and available technology. For conversational transcripts that is a high bar, because combinations of ordinary details frequently identify one person within a client’s own records.
What threshold should a scrubbing contract set?
A recall floor and a precision floor, stated separately and measured against identifiers seeded into live batches. A single accuracy figure covers operating points that differ by an order of magnitude in what escapes and in what is destroyed.
What happens to identifiers the pipeline misses?
They remain in the corpus, and at realistic volumes there will be some. Plan for it: access controls on the corpus, a route for identifiers discovered later, and a pre-agreed process for assessing notifiability. The residual is a design input rather than an exception.
Is differential privacy part of data scrubbing?
No. It is applied during model training, by perturbing gradients, and its guarantee attaches to the model rather than the dataset. A differential privacy claim without a stated privacy budget conveys nothing, since the budget is the whole of the guarantee.
How long does sanitising a million records take?
Two to four weeks is the figure commonly quoted, and it is only credible where automation carries the great majority and humans adjudicate a flagged minority. Ask for the split, because an auto-passed record and a human-reviewed one are not the same deliverable.
What does this work cost in the Philippines?
Hourly rates of roughly $8 to $14 for general annotation and sanitisation work, or per-record pricing depending on complexity. Where the corpus needs domain judgement to identify implicit identifiers, the applicable tier is higher and should be quoted separately.
Unlock cost-efficient growth with expert BPO guidance!
Partner with Cynergy BPO to connect with top outsourcing providers.
Streamline operations, cut costs, and scale your business with confidence.

Ralf Ellspermann is the Chief Strategy Officer (CSO) of Cynergy BPO and a globally recognized authority in business process and contact center outsourcing. With more than 25 years of experience advising enterprises and SMEs, he provides strategic guidance on vendor selection, CX optimization, and scalable outsourcing strategies across global markets. His expertise spans fintech, ecommerce and retail, healthcare, insurance, travel and hospitality, and technology (AI & SaaS) outsourcing.
A frequent speaker at leading industry conferences, Ralf is also a published contributor to The Times of India and CustomerThink, where he shares insights on outsourcing strategy, customer experience, and digital transformation.
