Image

What SLAs Should Enterprises Require for LLM Training Projects in the Philippines?

Image

By: Ralf Ellspermann
25-Year, Multi-Awarded BPO Veteran
Published: 24 September 2026

Image

Reviewed By: John Maczynski
Former EVP, World's Largest Contact Center
Updated: 24 September 2026

Quantitative targets for accuracy, inter-annotator agreement, throughput and security — each paired with the audit volume needed to measure it. A threshold expressed to a tenth of a point is a commitment to sample thousands of items a week, and a clause the audit cannot resolve penalises noise rather than performance.

Key Takeaways

  • Every threshold implies a sampling budget. Detecting a one-point shortfall below a 98.5% target needs roughly 2,400 audited items. Half a point needs about 8,500.
  • A 300-item weekly audit resolves almost nothing. The 95% interval around an observed 98.5% runs from 96.6% to 99.5%, so a batch delivered at 97% and one at 99% return the same verdict.
  • An unmeasurable clause hurts both parties. At 300 items a week, a vendor genuinely at 99.0% is credited against in about one week in five, while one at 97.5% passes in about one week in eight.
  • Cohen’s kappa is defined for two raters only. A clause written over a cohort needs Fleiss’ kappa or mean pairwise Cohen’s, plus a stated sampling protocol. Ambiguity inside a bonus-malus clause becomes a dispute.
  • Zero-tolerance security clauses reward weak detection. If any detected event triggers contract review, the cheapest compliance is to see less. Graduate the remedy by whether the event was self-reported.
  • The two-hour notification protects the buyer’s 72-hour clock. Philippine and EU law both require notification within 72 hours of knowledge. The contractual clock is shorter so the buyer has time to act on its own.

What Performance Metrics Belong in a Training SLA?

Annotation accuracy against a gold-standard key, inter-annotator agreement with the statistic named, throughput per specialist per shift, and security event handling. Each needs a target, a measurement method and a remedy — and the measurement method is the part most often left out.

The move away from call-centre metrics is right. Average handle time tells a buyer nothing about whether a preference ranking was correct, and enterprise agreements in this category properly specify accuracy thresholds, statistical consensus scores and operational definitions for drift, hallucination correction and rubric compliance.

Figure 1. Four SLA metrics, and what it takes to police each.

The third column is what converts a target into an obligation. A clause requiring at least 98.5% precision, audited weekly, sounds precise and is only as precise as the audit behind it. Two of the four metrics in Figure 1 are estimated from a sample and therefore carry sampling error; throughput is counted completely and does not; security events are not a sampling question at all, and treating them as one is a category error addressed later.

One drafting point before the arithmetic. Cohen’s kappa is defined for exactly two raters. A clause requiring a coefficient of 0.80 or above across a team cohort is underdetermined as written, and the correct statistic is Fleiss’ kappa or the mean of pairwise Cohen’s coefficients. This costs nothing to fix at drafting and is expensive to argue about once a credit is in dispute.

Can a Weekly Audit Actually Measure a 98.5% Threshold?

Not at the sample sizes typically used. A 300-item audit produces a 95% confidence interval running from about 96.6% to 99.5% around an observed 98.5%. Resolving a one-point shortfall requires roughly 2,400 items per audit; resolving half a point requires about 8,500.

Accuracy thresholds in this market are quoted to a tenth of a percentage point, and the audit cadence is quoted as a weekly batch review. Those two numbers have to be compatible, and usually they are not.

Figure 2. What a weekly audit of a given size can actually see.

The reason is that errors at these accuracy levels are rare events. At 98.5%, a 300-item audit is expected to contain about four and a half errors. Ordinary sampling variation in a count that small is large in proportional terms, so the estimated error rate swings widely from week to week even when nothing about the pod has changed. Precision in the estimate comes only from volume, and the volume required rises sharply as the margin being policed narrows.

This is not an argument for loose SLAs. It is an argument that the threshold and the audit plan are a single decision. A buyer who wants to enforce 98.5% to the decimal should budget for an audit of a few thousand items, which is a real and quotable cost. A buyer who can only afford a few hundred should write a wider tolerance band, or trigger on a rolling multi-week average, and say so in the contract.

Why the case study illustrates the problem

The engagement described later delivered 99.2% against a 99.0% threshold across 100,000 items. That margin is 200 items in 100,000. A 500-item weekly audit at a true 99.2% returns a 95% interval of roughly 97.96% to 99.69% — straddling the threshold from both directions. The programme was compliant, and the measurement apparatus described could not have demonstrated it in any given week.

Who Pays When the Audit Cannot Resolve the Threshold?

Both parties, in different weeks. Under a hard 98.5% invoice-credit trigger audited at 300 items, a vendor genuinely delivering 99.0% is penalised in about 18% of weeks, while a vendor delivering 97.5% passes in about 13% of weeks. The clause transfers sampling noise rather than performance risk.

A remedy clause exists to move risk to the party best able to control it. A clause enforced against an imprecise measurement does something different: it moves variance the vendor cannot control onto the vendor, and lets genuine failures through often enough to be unreliable protection for the buyer.

Figure 3. Outcomes of a hard threshold, by weekly audit size.

The commercial consequence is worse than the statistics suggest. A vendor that expects to be credited against in one week in five will price that expectation into its rate, so the buyer pays for the noise twice — once in the credit and once in the margin added to cover it. And a vendor disputing a credit it believes is unwarranted will usually be right, which corrodes the governance relationship the same agreement is trying to build.

Three fixes are available and all belong in the negotiation rather than in the remedy. Raise the audit volume to match the clause. Widen the clause to a tolerance band around the threshold, so a credit triggers only on a shortfall the audit can actually distinguish. Or trigger on a rolling four-week average, which quadruples the effective sample without quadrupling the weekly cost.

How Should Commercial Structures Allocate Risk?

By tying compensation to verified milestone completion rather than to seat count, while ensuring every metric in the bonus-malus structure is one the vendor improves by doing the work better rather than by managing the measurement. Fully loaded Philippine rates for specialised domain annotators run roughly $18 to $35 an hour.

The choice between full-time-equivalent pricing and managed-outcome contracting is a choice about where variance lands. FTE pricing gives stable capacity and leaves productivity optimisation with the buyer. Managed-outcome pricing carries a management margin and transfers accountability for delivery and quality to the provider. Hybrid structures with bonus-malus terms sit between them and are increasingly the default for this category.

Figure 4. What each remedy formulation actually rewards.

The test to apply to any proposed clause is simple: is the cheapest way for the vendor to satisfy this to improve the work, or to manage the measurement? A zero-tolerance security clause that triggers contract review on any detected event is satisfied most cheaply by tuning detection conservatively. A batch-level accuracy credit is satisfied partly by sampling luck. A throughput target running parallel to a quality gate is satisfied by speed, and the two obligations compete.

Each has a better formulation that preserves the intent. Security remedies should be graduated by whether the event was self-reported — severe for concealment, mild for disclosure — so that the vendor’s interest aligns with detecting and reporting rather than with not seeing. Accuracy credits should trigger on a rolling average within a tolerance band. Throughput should be conditional on the quality gate passing rather than a separate obligation, so there is no volume a vendor can deliver that offsets failing the accuracy clause.

How Should Security and Confidentiality Obligations Be Drafted?

As a combination of certification maintenance, specific technical prohibitions, and a notification clock shorter than the statutory one. Providers should be bound to maintain current ISO 27001 certification and SOC 2 Type II attestation, to prohibit unapproved storage media, and to notify the buyer within two hours of knowledge of an incident.

The technical provisions are standard and correctly stated in most agreements of this type: controlled desktop environments, no unapproved local storage, clean-desk policies across production floors. The clause worth thinking about carefully is the notification clock, because its purpose is frequently misunderstood.

Figure 5. How the contractual and statutory clocks nest.

The two-hour deadline is not a regulatory requirement. Under the Philippine Data Privacy Act, a personal information controller or processor must notify the National Privacy Commission and affected data subjects within 72 hours of knowledge of a notifiable breach, with a complete report following within five business days; GDPR sets the same 72-hour clock to the supervisory authority. The contractual two hours exists so that the buyer, who carries the statutory obligation, has the remaining seventy hours to assess, decide and file rather than discovering the incident on day three.

Both clocks start on knowledge, which is why detection capability rather than incident count is the thing to contract for. A provider whose monitoring detects late has satisfied every deadline in the agreement and left the buyer with no time. Requiring evidence of what the monitoring is configured to catch — and a periodic test exercise demonstrating it catching something — converts a clean record from an absence of reports into a demonstrated control.

On indemnification, the clause should cover data leakage, intellectual property infringement and regulatory non-compliance, and should flow down to any subcontractor. Note separately that certification does not discharge transfer obligations: the Philippines holds no EU adequacy decision, so EU personal data reaching a Philippine processor requires an Article 46 mechanism regardless of what the provider is certified to.

What Governance Keeps an SLA Working After Signature?

A joint committee of client data scientists and vendor operations managers reviewing performance weekly, senior quality re-sampling of completed batches before ingestion, retraining whenever rubrics change, and tracking of individual variance to catch fatigue and drift before they reach a batch.

  • Report the estimate and its precision together. An accuracy figure without the audit size behind it is not interpretable, and a governance meeting spent arguing about an unquantified number is wasted.
  • Put rubric changes on the agenda, not just results. A quality dip immediately after a rubric update is an onboarding cost, not a performance failure, and the two should not be remedied the same way.
  • Track individual variance separately from batch accuracy. Drift concentrated in two evaluators is a retraining problem. Drift spread evenly across the pod is usually a rubric problem.
  • Re-sample before ingestion, not after. A senior re-sample of completed batches is the last point at which a defect costs a review rather than a retraining cycle on the model.
  • Agree in advance what a missed target triggers. Whether the first response is investigation, remediation or credit should be settled at drafting, when neither party knows who will be arguing which side.

What Do Industry Leaders Say About Structuring AI Contracts?

That treating model training like routine data entry is the central procurement error, and that commercial terms should be tied to statistical consensus and domain accuracy rather than to seat count or turnaround time. Aligning vendor profitability with model quality changes performance.

The principle is right, and it raises the standard for how carefully the metrics themselves are specified.

The biggest mistake procurement teams make in artificial intelligence outsourcing is treating model training like routine data entry. SLAs that focus only on seat count or simple turnaround time guarantee substandard model performance. Enterprise buyers must tie commercial terms directly to statistical consensus and rigorous domain accuracy metrics. When you align vendor profitability with model intelligence, performance transforms completely.

— John Maczynski, CEO, Cynergy BPO

Aligning profitability with quality is exactly the objective, and it places a heavier burden on the metric than a seat-count contract ever does. When money moves on a number, that number needs to be measurable at the precision claimed, computed by a named statistic, drawn from a stated sample, and difficult to influence by any route other than doing the work well. A tightly drafted seat-count clause is merely unambitious. A loosely drafted quality clause actively misallocates money.

How Did One Healthcare Enterprise Structure Its Clinical SLA?

A healthcare technology enterprise required firm operational SLAs before approving offshore clinical annotation. A 50-person Manila pod under a managed-outcome agreement delivered 99.2% average accuracy across 100,000 medical instruction pairs over six months, reduced preparation cost by 45%, and passed all quarterly compliance audits.

The compliance committee’s position was the right one and is common in regulated healthcare work: no offshore engagement without contractual guarantees. The engagement delivered against them.

Figure 6. Reported outcomes from a 50-person clinical pod in Manila.

The throughput figure is worth drawing out because it is the clearest signal of what the programme actually bought. Fifty specialists over six months is on the order of 50,000 person-hours, and 100,000 items across that is roughly two items per person-hour, or about thirty minutes of attention each. That is deliberative clinical review rather than volume labelling, and it is the correct shape for a 99% accuracy target. It also means a throughput clause expressed as output per shift needed to be set against that reality rather than against back-office benchmarks.

The measurement question remains the useful lesson. Delivered accuracy of 99.2% against a 99.0% threshold is a margin of 200 items in 100,000, and no weekly audit of a few hundred items could have resolved it either way. Whatever the programme actually used to substantiate the 99.2% — a cumulative audit across the full engagement, a much larger sample, or a complete gold-set pass — is the detail a buyer replicating this result most needs, and it is the one the outcome summary does not record.

Why Do Organizations Work with Cynergy BPO on Contract Structuring?

Cynergy BPO is an independent, vendor-neutral outsourcing advisory firm headquartered in Manila, representing a vetted network of more than 100 Philippine providers. It maps requirements against performance data to produce a shortlist within days and manages competitive negotiation on the buyer’s behalf.

Who Is Cynergy BPO?

Cynergy BPO is an independent outsourcing advisory and consultancy firm headquartered in Manila, founded by industry veterans with more than 65 years of combined operational experience governing major global accounts. It specialises in connecting mid-market and enterprise organisations with vetted Philippine BPO providers across voice, back-office and AI data operations.

How Does Cynergy BPO Differ from Traditional Outsourcing Brokers?

Traditional brokers are transactional and are compensated by the providers they place, which shapes which provider is recommended. Cynergy BPO applies an advisory-led methodology, mapping exact technical, security and commercial requirements against performance data rather than against availability. In contract structuring in particular, an adviser paid by the placement has no reason to argue for a clause the provider will resist.

How Does Cynergy BPO’s Network of 100+ Vetted Philippine BPO Providers Benefit Organizations?

The network establishes what is actually achievable before a threshold is written into a contract. Knowing the accuracy levels comparable pods reach on comparable material, and the audit volumes providers will accept, is the difference between an SLA that holds and one that produces a credit dispute in month two.

How Does Cynergy BPO’s Advisory-Led Vendor Matching Process Work?

Requirements are mapped against operational and commercial criteria, a tailored shortlist of vetted providers is delivered within a few working days, and the firm then manages competitive proposal and negotiation processes on the buyer’s behalf. Service levels, audit protocols and the statistics behind them are settled as part of that process rather than after selection.

Why Do Organizations Use Cynergy BPO?

Because a service level agreement is the one part of an outsourcing engagement that is easy to write badly and expensive to renegotiate. Thresholds set without reference to what the market delivers, or without an audit plan capable of measuring them, create disputes that neither party wanted and that better drafting would have avoided.

Frequently Asked Questions

What accuracy threshold should an LLM training SLA require?

Enterprise agreements commonly specify 98.5% to 99.0% against a gold-standard key, and the figure should be chosen together with the audit volume that can measure it. A 98.5% threshold needs roughly 2,400 audited items to detect a one-point shortfall, so the threshold and the sampling budget are one decision.

How large should the audit sample be?

Large enough that the confidence interval is narrower than the shortfall you intend to act on. A few hundred items per week cannot distinguish 97% from 99%. If that volume is all the budget allows, trigger remedies on a rolling four-week average or a tolerance band rather than on a hard weekly threshold.

Which agreement statistic should the contract name?

Fleiss’ kappa where three or more raters judge the same items, Cohen’s kappa where exactly two do, or the mean of pairwise Cohen’s coefficients. Name one, and state how the overlap sample is drawn, since agreement rises when the sampled comparisons are easy.

What financial remedies are appropriate for non-compliance?

Proportional invoice credits, vendor-funded retraining, and termination rights for uncorrected failure are all standard. What matters more than the remedy is the trigger: a credit fired by a measurement that cannot resolve the threshold charges the vendor for sampling noise and will be disputed.

Should security clauses be zero tolerance?

Zero tolerance for concealment, yes. Zero tolerance for any detected event creates an incentive to detect less, since the cheapest route to zero reported events is conservative monitoring. Graduate the remedy so that self-reported events are treated mildly and concealed ones severely.

Why require two-hour breach notification when the law allows 72?

Because the 72 hours belongs to the buyer, not the vendor. Philippine and EU law both require notification to the regulator within 72 hours of knowledge. A short contractual clock leaves the buyer time to assess, decide and file within its own statutory deadline.

How often should SLA performance be reviewed?

Operational dashboards weekly, with formal governance monthly to review yield, drift and milestone velocity. Report every estimate with the sample size behind it, so the meeting can distinguish a real movement from ordinary variation.

What do specialised Philippine domain annotators cost?

Roughly $18 to $35 an hour fully loaded depending on the professional qualification, security overhead and contract model. Savings against onshore alternatives are usually quoted as programme figures rather than rate comparisons, so establish which basis a proposal is using before comparing two of them.

Share This
Jump to a Section

Unlock cost-efficient growth with expert BPO guidance!

Partner with Cynergy BPO to connect with top outsourcing providers.
Streamline operations, cut costs, and scale your business with confidence.

Book a Free Call
Image

Ralf Ellspermann is the Chief Strategy Officer (CSO) of Cynergy BPO and a globally recognized authority in business process and contact center outsourcing. With more than 25 years of experience advising enterprises and SMEs, he provides strategic guidance on vendor selection, CX optimization, and scalable outsourcing strategies across global markets. His expertise spans fintech, ecommerce and retail, healthcare, insurance, travel and hospitality, and technology (AI & SaaS) outsourcing.

A frequent speaker at leading industry conferences, Ralf is also a published contributor to The Times of India and CustomerThink, where he shares insights on outsourcing strategy, customer experience, and digital transformation.