Image

What Volume Discounts Can Enterprises Expect for Million-Row LLM Annotation Projects in the Philippines?

Image

By: Ralf Ellspermann
25-Year, Multi-Awarded BPO Veteran
Published: 28 September 2026

Image

Reviewed By: John Maczynski
Former EVP, World's Largest Contact Center
Updated: 28 September 2026

Published ladders run 15% to 38% off base rate, activating around 250,000 rows. But the base rate band spans 2.67 times, so provider choice moves price further than any discount — and a 38% discount off one list price is unverifiable against another. Normalise every bid to dollars per row after discount.

Key Takeaways

  • The discount is the smaller of the two levers, and the only one buyers pull. Base rates of $0.03 to $0.08 a row are a 2.67-times spread. The largest published discount is 38%. An undiscounted $0.03 beats a maximally discounted $0.08 by 39.5%.
  • A discount percentage carries no information across bids. It is quoted against a list price only that provider sets and no buyer sees. A 38% discount off $0.08 is arithmetically identical to a $0.0496 list price with nothing off.
  • The quoted rates and throughput imply $1.12 to $5.00 per person-hour. Against $12 to $22 quoted for Philippine AI annotation labour elsewhere in this category. At $15 an hour, $0.08 a row buys 19 seconds of attention — a simple classification price, not a complex-row price.
  • Retroactive volume brackets create cliffs worth exploiting. Above roughly 88% of any threshold, ordering up to the threshold costs less in total. At 920,000 rows you pay $34,500; at a full million you pay $33,000 and get 80,000 extra rows.
  • A flat per-row price makes the hard rows unprofitable to do properly. Shaving effort on the 10% of rows that need thought cuts delivery cost 19.2% — more than a whole published bracket — and moves no throughput or agreement metric at all.
  • The stated turnaround does not follow from the stated throughput. At 15,000 rows a day a million rows takes 13.3 weeks, not 6. Six weeks would need 33,333 a day, above the maximum quoted in the same paragraph.

How Do Philippine Providers Structure Volume Tiers?

As cumulative row brackets with a discount attached to each: roughly 0% to 10% below 250,000 rows, 12% to 20% to half a million, 22% to 28% to a million, and 30% to 38% above it. The brackets are conventional. What sits behind each one is where the money moves.

The structure reflects a real cost curve. Guideline development, calibration and pilot supervision are largely fixed, so spreading them over more rows genuinely lowers unit cost, and the first two brackets are mostly that amortisation showing up in the price. The upper brackets are a different thing: they are paid for by commitment, and it is worth being precise about who is conceding what.

Figure 1. The published discount ladder, and the question behind each rung.

The phrase to examine is in the top bracket. “Long-term labour locking” describes an obligation the buyer takes on — minimum volumes, notice periods, and often a shortfall clause — in exchange for the deepest discount. That may well be a good trade, and it should be priced as one. A 38% discount against a three-year minimum is a different instrument from a 38% discount on a single delivery, and only one of them is a discount in the ordinary sense.

The second thing to check at each rung is what funds it. A provider moving from 22% to 34% has to find the margin somewhere, and the two easiest places are the review tier and the time per row. Neither shows up in a throughput report. Asking for the quality-assurance ratio at each bracket — reviewers per annotator, secondary review percentage, adjudication headcount — converts an invisible trade-off into a contractual one.

Do the Quoted Per-Row Rates Cover the Work Being Described?

Not as stated. A base rate of $0.03 to $0.08 a row, combined with 50 annotators producing 15,000 to 25,000 rows a day, implies revenue of $1.12 to $5.00 per person-hour. Philippine AI annotation labour is quoted at $12 to $22 an hour in this same category.

This is the most important check a buyer can run on any per-row proposal, and it takes one line of arithmetic. Multiply the rate by the volume, divide by the person-hours the stated throughput implies, and see what hourly figure comes out the other end.

Figure 2. What the quoted prices pay for an hour of work.

Nothing about the conclusion is subtle. A price that returns $3.00 per person-hour will not staff an office-based graduate annotator in Manila, once facilities, supervision, quality assurance, benefits and margin are included. Two of the three quoted quantities have to move.

The reconciliation is almost certainly the task definition. At $15 an hour fully loaded, $0.08 a row buys 19 seconds of attention per row. Nineteen seconds is a realistic budget for simple binary classification or single-entity tagging — exactly the work the source paragraph names before it starts describing complex semantic parsing. Work backwards from realistic labour rates and complex rows price at roughly $0.11 to $0.35 each, between 1.4 and 4.4 times the top of the quoted band.

The practical consequence is that $0.03 to $0.08 is a usable benchmark for the simplest tier of annotation work and badly misleading as a planning figure for an LLM evaluation corpus. A buyer budgeting a million complex rows from the published band will be out by a factor of three to four, and will discover it at the point where the provider proposes a change order. Ask any per-row quote to name the task, the expected seconds per row, and the review depth it assumes, then run the same division before signing.

Is the Volume Discount the Biggest Lever on Price?

No. The base rate band spans 2.67 times while the largest published discount is 38%, so which provider is chosen moves the price further than how hard the discount is negotiated. An undiscounted $0.03 a row beats a maximally discounted $0.08 by 39.5%.

Procurement attention tends to concentrate on the discount because it is the number that feels negotiable. It is also the number that carries the least information.

Figure 3. Two levers, drawn to the same scale.

The structural problem is that a discount is quoted against a list price the provider sets and the buyer never sees. Provider A offering 38% off $0.08 and Provider B offering 10% off $0.055 land within a fraction of a cent of each other, and the one with the more impressive discount is marginally the more expensive. Nothing in the percentage reveals this, because the two list prices are not observable in the same market.

Three habits fix it. Normalise every proposal to dollars per row after all discounts, for an identical task definition and an identical quality regime — without those two controls the comparison is still meaningless. Ask each provider for its rate at the volume you actually intend to buy rather than for a ladder, which removes the list-price question entirely. And treat a large quoted discount as a prompt to check the list price rather than as evidence of a good deal.

Do the Bracket Boundaries Create Cliffs Worth Exploiting?

Yes, where the discount applies retroactively to the whole volume, which is how these tables are normally read. Above roughly 88% of any threshold, ordering up to the threshold costs less in total than ordering what you need.

This is the rare case where the arithmetic hands the buyer a free option, and it is worth checking before the order quantity is finalised rather than after.

Figure 4. Total cost against rows ordered, across the one-million boundary.

At a $0.05 base rate and midpoint discounts, 920,000 rows in the 25% bracket cost $34,500. A full million rows in the 34% bracket cost $33,000. The extra 80,000 rows are not merely free; they come with a $1,500 rebate. The break-even sits at 880,000 rows, and the same rule holds at each boundary — 88.4% of 250,000, 89.3% of 500,000, 88.0% of a million.

Two responses are available and a buyer should consider both. The first is to use the cliff: if the corpus is within 12% of a threshold, extend the scope to reach it, which is usually easy because there is always more data worth labelling. The second is to negotiate it away by asking for marginal-tier pricing, where each bracket’s discount applies only to the rows inside that bracket. Most providers will agree, because the cliff is an artefact of how the table was drawn rather than a deliberate commercial design, and it creates awkward incentives in both directions.

What should not happen is neither. A retroactive bracket table that nobody has read carefully will eventually produce an invoice where a smaller order cost more than a larger one, and that conversation is harder after the fact than before.

What Does a Flat Per-Row Price Do to the Hard Rows?

It makes them unprofitable to do properly. Every row pays the same regardless of difficulty, so the margin on rows requiring genuine judgement is negative and every escalation costs the provider money. Reducing effort on the hardest 10% cuts delivery cost 19.2%, more than a whole published bracket.

This is the quietest economics in the contract, and it runs directly against everything the quality section of the same proposal promises.

Figure 5. Two ways to reduce delivery cost, only one of which is negotiated.

Take a corpus where 10% of rows need roughly four times the effort of an easy one. Blended effort is 1.30 units per row. Cut the treatment of those hard rows to 1.5 units — a quick decision rather than an adjudication — and blended effort falls to 1.05, a 19.2% cost reduction. That is larger than the second published bracket and about the same as the third, and the buyer wins none of it.

What makes it durable is that no metric in the standard quality regime detects it. Rows per day rises. Inter-annotator agreement rises, because a quick default is more consistent than a considered judgement. Golden-set accuracy holds, because golden items are drawn to be representative and 90% of them are easy. Every dashboard improves while the specific 10% of the corpus that carries the dataset’s informational value gets worse.

The fix is to stop pricing difficulty as if it were uniform. Price adjudicated ambiguity separately — a per-item fee for rows escalated and resolved through a defined route — so that surfacing a hard row becomes revenue for the provider rather than a cost. Cap it, audit it, and expect to pay it: a corpus that generates no escalations at all is not an easy corpus, it is an unexamined one.

Scaling data annotation past the million-row mark in the Philippines is no longer just about raw headcount; it requires sophisticated workforce orchestration, automated pre-labeling integration, and rigid quality assurance frameworks that protect enterprise intellectual property while driving down unit costs.

— John Maczynski, CEO, Cynergy BPO

Automated pre-labelling is the lever worth examining most closely there, because it lowers unit cost through two different mechanisms that are easy to confuse. One is genuine: the model handles the rows that were never in doubt, and human attention concentrates where it is needed. The other is anchoring — an annotator shown a proposed label confirms it more often than they would have chosen it, which raises throughput and agreement simultaneously while lowering independent judgement. Both show up as a unit-cost improvement. Only the first is one. The way to tell them apart is a small blind sample worked without pre-labels, compared against the pre-labelled stream.

How Long Does a Million-Row Project Actually Take?

Eight to 13.3 weeks on the throughput normally quoted, not six to eight. At 25,000 rows a day, 50 annotators need exactly 40 working days. At 15,000 a day they need 66.7 working days. A six-week delivery would require 33,333 rows a day, above the stated maximum.

Turnaround claims are worth checking for the same reason as rate claims: the arithmetic is fixed once the throughput is stated, and the two figures usually appear in the same sentence.

Figure 6. Weeks to deliver a million rows at the stated throughput.

There is one reading on which six weeks works. Running seven days a week, 25,000 rows a day delivers a million in 5.7 calendar weeks — but that is not 50 full-time annotators, it is a shift structure with substantially more people, and the quoted headcount would not cover it. The distinction between working days and calendar days is worth settling explicitly before a timeline enters a milestone schedule, because the two readings differ by 40%.

These figures are also floors rather than forecasts. They exclude ramp, rework of rejected batches, guideline revisions mid-project, and public holidays, all of which are ordinary and none of which are small on a quarter-long engagement. A plan built on the arithmetic alone will be late; a plan that adds nothing for rework will be late by more.

What Should Buyers Negotiate on a Million-Row Contract?

Seven terms, most of which cost the provider little and all of which are easier to set before signature than after. They concern the unit of pricing, the treatment of difficulty, and what the discount is being exchanged for.

  • Dollars per row after all discounts, for a named task and a named review depth. The discount percentage is not comparable across bids. The normalised unit price, for identical work, is.
  • Expected seconds per row, stated by the provider. It converts the quote into an implied hourly rate in one division, and it surfaces a task-definition mismatch before the first batch rather than at the first change order.
  • Marginal-tier pricing, or a deliberate decision to use the cliffs. Retroactive brackets mean a smaller order can cost more than a larger one. Either remove the effect or exploit it, but decide which.
  • A separate price for adjudicated ambiguity. A per-item escalation fee makes difficulty profitable to surface. Under a flat rate it is profitable to bury.
  • The quality-assurance ratio at each bracket, written into the agreement. Reviewers per annotator and secondary review percentage are the first things thinned to fund a deeper discount, and the only ones a contract can protect.
  • What the top bracket commits the buyer to, priced as an option. Minimum volumes, notice periods and shortfall clauses are the consideration for the deepest discount. They belong in the comparison.
  • Whether quoted weeks are working days or calendar days. The two readings of the same throughput differ by about 40%, which is the difference between an on-time delivery and a missed quarter.

Why Do Organizations Work with Cynergy BPO on High-Volume Annotation?

Cynergy BPO is an independent, vendor-neutral outsourcing advisory firm headquartered in Manila, representing a vetted network of more than 100 Philippine providers. It maps requirements against performance data to produce a shortlist within days and manages competitive negotiation on the buyer’s behalf.

Who Is Cynergy BPO?

Cynergy BPO is an independent outsourcing advisory and consultancy firm headquartered in Manila, founded by industry veterans with more than 65 years of combined operational experience governing major global accounts. It specialises in connecting mid-market and enterprise organisations with vetted Philippine BPO providers across voice, back-office and AI data operations.

How Does Cynergy BPO Differ from Traditional Outsourcing Brokers?

Traditional brokers are transactional and are compensated by the providers they place, which shapes which provider is recommended. Cynergy BPO applies an advisory-led methodology, mapping exact technical, security and commercial requirements against performance data. On volume pricing, where the decisive comparison is a normalised unit price rather than a headline discount, that independence determines which number gets compared.

How Does Cynergy BPO’s Network of 100+ Vetted Philippine BPO Providers Benefit Organizations?

The network makes list prices comparable. A single buyer sees one quote and one discount ladder; a firm holding rates across more than 100 providers can tell whether a 38% discount lands above or below the market for the same task definition, review depth and volume commitment.

How Does Cynergy BPO’s Advisory-Led Vendor Matching Process Work?

Requirements are mapped against operational, security and commercial criteria, a tailored shortlist of vetted providers is delivered within a few working days, and the firm then manages competitive proposal and negotiation processes on the buyer’s behalf. Task definition, seconds per row, escalation pricing and bracket structure are normalised across bids during that process.

Why Do Organizations Use Cynergy BPO?

Because a per-row quote is not comparable to another per-row quote without the task definition, the review depth and the escalation treatment behind it. Making four proposals comparable is most of the work, and it is the part a buyer running a single procurement cannot do alone.

Frequently Asked Questions

What volume discount should a million-row project expect?

Published ladders reach 30% to 38% above a million rows. Treat that as a starting point rather than an outcome, and compare the resulting dollars per row against other bids for the same task — the spread between providers’ list prices is wider than the spread across the whole discount ladder.

Is $0.03 to $0.08 a row realistic?

For simple classification or single-entity tagging, yes. At a realistic fully loaded labour rate it corresponds to roughly 19 to 24 seconds per row. For complex semantic work requiring judgement, the implied hourly rate falls below what Philippine delivery costs, and realistic pricing runs closer to $0.11 to $0.35 a row.

How long does a million-row annotation project take?

At the throughput commonly quoted — 50 annotators at 15,000 to 25,000 rows a day — between 8 and 13.3 working weeks. Six weeks requires 33,333 rows a day on a five-day week, above the usual stated maximum, or seven-day running with more than 50 people.

Should volume discounts apply retroactively or marginally?

Marginally, in most cases. A retroactive bracket creates a cliff where ordering more costs less in total — above about 88% of a threshold, buying up to the threshold is cheaper outright. Marginal-tier pricing removes the anomaly and is rarely resisted.

What hidden costs appear on high-volume annotation contracts?

Guideline development and calibration billed outside the rate, platform licensing surcharges, re-labelling penalties, shift differentials for continuous running, and change orders where the delivered task proves harder than the quoted one. The last is the largest and the most avoidable, by fixing the task definition and expected seconds per row at signature.

Should a million-row project be fixed-price or time-and-materials?

Fixed price gives budget certainty and transfers volume risk to the provider. It does not transfer quality risk, and it creates a standing incentive against escalating ambiguous rows, since under a flat per-row price every escalation is a cost. A fixed price with a separately priced escalation route captures most of the predictability without that effect.

How should quality be protected while taking a deep discount?

Write the quality-assurance ratio into the agreement at each bracket, rather than the discount alone. Secondary review percentage and reviewers per annotator are what a deeper discount is usually funded from, and they are invisible in throughput reporting until a downstream model shows the damage.

What security and compliance standards apply to enterprise annotation data?

ISO 27001 and SOC 2 Type II are certifications a provider holds, and their scope statements matter more than their existence. GDPR and HIPAA are statutes, not certifications — nothing is certified against them, and the obligations fall on the buyer as controller. For EU personal data, note that the Philippines has no adequacy decision, so transfers require standard contractual clauses and a transfer impact assessment.

Share This
Jump to a Section

Unlock cost-efficient growth with expert BPO guidance!

Partner with Cynergy BPO to connect with top outsourcing providers.
Streamline operations, cut costs, and scale your business with confidence.

Book a Free Call
Image

Ralf Ellspermann is the Chief Strategy Officer (CSO) of Cynergy BPO and a globally recognized authority in business process and contact center outsourcing. With more than 25 years of experience advising enterprises and SMEs, he provides strategic guidance on vendor selection, CX optimization, and scalable outsourcing strategies across global markets. His expertise spans fintech, ecommerce and retail, healthcare, insurance, travel and hospitality, and technology (AI & SaaS) outsourcing.

A frequent speaker at leading industry conferences, Ralf is also a published contributor to The Times of India and CustomerThink, where he shares insights on outsourcing strategy, customer experience, and digital transformation.