

By: Ralf Ellspermann
25-Year, Multi-Awarded BPO Veteran
Published: 30 September 2026

Reviewed By: John Maczynski
Former EVP, World's Largest Contact Center
Updated: 30 September 2026
Through per-unit rates or dedicated seats at $1,800 to $2,800 a month, which reconciles exactly with the $10 to $16 hourly band quoted alongside it. The published unit table cannot be used as it stands, because its three rows are priced per object, per frame and per polygon — three denominators that cannot be compared.
Key Takeaways
- The three price rows use three different units, so no two can be compared. 2D boxes are quoted per object, 3D cuboids per point-cloud frame and segmentation per complex polygon. A frame contains many objects and a segmented frame many instances, so the table cannot be added up or budgeted from.
- As written, the 3D row prices below the 2D row. At $0.525 a frame and 20 annotated objects, LiDAR cuboids work out at $0.026 an object — roughly four times cheaper than the quoted 2D band. Past about ten objects a frame the row falls below the 2D band entirely.
- Restated per frame, the spread is closer to fifty times than to five. One frame at 20 objects and 30 segmentation instances costs about $2.00 in 2D boxes, $0.53 in cuboids as written, and $25.50 in segmentation. The FAQ’s “three to five times” understates the segmentation gap threefold.
- The seat and hourly bands, in contrast, reconcile precisely. $1,800 to $2,800 a month over 173 productive hours is $10.40 to $16.18 an hour, against the $10 to $16 quoted in the same article. That is a rare and welcome internal consistency.
- One question decides between a seat and a unit price: measured throughput. A $2,300 seat beats a $0.10 unit price above about 133 objects an hour, a $0.525 frame price above 25 frames an hour, and a $0.85 instance price above 16 instances an hour. The answer differs by tier.
- And neither model is more predictable; they fix different quantities. A seat fixes the bill and lets cost per unit float. Unit pricing fixes cost per unit and lets the bill float. Choose by which one the budget is actually exposed to.
What Do the Quoted Prices Actually Buy?
Plausible rates in three incompatible units. 2D bounding boxes at $0.05 to $0.15 per object, 3D cuboids at $0.30 to $0.75 per point-cloud frame, and semantic segmentation at $0.50 to $1.20 per complex polygon. The numbers are reasonable; the denominators make them unusable together.
Complexity tiering is the right structure for annotation pricing, and the tiers named are the right ones: geometric complexity, occlusion handling, temporal tracking across frames and edge precision genuinely drive time per item. The problem is arithmetic rather than commercial.

Figure 1. The published price table, with each row restated in a comparable unit.
A buyer reading this table cannot answer the question it exists to answer. Costing a corpus of, say, 50,000 frames requires knowing how many objects and instances each frame carries, and then converting every row into the same unit. Without that, the rows sit side by side implying a comparison that the units do not support.
The fix is a single sentence at the top of the table naming the unit, and then using it in every row. Per object is the most common convention and the easiest for a buyer to count. Per frame is defensible where scene density is stable. Either works; mixing them does not.
Why Does the 3D Row Price Below the 2D Row?
Because it is quoted per frame while the 2D row is quoted per object. At the midpoint of $0.525 a frame and 20 annotated objects, LiDAR cuboids work out at $0.026 an object, against a quoted 2D band of $0.05 to $0.15 — which would make 3D work the cheaper of the two.
This is the clearest symptom of the denominator problem and it is worth working through, because the conclusion is obviously wrong and the arithmetic is obviously right.

Figure 2. The cuboid price expressed per object, against the quoted 2D band.
3D cuboid annotation is harder than 2D boxing by every measure the article itself names: spatial reasoning in a sparse point cloud, occlusion handling, and identity tracking across consecutive frames. Nobody prices it below a rectangle drawn on an image. So when the published figures imply exactly that, the figures are not describing what the table’s column headings say they describe.
The reading that makes the row sensible is that the 3D price is meant per object, in which case $0.30 to $0.75 an object sits sensibly at three to five times the 2D band — which is also what the FAQ claims. That interpretation is almost certainly the intent, and it is not what the table says.
The check a buyer can run in a minute: take any provider’s published table, pick one representative frame from your own data, and price it end to end under each row. If two rows cannot both be applied to the same frame without ambiguity, the table needs restating before it can be negotiated against.
What Is the Real Spread Between Annotation Tiers?
Far wider than the FAQ suggests, once everything is expressed per frame. A frame carrying 20 annotated objects and 30 segmentation instances costs roughly $2.00 in 2D boxes and $25.50 in semantic segmentation — about thirteen times, against a stated three to five.
Getting this ratio right matters more than getting any individual rate right, because it is what decides how a mixed corpus is budgeted and which parts of it are worth segmenting at all.

Figure 3. The three tiers normalised to one frame.
The mechanism is instance count. A 2D box is one object; a segmentation frame carries every instance in the scene, and each one is a polygon to be traced rather than a rectangle to be dragged. That is why segmentation dominates any corpus that contains it, and why the decision about what fraction of the corpus genuinely needs pixel-level or point-level labelling is a larger cost lever than the negotiation.
A budget built on the FAQ’s three-to-five-times ratio will be roughly threefold short on the segmentation portion. On a corpus where segmentation is a tenth of frames, that is the difference between a comfortable estimate and a mid-project overrun, and it is entirely avoidable with one worked example.
The request that settles it is small: ask each provider to price one representative frame from your own data under each annotation scheme, showing the object and instance counts they assumed. That removes every assumption in this analysis and gives a directly comparable number across bids.
Unit Pricing or a Dedicated Seat?
It depends on measured throughput, and the threshold differs by tier. At $2,300 a month over 173 hours, a seat beats a $0.10 object price above about 133 objects an hour, a $0.525 frame price above 25 frames an hour, and a $0.85 instance price above 16 instances an hour.
The tradeoff is usually framed qualitatively — unit pricing for stable datasets, seats for iterative ones — which is sound as far as it goes and leaves the buyer without a number. The number is straightforward.

Figure 4. The throughput at which a seat becomes cheaper than a unit price.
Divide the monthly seat cost by the quoted unit price to get the units a seat must produce to break even, then divide by productive hours to get an hourly rate. Above that figure the dedicated seat is cheaper; below it the unit price is. The only input a buyer lacks is measured throughput, and that is the one thing a provider can supply from an existing engagement.
Two things follow. First, the answer genuinely differs by tier, so a single contract model applied across bounding boxes, cuboids and segmentation is a guess rather than a decision — a mixed programme may well want seats for the dense work and unit pricing for the simple. Second, the two rate cards in this article reconcile exactly: $1,800 to $2,800 a month over 173 hours is $10.40 to $16.18 an hour, against the $10 to $16 quoted alongside. That consistency is worth noting because it is not the norm.
Which Model Is Actually More Predictable?
Neither. A dedicated seat fixes the monthly bill and lets cost per unit float; unit pricing fixes cost per unit and lets the bill float. If throughput comes in 30% below plan, a seat holds at $2,300 while cost per object rises 43%, and unit pricing holds the unit cost while the invoice falls.
Predictability is described in this market as a property of the dedicated model. It is better understood as a choice about which of two quantities to hold still, because both models are perfectly predictable in one dimension and perfectly exposed in the other.

Figure 5. What each model fixes, and what it lets move.
The asymmetry that matters is which quantity a given programme’s budget is actually exposed to. A team with a fixed headcount budget and an elastic delivery target is exposed to the bill, and a seat protects it. A team with a defined corpus to label and a cost per unit written into a business case is exposed to the unit cost, and unit pricing protects that. Describing either as simply more predictable skips the question that decides it.
There is a second-order effect worth naming on iterative work, and it runs against the usual recommendation. An evolving ontology lowers throughput, because annotators re-learn and rework. Under a seat, that shows up as a silently rising cost per unit that never appears on an invoice; under unit pricing, it shows up as the provider absorbing it. So the model usually recommended for iterative projects is also the one that hides the cost of iteration, which is an argument for measuring throughput per sprint regardless of which model is chosen.
Annotation pricing is never just about the lowest per-item quote; it is about total dataset integrity and predictable cost scaling. When enterprises partner with Cynergy BPO, we match them with providers whose transparent pricing models eliminate hidden overhead and secure long-term value.
— John Maczynski, CEO, Cynergy BPO
Transparency in a pricing model has a precise test, and it is not the number of line items. It is whether two providers’ quotes can be reduced to the same unit and compared. A table using three denominators is not opaque by intention — it is simply not yet comparable, and making it so is a five-minute exercise that changes what a buyer is able to negotiate. The same applies to the quality figure: this article quotes accuracy above 98.5% in its takeaways and 99.2% in its case study, which are different numbers describing an unnamed metric.
What Should Buyers Specify in an Annotation Pricing Agreement?
Seven things, most of which cost a provider nothing to supply and all of which make competing quotes comparable. They concern the unit, the throughput behind it, and what happens to the work that does not fit the schema.
- One unit, named once and used in every row. Per object, per frame or per instance. A table mixing them cannot be added, compared across bids, or turned into a budget.
- A worked example frame from your own data, priced under each scheme. With the object and instance counts the provider assumed. It removes every assumption and takes an afternoon.
- Measured throughput per seat, per annotation type. It is the only input needed to decide between a seat and a unit price, and the threshold differs by tier.
- Which quantity the contract fixes, stated explicitly. A seat fixes the bill; unit pricing fixes the cost per unit. Choose by which one your budget is exposed to, not by which sounds steadier.
- A defined route and price for edge cases and ambiguous objects. Under unit pricing these are where a provider loses money, so an escalation fee makes difficulty profitable to surface rather than costly.
- Toolchain licensing stated separately, then bundled if you prefer. Bundled into a seat it is a fixed cost; bundled into a unit price it is amortised over volume, which is expensive at low volume.
- One quality threshold, with the metric it applies to named. “Above 98.5%” and “99.2%” appear in the same document. Per-class recall, intersection over union at a stated threshold, or something else — but one of them.
How Did One Startup Stabilise Its Annotation Spend?
An autonomous vehicle developer running 65% over budget on fragmented micro-task platforms, three months behind on model training, assessed three Philippine providers and moved to a 50-person squad on a fixed monthly model inside its own cloud toolchain — cutting labelling expenditure 58% and compressing delivery from fourteen days to forty-eight hours.

Figure 6. Reported outcomes from a 50-person annotation squad.
The diagnosis in the challenge is worth separating from the remedy in the solution. What drove costs 65% over budget was per-item change orders and labelling errors, and both of those are ontology problems: a schema that changed after work began, and rules ambiguous enough to produce inconsistent output. A fixed monthly model does not remove change orders; it stops them appearing as separate line items, which is a different thing.
What plausibly did remove them was consolidation. Moving from many micro-task suppliers to one governed squad working a single settled ontology eliminates the source of the variation, and that is available under either pricing model. Attributing the result to the fixed-fee structure rather than to consolidation would lead the next buyer to change the contract and keep the problem.
The 99.2% consistency figure also needs a metric before it can be reused, and the same document quotes above 98.5% elsewhere. On annotation work those could describe per-class recall, intersection over union at a threshold, or agreement against a gold set, and they are not interchangeable. Ask which was measured, on what sample, and how often.
Why Do Organizations Work with Cynergy BPO on Annotation Pricing?
Cynergy BPO is an independent, vendor-neutral outsourcing advisory firm headquartered in Manila, representing a vetted network of more than 100 Philippine providers. It maps requirements against performance data to produce a shortlist within days and manages competitive negotiation on the buyer’s behalf.
Who Is Cynergy BPO?
Cynergy BPO is an independent outsourcing advisory and consultancy firm headquartered in Manila, founded by industry veterans with more than 65 years of combined operational experience governing major global accounts. It specialises in connecting mid-market and enterprise organisations with vetted Philippine BPO providers across voice, back-office, engineering support and AI data operations.
How Does Cynergy BPO Differ from Traditional Outsourcing Brokers?
Traditional brokers are transactional and are compensated by the providers they place, which shapes which provider is recommended. Cynergy BPO applies an advisory-led methodology, mapping exact technical, security and commercial requirements against performance data. On annotation pricing, where the decisive step is reducing four quotes to one comparable unit, that independence determines which number gets compared.
How Does Cynergy BPO’s Network of 100+ Vetted Philippine BPO Providers Benefit Organizations?
The network supplies the missing denominator. A single buyer sees one provider’s rate card; a firm holding delivery data across more than 100 providers can establish what throughput is actually achieved per annotation type, which is what converts incomparable unit prices into a decision.
How Does Cynergy BPO’s Advisory-Led Vendor Matching Process Work?
Requirements are mapped against operational, security and commercial criteria, a tailored shortlist of vetted providers is delivered within a few working days, and the firm then manages competitive proposal and negotiation processes on the buyer’s behalf. Pricing unit, measured throughput, edge-case treatment and the quality metric are normalised across bids during that process.
Why Do Organizations Use Cynergy BPO?
Because two annotation quotes are rarely expressed in the same unit, and reducing them to one is most of the work of choosing between them. A buyer running a single procurement has no throughput baseline to do that conversion against.
Frequently Asked Questions
What does robotics annotation cost in the Philippines?
Published bands run $0.05 to $0.15 for 2D bounding boxes, $0.30 to $0.75 for LiDAR cuboids and $0.50 to $1.20 for semantic segmentation, alongside dedicated seats at $1,800 to $2,800 a month. Establish the unit each price applies to before comparing, since published tables frequently mix per object, per frame and per instance.
How much more expensive is 3D annotation than 2D?
On a consistent per-object basis, roughly three to five times. On a per-frame basis, segmentation can reach thirteen times a 2D frame because it prices every instance in the scene rather than one object. The ratio depends entirely on which unit is used, which is why the unit has to be fixed first.
Should a buyer choose unit pricing or a dedicated FTE model?
Compare measured throughput against the break-even. At $2,300 a month over 173 hours, a seat beats a $0.10 object price above about 133 objects an hour, a $0.525 frame price above 25 frames an hour, and a $0.85 instance price above 16 an hour. The answer differs by annotation type.
Which model gives more predictable costs?
Both, in different dimensions. A seat fixes the monthly bill and lets cost per unit float; unit pricing fixes cost per unit and lets the bill float. A 30% throughput shortfall raises cost per object 43% under a seat and leaves it unchanged under unit pricing. Choose by which quantity your budget is exposed to.
Do the monthly and hourly rate bands agree?
Yes, and precisely. $1,800 to $2,800 a month across 173 productive hours is $10.40 to $16.18 an hour, against a quoted band of $10 to $16. The two describe the same labour at the same price, which makes either usable as a cross-check on the other.
How should edge cases be handled under unit pricing?
With a defined escalation route and a separate per-item price. Under a flat unit rate the hardest items are where a provider loses money, so surfacing them is costly and burying them is profitable. Pricing adjudicated ambiguity separately reverses that incentive at modest cost.
Are annotation toolchain licences included in the quoted rate?
Usually, and it is worth seeing them separately before agreeing to bundle. Inside a seat they are a fixed monthly cost; inside a unit price they are amortised across volume, which makes low-volume work disproportionately expensive. Ask for the standalone figure, then decide where it sits.
How is sensitive annotation data protected?
Through data loss prevention controls, restricted egress, controlled environments and non-disclosure terms, with ISO 27001 and SOC 2 Type II scope statements confirmed to cover the delivery floor and tooling. Where scenes contain identifiable people the data is personal data, and the Philippines holds no EU adequacy decision, so European footage requires standard contractual clauses and a transfer impact assessment.
Unlock cost-efficient growth with expert BPO guidance!
Partner with Cynergy BPO to connect with top outsourcing providers.
Streamline operations, cut costs, and scale your business with confidence.

Ralf Ellspermann is the Chief Strategy Officer (CSO) of Cynergy BPO and a globally recognized authority in business process and contact center outsourcing. With more than 25 years of experience advising enterprises and SMEs, he provides strategic guidance on vendor selection, CX optimization, and scalable outsourcing strategies across global markets. His expertise spans fintech, ecommerce and retail, healthcare, insurance, travel and hospitality, and technology (AI & SaaS) outsourcing.
A frequent speaker at leading industry conferences, Ralf is also a published contributor to The Times of India and CustomerThink, where he shares insights on outsourcing strategy, customer experience, and digital transformation.
