

By: Ralf Ellspermann
25-Year, Multi-Awarded BPO Veteran
Published: 30 September 2026

Reviewed By: John Maczynski
Former EVP, World's Largest Contact Center
Updated: 30 September 2026
Through perception data annotation, simulation scenario building and remote assistance desks. The distinction that decides scope is the mode: Philippine centres run remote assistance, where a stopped vehicle awaits guidance, and cannot run direct remote driving, which needs under 100 milliseconds glass-to-glass and the Pacific costs 182.
Key Takeaways
- Remote assistance and remote driving are different products, and only one can run from Manila. Direct remote driving needs glass-to-glass latency under about 100 milliseconds. Manila to the US east coast floors at 182 milliseconds on physics alone, so the network round trip exceeds the whole budget before a camera has encoded a frame.
- Which is exactly the work the industry already places here. Waymo’s Chief Safety Officer told the US Senate that fleet response agents work from the Philippines, and that “they provide guidance, they do not remotely drive the vehicles.” The mode distinction is the selling point, not a limitation to work around.
- “Accuracy above 99%” is not a perception metric and hides the failure that matters. On an illustrative street scene, a dataset that misses every pedestrian, cyclist and rider still scores 98.2% pixel accuracy and 0.947 mean IoU. Recall on those classes is zero.
- So contract on per-class recall by distance band, with IoU thresholds set per class. The field’s own metrics for 3D boxes are IoU at a threshold plus translation, scale and orientation error. Aggregate accuracy corresponds to none of them and is unenforceable.
- European coverage from Manila is an afternoon shift, not a graveyard. Central European business hours land at 15:00 to 23:00 Manila time; US Eastern lands at 21:00 to 05:00. The two are not the same staffing proposition and should not carry the same differential.
- The case study’s real result is capacity, not cost. Engineers spent 40% of the week on routine annotation. Removing it lifts core development from 60% of their time to all of it — a 67% increase in the capacity that moves a launch date, against a 58% cut on the annotation line.
What Do the Comparison Metrics Actually Commit To?
Rates of $10 to $16 an hour against $45 to $75 in-house, an eight to ten week ramp for 50 seats, multi-shift throughput and an operating-expense model. Each line is defensible. What turns them into a contract is the question sitting behind each one.
Comparison tables of this shape are the standard opening move in technical outsourcing proposals, and this one is more honest than most: the rate band is narrow enough to be useful, and the capital-to-operating expense shift is a real structural advantage rather than a rhetorical one.

Figure 1. The comparison table, and the question behind each row.
Two rows deserve particular attention. The ramp comparison is only fair if both sides start from the same place: sixteen to twenty-four weeks for an in-house build includes recruiting into a Western engineering organisation, while eight to ten weeks offshore usually assumes a provider with a standing bench. Ask which situation applies, because a provider recruiting specialists for your engagement is quoting a forecast rather than a plan.
The quality row is the one that matters most and is addressed last in most proposals. A single aggregate accuracy figure cannot express the failure modes that perception datasets exist to prevent, which the fourth section of this article takes up in detail. It is also, almost always, the figure that ends up written into the statement of work.
Can Teleoperation Really Run from Manila?
Remote assistance, yes. Direct remote driving, no. Published thresholds put safe remote driving at under about 100 milliseconds glass-to-glass, and the Manila-to-US network round trip alone floors near 182 milliseconds on the speed of light in fibre. That is physics, not a carrier problem.
This is the single distinction that determines what can honestly be sold, and it is worth drawing precisely rather than softening. “Ultra-low-latency command-and-control” describes one of two very different products, and the Philippines is well positioned for one of them and structurally excluded from the other.

Figure 2. Glass-to-glass latency against the remote driving threshold.
Direct remote driving means a human closing a continuous control loop: steering, braking and accelerating a moving vehicle from a console. The operator’s inputs must arrive while the scene they were made for still exists. Published guidance puts the requirement at roughly 100 milliseconds glass to glass, with operator performance degrading noticeably beyond 200 to 250. A Manila console reaches around 320 milliseconds before a human has reacted at all, and a vehicle at 40 km/h covers 2.4 metres in the network round trip alone.
Remote assistance is a different job. The vehicle detects a situation it cannot resolve, brings itself to a safe state, and asks a human for context — is that a construction zone or a stalled car, is this path passable, which of these two routes. The software retains the driving task throughout. Because nothing is moving while the human considers, a round trip of several hundred milliseconds is immaterial, and the work is bounded by judgement quality rather than by reaction time.
That is not a consolation prize. It is what the industry’s most advanced operator already places here: in Senate testimony, Waymo’s Chief Safety Officer confirmed that fleet response agents work from the Philippines, stating that “they provide guidance, they do not remotely drive the vehicles” and that the Waymo Driver “is in control of the vehicle at all times.” Philippine capability in this space has an unusually strong public proof point, and the way to use it is to name the mode rather than to claim latency the Pacific will not permit.
The practical rule for scoping: any workload where a human closes a control loop on a moving machine belongs in-region. Any workload where a human supplies context to a stopped or safely-operating machine offshores cleanly, and the further advantage is that time-zone spread becomes a feature rather than a constraint.
How Should a Remote Assistance Desk Be Sized?
From the share of operating time a vehicle spends in an assistance state, not from fleet size. At 3% and a target desk utilisation of 70%, one operator supports roughly 23 vehicles. At 6% that falls to about 12, and at 2% it rises to 35.
Desk sizing is where a remote operations contract is priced, and it is usually done by ratio rather than by load. Treating it as a queueing problem produces a defensible number and, more usefully, makes the sensitivity visible.

Figure 3. Vehicles per operator, by intervention rate.
The offered load is the fleet size multiplied by the share of operating time spent awaiting guidance. Published estimates put that at roughly 2% to 4% in complex urban environments, though it varies enormously with route difficulty, weather and the maturity of the autonomy stack. Sizing the desk at around 70% utilisation leaves headroom for the clustering that matters here: interventions are not independent, because the conditions that cause one — a closed street, an unusual vehicle, heavy rain — cause several at once.
Two contract terms follow from the shape of the curve. First, the intervention rate should be measured and reported, because it is the variable that determines cost and it improves as the autonomy stack matures — a desk sized on year-one telemetry will be overstaffed by year two, and whoever captures that saving should be settled in advance. Second, the service level should be stated as time to first response under a defined queue depth, not as an operator-to-vehicle ratio, since the ratio is an output of the model rather than an input to it.
Capture every intervention as data or you will pay for the same lesson twice. The whole point of a good remote-ops partner is that today’s exception becomes tomorrow’s autonomous behavior.
— John Maczynski, CEO, Cynergy BPO
This is the correct principle and it has a specific implementation. An intervention becomes training data only if what the operator saw, what they decided and why is captured in a structured form at the moment of the decision — not reconstructed later from a ticket. That means the assistance console writes a labelled record as a by-product of the workflow, with the scene, the options presented, the choice made and a reason code. It also means the intervention rate should be expected to fall on the routes the desk has already seen, which is both the measure of whether the loop is working and the reason the commercial model needs to anticipate its own success.
Is “Accuracy Above 99%” a Meaningful Perception Metric?
No, and on safety-critical data it is actively misleading. Because road, building and vegetation dominate a street scene, a dataset that misses every pedestrian, cyclist and rider still scores 98.2% pixel accuracy and 0.947 mean IoU. Recall on those classes is zero.
Perception datasets exist to prevent a specific class of failure: a machine not seeing a person. An aggregate accuracy figure is structurally incapable of detecting that failure, which makes it the wrong number to write into a statement of work and an unfortunately common one.

Figure 4. Three scores for one dataset, and what each one sees.
The arithmetic is driven by class imbalance. On an illustrative street scene, road, building, vegetation and car account for roughly three quarters of all pixels, while people, riders, bicycles and motorcycles together account for under 2%. A dataset that handles the dominant classes perfectly and misses every vulnerable road user therefore loses less than two points of pixel accuracy. Averaging IoU across classes helps less than it appears to: one class at zero among nineteen still leaves a mean above 0.94.
The metrics that do detect it are per-class and conditional. Recall on each vulnerable road user class, reported separately. The same recall banded by distance, because a pedestrian at 60 metres is both harder to annotate and the one that determines stopping distance. IoU thresholds set per class rather than globally, since the field’s own benchmarks use a stricter threshold for vehicles than for pedestrians precisely because small objects are harder to localise. And for 3D boxes, the translation, scale and orientation errors that benchmark suites report alongside detection scores.
None of this makes a Philippine provider a worse choice — the work is done well here, and the same critique applies to an in-house team measuring itself the same way. It changes what a buyer asks for. A provider that can report per-class recall by distance band from an existing engagement is demonstrating a quality system built for perception data. A provider that reports a single accuracy figure is reporting a number that would look identical whether or not the dataset is fit for purpose.
Which Shifts Does Offshore Coverage Actually Require?
It depends entirely on which client time zone is being covered, and the answer is counterintuitive. Central European business hours fall at 15:00 to 23:00 in Manila — an afternoon shift. US Eastern hours fall at 21:00 to 05:00, and US Pacific at midnight to 08:00.
Offshore delivery is assumed to mean night work, and for North American coverage it does. For European coverage it does not, and the difference is large enough to change both staffing cost and attrition.

Figure 5. Client business hours, expressed in Manila local time.
This matters because shift allocation drives retention, and retention drives cost on exactly the roles that take longest to replace. A specialist covering European hours works mid-afternoon to late evening in Manila, which sits at the favourable end of the published attrition bands. The same specialist covering the US west coast works midnight to 08:00, which sits at the punishing end. Paying an identical shift differential for both, which is the common practice, systematically underpays one and overpays the other.
It also suggests a sequencing strategy for a first engagement. A European-hours pilot proves the delivery model on the easier shift, with lower churn and a smaller differential, before a North American desk is stood up. Where a programme needs both, the specialist tiers with the longest replacement pipelines should sit on the European window and the shorter-pipeline roles on the North American one, which is the opposite of how pods are usually allocated.
One scheduling detail worth writing into the contract: Manila is UTC+8 and observes no daylight saving, so every one of these windows moves an hour later when northern-hemisphere clocks go back. A coverage commitment expressed in client local time changes the Manila roster twice a year; expressed in Manila time, it silently drops an hour of client coverage.
What Should Buyers Specify in a Perception and Robotics Contract?
Seven things, most of which cost the provider nothing and none of which appear in a standard proposal. They concern how quality is measured, which teleoperation mode is in scope, and where the data may lawfully go.
- Per-class recall by distance band, not aggregate accuracy. A dataset missing every vulnerable road user scores 98% on the aggregate. Recall on those classes is the only measure that detects it.
- IoU thresholds set per class, with translation, scale and orientation error for 3D boxes. These are the field’s own metrics. A contract written in them is enforceable; one written on “accuracy” is not.
- Which teleoperation mode is in scope, stated explicitly. Remote assistance against a stopped vehicle offshores cleanly. Direct remote driving of a moving vehicle does not, and no network contract changes that.
- The intervention rate, measured and reported, with desk size derived from it. Ratios are outputs. Specify time to first response under a defined queue depth, and agree who captures the saving as the rate falls.
- A structured intervention record written at the moment of decision. Scene, options, choice and reason code, captured as a by-product of the workflow. Reconstructed tickets do not become training data.
- Redaction of faces and plates before transfer, or a stated legal basis where the task needs them. Street-scene data is personal data. The Philippines holds no EU adequacy decision, so European footage needs standard contractual clauses and a transfer impact assessment.
- Business continuity across genuinely decorrelated sites. Typhoon tracks commonly cross both Luzon and the Visayas, so a two-site plan spanning only those is less independent than it looks.
How Did One Industrial Enterprise Scale Robotics Data Operations?
An industrial automation firm with a LiDAR point-cloud backlog and local hiring costs 60% over budget assessed three specialised Philippine providers and stood up a 70-person engineering support unit on a staggered 24/7 model, working inside the client’s own cloud environment — cutting annotation cost 58% and compressing iteration cycles from weeks to days.

Figure 6. Reported outcomes from a 70-person engineering support unit.
The headline the engagement buries is the capacity release. Internal engineers were spending 40% of each week on routine annotation, which means 60% remained for core algorithm development. Removing the annotation load lifts that to 100% — a 67% increase in the capacity that actually moves a launch date. That is a larger effect than the 58% saving on the annotation line, and it is the far likelier explanation of a four-month pull-forward in the deployment schedule.
Stating it that way also changes how the business case should be built. A cost comparison between $10 to $16 an hour offshore and $45 to $75 in-house understates the value, because the in-house figure is the cost of the engineer’s time rather than the value of what that engineer would otherwise have built. For a team whose bottleneck is algorithm development rather than headcount budget, the capacity argument is both larger and easier to defend.
The two-week parallel validation phase is the implementation detail most worth copying. Running both teams against the same data before the in-house one stands down is where an undocumented workflow becomes a transferable one, and it is the concrete form of the engagement’s stated lesson about front-loading procedure documentation. Working inside the client’s own cloud environment is the second: it keeps the data and the toolchain on the buyer’s side of the boundary by construction rather than by contract.
Why Do Organizations Work with Cynergy BPO on Technical Outsourcing?
Cynergy BPO is an independent, vendor-neutral outsourcing advisory firm headquartered in Manila, representing a vetted network of more than 100 Philippine providers. It maps requirements against performance data to produce a shortlist within days and manages competitive negotiation on the buyer’s behalf.
Who Is Cynergy BPO?
Cynergy BPO is an independent outsourcing advisory and consultancy firm headquartered in Manila, founded by industry veterans with more than 65 years of combined operational experience governing major global accounts. It specialises in connecting mid-market and enterprise organisations with vetted Philippine BPO providers across voice, back-office, engineering support and AI data operations.
How Does Cynergy BPO Differ from Traditional Outsourcing Brokers?
Traditional brokers are transactional and are compensated by the providers they place, which shapes which provider is recommended. Cynergy BPO applies an advisory-led methodology, mapping exact technical, security and commercial requirements against performance data. On perception work, where the decisive question is which quality metrics a provider can actually report, that independence determines what gets asked.
How Does Cynergy BPO’s Network of 100+ Vetted Philippine BPO Providers Benefit Organizations?
The network establishes which providers have worked specific sensor stacks rather than generic imagery, which can evidence per-class recall from prior engagements, and which already run assistance desks on client telemetry — distinctions invisible in a capability deck and decisive on a safety-critical dataset.
How Does Cynergy BPO’s Advisory-Led Vendor Matching Process Work?
Requirements are mapped against operational, security and commercial criteria, a tailored shortlist of vetted providers is delivered within a few working days, and the firm then manages competitive proposal and negotiation processes on the buyer’s behalf. Quality metrics, teleoperation mode, shift windows and the data transfer mechanism are settled during that process rather than after selection.
Why Do Organizations Use Cynergy BPO?
Because the questions that determine whether a perception dataset is fit for purpose — which metric governs, which classes are reported separately, which teleoperation mode is in scope — are not in any proposal, and a single procurement rarely knows to ask them before the first delivery.
Frequently Asked Questions
Can Philippine teams perform teleoperation for autonomous vehicles?
They can perform remote assistance, where a vehicle brings itself to a safe state and asks a human for context while the software keeps the driving task. Waymo has confirmed it staffs fleet response agents in the Philippines on exactly that basis. Direct remote driving of a moving vehicle requires under about 100 milliseconds glass to glass and is not achievable across the Pacific.
Why can a leased line not solve teleoperation latency?
Because the constraint is distance. Manila to the US east coast floors near 182 milliseconds round trip at the speed of light in fibre, and observed links run 200 to 234. A leased line removes contention and jitter, which matter, but no commercial arrangement shortens the path.
What annotation quality metrics should a perception contract use?
Per-class recall on vulnerable road users, banded by distance; IoU thresholds set per class; and for 3D boxes the translation, scale and orientation errors that benchmark suites report. A single accuracy figure is not one of the field’s metrics and cannot detect a dataset that systematically misses rare classes.
How is a remote assistance desk sized?
From offered load rather than fleet size. Multiply fleet size by the share of operating time spent awaiting guidance — published estimates run 2% to 4% in complex urban settings — and staff to roughly 70% utilisation. At 3% that is about 23 vehicles per operator, with headroom for the clustering that weather and road closures produce.
Which shift does offshore coverage actually require?
It depends on the client’s time zone. Central European business hours fall at 15:00 to 23:00 Manila time, an afternoon shift at the favourable end of the attrition bands. US Eastern falls at 21:00 to 05:00 and US Pacific at midnight to 08:00. Manila observes no daylight saving, so each window moves an hour when northern clocks change.
What are the data protection implications of sending street-scene data offshore?
Street scenes contain faces and licence plates, which are personal data. Redact before transfer wherever the annotation task permits it. Where it does not, note that the Philippines holds no EU adequacy decision, so European footage requires standard contractual clauses and a transfer impact assessment — routine to arrange and expensive to discover late.
How robust is Philippine business continuity against typhoons?
Established centres maintain generators, carrier-diverse fibre and tested work-from-home failover, all of which are real. The point to probe is site independence: typhoon tracks commonly cross both Luzon and the Visayas, so a two-site plan spanning only those regions is less decorrelated than it appears. Ask what the second site is genuinely independent of.
What hourly rates apply to Philippine engineering support work?
Bands of $10 to $16 fully loaded are quoted for technical support and annotation roles, against $45 to $75 and above for in-house equivalents. Rates quoted across Philippine AI and technical work in this category span $10 to $35 depending on seniority, so a figure needs a named role and a stated definition of fully loaded cost before it is comparable.
Unlock cost-efficient growth with expert BPO guidance!
Partner with Cynergy BPO to connect with top outsourcing providers.
Streamline operations, cut costs, and scale your business with confidence.

Ralf Ellspermann is the Chief Strategy Officer (CSO) of Cynergy BPO and a globally recognized authority in business process and contact center outsourcing. With more than 25 years of experience advising enterprises and SMEs, he provides strategic guidance on vendor selection, CX optimization, and scalable outsourcing strategies across global markets. His expertise spans fintech, ecommerce and retail, healthcare, insurance, travel and hospitality, and technology (AI & SaaS) outsourcing.
A frequent speaker at leading industry conferences, Ralf is also a published contributor to The Times of India and CustomerThink, where he shares insights on outsourcing strategy, customer experience, and digital transformation.
