

By: Ralf Ellspermann
25-Year, Multi-Awarded BPO Veteran
Published: 30 September 2026

Reviewed By: John Maczynski
Former EVP, World's Largest Contact Center
Updated: 30 September 2026
By aligning LiDAR, radar and camera streams into one object identity per frame. The methodology is sound; the precision is misplaced. Millisecond cross-sensor sync sits inside a LiDAR sweep lasting 50 to 200 milliseconds, during which the vehicle moves one to four metres — sixty-seven times the sync error.
Key Takeaways
- Millisecond timestamp alignment is precision at the wrong layer. A spinning LiDAR captures a frame over a full revolution, so at 10 Hz and 20 m/s the ego vehicle moves 2.00 metres during one sweep. A 1 ms cross-sensor sync error at 30 m/s closing is 0.03 metres — sixty-seven times smaller.
- So specify motion-compensated clouds, not tighter sync. Production autonomous driving stacks ship a distortion corrector that undistorts each sweep from per-point timestamps and ego velocity. Ask whether the clouds handed to annotators have been through it, and at which instant of the sweep the frame is stamped.
- Radar labour is association, not annotation. Cross-range uncertainty is 1.75 metres at 100 metres even for a 1° imaging radar, and 7 metres for a typical 4° forward unit. A human cannot derive object extent from radar, so it should be priced as verification rather than as labelling.
- Per-modality accuracy does not survive fusion. A fused object is correct only where every layer agrees. At 99% per modality across three layers, identity consistency falls to 97.0%; at 98.5% it falls to 95.6% — below the threshold the contract names.
- Which means the metric has to be measured at the fused level. The share of objects carrying one consistent identity across every sensor layer is a single query, and it is the only figure that measures what the engagement exists to produce.
- Run the cross-sensor consistency check before human review, not after. Automated checks that flag roughly 8% of objects can contain around 70% of true errors, finding about ten times more errors per hour of review — and a reviewer shown a disagreement resolves it rather than confirming a box already drawn.
What Does Sensor-Fusion Annotation Actually Involve?
Three different jobs, not one. LiDAR is where a human derives 3D extent and identity from geometry. Camera resolves what an object is rather than where. Radar supplies velocity and confirmation. Only the first supports independent labelling, and the deliverable is none of the three.
The modality table in general use is accurate about what each sensor is hard at. What it does not say is that the three present fundamentally different tasks to a human annotator, and that treating them as one priced activity both overpays for the easiest and underspecifies the output.

Figure 1. Three modalities, three different jobs.
The fourth row is the one that matters commercially. A perception model does not consume a LiDAR annotation, a camera annotation and a radar annotation; it consumes a fused object with one identity, one extent and one velocity. That artefact is produced by the association work between layers, which is precisely the part no row of the standard table describes and no per-modality quality figure measures.
Scoping follows from this. Price LiDAR cuboid and segmentation work as labelling, camera work as refinement and classification, radar work as verification, and specify the fused object’s identity consistency as the acceptance criterion. Four line items rather than one blended rate, with the fourth carrying the acceptance test.
Is Millisecond Timestamp Alignment the Right Precision?
It is finer than it needs to be, and it sits inside a much larger error nobody mentions. A spinning LiDAR builds a frame over one full revolution — 50 to 200 milliseconds — during which the vehicle travels one to four metres at highway speed. Cross-sensor sync of 1 ms covers 3 centimetres.
Aligning timestamps to the millisecond is the kind of claim that signals rigour, and taken on its own it is correct and worth doing. The difficulty is that it addresses the smaller of two timing errors, and the larger one determines whether an annotator can place a box accurately at all.

Figure 2. Ego displacement during one sweep, against the sync error it contains.
A rotating LiDAR does not photograph a scene; it sweeps one. Points recorded at the start of a revolution and points recorded at the end are separated by the full sweep period, and if the vehicle has moved between them the returns no longer describe a single geometric moment. At 10 Hz and 20 metres per second the displacement across one sweep is 2.00 metres. An annotator fitting a cuboid to that cloud is fitting a smear.
This is a solved problem in production, which is what makes it worth asking about. Open-source autonomous driving stacks ship a distortion corrector that reconstructs where each point should have been, using per-point timestamps together with ego velocity from the vehicle’s twist and optionally an inertial measurement unit. The question for a provider is not whether they can align timestamps but whether the clouds their annotators see have already been through that correction — and if not, whose job it is.
A second question follows from the same geometry. When a frame carries a single timestamp, which instant of the sweep does it refer to: the start, the middle or the end? At 10 Hz those differ by 100 milliseconds, which is 2 metres of ego motion at highway speed and is more than enough to misalign a camera projection against a LiDAR box. Both questions are answerable in a sentence by a provider who has done this work and are diagnostic of one who has not.
What Can a Human Actually Label from Radar?
Velocity and presence, not extent. Automotive radar azimuth resolution ranges from about 1 degree on large imaging arrays to tens of degrees on compact parts, which puts cross-range uncertainty at 1.75 metres at 100 metres in the best case and 7 metres for a typical forward unit.
The standard table describes radar’s challenge as low angular resolution and velocity noise, and proposes velocity overlay and cross-verification as the solution. That is the right answer, and it implies something about the nature of the work that is worth making explicit.

Figure 3. Cross-range uncertainty against range.
Cross-range uncertainty is twice the range times the tangent of half the azimuth resolution. Even a high-end imaging radar at 1 degree blurs a target across 1.75 metres at 100 metres — approximately a car’s width. A typical 4-degree forward unit blurs it across 7 metres, which is two lanes. Published resolution figures also assume static point targets at identical range under classic beamforming, so real-world separation is worse than the specification.
The consequence is that a human working radar returns is not deriving an object’s boundary; there is no boundary to derive. They are associating returns with objects already identified in LiDAR and camera, and checking that the Doppler velocity is consistent with the tracked motion. That is genuinely valuable work — Doppler is the one measurement neither camera nor LiDAR provides directly, and it is what separates two objects sharing a range and bearing but moving differently — and it is a different and faster task than labelling.
Commercially this means radar should not sit in the same rate tier as segmentation, and its quality metric should be association correctness rather than boundary precision. A provider quoting one blended rate across all three modalities is either overpricing the radar work or has not separated the tasks.
Does 99% Per-Modality Accuracy Survive Fusion?
No. A fused object is correct only where every layer agrees on its identity, so per-layer accuracies compound. At 99% per modality across three layers, fused identity consistency is 97.0%; at 98.5% it is 95.6% — below the threshold the same document specifies.
This is the sensor-fusion-specific version of a measurement problem, and it is the one that matters most here because identity consistency across layers is the entire point of fusion annotation.

Figure 4. Per-modality accuracy against fused identity consistency.
The arithmetic is unforgiving in a way that is easy to miss when each layer is inspected separately. Three layers at 99% leave 97.0% of fused objects fully consistent; the missing three points are objects where one sensor layer disagrees about what the object is or which track it belongs to. Those are not small errors distributed evenly — they are exactly the identity swaps and association failures that produce a model confidently tracking the wrong thing.
Errors are also correlated in practice, through shared occlusion and shared calibration, which means the independent-error curves should be read as a floor on the gap rather than a point estimate. The direction is robust regardless: whatever the per-layer figure, the fused figure is lower, and the contract names the higher one.
The fix is cheap and specific. Report the share of objects carrying one consistent identity across every sensor layer, per sequence. It is a single query against the delivered data, it measures the deliverable rather than its components, and it is the only figure that would have caught the failure mode this work exists to prevent. Three different thresholds appear in the same source document — 99% and above, 98.5% and above, and a 99.3% pilot baseline — so this is also the moment to settle on one number and name the metric it applies to.
Sensor-fusion annotation is not a volume numbers game; it is an exercise in data integrity. Matching an enterprise with an outsourcing partner that possesses precise toolchain expertise is the difference between a stalled AI model and a rapid market deployment.
— John Maczynski, CEO, Cynergy BPO
Data integrity is the right frame, and in fusion work it has a precise definition: one object, one identity, consistent across every layer and every frame of a sequence. That is measurable, and measuring it is what separates an integrity claim from an assertion. On toolchain expertise, the caution worth adding is that platform-specific fluency depreciates when platforms change hands or change price, whereas the ontology and the accumulated record of how ambiguous associations were resolved transfer between tools. Match on the second and treat the first as a convenience.
Where Should the Quality Check Sit in the Pipeline?
Before the human, not after. The described hierarchy — junior annotators bound, mid-level specialists cross-validate, seniors audit edge cases — puts cross-sensor validation after a label has already been committed, which turns review into confirmation rather than derivation.
The tiered structure is sound and the roles are right. The ordering is what wastes the most valuable signal in the pipeline.

Figure 5. Review effort against errors found, for two orderings.
Cross-sensor disagreement is the cheapest high-yield error detector available on this workload, and most of it can be computed rather than reviewed. Does the LiDAR cuboid contain the radar return? Is the Doppler velocity consistent with the track’s own motion? Does the camera projection of the 3D box overlap the corresponding image region? Each of these is a few lines of geometry, runs on every object, and needs no human at all.
Running those checks first changes what the human tier does. On illustrative figures — 8% of objects flagged, containing 70% of true errors — reviewing only the flagged set finds roughly ten times more errors per hour of review than examining everything. The exact capture rate depends on which checks are implemented, but the shape holds: scarce senior attention concentrated on disagreements beats the same attention spread across objects that already agree.
The second reason to reorder is behavioural rather than arithmetic. A reviewer shown an existing bounding box evaluates that box; a reviewer shown a conflict between two sensor layers has to resolve it independently. The first invites confirmation, the second requires judgement, and on subjective association calls that difference is substantial.
What Should Buyers Specify in a Sensor-Fusion Contract?
Seven things, most of which a capable provider can confirm in a sentence and none of which appear in a standard proposal. They concern the state of the data handed to annotators, what each modality’s labour actually is, and where quality is measured.
- Motion-compensated point clouds, with per-point timestamps preserved. Ego displacement across one sweep is one to four metres at highway speed. A box fitted to an uncorrected cloud is fitted to a smear.
- Which instant of the sweep a frame timestamp refers to. Start, middle and end differ by a full revolution — 100 ms at 10 Hz, or 2 metres of ego motion. It decides whether the camera projection lines up.
- Identity consistency measured at the fused level, per sequence. Three layers at 99% each give 97% fused. The contract should name the fused figure, because that is what the model consumes.
- Radar scoped and priced as association, not as labelling. Extent cannot be derived from returns with metres of cross-range uncertainty. Its metric is association correctness and its rate tier should reflect a faster task.
- Automated cross-sensor consistency checks ahead of human review. Geometry that runs on every object, flagging disagreements. It multiplies the yield of the senior tier roughly tenfold.
- One accuracy threshold, with the metric it applies to named. Three different figures appear across a single proposal in this market. Pick one, state whether it is precision, recall, IoU or identity consistency, and report it per class.
- The ontology handbook as a buyer-owned deliverable. It is what holds labels consistent across a rotation where no two shifts share a supervisor, and it survives a change of provider, platform or sensor generation.
How Did One Autonomous Delivery Startup Optimise Its Annotation?
An autonomous delivery vehicle startup running 60% over budget on fragmented domestic contractors, with engineers losing 38% of each sprint to fixing labels, assessed three Philippine providers and deployed a 60-person squad on a 24/7 staggered schedule inside its own toolchain — cutting annotation spend 59% and compressing iteration from fourteen days to forty-eight hours.

Figure 6. Reported outcomes from a 60-person sensor-fusion squad.
The capacity figure is the one worth deriving, because the engagement reports it as a problem rather than as a result. Engineers were spending 38% of each sprint fixing labels, which left 62% for the work they were hired to do. Ending the rework lifts that to all of it — a 61% increase in effective engineering capacity, which is a larger effect than the 59% spend reduction and a far more plausible explanation of four months of schedule returned.
The mechanism sits in the lessons learned. A standardised ontology handbook and a two-week parallel pilot locked in 99.3% before volume handoff, which is what stopped the rework rather than any change in hourly rate. This generalises with an edge: a buyer who moves to a cheaper provider without fixing the ontology relocates the rework instead of removing it, and will see the cost saving without the schedule benefit.
The ontology handbook also deserves attention as the artefact rather than the process. On a 24/7 staggered schedule no two shifts share a supervisor, so consistency cannot come from oversight; it has to come from a written standard that resolves the association questions in advance. That document is the thing that makes a pilot accuracy figure hold at volume, and it is the one deliverable from this engagement that would survive a change of provider, platform or sensor generation.
Why Do Organizations Work with Cynergy BPO on Sensor-Fusion Sourcing?
Cynergy BPO is an independent, vendor-neutral outsourcing advisory firm headquartered in Manila, representing a vetted network of more than 100 Philippine providers. It maps requirements against performance data to produce a shortlist within days and manages competitive negotiation on the buyer’s behalf.
Who Is Cynergy BPO?
Cynergy BPO is an independent outsourcing advisory and consultancy firm headquartered in Manila, founded by industry veterans with more than 65 years of combined operational experience governing major global accounts. It specialises in connecting mid-market and enterprise organisations with vetted Philippine BPO providers across voice, back-office, engineering support and AI data operations.
How Does Cynergy BPO Differ from Traditional Outsourcing Brokers?
Traditional brokers are transactional and are compensated by the providers they place, which shapes which provider is recommended. Cynergy BPO applies an advisory-led methodology, mapping exact technical, security and commercial requirements against performance data. On fusion work, where the decisive question is whether a provider measures anything at the fused level, that independence determines what gets asked.
How Does Cynergy BPO’s Network of 100+ Vetted Philippine BPO Providers Benefit Organizations?
The network establishes which providers have handled synchronised multi-sensor data rather than single-modality imagery, which receive motion-compensated clouds as a matter of course, and which can report identity consistency from prior engagements — three distinctions that decide dataset quality and appear in no capability deck.
How Does Cynergy BPO’s Advisory-Led Vendor Matching Process Work?
Requirements are mapped against operational, security and commercial criteria, a tailored shortlist of vetted providers is delivered within a few working days, and the firm then manages competitive proposal and negotiation processes on the buyer’s behalf. Data preparation state, per-modality scoping, the fused acceptance metric and ontology ownership are normalised across bids during that process.
Why Do Organizations Use Cynergy BPO?
Because the questions that determine whether a fusion dataset is usable — has the cloud been de-skewed, what is the frame timestamp anchored to, is identity measured at the fused level — are answerable in a sentence by a provider who has done the work, and are rarely asked by a buyer running a single procurement.
Frequently Asked Questions
How precisely do sensor streams need to be synchronised?
Millisecond cross-sensor alignment is more than sufficient — at 30 m/s closing speed it corresponds to 3 centimetres. The larger error is intra-sweep: a LiDAR frame is captured over a full revolution, so at 10 Hz and 20 m/s the vehicle moves 2 metres during one sweep. Ask for motion-compensated clouds.
What is point-cloud motion distortion and why does it matter for annotation?
A rotating LiDAR records points sequentially, so if the vehicle moves during the revolution the returns no longer describe one geometric instant and the scene smears. Production stacks correct this using per-point timestamps and ego velocity. An annotator fitting a cuboid to an uncorrected cloud is fitting distorted geometry.
Can annotators label objects directly from radar?
Not their extent. Azimuth resolution of 1 to several degrees puts cross-range uncertainty at 1.75 metres or more at 100 metres, which exceeds a car’s width. Radar work is association — linking returns to objects identified in LiDAR and camera and checking Doppler consistency — and should be scoped and priced as verification.
What accuracy threshold should a sensor-fusion contract specify?
One threshold, applied at the fused level, with the metric named. Per-modality accuracy compounds: three layers at 99% each give 97.0% fused identity consistency, and 98.5% each gives 95.6%. Reporting the share of objects with one consistent identity across every layer measures the actual deliverable.
How should quality assurance be sequenced across tiers?
Run automated cross-sensor consistency checks before human review rather than after. Geometry checks flag disagreements on every object at almost no cost, concentrating senior review where it pays. Reviewing after a box is drawn also invites confirmation, whereas resolving a flagged conflict requires independent judgement.
What file formats do Philippine sensor-fusion teams work with?
PCD, PLY and LAS for point clouds, alongside standard video and synchronised image sequences. The more consequential question is not the container but whether per-point timestamps and ego pose travel with the data, since both are required for motion compensation and for verifying frame alignment.
How long does onboarding a fusion annotation team take?
Six to eight weeks is commonly quoted, covering recruitment, toolchain integration, ontology training and parallel validation. That assumes a provider with an existing technical bench; where specialists are being recruited for the engagement, the specialist tier alone typically runs longer than the whole quoted window.
How is proprietary autonomous vehicle data protected?
ISO 27001 and SOC 2 Type II are certifications a provider holds, and the scope statement matters more than the certificate — confirm the delivery floor and tooling sit inside it. Data protection statutes are laws rather than certifications. Where scenes contain identifiable faces or plates they are personal data, and the Philippines holds no EU adequacy decision, so European footage requires standard contractual clauses and a transfer impact assessment.
Unlock cost-efficient growth with expert BPO guidance!
Partner with Cynergy BPO to connect with top outsourcing providers.
Streamline operations, cut costs, and scale your business with confidence.

Ralf Ellspermann is the Chief Strategy Officer (CSO) of Cynergy BPO and a globally recognized authority in business process and contact center outsourcing. With more than 25 years of experience advising enterprises and SMEs, he provides strategic guidance on vendor selection, CX optimization, and scalable outsourcing strategies across global markets. His expertise spans fintech, ecommerce and retail, healthcare, insurance, travel and hospitality, and technology (AI & SaaS) outsourcing.
A frequent speaker at leading industry conferences, Ralf is also a published contributor to The Times of India and CustomerThink, where he shares insights on outsourcing strategy, customer experience, and digital transformation.
