Compare cost per accepted label, not the headline rate
Data annotation outsourcing cost is the full amount required to turn raw assets into labels your model team can accept and use. It includes task design, tooling, annotator time, quality review, adjudication, rework, security controls, project management, and delivery. A low price per box or document can become expensive when the quote excludes difficult cases or when your engineers must repair the export.
The fastest way to make quotes comparable is to give every vendor the same representative sample, guideline version, output format, acceptance rules, security requirements, and volume assumptions. Then calculate the cost per accepted unit after quality control, not only the cost per submitted unit. This guide explains the main pricing models, modality-specific drivers, hidden costs, and a practical quote-normalization method for procurement, product, and ML operations teams.
Four pricing models and when each fits
Per-task or per-object pricing charges for an image, bounding box, text span, audio minute, video frame, or another defined unit. It works when the unit is stable and vendors count it the same way. The contract must say whether skipped, rejected, empty, or multi-object assets are billable and whether review is included.
Hourly pricing fits exploratory work, complex research, changing taxonomies, and tasks whose effort cannot be predicted from volume alone. It makes expert time visible, but the buyer needs role-based rates, time-reporting rules, productivity evidence from a pilot, and a cap or review point.
Dedicated-team pricing reserves capacity for a month or another period. It can suit continuous pipelines with frequent guideline changes and close collaboration. Compare productive capacity after training, holidays, supervision, QA, and expected utilization rather than multiplying headcount by a headline rate.
Outcome-based or managed-service pricing covers an agreed delivery, such as a validated dataset that passes defined acceptance criteria. This can simplify budgeting when scope is stable, but the quote must state assumptions, exclusions, change control, and who owns remediation. No pricing model is automatically cheaper. Choose the model that makes the project’s largest uncertainty measurable.
Public cloud labeling services illustrate why unit definitions matter: AWS describes object pricing around dataset objects reviewed, while its Ground Truth documentation distinguishes human labeling, consolidation, task types, and output metadata. Vendor proposals should be at least as explicit about what creates a billable event.
Build a cost equation before comparing vendors
Use a simple total-cost model:
- one-time setup: sample analysis, taxonomy and guideline work, workflow configuration, integrations, security onboarding, training, and calibration;
- production: billable units multiplied by the agreed unit or labor rate;
- quality: second-pass review, gold checks, overlap, adjudication, audit sampling, and reporting;
- exceptions: specialist review, difficult assets, privacy handling, source cleanup, and approved scope changes;
- buyer-side cost: your team’s preparation, clarifications, acceptance testing, engineering, and rework.
Ask every vendor to separate one-time and recurring charges. A setup fee is not inherently a problem: a well-designed pilot can prevent repeated production errors. The risk is an undefined fee or a cheap setup that moves taxonomy decisions back to the buyer after work begins.
Record the denominator for every rate. “Per image” means little if one image contains two objects in the pilot and twenty in production. For text, count documents, tokens, spans, or judgments consistently. For speech, distinguish media minutes from human working minutes. For video, state whether the unit is a clip, sampled frame, tracked object, keyframe, or annotated event.
Cost drivers by modality
Image annotation. Classification is usually simpler than dense object detection or segmentation, but cost still depends on objects per image, occlusion, image quality, class count, boundary rules, minimum object size, and required attributes. Polygon work increases with contour complexity; bounding-box work increases with object density and ambiguous overlap. Ask for pilot productivity by complexity band rather than one blended rate.
Video annotation. Duration alone is a poor estimator. Frame rate, sampling interval, number of tracks, shot changes, occlusion, interpolation, temporal consistency, event definitions, and whether reviewers watch the full sequence all matter. A one-minute static security clip and a one-minute crowded traffic sequence are not equivalent units.
Text and document annotation. Cost changes with document length, language, domain, taxonomy depth, context needed to decide a label, overlapping spans, entity relationships, and whether annotators must research external material. Documents that require layout context or contain tables may take longer than plain text with the same word count.
Audio and speech. Rate drivers include duration, noise, number of speakers, overlap, accent, language availability, transcription depth, timestamps, diarization, non-speech events, and privacy redaction. “Per audio minute” should specify whether the minute includes transcription, labels, review, and difficult-audio surcharges.
3D, sensor, and specialist data. Point clouds, medical images, geospatial data, legal documents, and scientific material can require specialized tools and qualified reviewers. Tool performance, rendering latency, coordinate systems, multi-sensor synchronization, and domain decision rules affect throughput. Do not compare a specialist task with a general crowd rate.
Quality cost belongs inside the quote
Quality is not a final inspection line added after labeling. The quote should explain training, qualification, calibration, in-process checks, reviewer coverage, adjudication, correction, and acceptance reporting. The NIST AI Risk Management Framework calls for documented metrics, test methods, measurement results, and relevant expertise; that is a useful procurement principle even when a project is not claiming formal compliance.
Define the quality unit and the error taxonomy before setting a target. A class label, box, polygon, text span, transcript segment, and tracked object need different measurements. State how severity, abstentions, partial correctness, empty assets, and ambiguous items are handled. A single “accuracy” percentage without its denominator and sampling method cannot normalize quotes.
Overlap is a design choice, not a universal requirement. Sending every item to three annotators may waste budget on clear cases, while single-pass labeling may be insufficient for subjective or high-risk decisions. A practical design can combine single annotation, targeted overlap, hidden gold tasks, risk-based review, and expert adjudication. Price each layer separately and define who pays when the guideline—not the annotator—caused the disagreement.
Hidden costs that change the accepted-unit price
The largest hidden cost is rework. Causes include incomplete instructions, inconsistent source data, late taxonomy changes, tool exports that do not match the training pipeline, and vendor corrections that cover only the audited sample. The contract should distinguish vendor error, buyer scope change, source-data defect, and genuinely ambiguous cases, with a remedy for each.
Tooling can add license fees, storage, compute, integration, custom interface work, data transfer, and export conversion. Confirm whether the vendor uses your platform, provides its own, or charges both a managed-service fee and a platform fee. Test a delivered file in the real training pipeline during the pilot; visual correctness in the annotation UI is not enough.
Management costs include onboarding, daily coordination, questions, reporting, change control, and staff replacement. These may be bundled or explicit. Ask who owns the guideline, who answers edge cases, how quickly decisions reach the workforce, and whether project management hours have a limit.
Security and privacy can affect price through restricted work environments, device controls, access reviews, background checks, data localization, logging, redaction, and secure deletion. These are not optional add-ons when the data requires them. The UK ICO’s current data protection by design guidance emphasizes deciding necessary personal-data use, access, and safeguards from the design stage. Buyers should confirm applicable requirements with their own privacy and legal owners.
A normalized quote comparison example
Consider a hypothetical pilot for 100,000 images. Each quote uses the same sample, taxonomy, output schema, and acceptance test. Vendor A quotes $0.08 per submitted image, a $2,000 setup fee, and excludes review. Vendor B quotes $0.11 per image including sampled review and a $3,000 setup fee. These figures are illustrative, not market benchmarks.
Vendor A’s production line is $8,000, producing a headline total of $10,000. Suppose the pilot indicates that 82% of units pass the buyer’s acceptance test without buyer-side repair, and correcting the remaining units is estimated at $2,700. The comparable cost is $12,700, or about $0.155 per accepted unit when divided by 82,000 accepted units.
Vendor B’s headline total is $14,000. Suppose 96% pass, included remediation brings the approved delivery to 98,000 units, and buyer-side acceptance work costs $600. The comparable total is $14,600, or about $0.149 per accepted unit. Vendor B has the higher unit rate but the lower normalized accepted-unit cost in this scenario. Change the assumptions, and the result may reverse—so require pilot evidence and sensitivity ranges rather than treating this example as a promise.
Also compare schedule, capacity, security fit, documentation, and correction turnaround. The financially preferred vendor is the one with the best risk-adjusted fit, not automatically the lowest or highest rate.
Questions to ask each vendor
- What exactly is the billable unit, and how are empty, skipped, rejected, duplicate, or multi-object items counted?
- Which setup, platform, storage, transfer, management, reviewer, expert, and remediation fees are included?
- What complexity assumptions support the rate, and what triggers a surcharge or change order?
- Which roles label, review, adjudicate, manage, and approve the delivery? What qualifications are required?
- How are annotators trained and calibrated when instructions change?
- Which quality metrics, denominators, sample sizes, severity levels, and acceptance thresholds appear in the report?
- Who pays for rework caused by vendor error, unclear guidelines, defective source data, or changed scope?
- What throughput and acceptance evidence came from the representative pilot, and what range should the buyer budget for?
- What access, logging, retention, deletion, confidentiality, and incident-handling controls apply?
- Can the vendor deliver the required schema, record identifiers, lineage, issue log, correction history, and final QA evidence?
Budget checklist before requesting quotes
- Freeze the task definition, taxonomy version, supported languages, modalities, and output schema.
- Provide a representative sample with easy, typical, difficult, ambiguous, and invalid items.
- State expected volume by complexity band, not only one total count.
- Define the billable unit and the accepted unit separately.
- Set quality measures, sampling or overlap rules, severity levels, adjudication, and release criteria.
- List security, privacy, access, retention, and location requirements before pricing.
- Identify buyer and vendor owners for guideline questions and change approval.
- Request separate setup, production, QA, specialist, platform, management, and contingency lines.
- Estimate buyer-side engineering, acceptance testing, and correction work.
- Model low, expected, and high cases for complexity, throughput, pass rate, and change volume.
Use a representative sample to obtain a defensible estimate
A pricing pilot should reproduce production difficulty. Include rare classes, crowded scenes, poor audio, long documents, multiple languages, edge cases, and records that should be rejected. Run the intended workflow through export and buyer acceptance. Track active labeling time, review time, question volume, pass rate, error severity, correction cycles, and excluded units.
For related planning, see our guides to AI data collection cost, data annotation quality control, and evaluating AI data service companies. Together they help separate collection constraints, labeling quality, and supplier fit.
Smart Language Service can scope a representative annotation sample across image, video, text, or audio workflows. Request a scoped annotation estimate using a representative sample, and we will help define the billable unit, quality layers, exception process, delivery format, and acceptance evidence before production volume is committed.

