Personal computing leader lands 140K speech and video assets

140,000+ multimodal assets captured on 8 client-branded devices across multiple countries and languages ” Firstsource-delivered in-facility DC for a Personal Computing Leader.
Personal computing leader lands 140K speech and video assets

Overview

When a global PC maker is training on-device AI to see, hear, and respond, the data has to be captured under the brand's own quality bar, and on its own hardware.

A Personal Computing Leader needed in-facility moderated speech and video data collected using 8 client-branded devices, across multiple countries and languages; to train eyes, hands, and voice models on the actual devices customers would use.

Firstsource captured 140,000+ assets in 24 weeks, speech and video collected, annotated, and tested end-to-end with 100% brand safety and 100% Speech and Video QC pass.

This was Intelligence that Operates: in-facility multimodal capture run on the client's own devices, on a 24-week program clock.

Challenges

  • Eyes, hands, and voice training data can't be patched together from separate vendors. On-device AI ships when all three modalities have been trained on the same hardware under the same quality discipline. Three vendors produce three signal mismatches.
  • 8 client-branded devices add a compliance dimension. Capturing on the client's own hardware brings IP, brand safety, and device-handling requirements that don't show up in third-party-hardware programs.
  • Language and locale coverage is a recruitment, not a translation, problem. Voice and gesture data per language requires in-country participants, not script-readers, and the moderation discipline has to hold consistently across all of them.

How We Made It Happen

We ran in-facility capture, annotation, and testing as one connected program ” on the client's own devices.

  • In-facility moderated capture on 8 client-branded devices. Speech, video, and on-device interaction captured in controlled environments with full brand-safety and device-handling protocols.
  • OTS Data delivered with annotation and testing in the same envelope. Eyes, hands, and voice training data emerged from one pipeline, not three.
  • Participant recruitment matched to locales. Recruitment matched the locales the model would ship in ” not the locales the vendor could reach.

Conclusion

On-device AI only ships when the training data lives on the device it'll run on. Firstsource ran in-facility multimodal capture on the client's hardware at a global scale, turning device-grade data collection into Intelligence That Operates.

Outcomes

The partnership delivered measurable financial, operational, and customer engagement results:

140,000+ assets in 24 weeks

speech and video captured, annotated, and tested end-to-end.

Multiple countries and languages

participants recruited in-country across the model's target locales.

100% brand safety, 100% QC pass

captured on 8 client-branded devices, <5% rejection rate.

Most recent

Delivering exceptional customer outcomes while managing surge and helping client collections team overcome cost of living challenges

Managing surge, exceptional CX outcomes | Firstsourc

Transforming debt collection: a digital solution for enhanced customer engagement and operational efficiency

Digital debt collection engagement | Firstsource

How a digital-first collections model delivered top-ranked recovery for a smart home technology provider

Digital collections model tops recovery | Firstsource

Technology