Untagged Mentions: Why Social Listening Never Collects Them
Author :
Luke Bae
Published :

TL;DR: Social listening tools miss brand mentions in video because they retrieve content by matching text, not by watching video. Platform APIs hand over the caption, the hashtags and the handle tags — so a brand that is spoken aloud, written on screen, or simply held up to the camera is never pulled into the dataset at all. Untagged mentions are not an analysis failure that better AI fixes downstream. They are a collection failure, and it happens before any analysis runs.
The hole in your listening dashboard is not in the analysis layer. It is in the query.
Most marketers assume a mention exists if it happened, and that the platform's job is to find it and score it. That assumption held while brand conversation was typed. It broke when the conversation moved into short-form video: the typical TikTok user now spends 1 hour and 37 minutes a day inside the app, roughly 14% longer than the typical YouTube user spends inside YouTube's (Source: DataReportal / We Are Social, Digital 2026, on Similarweb App Intelligence data).
Call them invisible mentions: the ones that live in the content but never in the text. Nobody has credibly counted them, which is revealing in itself. What does exist is documentation — vendors describing, in their own help centers, where their collection stops. A companion piece enumerates ten types of missed mention, each with the detection layer it needs. This piece walks that paper trail and what it does to every number you compute from mention volume.
What untagged mentions actually are
An untagged mention is a brand mention that is present in the video but absent from its text. It is not a mention the platform saw and misjudged — it is a mention the platform never retrieved.
Untagged mention: a brand mention that keyword search cannot reach — spoken out loud, written on screen, or simply shown, but never typed into a caption, hashtag, or @handle.
They arrive in three forms, and each needs a different reader:
Spoken — a creator says the brand name mid-review. Recoverable with speech-to-text.
On-screen text — the name appears in a caption overlay, a price card, or a comparison chart burned into the frame. Recoverable with frame-level optical character recognition.
Shown — the product, packaging or logo is visible and never named at all. Recoverable with visual detection.
Which layer matters most depends on the category; the audio versus visual matrix sorts that out for B2C brands.
This is where our definition parts company with the industry's. Hootsuite's guide treats an untagged mention as a brand name written in a caption or comment without the @ symbol — still text, just missing a symbol (Source: Hootsuite, 2026). Sprout Social's brand-mentions guide splits mentions into direct and indirect and lists detection methods that are all keyword-based; its published guidance does not describe video transcription, on-screen text extraction or logo detection (Source: Sprout Social).
Neither taxonomy has a slot for a mention that was never typed. Ours does, and the full definition with worked examples lives on our untagged mentions page.
Why text-based social listening can't see inside video
Because the pipeline runs query-first. A listening platform sends a keyword to a platform API, the API matches that keyword against text fields, and what comes back is metadata — not media. That ordering is the whole difference between text-era and video-era social listening.
TikTok's Research API is the clearest published example. Its returnable field list includes video_description, hashtag_names, username, engagement counts and voice_to_text — but no video file and no audio stream (Source: TikTok for Developers, 2026).
That voice_to_text field deserves an honest note: speech data is not universally absent. TikTok's own documentation describes it as text shown only "for videos that have voice to text features on" — auto-caption text, opt-in, on a research-only endpoint that commercial listening vendors do not query. Partial, conditional, and out of reach for the products brands actually buy.
Then comes the ordering problem. Mentionlytics documents its own sequence plainly: its AI runs on the mentions it has "already collected" by text keyword, and the platform "doesn't scan the online space for videos and audios associated with the keywords you track" (Source: Mentionlytics Help Center, 2026). Video understanding, where a vendor has it, is applied second — to whatever the text query already returned.
Collection gap: the mentions a listening platform never retrieves, because its keyword query matched only text fields. Downstream video AI cannot recover them. They were never in the dataset.
And the excuse that the technology is not ready has expired. On the Open ASR Leaderboard, the leading open-source model transcribes clean English speech at 1.6% word error rate and the harder, noisier split at 3.1%, holding 2.41% even at 10 dB signal-to-noise (Source: Northflank, 2026). MLPerf's Whisper benchmark cut the prior reference ASR model's word error rate by more than 72% on more challenging audio (Source: MLCommons, 2025).
The transcription is not the constraint. The pipeline was simply never pointed at the video.
Where the brand match happens Video retrieved with no brand text? Documented limit Text-first listening (default config) Caption, hashtags, @handle No Sprinklr: no general keyword listening on TikTok Text-first plus video AI bolt-on Text query first, video read second No Mentionlytics: analyzes already-collected mentions only Visual or logo analytics add-on Images and keyframes, after collection Only inside the already-collected set Collection method not disclosed Video-native listening Speech, on-screen text and objects in the video itself Yes —
What the tools themselves say about their coverage
The strongest evidence for the collection gap is not our analysis. It is the vendors' own help centers, which are considerably more candid than their marketing pages.
Sprinklr's help center states that TikTok listening covers @mentions of authenticated business accounts and registered brand hashtags, and that "general keyword or competitive listening is not supported." The same page notes that hashtag listening "includes video captions only," excludes comments, and retrieves only the top 1,000 hashtag posts by likes inside a rolling 90-day window (Source: Sprinklr Help Center, 2026). Read that carefully: on the platform where most short-form brand conversation happens, an enterprise listening product documents that you cannot run a keyword search at all.
Mentionlytics documents the order of operations. Hootsuite's guide defines untagged in a way that excludes the video case entirely, and Sprout's guide is silent on video.
Now the fairness clause, because it matters and because the opposite claim is false and checkable: legacy platforms can analyze video. Talkwalker markets logo recognition across "50,000,000+ videos analyzed every day" and search across "30,000 brands, scenes, objects and products" (Source: Talkwalker). Brandwatch has shipped image and logo analysis since 2017, picking up marks wherever they appear in an image (Source: Brandwatch). YouScan, Meltwater and Sprinklr all ship some visual or audio capability too.
The argument here is not about capability. It is about collection order. A visual recognition engine pointed at a dataset assembled by caption matching will faithfully analyze the wrong sample. That is why the fix has to move upstream, into how video gets retrieved in the first place.
How much brand conversation happens without a tag
Nobody has measured it properly, and we are not going to pretend otherwise. Several confident-sounding percentages circulate in this category; each one traces back to a vendor blog with no stated methodology, no sample size and no fieldwork date. None is safe to plan against, so we are not repeating them here.
The absence is structural. The tools that would run the study are the tools that cannot see the thing being measured.
What can be documented is the ceiling. On TikTok, per Sprinklr's help center, the measured universe is @mentions of your own authenticated account plus hashtags you registered — captions only. That is the denominator a great many brands are quietly reporting from.
YouScan comes closest to publishing a share for YouTube, and hedges it inside its own sentence as an approximation rather than a measurement — no methodology, no sample, no fieldwork date (Source: YouScan, 2026). Our own figure is equally ours: Syncly observes 3-4x more brand mentions captured when collection shifts from text matching to video-native retrieval (Source: Syncly, 2026). That is our number from our data, not an industry benchmark, and platform-level gaps show up clearly when you compare TikTok coverage across products.
Sizing your own gap is procedural, and how to measure it on TikTok, Reels and Shorts is its own walkthrough.
What changes when you start counting untagged mentions
Every metric built on mention volume moves, and share of voice moves most. The foundational excess share of voice work — Binet and Field's analysis of the IPA Databank in The Long and the Short of It (IPA, 2013) — found that each 10 points of share of voice above share of market is associated with roughly 0.5% annual market-share growth.
That turns a tooling complaint into a budget error. If share of voice is an input to planning, and share of voice is computed from a caption-only denominator, the plan is not slightly imprecise — it is wrong by construction, and it is wrong in a direction you cannot see.
There is a second, quieter distortion. Sprinklr's engagement-ranked hashtag cap means that even inside the tagged set, the sample tilts toward high-engagement posts. Early-stage conversation, small creators and slow-building complaints are structurally underweighted — exactly the material that would have given you warning.
So expect to re-baseline rather than compare. Untagged volume is a new denominator: sentiment mix shifts, because spoken and shown mentions skew toward unprompted opinion; crisis thresholds set on the old volume fire late or not at all; and competitive share needs recomputing for every brand in the set, not just yours.
Teams that want the mention data inside their own models pull it through the API rather than a dashboard, and the broader method sits in our video listening guide.
Key Takeaways
The miss is a collection failure, not an analysis failure. A video whose brand appears only in speech, on screen, or in the product itself is never retrieved, so no downstream video AI can recover it.
Legacy platforms can analyze video. Talkwalker, Brandwatch, YouScan, Meltwater and Sprinklr all ship visual or audio capability. The problem is that the text query runs first and decides the sample.
The industry definition of untagged means a brand name typed without an @. Spoken, on-screen and shown mentions are a strictly larger set that the incumbent taxonomy cannot express.
No independent study quantifies the untagged share. Ignore the circulating percentages and reason from documented coverage limits instead.
Share of voice computed from a caption-only denominator is wrong by construction, and share of voice is an input to budget planning.
Fix the query, not the dashboard. Better sentiment models and smarter alerting all operate on a sample some keyword already selected, and no amount of analytical polish adds back a video that was never collected.
Which reframes what you should be asking vendors. Not "can you detect a logo?" — most of them can. Ask "was the video retrieved at all when my brand appeared nowhere in its text?"
That one question separates architecture from feature list, and the category's own help centers already answer it. We compare on it head to head with YouScan.
See what your current tool never collected. Book a Syncly demo →