Skip to main content
Monitoring: queries, data & reliability

Duplicate mention detection: separating repeats

Duplicates distort time, numbers and decisions. How to separate duplicate posts and repeated mentions, and count them correctly in reports.

October 9, 20266 min read

Short answer

To separate duplicate posts and repeated mentions, first recognise the kind of duplicate: the same page under different addresses, news republished on other sites, the same post shared in several places, a quote or repost, an edited post, and near-identical copies. Then merge them with one of three keys: the platform's post ID, the canonical URL, or a fingerprint of the text. In reports, count unique items and unique authors, not posts. Do not delete duplicates — link them to one result, because how often something is repeated can itself be a signal.

Why duplicates are a problem

Duplicates do three kinds of damage. First, time: the team reads the same news story ten times on ten sites. Second, the numbers: when a press release is republished on fifteen sites, the mention count rises fifteen-fold and the report lies. Third, decisions: when one customer's complaint is shared on three forums, it looks like "three complaints" and priorities go wrong. Separating duplicates is monitoring's least visible but most useful cleaning job.

Six kinds of duplicate

  • Same page, different address — tracking parameters, mobile version, with and without "www".
  • Republished news — a press release or agency story on several sites with the same text.
  • Cross-post — the same author places the same text in several groups or forums.
  • Quote and repost — sharing someone else's post with or without a comment.
  • Edited post — the same post whose text changed later, or which gained new comments.
  • Near-identical copies — texts copied with a word or two changed, including spam.

Three dedup keys

  1. Platform IDIf the platform gives each post a unique ID, that is the most reliable key: same ID, same post.
  2. Canonical URLWithout an ID, the address is normalised: tracking parameters, "www" and a trailing "/" are removed, and the page's own canonical link is used if it has one.
  3. Text fingerprintFor the same text on different sites: the text is normalised and a short fingerprint computed from it; same fingerprint, same content.

Within a search and across searches

There are two different questions. First: among today's results, does the same thing appear twice? That is dedup within a search. Second: did we already see today's post yesterday? That is dedup across searches, and it matters more for daily work. A post reviewed yesterday and unchanged should not be shown again today. But if its text has changed — new comments in a forum thread, say — it should come back as "updated", because the situation has changed.

How to count in reports

  • Unique items — the count left after duplicates are merged.
  • Unique authors — ten posts by one person count as one author.
  • Number of sources — on how many sites a unique item appeared; a measure of spread.
  • Total mention count — only with context and next to the "unique" figure; on its own it misleads.

Reach (how many people saw something) is not a mention count and cannot be measured precisely from public sources. Put such a figure in a report only when the platform's official data exists.

When duplicates are a signal

Sometimes the repetition itself is information. If a complaint is shared in many forums and groups within a short time, that is a spread signal and may call for a crisis check. The same text posted from different accounts at the same time is a sign of a coordinated campaign or spam — such posts should not count as the customer's voice and should be flagged separately. So when merging duplicates, keep the "in how many sources" and "over what time" information.

Illustrative example

This is an illustrative example. A bank issues a press release about a new card, and 42 mentions appear in one day. After dedup the picture is: the press release on 18 news sites with the same text (one unique item, 18 sources), the bank's own post shared in 9 groups (one unique item), 6 customer questions and 2 complaints — one of them cross-posted on three forums. Result: not 42 mentions but 10 unique items and 9 unique customer authors. The report gives these two figures, and 42 only in the context of spread.

Edited posts

When a post is edited, two mistakes are possible. One is counting it as a new post — the mention count rises artificially. The other is hiding it as an old post — the text has changed, and perhaps the author strengthened the complaint or wrote "resolved". The right way: keep it as the same post but flag the change. When the fingerprint is calculated only on the part that concerns the brand, changes elsewhere on the page — a news feed refreshing, say — do not create a false signal.

Limitations

No dedup rule catches every case. Copies with a word or two changed can produce a different fingerprint; different people happening to write the same short sentence can cause a wrong merge. A comment added to a quote or repost is new information — treating it as a pure duplicate is a mistake. Important results need a person to look.

Common mistakes

  • Presenting the total mention count as the main measure.
  • Deleting duplicates and losing the spread information.
  • Counting an edited post as a new post.
  • Rereading yesterday's posts every day.
  • Counting coordinated spam as the customer's voice.

Dedup in Vexvon Monitoring

Vexvon Monitoring never records the same result twice within a search: it merges by platform ID where one exists, otherwise by canonical URL and keyword. Across searches, a content fingerprint is kept for each result: a post shown before and unchanged is marked as a repeat and hidden by default, while a post whose text changed is shown again as "updated". On web pages the fingerprint is calculated from the text around the keyword, so a change elsewhere on the page does not create a false "updated". Posts by the same author with the same text within a short time are reduced to one result. More on the Vexvon Monitoring page.

Next step

Take last week's results and calculate three numbers: total mentions, unique items and unique authors. If the gap is large, switch your report to the unique figures from this week. Other articles are in the monitoring queries, data and reliability section.

Live demo

Ready? Let's start

See Vexvon live in a 10-minute demo.

  • A scenario built for your business
  • A live sample call
  • A tour of the platform
Get a demoorBook a meeting

Your details are used only for the demo and to get in touch.