light, points, social, media, network, networking, social network, social networking, symbols, internet, digital, web, facebook, icon, button, applications, twitter, global, togeth. Social media slang statistics 2027: practical details
Photo by geralt on Pixabay

Industry

Part of Social media slang: what beginners should know

Social media slang statistics 2027: practical details

Social media slang statistics count a corpus, not a language: the spelling problem, why repetition is not use, and what can honestly be said instead.

There is no honest number for how many people use a word. Not approximately, not with error bars, not for a single term in a single year. This page explains why, and what can be said instead, because "we cannot count it" is not the same as "we know nothing".

No figures appear anywhere on this page. That is deliberate, and the reasoning is the subject.

What to take away

  • Any count of a word online is a count of a corpus: a pile of text somebody assembled. The pile is not the language.
  • The pile systematically over-represents public, indexed, long-lived, text-first writing, and speech is missing entirely.
  • Nonstandard spelling defeats ordinary counting, and the standardizing fixes quietly delete the variation that was the interesting part.
  • Copying, quoting and automated reposting inflate counts in ways that cannot be separated out after the fact.

What a corpus is, and what it is not

A corpus is a collection of text, chosen by someone, for some purpose. Everything a count tells you is a fact about that collection. The discipline built around this, corpus linguistics, spends most of its effort on the design of the collection precisely because the design decides the answer.

An online corpus is assembled from what could be collected. That means:

  • Public posts, not private messages, and most conversation is private.
  • Text, not speech, and most language is spoken.
  • Whatever survived to the moment of collection, which excludes everything deleted, expired or moved.
  • Whatever was reachable by whatever tool did the collecting, which excludes anything behind a login or built by script in the reader's browser.

Each exclusion removes a different population, and the removed populations are not random with respect to vocabulary. Younger speech, in-group speech and the earliest uses of anything are concentrated exactly in the parts that are hardest to collect. This is the general problem of sampling bias, in a setting where the sampling frame cannot be described, let alone corrected.

Why the spelling problem is worse than it looks

Counting a word means deciding what counts as that word. For standard vocabulary this is dull. For online vocabulary it is the whole task, because variation in spelling is not noise. It carries register, community and stance, which is to say it carries most of the meaning.

Decision If you merge variants If you keep them apart
Deliberate misspellings You count a hostile use and an affectionate one as the same word You count one word as several and each looks rarer than it is
Repeated letters marking stress You lose the intensity distinction entirely The counts fragment across an unbounded set of forms
Character substitution You silently merge filter evasion with ordinary use You miss most uses, because the substituted form may be the common one
Capitalization You lose a tone marker that speakers rely on You split every term in two

There is no correct answer. There is only a decision, and any published figure has made one without telling you. Two studies that made different decisions will disagree, and the disagreement will be reported as a finding about language rather than as a difference in method. That is the operationalization problem in its plainest form.

Repetition is not use

The other half of the problem is the numerator. A count of occurrences counts occurrences, and an occurrence is not a speaker choosing a word.

  • A quoted post carries the word again.
  • A screenshot carries it invisibly, so it is missed.
  • Automated reposting multiplies a single utterance without limit.
  • A single conversation between two people who use the term constantly produces a great many occurrences from two speakers.

Deduplicating fixes none of this reliably, because the copies are not identical and the near-copies are exactly what you cannot classify. The same difficulty in the neighboring case of measuring attention rather than words is worked through under benefits of social media.

What can honestly be said

Quite a lot, as long as it is stated as a claim about the corpus.

  • Relative movement inside one collection. If the same collection, gathered the same way, shows a form appearing more often over a period, that is a real observation about that collection.
  • Presence and absence. Finding a form in a collection proves it existed there. Not finding it proves very little, and saying so is the honest form.
  • Co-occurrence. What a term appears next to is more stable than how often it appears, because it does not depend on the denominator.
  • Order of appearance within one collection. Weaker than it sounds, but stronger than a frequency.

None of these support a sentence of the form "X percent of people say this". Nothing supports that sentence, which is why you should distrust it wherever you meet it. The way an unsupported figure survives correction, and why the correction never travels as far, is set out under memes and virality.

Reading a claim about how common a word is

Ask four questions, in this order.

  1. What text was in the collection, and how did it get there?
  2. What counted as the word? Which variants were merged?
  3. What is the denominator: posts, accounts, sessions, or something the writer has not defined?
  4. Were repeats removed, and by what rule?

If the source cannot answer the first two, the figure is not a measurement. If it cannot answer the third, the percentage has no meaning at all, since a percentage without a denominator is a number with a symbol after it. The reason so little of the older record can answer any of them is set out under internet culture archives, and the reason nobody was counting at the start is covered under early social networks.

Common questions

Is a large collection better than a careful one?

No. Size does not repair a biased frame; it makes the bias more precisely measured. A large collection of public posts is a very good measurement of public posts.

What about asking people directly?

Self-report about your own vocabulary is unreliable in a specific direction: people report the forms they think are correct. It is useful for attitudes and poor for frequency.

Could a service count its own users accurately?

It could count events in its own system, which is a genuine number about that system. It still faces every definitional problem above, and it is not a measurement of the language.

So what should a writer do?

Describe the mechanism and skip the magnitude. Almost every interesting claim about vocabulary is about direction and process, and neither requires a number.

More in Industry

Latest from Reporting Desk