<- Back to laws

Empirical rank-frequency regularity

Zipf's
Law

Rank the words in a large corpus from most to least frequent. Their frequencies often decline approximately as a power of rank: a small vocabulary forms the head, while an enormous number of uncommon words creates the long tail.

THEOFANDTOAINTHATLANGUAGEUNCOMMONHAPAXf(r) ~ 1/r
ObjectRanked word frequencies
Classical formf(r) = C / rs
Typical exponents near 1
Scientific statusApproximate empirical law
Critical choiceTokenization
Common misuseStraight line = proof
INTERACTIVE 01 / LANGUAGE DATA LAB

Turn a text into a ranked vocabulary.

Analyze the words in this guide, compare controlled counterexamples, or paste your own English text. All calculations run locally in the page.

DESCRIPTIVE LABThe displayed slope uses ordinary least squares on log ranks for teaching. It is not a valid power-law goodness-of-fit test and does not compare alternative distributions.
RANK-FREQUENCY PROFILE
Tokens0all counted word occurrences
Word types0distinct normalized forms
Fitted exponent s--descriptive head-to-rank 100
Hapax types0types appearing exactly once
Interactive visual model for Zipf's Law.
ObservedIdeal s = 1 guideR-squared: --
ANALYZING CORPUS

The page will summarize the distribution after tokenization.

RANKWORDCOUNTSHARE
HEADfew typesmany tokens
TORSOtopic + grammarstructured deviations
LONG TAILmany rare typesnames, terms, innovations
01 / MEANING

Frequency falls as rank rises.

Count every word token in a corpus, group identical tokens into word types, and order the types by decreasing frequency. The classical form says that frequency f at rank r is proportional to r raised to a negative exponent s. When s is near one, rank 2 has roughly half the frequency of rank 1, rank 10 roughly one tenth, and rank 100 roughly one hundredth.

GENERAL FORMf(r) = C / rs
rfrequency rank
f(r)count or relative frequency
sscaling exponent
RANK 1100%
RANK 250%
RANK 520%
RANK 1010%

This is an approximate scaling relation, not a word-by-word prediction. Natural corpora contain systematic bends and local structure beyond one exponent. Piantadosi's critical review concludes that large-scale word frequencies are robustly Zipfian while also being reliably more complex than the classic formula.[3]

Zipf's Law compresses the silhouette of a vocabulary. It does not erase grammar, meaning, genre, history, or the corpus that produced it.
02 / READING THE CURVE

Linear and logarithmic views tell different stories.

LINEAR AXES

The head dominates.

A handful of function words towers over everything else. Most ranks collapse visually near zero.

LOG-LOG AXES

Scaling becomes visible.

A pure power law becomes a straight line with slope -s. Deviations in the head and tail are easier to inspect.

HEADFunction words and corpus conventions

Very common words may bend away from a single fitted line.

MIDDLEBroad scaling region

The rank-frequency relation often looks most regular over an intermediate span.

TAILRare words and finite samples

Counts become discrete, noisy, and dominated by words seen once or twice.

Visual cautionA line on log-log axes is suggestive, not conclusive. Lognormal, stretched-exponential, Yule-Simon, and truncated distributions can resemble power laws over limited ranges.
03 / CORPUS PIPELINE

Before counting words, define what a word is.

Rank-frequency curves depend on editorial and computational choices. A result without its tokenization rules is not reproducible.

01Select corpus

Language, speaker, genre, date, topic, medium, and corpus length shape the distribution.

->
02Segment tokens

Whitespace is insufficient for every language. Punctuation, apostrophes, compounds, emoji, and scripts need rules.

->
03Normalize

Decide case folding, spelling variants, numbers, markup, and whether contractions split.

->
04Choose type

Surface forms, lemmas, morphemes, characters, and subword tokens answer different questions.

->
05Rank and model

Report ties, frequency units, fitting range, uncertainty, and alternative models.

RAW"Dogs, dog's, DOGS, and dog-like."
SURFACE FORMSdogs / dog's / dogs / and / dog / like
POSSIBLE LEMMASdog / dog / dog / and / dog / like
STOP WORDS

Keep or remove?

Keeping them reveals the natural high-frequency head. Removing them may help topic analysis but changes the law being measured.

LEMMATIZATION

Merge inflections?

Combining runs, ran, and running changes ranks and can alter fitted parameters. It is a linguistic model, not neutral cleanup.

SUBWORD MODELS

Tokens are learned.

Modern NLP vocabularies split words into pieces. Their frequency distribution is determined partly by the tokenizer's training objective.

MULTILINGUAL TEXT

Segmentation differs.

Chinese, Japanese, Thai, agglutinative languages, and mixed scripts make English-style word boundaries inappropriate.

04 / EMPIRICAL EVIDENCE

Robust at large scale, structured in detail.

Zipf-like rank-frequency structure appears across natural languages and many corpus types, but real curves are not exact 1/r laws. Cross-language work reports recurring multi-segment patterns, while genre and preprocessing shift local slopes.[9]

ROBUSTExtreme frequency inequality

A small set of words accounts for a large share of tokens, and most types are rare.

ROBUSTApproximate scaling

Frequency falls roughly as a power of rank over substantial regions of large corpora.

VARIABLEExponent and curvature

There is no single exact slope shared by every language, corpus, rank range, or token definition.

VARIABLEHead and tail shape

Function words, topic mixture, finite sample size, and vocabulary growth create systematic departures.

TOKENSN

Every occurrence: "the" counted 8,000 times contributes 8,000 tokens.

TYPESV

Distinct forms: all 8,000 occurrences of "the" contribute one type.

HAPAX LEGOMENAf = 1

Types seen once. They are evidence of the large tail and of incomplete vocabulary sampling.

05 / WHY DOES IT EMERGE?

Many mechanisms can produce a similar curve.

The rank-frequency pattern is not a unique fingerprint of one causal theory. Explanations operate at different levels and may be complementary rather than exclusive.

COMMUNICATIVE TRADEOFF

Least effort

Speakers favor reusable, accessible forms; hearers benefit from distinct, informative signals. Zipf proposed that language balances competing pressures, and later models formalized this tradeoff.[6]

GROWTH PROCESS

Preferential reuse

Existing words are likely to recur while new words enter at a lower rate. Simon showed how such stochastic growth can generate highly skewed distributions.[5]

NULL MODEL

Random typing

Random character strings separated by spaces can create Zipf-like ranks. This proves that the shape alone does not demonstrate a realistic language mechanism.

INFORMATION / COST

Compression

Short, frequent forms and longer, rarer forms can arise when coding cost and information are jointly constrained.

CORPUS STRUCTURE

Mixture of topics

Combining speakers, documents, genres, and contexts creates heterogeneity and can strengthen heavy-tailed frequency structure.

COGNITIVE SYSTEM

Learning and memory

Exposure, accessibility, prediction, production, and lexical innovation interact across development and use.

06 / PROFESSIONAL TESTING

Fit, challenge, and compare the model.

01State the hypothesis

Specify the rank range, token definition, corpus, and whether the exponent is fixed at one or estimated.

02Inspect without binning

Plot counts, complementary cumulative distributions, and residuals. Preserve the discrete tail.

03Estimate appropriately

Least squares on logged values is biased. Use likelihood-based methods suited to discrete data.

04Test fit

Use simulated goodness-of-fit procedures and quantify uncertainty in the lower cutoff and exponent.

05Compare alternatives

Evaluate lognormal, exponential, stretched-exponential, Yule-Simon, and truncated models using likelihood ratios.

06Replicate across corpora

Check sensitivity to genre, length, preprocessing, language, time period, and sampling unit.

THE TRAPlog(count) = a - s log(rank)

A high R-squared after transformation can coexist with systematic curvature, correlated errors, discrete tail effects, and a better alternative model.

R-squaredis notpower-law probabilityand notcausal evidence

Clauset, Shalizi, and Newman provide a widely used framework combining maximum-likelihood estimation, Kolmogorov-Smirnov goodness-of-fit, simulation, and likelihood-ratio comparisons. They show why visual inspection and ordinary least squares are insufficient.[4]

07 / HISTORY

From word counts to statistical language science.

1910s-30sEarly frequency studies

Researchers including Jean-Baptiste Estoup documented ranked word-frequency patterns before Zipf's mature synthesis.

1932-35George Kingsley Zipf

Zipf systematically analyzed relative word frequencies in language and connected rank to frequency.

1949Principle of least effort

Zipf placed linguistic distributions inside a broader theory of human behavior and competing effort.[1]

1953-55Mandelbrot and Simon

Information-theoretic and stochastic-growth models supplied distinct mathematical routes to skewed frequency distributions.

1990s-nowCorpus era

Massive multilingual corpora expose stable large-scale structure, detailed deviations, sampling effects, and stronger statistical tests.

08 / APPLICATIONS

The head and tail demand different engineering strategies.

SEARCH

Frequent and rare queries

Cache and optimize the head, but design retrieval, spelling, and fallback systems for the enormous query tail.

NATURAL LANGUAGE PROCESSING

Imbalanced training data

Common tokens receive abundant updates while rare words, names, dialects, and specialist terms remain data poor.

COMPRESSION

Short codes for common events

Frequency-skewed symbols motivate variable-length coding, though linguistic structure is richer than isolated token counts.

LEXICOGRAPHY

Coverage and vocabulary growth

More text keeps revealing new types. Corpus size must accompany any claim about vocabulary coverage.

PRODUCT ANALYTICS

Head-tail allocation

A few actions dominate volume while rare workflows create support, accessibility, and reliability challenges.

EVALUATION

Do not average away the tail

Aggregate accuracy can look strong while performance fails on rare terms, languages, intents, or users.

10 / REFERENCES

Sources and further reading.

Original books, foundational models, critical reviews, and rigorous statistical methods are prioritized.

  1. George Kingsley Zipf (1949) - Human Behavior and the Principle of Least EffortZipf's mature synthesis connecting linguistic frequency to a broader least-effort principle.Public-domain scan via Wikimedia Commons
  2. George Kingsley Zipf (1935) - The Psycho-Biology of LanguageAn early systematic presentation of relative frequency patterns in language.Internet Archive catalog
  3. Steven T. Piantadosi (2014) - Zipf's Word Frequency Law in Natural LanguageA critical review emphasizing robust large-scale structure and reliable deviations beyond the classical law.doi.org/10.3758/s13423-014-0585-6
  4. Clauset, Shalizi, and Newman (2009) - Power-Law Distributions in Empirical DataLikelihood-based fitting, goodness-of-fit testing, and comparison against alternative distributions.SIAM Review 51(4), 661-703
  5. Herbert A. Simon (1955) - On a Class of Skew Distribution FunctionsA foundational stochastic-growth explanation for highly skewed empirical distributions.Biometrika 42(3/4), 425-440
  6. Ferrer i Cancho and Sole (2003) - Least Effort and the Origins of Scaling in Human LanguageA formal communicative tradeoff model producing Zipf-like scaling.PNAS 100(3), 788-791
  7. Benoit Mandelbrot (1953) - An Informational Theory of the Statistical Structure of LanguageAn information-theoretic account extending the classical rank-frequency relation.Communication Theory, 486-502
  8. Altmann and Gerlach (2016) - Statistical Laws in LinguisticsA critical discussion of fitting, falsification, correlations, and fluctuations in linguistic laws.arXiv:1502.03296
  9. Yu, Xu, and Liu (2018) - Zipf's Law in 50 LanguagesCross-language evidence for recurring multi-segment structure and tail deviations.arXiv:1807.01855
  10. Petersen et al. (2012) - Languages Cool as They ExpandLarge-corpus analysis linking vocabulary growth, word birth rates, and language expansion.Scientific Reports 2, 943
CONTINUE EXPLORING

Related laws, with the relationship made explicit.

These are editorial connections, not claims that the laws are mathematically equivalent.

CONTINUE READING

Place this law inside the collection.

LAW 007 / 100 PUBLISHED