Text

Word Frequency Counter

Count how often each word appears in a text, with stop-word handling and the caveats that make raw counts misleading.

Last reviewed by the Radiatus Cloud team

Need this done properly for your business?

Radiatus delivers secure cloud, DevOps & compliance engineering.

Book a free consult

Raw counts are dominated by function words

Run any English text through a naive counter and the top ten will be the, of, and, to, a, in, that, is, it, for. These carry almost no topical information. Filtering stop words is what makes the output useful, though the right stop-word list depends on the task: for authorship analysis, function words are the signal, not the noise, because their frequencies are habitual and hard to disguise.

Zipf's law and why the tail matters

Word frequencies follow a power law: the most common word appears roughly twice as often as the second, three times as often as the third, and so on. The practical consequence is that a small handful of words accounts for most of the text, while the words that actually distinguish this document from any other sit far down the list. Reading only the top twenty tells you the text is in English, not what it is about.

Stemming changes the answer

Are run, runs, running and ran one word or four? Stemming crudely chops suffixes; lemmatisation maps words to dictionary forms properly. Counting without either scatters a concept across several entries and understates it. Counting with aggressive stemming merges words that should stay apart. Neither is correct in general — the choice depends on whether you care about concepts or about exact forms.

Keyword density is not an SEO strategy

Frequency analysis is often reached for to hit a target keyword density. There is no such target: Google has not used keyword density as a ranking signal for many years, and writing to a percentage reliably produces text that reads badly. Frequency analysis is genuinely useful for the opposite purpose — spotting a word you have unconsciously repeated forty times, or confirming a page never actually says the thing it is about.

Where it does real work

Finding filler and crutch words in your own writing. Checking that a translation preserved terminology consistently. Comparing two documents' vocabularies to see what is distinctive to each. Building a glossary from a corpus. In every case the interesting output is a comparison, not an absolute count.

Tokenisation is where errors enter

Hyphenated words, apostrophes, numbers, URLs and non-Latin scripts all need decisions. Is "don't" one token or two? Is "state-of-the-art" one or four? Different tools answer differently, which is why the same text counted by two tools produces two different totals. Consistency within one analysis matters more than which convention you pick.

Related tools

  • Emoji Picker & Copier — Browse, search, and copy emojis by category. Collect multiple emojis and copy them all at once. Recently used emojis saved locally.
  • Text to Speech Player — Convert text to natural speech using the Web Speech API. Adjustable voice, speed, pitch, and volume with playback controls.
  • Speech to Text Transcriber — Transcribe speech to text using your microphone. Supports multiple languages with continuous listening, copy, and download features.
  • Instagram Caption Generator — Generate Instagram caption ideas from a topic, tone, and optional CTA. Fast, template-based suggestions.

Frequently Asked Questions

Why are the top words always the, of and and?

Because function words dominate any English text and carry almost no topical information. Filtering stop words is what makes the output useful — except in authorship analysis, where those frequencies are the signal.

What is Zipf's law?

Word frequencies follow a power law where the most common word appears about twice as often as the second. It means a few words account for most of the text and the distinguishing words sit far down the list.

Should I stem words before counting?

It depends on whether you care about concepts or exact forms. Without stemming, run and running count separately and understate the concept; with aggressive stemming, distinct words get merged.

Is there an ideal keyword density?

No. Google has not used keyword density as a ranking signal for years, and writing to a percentage produces bad text. Frequency analysis is better used to catch unconscious repetition.

Why do two tools give different word counts?

Tokenisation choices differ: whether "don't" is one token or two, whether hyphenated words split, how numbers and URLs are handled. Consistency within one analysis matters more than the convention.

Privacy & Security

All processing happens locally in your browser — nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Paste text to see a ranked list of word frequencies. Toggle case and minimum length.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.