N-gram Keyword Extractor
Extract the most frequent one, two, three and four word phrases from a text, filtered for stop words, so you can see which topics the content actually covers rather than which it intends to.
Last reviewed by the Radiatus Cloud team
Need this done properly for your business?
Radiatus delivers secure cloud, DevOps & compliance engineering.
Single word counts hide the topic
Counting individual words tells you a page mentions "running" forty times and "shoes" thirty five times. It does not tell you whether the page is about running shoes, shoe running, or two unrelated subjects sharing a page. Phrases of two, three and four words carry the meaning that individual words lose, and extracting them shows what a piece of content is actually about rather than what its author believes it is about. The gap between those two is frequently the whole problem.
Stop words have to go, but only from the edges
Removing every occurrence of "of" and "the" destroys genuine phrases such as "rate of return" and "the state of the art". The workable rule is to discard phrases that begin or end with a stop word while keeping those where one sits in the middle. That keeps meaningful multi-word terms intact while removing the enormous number of grammatical fragments that would otherwise dominate any frequency list.
Frequency is a description, not a target
Knowing which phrases repeat is diagnostic. It is not an instruction to repeat them more. Keyword density as a ranking factor has not been meaningful for many years, and writing to hit a density figure produces text that reads badly to the only audience that matters. The useful readings are different: a phrase you expected to appear and do not, a phrase dominating that you did not intend, and the vocabulary a page uses compared with one that already ranks.
Related tools
- Meta Tag Analyzer — Check a page's meta tags for the problems that actually affect indexing and click-through.
- Redirect Chain Checker — Trace the full path of redirects (301/302) to find loops or lost link juice.
- Robots.txt Validator — Validate robots.txt syntax and check whether a specific URL is allowed or blocked for a given crawler.
- Sitemap Generator (Lite) — Generate a valid XML sitemap with lastmod dates and correct structure, and learn which URLs belong in it and which do not.
Frequently Asked Questions
What is an n-gram?
A sequence of n consecutive words. A 2-gram is a two word phrase, a 3-gram is three. Longer phrases carry the meaning that individual word counts lose, which is why they show what a page is about.
Why filter stop words at the edges only?
Because removing them entirely destroys real phrases such as "rate of return". Discarding phrases that start or end with a stop word removes grammatical fragments while keeping meaningful terms with a stop word in the middle.
Should I write to a keyword density target?
No. Density as a ranking factor has not been meaningful for many years, and writing to a number produces text that reads badly to readers, who are the audience that matters. Use the output diagnostically instead.
What is the most useful thing to look for?
A phrase you expected and do not see, or one dominating that you did not intend. Both indicate the content covers something different from what was planned, which no amount of optimisation fixes.
How does this compare with keyword density tools?
Density tools report a percentage for single words. This reports frequency for phrases of one to four words, which is a different and generally more informative view, since topics are expressed in phrases rather than in isolated words.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Paste your content to extract the phrases it actually repeats, by length.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
Meta Tag Analyzer
SEOCheck a page's meta tags for the problems that actually affect indexing and click-through.
Redirect Chain Checker
SEOTrace the full path of redirects (301/302) to find loops or lost link juice.
Robots.txt Validator
SEOValidate robots.txt syntax and check whether a specific URL is allowed or blocked for a given crawler.