GoRhyme Credits
GoRhyme’s rhyme, meaning, and word tools are moving to GoRhyme’s own word engine, one tool at a time. The engine is built from open data that researchers and volunteers made and share, and could not exist without them. This page credits each source, says how GoRhyme uses it, and gives the license it is shared under.
Some of this data also reaches people directly. GoRhyme’s word service answers the requests made by GoRhyme’s tool pages, and anyone else who calls the service receives the same answers, each with a link back to this page. Where the service passes along a source’s data, the notes below note the usage.
None of the people or organizations named here is affiliated with GoRhyme, and none of them endorses GoRhyme or its tools.
Table of Contents
Word frequency: wordfreq
wordfreq was created by Robyn Speer. It estimates how often words appear across many kinds of writing and speech, from books and subtitles to websites.
How GoRhyme uses it: word frequency helps order the results in GoRhyme’s tools, so familiar words come before rare ones. GoRhyme’s word service also shares a frequency number for each word it returns.
License: wordfreq’s data is shared under the Creative Commons Attribution-ShareAlike 4.0 International license (CC BY-SA 4.0), https://creativecommons.org/licenses/by-sa/4.0/. GoRhyme’s frequency numbers are derived from wordfreq’s data and have been changed from the original. Because of the ShareAlike terms, the frequency numbers GoRhyme shares are themselves offered under CC BY-SA 4.0, and you are free to reuse them under that license.
wordfreq: https://github.com/rspeer/wordfreq
Sources wordfreq was built from
wordfreq’s data comes from Exquisite Corpus, by Luminoso, which combines many sources of text. GoRhyme thanks these sources as well:
The SUBTLEX word lists, by Marc Brysbaert and colleagues, used in wordfreq with their permission. SUBTLEX is freely available data: http://crr.ugent.be/programs-data/subtitle-frequencies
OpenSubtitles, through OPUS OpenSubtitles 2018, https://www.opensubtitles.org
Google Books Ngrams and Google Books Syntactic Ngrams, https://books.google.com/ngrams
Wikipedia, NewsCrawl, GlobalVoices, OSCAR, ParaCrawl, the Leeds Internet Corpus (University of Leeds Centre for Translation Studies), Twitter, and Reddit
Words in poetry: the Gutenberg Poetry Corpus
The Gutenberg Poetry Corpus was created by Allison Parrish. It gathers more than three million lines of poetry from public-domain books digitized by Project Gutenberg’s volunteers.
How GoRhyme uses it: GoRhyme counted how often poets used each word, and those counts help set the order of most rhyme results, so words that poets have actually used come first. On longer lists, that order also decides which words make the list. The corpus also helps GoRhyme tell names from ordinary words. In the meaning tools, the poems help GoRhyme find related words, words that tend to appear together, and adjective and noun pairs. The poetry counts themselves are never shared.
License: the corpus is dedicated to the public domain under Creative Commons CC0 1.0 Universal, https://creativecommons.org/publicdomain/zero/1.0/, and the poems themselves are in the public domain. No credit is required, and GoRhyme gives it with thanks.
Gutenberg Poetry Corpus: https://github.com/aparrish/gutenberg-poetry-corpus
Pronunciations: the CMU Pronouncing Dictionary
The CMU Pronouncing Dictionary (CMUdict) was made at Carnegie Mellon University. It gives the pronunciation of more than 130,000 English words, sound by sound.
How GoRhyme uses it: CMUdict is where GoRhyme learns how words sound. Rhyme matches, syllable counts, and sound lists are built from it. GoRhyme’s word service also shares pronunciation codes taken from it.
License: CMUdict is shared under a BSD-style license. Its full notice appears at the bottom of this page, as the license requires.
CMUdict: https://github.com/cmusphinx/cmudict
Word meanings: WordNet
WordNet is a large database of English created at Princeton University. It groups words by meaning and records how those meanings relate to one another.
How GoRhyme uses it: WordNet supplies the definitions, parts of speech, synonyms, opposites, and “kind of” lists in GoRhyme’s meaning tools, and many of the multi-word phrases that appear among rhyme results. GoRhyme’s word service also shares these definitions and lists.
License: WordNet 3.0 is shared under the WordNet 3.0 license. Its full notice appears at the bottom of this page, as the license requires.
WordNet: https://wordnet.princeton.edu
Citation: George A. Miller (1995). WordNet: A Lexical Database for English. Communications of the ACM, 38(11), 39-41.
Word pairs: Google Books Ngrams
Google Books Ngrams counts how often words appear together across millions of published books. GoRhyme uses the English two-word data, version 20120701.
How GoRhyme uses it: the adjective and noun pair lists, drawn partly from this data, show which adjectives and nouns writers actually put side by side. GoRhyme’s word service also shares these pairs.
License: the data is shared under the Creative Commons Attribution 3.0 Unported license (CC BY 3.0), https://creativecommons.org/licenses/by/3.0/. The original work has been modified: GoRhyme counted, filtered, and ranked the word pairs.
Credit: Google Books Ngram Viewer, https://books.google.com/ngrams
Related words: GloVe
GloVe was created by Stanford University’s natural language processing group. It maps words by the company they keep, so words used in similar ways sit close together. GoRhyme uses the 2024 Wikipedia and Gigaword vectors.
How GoRhyme uses it: GloVe helps rank the related-word suggestions in GoRhyme’s meaning tools. The GloVe data itself is never shared.
License: the GloVe vectors are dedicated to the public domain under the Open Data Commons Public Domain Dedication and License (PDDL 1.0), https://opendatacommons.org/licenses/pddl/1-0/. No credit is required, and GoRhyme gives it gladly.
GloVe: https://nlp.stanford.edu/projects/glove/
First names: the Names Corpus
The Names Corpus was compiled by Mark Kantrowitz, with additions by Bill Ross, and is distributed with the Natural Language Toolkit (NLTK).
How GoRhyme uses it: the list helps GoRhyme recognize first names, so they rank lower in rhyme results and stay out of the meaning lists. The list itself is never shared.
License: the list may be used freely with credit. Copyright (C) 1991 Mark Kantrowitz. Additions by Bill Ross.
Names and abbreviations: public-domain lists
Two public-domain lists help GoRhyme tell names and abbreviations from ordinary words.
The US Census Bureau’s list of frequently occurring surnames from the 2010 Census helps GoRhyme recognize last names: https://www.census.gov/topics/population/genealogy/data/2010_surnames.html
The word list from Webster’s Second International Dictionary (1934), included with macOS, helped build GoRhyme’s lists of names and abbreviations.
License: both lists are in the public domain, so no credit is required, and GoRhyme gives it with thanks. Neither list is shared.
License notices
The following notices are reproduced in full, as their licenses require.
WordNet 3.0 license
WordNet Release 3.0
This software and database is being provided to you, the LICENSEE, by Princeton University under the following license. By obtaining, using and/or copying this software and database, you agree that you have read, understood, and will comply with these terms and conditions.:
Permission to use, copy, modify and distribute this software and database and its documentation for any purpose and without fee or royalty is hereby granted, provided that you agree to comply with the following copyright notice and statements, including the disclaimer, and that the same appear on ALL copies of the software, database and documentation, including modifications that you make for internal use or for distribution.
WordNet 3.0 Copyright 2006 by Princeton University. All rights reserved.
THIS SOFTWARE AND DATABASE IS PROVIDED “AS IS” AND PRINCETON UNIVERSITY MAKES NO REPRESENTATIONS OR WARRANTIES, EXPRESS OR IMPLIED. BY WAY OF EXAMPLE, BUT NOT LIMITATION, PRINCETON UNIVERSITY MAKES NO REPRESENTATIONS OR WARRANTIES OF MERCHANTABILITY OR FITNESS FOR ANY PARTICULAR PURPOSE OR THAT THE USE OF THE LICENSED SOFTWARE, DATABASE OR DOCUMENTATION WILL NOT INFRINGE ANY THIRD PARTY PATENTS, COPYRIGHTS, TRADEMARKS OR OTHER RIGHTS.
The name of Princeton University or Princeton may not be used in advertising or publicity pertaining to distribution of the software and/or database. Title to copyright in this software, database and any associated documentation shall at all times remain with Princeton University and LICENSEE agrees to preserve same.
CMU Pronouncing Dictionary license
Copyright (C) 1993-2015 Carnegie Mellon University. All rights reserved.
Redistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer. The contents of this file are deemed to be source code.
2. Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution.
This work was supported in part by funding from the Defense Advanced Research Projects Agency, the Office of Naval Research and the National Science Foundation of the United States of America, and by member companies of the Carnegie Mellon Sphinx Speech Consortium. We acknowledge the contributions of many volunteers to the expansion and improvement of this dictionary.
THIS SOFTWARE IS PROVIDED BY CARNEGIE MELLON UNIVERSITY “AS IS” AND ANY EXPRESSED OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL CARNEGIE MELLON UNIVERSITY NOR ITS EMPLOYEES BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.