skip to navigation
skip to content

Not Logged In

google-ngram-downloader 3.0

The streaming access to the Google ngram data.

Latest Version: 3.1.1

https://travis-ci.org/dimazest/google-ngram-downloader.png?branch=master https://coveralls.io/repos/dimazest/google-ngram-downloader/badge.png?branch=master

The Google Books Ngram Viewer dataset is a freely available resource under a Creative Commons Attribution 3.0 Unported License which provides ngram counts over books scanned by Google.

The data is so big, that storing it is almost impossible. However, sometimes you need an aggregate data over the dataset. For example to build a co-occurrence matrix.

This package provides an iterator over the dataset stored at Google. It decompresses the data on the fly and provides you the access to the underlying data.

Example use

>>> from google_ngram_downloader import readline_google_store
>>>
>>> fname, url, records = next(readline_google_store(ngram_len=5))
>>> fname
'googlebooks-eng-all-5gram-20120701-0.gz'
>>> url
'http://storage.googleapis.com/books/ngrams/books/googlebooks-eng-all-5gram-20120701-0.gz'
>>> next(records)
Record(ngram=u'0 " A most useful', year=1860, match_count=1, volume_count=1)

Installation

pip intall google-ngram-downloader

The command line tool

It also provides a simple command line tool to download the ngrams called google-ngram-downloader.

Changes

Version 3.0

  • download, readile and cooccurrence subcommands.
  • readline_google_store transforms lines to Record in several processes.
 
File Type Py Version Uploaded on Size
google-ngram-downloader-3.0.tar.gz (md5) Source 2013-12-11 8KB
  • Downloads (All Versions):
  • 15 downloads in the last day
  • 82 downloads in the last week
  • 489 downloads in the last month