sumy · PyPI

Module for automatic summarization of text documents and HTML pages.

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

https://api.travis-ci.org/miso-belica/sumy.png?branch=master

Here are some other summarizers:

https://github.com/thavelick/summarize/ - Python, TF (very simple)
Reduction - Python, TextRank (simple)
Open Text Summarizer - C, TF without normalization
Simple program that summarize text - Python, TF without normalization
Intro to Computational Linguistics - Java, LexRank
Sumtract: Second project for UW LING 572 - Python
TextTeaser - Scala
PyTeaser - TextTeaser port in Python
Automatic Document Summarizer - Java, Bipartite HITS (no sources)
Pythia - Python, LexRank & Centroid
SWING - Ruby
Topic Networks - R, topic models & bipartite graphs
Almus: Automatic Text Summarizer - Java, LSA (without source code)
Musutelsa - Java, LSA (always freezes)
http://mff.bajecni.cz/index.php - C++
MEAD - Perl, various methods + evaluation framework

Installation

Make sure you have Python 2.6+/3.2+ and pip (Windows, Linux) installed. Run simply (preferred way):

$ [sudo] pip install sumy

Or for the fresh version:

$ [sudo] pip install git+git://github.com/miso-belica/sumy.git

Or if you have to:

$ wget https://github.com/miso-belica/sumy/archive/master.zip # download the sources
$ unzip master.zip # extract the downloaded file
$ cd sumy-master/
$ [sudo] python setup.py install # install the package

Usage

Sumy contains command line utility for quick summarization of documents.

$ sumy lex-rank --length=10 --url=http://en.wikipedia.org/wiki/Automatic_summarization # what's summarization?
$ sumy luhn --language=czech --url=http://www.zdrojak.cz/clanky/automaticke-zabezpeceni/
$ sumy edmundson --language=czech --length=3% --url=http://cs.wikipedia.org/wiki/Bitva_u_Lipan
$ sumy --help # for more info

Various evaluation methods for some summarization method can be executed by commands below:

$ sumy_eval lex-rank reference_summary.txt --url=http://en.wikipedia.org/wiki/Automatic_summarization
$ sumy_eval lsa reference_summary.txt --language=czech --url=http://www.zdrojak.cz/clanky/automaticke-zabezpeceni/
$ sumy_eval edmundson reference_summary.txt --language=czech --url=http://cs.wikipedia.org/wiki/Bitva_u_Lipan
$ sumy_eval --help # for more info

Python API

Or you can use sumy like a library in your project.

# -*- coding: utf8 -*-

from __future__ import absolute_import
from __future__ import division, print_function, unicode_literals

from sumy.parsers.html import HtmlParser
from sumy.parsers.plaintext import PlaintextParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer as Summarizer
from sumy.nlp.stemmers import Stemmer
from sumy.utils import get_stop_words


LANGUAGE = "czech"
SENTENCES_COUNT = 10


if __name__ == "__main__":
    url = "http://www.zsstritezuct.estranky.cz/clanky/predmety/cteni/jak-naucit-dite-spravne-cist.html"
    parser = HtmlParser.from_url(url, Tokenizer(LANGUAGE))
    # or for plain text files
    # parser = PlaintextParser.from_file("document.txt", Tokenizer(LANGUAGE))
    stemmer = Stemmer(LANGUAGE)

    summarizer = Summarizer(stemmer)
    summarizer.stop_words = get_stop_words(LANGUAGE)

    for sentence in summarizer(parser.document, SENTENCES_COUNT):
        print(sentence)

Tests

Run tests via

$ nosetests-2.6 && nosetests-3.2 && nosetests-2.7 && nosetests-3.3

Changelog

0.3.0 (2014-06-07)

Added possibility to specify format of input document for URL & stdin. Thanks to @Lucas-C.
Added possibility to specify custom file with stop-words in CLI. Thanks to @Lucas-C.
Added support for French language (added stopwords & stemmer). Thanks to @Lucas-C.
Function sumy.utils.get_stop_words raises LookupError instead of ValueError for unknown language.
Exception LookupError is raised for unknown language of stemmer instead of falling silently to null_stemmer.

0.2.1 (2014-01-23)

Fixed installation of my own readability fork. Added breadability to the dependencies instead of it #8. Thanks to @pratikpoddar.

0.2.0 (2014-01-18)

Removed dependency on SciPy #7. Use numpy.linalg.svd implementation. Thanks to Shantanu.

0.1.0 (2013-10-20)

First public release.

Project details

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

0.11.0

Oct 23, 2022

0.10.0

Apr 21, 2022

0.9.0

Oct 21, 2021

0.8.1

May 19, 2019

0.8.0

May 18, 2019

0.7.0

Jul 22, 2017

0.6.0

Mar 5, 2017

0.5.1

Nov 17, 2016

0.5.0

Nov 12, 2016

0.4.1

Mar 6, 2016

0.4.0

Dec 6, 2015

This version

0.3.0

Jun 7, 2014

0.2.1

Jan 25, 2014

0.2.0

Jan 18, 2014

0.1.0

Oct 20, 2013

0.0.1

Oct 19, 2013

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sumy-0.3.0.zip (45.7 kB view hashes)

Uploaded Jun 7, 2014 Source

Built Distribution

sumy-0.3.0-py2.py3-none-any.whl (42.1 kB view hashes)

Uploaded Jun 7, 2014 Python 2 Python 3

Hashes for sumy-0.3.0.zip

Hashes for sumy-0.3.0.zip
Algorithm	Hash digest
SHA256	`f0755f044118fe95a7c5e01dae973a2a894b1a5975b7bb7a7e73e63faad8ab9c`
MD5	`20cae178a7edb1e6499aece5704aa743`
BLAKE2b-256	`003b0cb489965c06955f8ca42b6922b293fe2ad6449f8be7bf2d12856f083a0f`

Hashes for sumy-0.3.0-py2.py3-none-any.whl

Hashes for sumy-0.3.0-py2.py3-none-any.whl
Algorithm	Hash digest
SHA256	`ac5a81e4b169a8d2549dcd1f0b6ca286698825b7f38f4dd7750dbbcf27bff338`
MD5	`20825d296e6ffda2f59435d7bea531e1`
BLAKE2b-256	`9ac27b9fe97308353b62b58b9beea7390f9b42e2080d5b5ab23d591c45ab1b8b`