sumy · PyPI

Module for automatic text summarization of HTML documents.

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

https://api.travis-ci.org/miso-belica/sumy.png?branch=master

Here are some other summarizers:

https://github.com/thavelick/summarize/ - Python, TF (very simple)
Reduction - Python, TextRank (simple)
Open Text Summarizer - C, TF without normalization
Simple program that summarize text - Python, TF without normalization
Intro to Computational Linguistics - Java, LexRank
Sumtract: Second project for UW LING 572 - Python
TextTeaser - Scala
Automatic Document Summarizer - Java, Bipartite HITS (no sources)
Pythia - Python, LexRank & Centroid
SWING - Ruby
Topic Networks - R, topic models & bipartite graphs
Almus: Automatic Text Summarizer - Java, LSA (without source code)
Musutelsa - Java, LSA (always freezes)
http://mff.bajecni.cz/index.php - C++
MEAD - Perl, various methods + evaluation framework

Installation

Currently only from git repo (make sure you have Python installed)

$ wget https://github.com/miso-belica/sumy/archive/master.zip # download the sources
$ unzip master.zip # extract the downloaded file
$ cd sumy-master/
$ [sudo] python setup.py install # install the package

Or simply run:

$ [sudo] pip install git+git://github.com/miso-belica/sumy.git

Usage

Sumy contains command line utility for quick summarization of documents.

$ sumy lex-rank --length=10 --url=http://en.wikipedia.org/wiki/Automatic_summarization # what's summarization?
$ sumy luhn --language=czech --url=http://www.zdrojak.cz/clanky/automaticke-zabezpeceni/
$ sumy edmundson --language=czech --length=3% --url=http://cs.wikipedia.org/wiki/Bitva_u_Lipan
$ sumy --help # for more info

Various evaluation methods for some summarization method can be executed by commands below:

$ sumy_eval lex-rank reference_summary.txt --url=http://en.wikipedia.org/wiki/Automatic_summarization
$ sumy_eval lsa reference_summary.txt --language=czech --url=http://www.zdrojak.cz/clanky/automaticke-zabezpeceni/
$ sumy_eval edmundson reference_summary.txt --language=czech --url=http://cs.wikipedia.org/wiki/Bitva_u_Lipan
$ sumy_eval --help # for more info

Python API

Or you can use sumy like a library in your project.

# -*- coding: utf8 -*-

from __future__ import absolute_import
from __future__ import division, print_function, unicode_literals

from sumy.parsers.html import HtmlParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer
from sumy.nlp.stemmers.czech import stem_word
from sumy.utils import get_stop_words


if __name__ == "__main__":
    url = "http://www.zsstritezuct.estranky.cz/clanky/predmety/cteni/jak-naucit-dite-spravne-cist.html"
    parser = HtmlParser.from_url(url, Tokenizer("czech"))

    summarizer = LsaSummarizer(stem_word)
    summarizer.stop_words = get_stop_words("czech")

    for sentence in summarizer(parser.document, 20):
        print(sentence)

Tests

Run tests via

$ nosetests-2.6 && nosetests-3.2 && nosetests-2.7 && nosetests-3.3

Changelog

0.1.0 (2013-MM-DD)

First public release.

Project details

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

0.11.0

Oct 23, 2022

0.10.0

Apr 21, 2022

0.9.0

Oct 21, 2021

0.8.1

May 19, 2019

0.8.0

May 18, 2019

0.7.0

Jul 22, 2017

0.6.0

Mar 5, 2017

0.5.1

Nov 17, 2016

0.5.0

Nov 12, 2016

0.4.1

Mar 6, 2016

0.4.0

Dec 6, 2015

0.3.0

Jun 7, 2014

0.2.1

Jan 25, 2014

0.2.0

Jan 18, 2014

0.1.0

Oct 20, 2013

This version

0.0.1

Oct 19, 2013

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sumy-0.0.1.zip (43.4 kB view hashes)

Uploaded Oct 19, 2013 Source

Hashes for sumy-0.0.1.zip

Hashes for sumy-0.0.1.zip
Algorithm	Hash digest
SHA256	`a1cd97309d1bb72f6ab44dde8b4015c67a60c223803f72b0cef287d0c3dadc35`
MD5	`68a5dd9a3da7b4951369e4447f741e5f`
BLAKE2b-256	`071b32d6a040cb7695434270ce4757a8352ff2ccb9cf0dfd9accbdcaeaff90f4`