scrapelib

a library for scraping things

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Development Status
- 4 - Beta
Intended Audience
- Developers
License
- OSI Approved :: BSD License
Natural Language
- English
Operating System
- OS Independent
Programming Language
- Python
Topic
- Software Development :: Libraries :: Python Modules

Project description

A Python library for scraping things.

Features include:

HTTP, HTTPS, FTP requests via an identical API

HTTP caching, compression and cookies

redirect following

request throttling

robots.txt compliance (optional)

robust error handling

Written by Michael Stephens <mstephens@sunlightfoundation.com> and James Turk <jturk@sunlightfoundation.com>.

Source is available at http://github.com/sunlightlabs/scrapelib.

Requirements

python >= 2.6

httplib2 optional but highly recommended.

Installation

scrapelib is available on PyPI and thus can be downloaded installed via pip install scrapelib or easy_install scrapelib.

To install from a source distribution run python setup.py install.

Example Usage

import scrapelib
s = scrapelib.Scraper(requests_per_minute=10, allow_cookies=True,
                      follow_robots=True)

# Grab Google front page
s.urlopen('http://google.com')

# Will raise RobotExclusionError
s.urlopen('http://google.com/search')

# Will be throttled to 10 HTTP requests per minute
while True:
    s.urlopen('http://example.com')

Project details

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Development Status
- 4 - Beta
Intended Audience
- Developers
License
- OSI Approved :: BSD License
Natural Language
- English
Operating System
- OS Independent
Programming Language
- Python
Topic
- Software Development :: Libraries :: Python Modules

Release history Release notifications | RSS feed

2.3.0

Dec 15, 2023

2.2.0

May 18, 2023

2.1.0

Nov 7, 2022

2.0.7

Jul 6, 2022

2.0.6

Jun 23, 2021

2.0.5

Jun 15, 2021

2.0.4

Apr 13, 2021

2.0.3

Apr 13, 2021

2.0.2

Apr 9, 2021

2.0.1

Apr 9, 2021

2.0.0

Apr 9, 2021

1.2.0

Nov 13, 2018

1.1.1

Apr 16, 2018

1.1.0

Jun 6, 2017

1.0.2

Apr 16, 2017

1.0.1

Apr 16, 2017

1.0.0

Mar 20, 2015

0.10.1

Jan 22, 2015

0.10.0

Jul 15, 2014

0.9.1

Mar 28, 2014

0.9.0

May 22, 2013

0.8.0

Mar 19, 2013

0.7.4

Dec 21, 2012

0.7.3

Jun 21, 2012

0.7.2

May 9, 2012

0.7.1

Apr 27, 2012

0.7.0

Apr 23, 2012

0.6.2

Apr 20, 2012

0.6.1

Apr 19, 2012

0.6.0

Apr 19, 2012

0.5.8

Apr 18, 2012

0.5.7

Feb 2, 2012

0.5.6

Nov 9, 2011

0.5.5

Sep 27, 2011

0.5.4

Jun 7, 2011

0.5.3

Jun 7, 2011

0.5.2

May 16, 2011

0.5.0

Mar 21, 2011

This version

0.4.3

Feb 11, 2011

0.4.2

Feb 8, 2011

0.4.1

Dec 7, 2010

0.4.0

Nov 8, 2010

0.3.0

Oct 5, 2010

0.2.0

Jul 13, 2010

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapelib-0.4.3.tar.gz (9.5 kB view hashes)

Uploaded Feb 11, 2011 Source

Hashes for scrapelib-0.4.3.tar.gz

Hashes for scrapelib-0.4.3.tar.gz
Algorithm	Hash digest
SHA256	`7e3f9a345b5e1b6c2801842f6f0fcb5f6cc28ee286c931697240a15e2ab1be74`
MD5	`8b37ee2b52f93939f2c4dcb7afc57725`
BLAKE2b-256	`e9b7bf6aaea4111ff555e51a238e74b9baa125f9569441dd6f2fff42c5d95583`