mediawiki-dump

Python package for working with MediaWiki XML content dumps

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

# mediawiki-dump
[![Build Status](https://travis-ci.org/macbre/mediawiki-dump.svg?branch=master)](https://travis-ci.org/macbre/mediawiki-dump)

```
pip install mediawiki_dump
```

[Python3 package](https://pypi.org/project/mediawiki_dump/) for working with [MediaWiki XML content dumps](https://www.mediawiki.org/wiki/Manual:Backing_up_a_wiki#Backup_the_content_of_the_wiki_(XML_dump)).

Wikipedia (bz2 compressed) and Wikia (7zip) content dumps are supported.

## Dependencies

In order to read 7zip archives (used by Wikia's XML dumps) you need to install [`libarchive`](http://libarchive.org/):

```
sudo apt install libarchive-dev
```

## API

### Tokenizer

Allows you to clean up the wikitext:

```python
from mediawiki_dump.tokenizer import clean
clean('[[Foo|bar]] is a link')
'bar is a link'
```

And then tokenize the text:

```python
from mediawiki_dump.tokenizer import tokenize
tokenize('11. juni 2007 varð kunngjørt, at Svínoyar kommuna verður løgd saman við Klaksvíkar kommunu eftir komandi bygdaráðsval.')
['juni', 'varð', 'kunngjørt', 'at', 'Svínoyar', 'kommuna', 'verður', 'løgd', 'saman', 'við', 'Klaksvíkar', 'kommunu', 'eftir', 'komandi', 'bygdaráðsval']
```

### Dump reader

Fetch and parse dumps (using a local file cache):

```python
from mediawiki_dump.dumps import WikipediaDump
from mediawiki_dump.reader import DumpReader

dump = WikipediaDump('fo')
pages = DumpReader().read(dump)

[title for _, _, title, *rest in pages][:10]

['Main Page', 'Brúkari:Jon Harald Søby', 'Forsíða', 'Ormurin Langi', 'Regin smiður', 'Fyrimynd:InterLingvLigoj', 'Heimsyvirlýsingin um mannarættindi', 'Bólkur:Kvæði', 'Bólkur:Yrking', 'Kjak:Forsíða']
```

`read` method yields the following per-revision information: `namespace`, `page_id`, `title`, `content`, `revision_id`, `timestamp`, `contributor` (`None` for anonymous edits).

By using `DumpReaderArticles` class you can read article pages only:

```python
import logging; logging.basicConfig(level=logging.INFO)

from mediawiki_dump.dumps import WikipediaDump
from mediawiki_dump.reader import DumpReaderArticles

dump = WikipediaDump('fo')
pages = DumpReaderArticles().read(dump)

print([title for _, _, title, *rest in pages][:25])
```

Will give you:

```
INFO:DumpReaderArticles:Parsing XML dump...
INFO:WikipediaDump:Checking /tmp/wikicorpus_62da4928a0a307185acaaa94f537d090.bz2 cache file...
INFO:WikipediaDump:Fetching fo dump from <https://dumps.wikimedia.org/fowiki/latest/fowiki-latest-pages-meta-current.xml.bz2>...
INFO:WikipediaDump:HTTP 200 (14105 kB fetched)
INFO:WikipediaDump:Cache set
...
['WIKIng', 'Føroyar', 'Borðoy', 'Eysturoy', 'Fugloy', 'Forsíða', 'Løgmenn í Føroyum', 'GNU Free Documentation License', 'GFDL', 'Opið innihald', 'Wikipedia', 'Alfrøði', '2004', '20. juni', 'WikiWiki', 'Wiki', 'Danmark', '21. juni', '22. juni', '23. juni', 'Lívfrøði', '24. juni', '25. juni', '26. juni', '27. juni']
```

## Reading Wikia's dumps

```python
import logging; logging.basicConfig(level=logging.INFO)

from mediawiki_dump.dumps import WikiaDump
from mediawiki_dump.reader import DumpReaderArticles

dump = WikiaDump('plnordycka')
pages = DumpReaderArticles().read(dump)

print([title for _, _, title, *rest in pages][:25])
```

Will give you:

```
INFO:DumpReaderArticles:Parsing XML dump...
INFO:WikiaDump:Checking /tmp/wikicorpus_f7dd3b75c5965ee10ae5fe4643fb806b.7z cache file...
INFO:WikiaDump:Fetching plnordycka dump from <https://s3.amazonaws.com/wikia_xml_dumps/p/pl/plnordycka_pages_current.xml.7z>...
INFO:WikiaDump:HTTP 200 (129 kB fetched)
INFO:WikiaDump:Cache set
INFO:WikiaDump:Reading wikicorpus_f7dd3b75c5965ee10ae5fe4643fb806b file from dump
...
INFO:DumpReaderArticles:Parsing completed, entries found: 615
['Nordycka Wiki', 'Strona główna', '1968', '1948', 'Ormurin Langi', 'Mykines', 'Trollsjön', 'Wyspy Owcze', 'Nólsoy', 'Sandoy', 'Vágar', 'Mørk', 'Eysturoy', 'Rakfisk', 'Hákarl', '1298', 'Sztokfisz', '1978', '1920', 'Najbardziej na północ', 'Svalbard', 'Hamferð', 'Rok w Skandynawii', 'Islandia', 'Rissajaure']
```

## Fetching full history

Pass `full_history` to `BaseDump` constructor to fetch the XML content dump with full history:

```python
import logging; logging.basicConfig(level=logging.INFO)

from mediawiki_dump.dumps import WikiaDump
from mediawiki_dump.reader import DumpReaderArticles

dump = WikiaDump('macbre', full_history=True) # fetch full history, including old revisions
pages = DumpReaderArticles().read(dump)

print('\n'.join(['%s %s (%s)' % (str(timestamp), title, author) for _, _, title, _, _, timestamp, author in pages]))
```

Will give you:

```
INFO:DumpReaderArticles:Parsing completed, entries found: 384
2016-10-12 19:51:06+00:00 Macbre Wiki (Default)
2016-10-12 19:51:05+00:00 Macbre Wiki (Wikia)
2016-11-04 10:33:20+00:00 Macbre Wiki (Macbre)
2016-11-04 10:37:17+00:00 Macbre Wiki (FandomBot)
2017-01-25 14:47:37+00:00 Macbre Wiki (FandomBot)
2017-04-10 11:20:25+00:00 Macbre Wiki (Ryba777)
2017-04-10 11:21:20+00:00 Macbre Wiki (Ryba777)
2018-03-07 12:51:12+00:00 Macbre Wiki (Macbre)
2016-10-12 19:51:05+00:00 Main Page (Wikia)
2016-11-08 10:15:33+00:00 FooBar (None)
2016-11-08 10:15:49+00:00 FooBar (None)
...
2018-06-05 11:45:44+00:00 YouTube tag (FANDOMbot)
2018-06-06 08:51:24+00:00 Maps (Macbre)
2018-06-07 08:17:13+00:00 Maps (Macbre)
2018-06-07 08:17:36+00:00 Maps (Macbre)
2018-07-24 14:52:20+00:00 Scary transclusion (Macbre)
2018-09-11 14:04:15+00:00 Lua (Macbre)
2018-09-11 14:14:24+00:00 Lua (Macbre)
2018-09-11 14:14:37+00:00 Lua (Macbre)
```

Project details

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

1.3.0

Apr 24, 2024

1.2.1

Apr 17, 2024

1.2.0

Sep 26, 2023

1.1.0

Mar 15, 2023

1.0.0

Aug 30, 2021

0.8.0

Jun 15, 2021

0.7.0

Sep 16, 2020

0.6.7

Jul 28, 2019

0.6.6

Jul 22, 2019

0.6.5

Jun 13, 2019

0.6.4

Apr 1, 2019

0.6.3

Mar 25, 2019

0.6.2

Nov 24, 2018

0.6.1

Nov 24, 2018

0.6

Nov 24, 2018

0.5

Nov 22, 2018

0.4

Nov 12, 2018

This version

0.3

Oct 30, 2018

0.2

Oct 26, 2018

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mediawiki_dump-0.3.tar.gz (8.3 kB view hashes)

Uploaded Oct 30, 2018 Source

Hashes for mediawiki_dump-0.3.tar.gz

Hashes for mediawiki_dump-0.3.tar.gz
Algorithm	Hash digest
SHA256	`5ad07d89a24c001cb17b3fea2542c1c615b158fa690b33ff986e64a4035d6e2f`
MD5	`ffd6242bd8020f33d529dc7ef96e7939`
BLAKE2b-256	`bc8a1ac5528ca475dcc9943952e0e02f9a63481a5ff11186d64e826e64bece72`