pdpipe

Easy pipelines for pandas.

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

Easy pipelines for pandas.

>>> df = pd.DataFrame(
        data=[[4, 165, 'USA'], [2, 180, 'UK'], [2, 170, 'Greece']],
        index=['Dana', 'Jack', 'Nick'],
        columns=['Medals', 'Height', 'Born']
    )
>>> pipeline = pdp.Coldrop('Medals').Binarize('Born')
>>> pipline(df)
            Height  Born_UK  Born_USA
    Dana     165        0         1
    Jack     180        1         0
    Nick     170        0         0

1 Installation

Install pdpipe with:

pip install pdpipe

Some stages require scikit-learn; they will simply not be loaded if scikit-learn is not found on the system, and pdpipe will issue a warning.

2 Use

2.1 Creating Pipline Stages

Create stages with the following syntax:

import pdpipde as pdp
drop_name = pdp.ColDrop("Name")

By default, pipeline stages raise an exception if a DataFrame not meeting their precondition is piped through. This behaviour can be set per-stage by assigning exraise with a bool in a constructor call:

drop_name = pdp.ColDrop("Name", exraise=False)

2.2 Creating Piplines

Pipelines can be created by supplying a list of pipeline stages:

pipeline = pdp.Pipeline([pdp.ColDrop("Name"), pdp.Binarize("Label")])

Alternatively, you can add pipeline stages together:

pipeline = pdp.ColDrop("Name") + pdp.Binarize("Label")

Or even by adding pipelines together or pipelines to pipeline stages:

pipeline = pdp.ColDrop("Name") + pdp.Binarize("Label")
pipeline += pdp.MapColVals("Job", {"Part": True, "Full":True, "No": False})
pipeline += pdp.Pipeline([pdp.ColRename({"Job": "Employed"})])

Pipline stages can also be chained to other stages to create pipelines:

pipeline = pdp.ColDrop("Name").Binarize("Label").ValDrop([-1], "Children")

2.3 Applying Pipelines Stages

You can apply a pipeline stage to a DataFrame using its apply method:

res_df = pdp.ColDrop("Name").apply(df)

Pipeline stages are also callables, making the following syntax equivalent:

drop_name = pdp.ColDrop("Name")
res_df = drop_name(df)

The initialized exception behaviour of a pipeline stage can be overriden on a per-application basis:

drop_name = pdp.ColDrop("Name", exraise=False)
res_df = drop_name(df, exraise=True)

2.4 Applying Pipelines

Pipelines are pipeline stages themselves, and can be applied to DataFrame using the same syntax, applying each of the stages making them up, in order:

pipeline = pdp.ColDrop("Name") + pdp.Binarize("Label")
res_df = pipeline(df)

Assigning the exraise paramter to a pipeline apply call with a bool set or unsets exception raising on failed preconditions for all contained stages:

pipeline = pdp.ColDrop("Name") + pdp.Binarize("Label")
res_df = pipeline.apply(df, exraise=True)

3 Pipeline Stages

3.1 Basic Stages

ColDrop - Drop columns by name.
ValDrop - Drop rows by by their value in specific or all columns.
ValKeep - Keep rows by by their value in specific or all columns.
ColRename - Rename columns.
Bin - Convert a continous valued column to categoric data using binning.
Binarize - Convert a categorical column to the several binary columns corresponding to it.
MapColVals - Convert column values using a mapping.

3.2 Scikit-learn-dependent Stages

Encode - Encode a categorical column to corresponding number values.

4 Credits

Created by Shay Palachy (shay.palachy@gmail.com).

Project details

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

0.3.2

Sep 19, 2022

0.3.1

Aug 9, 2022

0.3.0

Jul 4, 2022

0.2.8

Jun 23, 2022

0.2.7

Jun 22, 2022

0.2.6

Jun 22, 2022

0.2.5

Jun 7, 2022

0.2.4

May 12, 2022

0.2.3

Mar 13, 2022

0.2.2

Mar 10, 2022

0.2.1

Feb 23, 2022

0.2.0

Feb 14, 2022

0.1.6

Jan 30, 2022

0.1.5

Jan 29, 2022

0.1.4

Jan 29, 2022

0.1.3

Jan 26, 2022

0.1.2

Jan 23, 2022

0.1.0

Jan 23, 2022

0.0.72

Jan 19, 2022

0.0.71

Dec 26, 2021

0.0.70

Dec 19, 2021

0.0.69

Dec 10, 2021

0.0.68

Dec 8, 2021

0.0.67

Nov 15, 2021

0.0.66

Nov 8, 2021

0.0.65

Nov 8, 2021

0.0.64

Nov 8, 2021

0.0.63

Nov 3, 2021

0.0.62

Oct 27, 2021

0.0.61

Oct 27, 2021

0.0.60

Sep 29, 2021

0.0.59

Aug 30, 2021

0.0.58

Aug 28, 2021

0.0.57

Aug 25, 2021

0.0.56

Aug 18, 2021

0.0.55

Aug 18, 2021

0.0.54

Aug 18, 2021

0.0.53

Nov 9, 2020

0.0.52

Oct 30, 2020

0.0.51

Oct 1, 2020

0.0.50

Aug 27, 2020

0.0.49

May 5, 2020

0.0.48

May 5, 2020

0.0.46

Feb 26, 2020

0.0.45

Feb 24, 2020

0.0.44

Feb 24, 2020

0.0.43

Feb 17, 2020

0.0.42

Feb 5, 2020

0.0.41

Feb 3, 2020

0.0.40

Feb 3, 2020

0.0.39

Jan 26, 2020

0.0.38

Jan 20, 2020

0.0.37

Jan 7, 2020

0.0.35

Dec 21, 2019

0.0.33

Dec 7, 2019

0.0.32

Dec 3, 2019

0.0.31

Jun 27, 2019

0.0.30

Jun 14, 2019

0.0.29

Jun 14, 2019

0.0.27

May 28, 2018

0.0.26

May 28, 2018

0.0.25

May 9, 2018

0.0.24

May 2, 2018

0.0.23

Apr 22, 2018

0.0.22

Apr 16, 2018

0.0.21

Apr 16, 2018

0.0.20

Apr 8, 2018

0.0.19

Apr 8, 2018

0.0.18

Mar 20, 2018

0.0.17

Mar 12, 2018

0.0.16

Mar 11, 2018

0.0.15

Mar 7, 2018

0.0.14

Mar 7, 2018

0.0.13

Mar 3, 2018

0.0.12

Feb 12, 2018

0.0.11

Feb 12, 2018

0.0.10

Feb 12, 2018

0.0.9

Feb 5, 2018

0.0.8

Feb 5, 2018

0.0.7

Jan 30, 2018

0.0.6

Jan 14, 2018

0.0.5

May 24, 2017

0.0.4

May 5, 2017

0.0.3

May 5, 2017

This version

0.0.2

Mar 17, 2017

0.0.1

Mar 16, 2017

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pdpipe-0.0.2.tar.gz (28.0 kB view hashes)

Uploaded Mar 17, 2017 Source

Hashes for pdpipe-0.0.2.tar.gz

Hashes for pdpipe-0.0.2.tar.gz
Algorithm	Hash digest
SHA256	`05caa59c03267f84dc76d4dbf4eae706056eb052e67822df2360ec82583853a7`
MD5	`ac35089648ee28bcb4cc3c93efe2da87`
BLAKE2b-256	`8bea6b3625737d733148860249cee556146d715feb1eff409b2cd6ec911a9dde`