A multi-lingual approach to AllenNLP CoReference Resolution, along with a wrapper for spaCy.

These details have not been verified by PyPI

Project links

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

Crosslingual Coreference

Coreference is amazing but the data required for training a model is very scarce. In our case, the available training for non-English languages also proved to be poorly annotated. Crosslingual Coreference, therefore, uses the assumption a trained model with English data and cross-lingual embeddings should work for languages with similar sentence structures.

Install

pip install crosslingual-coreference

Quickstart

from crosslingual_coreference import Predictor

text = (
    "Do not forget about Momofuku Ando! He created instant noodles in Osaka. At"
    " that location, Nissin was founded. Many students survived by eating these"
    " noodles, but they don't even know him."
)

# choose minilm for speed/memory and info_xlm for accuracy
predictor = Predictor(
    language="en_core_web_sm", device=-1, model_name="minilm"
)

print(predictor.predict(text)["resolved_text"])
# Output
#
# Do not forget about Momofuku Ando!
# Momofuku Ando created instant noodles in Osaka.
# At Osaka, Nissin was founded.
# Many students survived by eating instant noodles,
# but Many students don't even know Momofuku Ando.

Chunking/batching to resolve memory OOM errors

from crosslingual_coreference import Predictor

predictor = Predictor(
    language="en_core_web_sm",
    device=0,
    model_name="minilm",
    chunk_size=2500,
    chunk_overlap=2,
)

Use spaCy pipeline

import spacy

import crosslingual_coreference

text = (
    "Do not forget about Momofuku Ando! He created instant noodles in Osaka. At"
    " that location, Nissin was founded. Many students survived by eating these"
    " noodles, but they don't even know him."
)


nlp = spacy.load("en_core_web_sm")
nlp.add_pipe(
    "xx_coref", config={"chunk_size": 2500, "chunk_overlap": 2, "device": 0}
)

doc = nlp(text)
print(doc._.coref_clusters)
# Output
#
# [[[4, 5], [7, 7], [27, 27], [36, 36]],
# [[12, 12], [15, 16]],
# [[9, 10], [27, 28]],
# [[22, 23], [31, 31]]]
print(doc._.resolved_text)
# Output
#
# Do not forget about Momofuku Ando!
# Momofuku Ando created instant noodles in Osaka.
# At Osaka, Nissin was founded.
# Many students survived by eating instant noodles,
# but Many students don't even know Momofuku Ando.

Available models

As of now, there are two models available "info_xlm", "xlm_roberta", "minilm", which scored 77, 74 and 74 on OntoNotes Release 5.0 English data, respectively.

More Examples

Project details

These details have not been verified by PyPI

Project links

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

0.3.1

Jun 19, 2023

0.3

Apr 5, 2023

0.2.9

Sep 24, 2022

0.2.8

Jul 14, 2022

0.2.7

Jul 14, 2022

0.2.6

Jun 8, 2022

0.2.5

May 25, 2022

0.2.4

May 10, 2022

This version

0.2.3

May 5, 2022

0.2.2

May 5, 2022

0.2.1

Apr 13, 2022

0.2.0

Apr 3, 2022

0.1.5

Mar 31, 2022

0.1.4

Mar 30, 2022

0.1.3

Mar 29, 2022

0.1.2

Mar 29, 2022

0.1.1

Mar 28, 2022

0.1.0

Mar 28, 2022

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

crosslingual-coreference-0.2.3.tar.gz (10.3 kB view hashes)

Uploaded May 5, 2022 Source

Built Distribution

crosslingual_coreference-0.2.3-py3-none-any.whl (11.6 kB view hashes)

Uploaded May 5, 2022 Python 3

Hashes for crosslingual-coreference-0.2.3.tar.gz

Hashes for crosslingual-coreference-0.2.3.tar.gz
Algorithm	Hash digest
SHA256	`c6bb56dfdca24a4d667c5c41fee1e562f9d3c1cc16bcbe8525990cd4c3114f19`
MD5	`9768435258f415327800f1c0c35768db`
BLAKE2b-256	`5d60565f342532d3d632b4ffe94521a367170b6c3d33a6dfd76111ded52daa02`

Hashes for crosslingual_coreference-0.2.3-py3-none-any.whl

Hashes for crosslingual_coreference-0.2.3-py3-none-any.whl
Algorithm	Hash digest
SHA256	`b216bd9591bcb91173fcaf945b591d47d0384582d6ab6dab7a37113a84f2b594`
MD5	`f740807a1164b2f10c103b5740bd28a0`
BLAKE2b-256	`e09c352141e6f10d3958e7338d23e112f3bc8fd57a9b2ca689fe74b56cbc0bf6`