English NER in Flair (4-class)

This is the multilingual 4-class NER model for Flair.

Our model can predict entities in any language text. However, for some languages you should use a language-specific tokenizer (see examples below).

Predicts 4 tags:

tag meaning
PER person name
LOC location name
ORG organization name
MISC other name

⚠️ Default license: noncommercial use only. This model is released under the Flukes NC 1.0 License. Commercial use — including using this model's predictions in a commercial product or service — requires a separate license. Contact alan.akbik@gmail.com.


Example 1: Named Entity Recognition in Spanish

Requires: Flair (pip install flair)

from flair.data import Sentence
from flair.models import SequenceTagger

# load tagger
tagger = SequenceTagger.load("flair/entity-multilingual-8class")

# make example sentence with Spanish text
sentence = Sentence("La selección española de fútbol, dirigida por Vicente del Bosque, ganó en Johannesburgo la final del Mundial 2010 contra los Países Bajos.")

# predict NER tags
tagger.predict(sentence)

# print sentence
print(sentence)

# print predicted NER spans
print('The following NER tags are found:')
# iterate over entities and print
for entity in sentence.get_spans('ner'):
    print(entity)

This yields the following output:

Span[1:5]: "selección española de fútbol" → ORG (1.0000)
Span[8:11]: "Vicente del Bosque" → PER (1.0000)
Span[14:15]: "Johannesburgo" → LOC (1.0000)
Span[18:20]: "Mundial 2010" → MISC (1.0000)
Span[22:24]: "Países Bajos" → ORG (1.0000)

So, the entities "selección española de fútbol" and "Países Bajos" are labeled as a organization (soccer teams), "Vicente del Bosque" is labeled as a person, "Johannesburgo" is labeled as a location and "Mundial 2010", labeled as other entity (MISC).


Example 2: Named Entity Recognition in French

Flair by default uses SegTok for tokenization, which does not work so well for French, as it does not split contractions with elision like "L'équipe". In the newest Flair version, you can add additional split tokens, which alleviates this issue.

from flair.data import Sentence
from flair.models import SequenceTagger

# load tagger
tagger = SequenceTagger.load("flair/entity-multilingual-8class")

# make a tokenizer for French that splits on contraction characters
tokenizer = SegtokTokenizer(additional_split_characters=["'"])

# make example sentence with French text, pass the special tokenizer
sentence = Sentence("L'équipe de France de football, entraînée par Aimé Jacquet, a remporté à Saint-Denis la finale de la Coupe du monde 1998 face au Brésil.", use_tokenizer=tokenizer)

# predict NER tags
tagger.predict(sentence)

# print sentence
print(sentence)

# print predicted NER spans
print('The following NER tags are found:')
# iterate over entities and print
for entity in sentence.get_spans('ner'):
    print(entity)

This yields the following output:

Span[2:5]: "équipe de France" → ORG (1.0000)
Span[10:12]: "Aimé Jacquet" → PER (1.0000)
Span[16:17]: "Saint-Denis" → LOC (1.0000)
Span[21:25]: "Coupe du monde 1998" → MISC (1.0000)
Span[27:28]: "Brésil" → ORG (1.0000)

So, the entities "équipe de France" and "Brésil" are labeled as a organization (soccer teams), "Aimé Jacquet" is labeled as a person, "Saint-Denis" is labeled as a location and "Coupe du monde 1998" is labeled as other entity (MISC).


Example 3: Named Entity Recognition in Japanese

Flair ships external tokenizers for Japanese.

from flair.data import Sentence
from flair.models import SequenceTagger

# load tagger
tagger = SequenceTagger.load("flair/entity-multilingual-8class")

# make a tokenizer for Japanese
tokenizer = JapaneseTokenizer("janome")

# make example sentence with Japanese text, pass the special tokenizer
sentence = Sentence("森保一監督率いる日本代表は、カタールで開催された2022年のワールドカップで、元世界王者のスペインを破る歴史的な勝利を収めました。", use_tokenizer=tokenizer)

# predict NER tags
tagger.predict(sentence)

# print sentence
print(sentence)

# print predicted NER spans
print('The following NER tags are found:')
# iterate over entities and print
for entity in sentence.get_spans('ner'):
    print(entity)

This yields the following output:

Span[0:2]: "森保一" → PER (1.0000)
Span[4:5]: "日本" → ORG (1.0000)
Span[8:9]: "カタール" → LOC (1.0000)
Span[17:18]: "ワールドカップ" → MISC (1.0000)
Span[24:25]: "スペイン" → ORG (1.0000)

So, the entities "日本" and "スペイン" are labeled as a organization (soccer teams), "森保一" is labeled as a person, "カタール" is labeled as a location and "ワールドカップ" is labeled as other entity (MISC).


Cite

Please cite the following paper when using this model.

@inproceedings{akbik2019flair,
    title={{FLAIR}: An easy-to-use framework for state-of-the-art {NLP}},
    author={Akbik, Alan and Bergmann, Tanja and Blythe, Duncan and Rasul, Kashif and Schweter, Stefan and Vollgraf, Roland},
    booktitle={{NAACL} 2019, 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations)},
    pages={54--59},
    year={2019}
}

License

Model weights: Flukes Noncommercial License 1.0. Personal, academic, and other noncommercial use permitted. Commercial use requires a separate license — contact alan.akbik@gmail.com.

Downloads last month
33
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including flair/entity-multilingual-4class