bert portuguese ner

Providerlfcc
Categorytoken-classification
Licensemit
Downloads508.8K
Stars0

Overview

The BERT Portuguese NER model is a specialized token-classification engine fine-tuned specifically for Named Entity Recognition within Portuguese text. For developers building NLP pipelines, this model provides a reliable way to extract structured entities—such as names, organizations, and locations—from unstructured Lusophone data. Unlike general-purpose LLMs, this BERT-based architecture is optimized for sequence labeling, offering lower latency and higher precision for specific extraction tasks. It integrates seamlessly into standard Hugging Face or PyTorch workflows, making it an efficient choice for automating data labeling, enhancing search indexing, or building entity-aware chatbots for the Brazilian and Portuguese markets.

Highlights

  • Optimized for Portuguese Named Entity Recognition tasks
  • Efficient token-classification for low-latency production environments
  • Seamless integration with standard NLP frameworks
  • Permissive MIT license for commercial flexibility

Usage

Install
# Install Hugging Face transformers
pip install transformers torch
SDK Usage
# Load model with transformers
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("lfcc/bert-portuguese-ner")
tokenizer = AutoTokenizer.from_pretrained("lfcc/bert-portuguese-ner")

Hugging Face Download

We recommend downloading the model via the Hugging Face CLI or Hub SDK.

Guidance:Before downloading, install huggingface_hub with:

Guidance
pip install -U huggingface_hub

CLI Download

Download the full repository

Download the full repository
huggingface-cli download lfcc/bert-portuguese-ner

Download a single file to a local folder (e.g. config.json into ./dir)

Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download lfcc/bert-portuguese-ner config.json --local-dir ./dir

See the official docs for more CLI options

SDK Download

SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('lfcc/bert-portuguese-ner')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://huggingface.co/lfcc/bert-portuguese-ner

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/lfcc/bert-portuguese-ner

Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.

PyTorch / Transformers Usage

Install Transformers

Install Transformers
pip install -U transformers torch

Load the model and run inference

Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('lfcc/bert-portuguese-ner')
tokenizer = AutoTokenizer.from_pretrained('lfcc/bert-portuguese-ner')

Full Documentation

来源: HuggingFace

---
license: mit
tags:

  • generated_from_trainer

metrics:
  • precision

  • recall

  • f1

  • accuracy

model_index:
  • name: bert-portuguese-ner-archive

results:
- task:
name: Token Classification
type: token-classification
metric:
name: Accuracy
type: accuracy
value: 0.9700325118974698
---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You
should probably proofread and complete it, then remove this comment. -->

bert-portuguese-ner

This model is a fine-tuned version of neuralmind/bert-base-portuguese-cased
It achieves the following results on the evaluation set:

  • Loss: 0.1140

  • Precision: 0.9147

  • Recall: 0.9483

  • F1: 0.9312

  • Accuracy: 0.9700

Model description

This model was fine-tunned on token classification task (NER) on Portuguese archival documents. The annotated labels are: Date, Profession, Person, Place, Organization

Datasets

All the training and evaluation data is available at: http://ner.epl.di.uminho.pt/

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05

  • train_batch_size: 16

  • eval_batch_size: 16

  • seed: 42

  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08

  • lr_scheduler_type: linear

  • num_epochs: 4

Training results

| Training Loss | Epoch | Step | Validation Loss | Precision | Recall | F1 | Accuracy |
|:-------------:|:-----:|:----:|:---------------:|:---------:|:------:|:------:|:--------:|
| No log | 1.0 | 192 | 0.1438 | 0.8917 | 0.9392 | 0.9148 | 0.9633 |
| 0.2454 | 2.0 | 384 | 0.1222 | 0.8985 | 0.9417 | 0.9196 | 0.9671 |
| 0.0526 | 3.0 | 576 | 0.1098 | 0.9150 | 0.9481 | 0.9312 | 0.9698 |
| 0.0372 | 4.0 | 768 | 0.1140 | 0.9147 | 0.9483 | 0.9312 | 0.9700 |

Framework versions

  • Transformers 4.10.0.dev0
  • Pytorch 1.9.0+cu111
  • Datasets 1.10.2
  • Tokenizers 0.10.3

Citation

bibtex
@Article{make4010003,
AUTHOR = {Cunha, Luís Filipe and Ramalho, José Carlos},
TITLE = {NER in Archival Finding Aids: Extended},
JOURNAL = {Machine Learning and Knowledge Extraction},
VOLUME = {4},
YEAR = {2022},
NUMBER = {1},
PAGES = {42--65},
URL = {https://www.mdpi.com/2504-4990/4/1/3},
ISSN = {2504-4990},
ABSTRACT = {The amount of information preserved in Portuguese archives has increased over the years. These documents represent a national heritage of high importance, as they portray the country&rsquo;s history. Currently, most Portuguese archives have made their finding aids available to the public in digital format, however, these data do not have any annotation, so it is not always easy to analyze their content. In this work, Named Entity Recognition solutions were created that allow the identification and classification of several named entities from the archival finding aids. These named entities translate into crucial information about their context and, with high confidence results, they can be used for several purposes, for example, the creation of smart browsing tools by using entity linking and record linking techniques. In order to achieve high result scores, we annotated several corpora to train our own Machine Learning algorithms in this context domain. We also used different architectures, such as CNNs, LSTMs, and Maximum Entropy models. Finally, all the created datasets and ML models were made available to the public with a developed web platform, NER@DI.},
DOI = {10.3390/make4010003}
}
Join our Telegram