bert large portuguese cased
Overview
Highlights
- Optimized specifically for high-accuracy Portuguese language processing
- Case-sensitive tokenization for better proper noun recognition
- Ideal for NER and sentiment analysis fine-tuning
- Seamless integration with Hugging Face Transformers library
- MIT licensed for flexible commercial and private deployment
Usage
# Install Hugging Face transformers
pip install transformers torch
# Load model with transformers
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("neuralmind/bert-large-portuguese-cased")
tokenizer = AutoTokenizer.from_pretrained("neuralmind/bert-large-portuguese-cased")
Hugging Face Download
We recommend downloading the model via the Hugging Face CLI or Hub SDK.
Guidance:Before downloading, install huggingface_hub with:
pip install -U huggingface_hub
CLI Download
Download the full repository
huggingface-cli download neuralmind/bert-large-portuguese-cased
Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download neuralmind/bert-large-portuguese-cased config.json --local-dir ./dir
See the official docs for more CLI options
SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('neuralmind/bert-large-portuguese-cased')
Git Download
Make sure git-lfs is installed first
git lfs install
git clone https://huggingface.co/neuralmind/bert-large-portuguese-cased
To skip LFS large-file downloads, use:
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/neuralmind/bert-large-portuguese-cased
Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.
PyTorch / Transformers Usage
Install Transformers
pip install -U transformers torch
Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('neuralmind/bert-large-portuguese-cased')
tokenizer = AutoTokenizer.from_pretrained('neuralmind/bert-large-portuguese-cased')
Full Documentation
---
language: pt
license: mit
tags:
- bert
- pytorch
datasets:
- brWaC
---
BERTimbau Large (aka "bert-large-portuguese-cased")
Introduction
BERTimbau Large is a pretrained BERT model for Brazilian Portuguese that achieves state-of-the-art performances on three downstream NLP tasks: Named Entity Recognition, Sentence Textual Similarity and Recognizing Textual Entailment. It is available in two sizes: Base and Large.
For further information or requests, please go to BERTimbau repository.
Available models
| Model | Arch. | #Layers | #Params |
| ---------------------------------------- | ---------- | ------- | ------- |
| neuralmind/bert-base-portuguese-cased | BERT-Base | 12 | 110M |
| neuralmind/bert-large-portuguese-cased | BERT-Large | 24 | 335M |
Usage
from transformers import AutoTokenizer # Or BertTokenizer
from transformers import AutoModelForPreTraining # Or BertForPreTraining for loading pretraining heads
from transformers import AutoModel # or BertModel, for BERT without pretraining heads
model = AutoModelForPreTraining.from_pretrained('neuralmind/bert-large-portuguese-cased')
tokenizer = AutoTokenizer.from_pretrained('neuralmind/bert-large-portuguese-cased', do_lower_case=False)
Masked language modeling prediction example
from transformers import pipeline
pipe = pipeline('fill-mask', model=model, tokenizer=tokenizer)
pipe('Tinha uma [MASK] no meio do caminho.')
[{'score': 0.5054386258125305,
'sequence': '[CLS] Tinha uma pedra no meio do caminho. [SEP]',
'token': 5028,
'token_str': 'pedra'},
{'score': 0.05616172030568123,
'sequence': '[CLS] Tinha uma curva no meio do caminho. [SEP]',
'token': 9562,
'token_str': 'curva'},
{'score': 0.02348282001912594,
'sequence': '[CLS] Tinha uma parada no meio do caminho. [SEP]',
'token': 6655,
'token_str': 'parada'},
{'score': 0.01795753836631775,
'sequence': '[CLS] Tinha uma mulher no meio do caminho. [SEP]',
'token': 2606,
'token_str': 'mulher'},
{'score': 0.015246033668518066,
'sequence': '[CLS] Tinha uma luz no meio do caminho. [SEP]',
'token': 3377,
'token_str': 'luz'}]
For BERT embeddings
import torch
model = AutoModel.from_pretrained('neuralmind/bert-large-portuguese-cased')
input_ids = tokenizer.encode('Tinha uma pedra no meio do caminho.', return_tensors='pt')
with torch.no_grad():
outs = model(input_ids)
encoded = outs[0][0, 1:-1] # Ignore [CLS] and [SEP] special tokens
encoded.shape: (8, 1024)
tensor([[ 1.1872, 0.5606, -0.2264, ..., 0.0117, -0.1618, -0.2286],
[ 1.3562, 0.1026, 0.1732, ..., -0.3855, -0.0832, -0.1052],
[ 0.2988, 0.2528, 0.4431, ..., 0.2684, -0.5584, 0.6524],
...,
[ 0.3405, -0.0140, -0.0748, ..., 0.6649, -0.8983, 0.5802],
[ 0.1011, 0.8782, 0.1545, ..., -0.1768, -0.8880, -0.1095],
[ 0.7912, 0.9637, -0.3859, ..., 0.2050, -0.1350, 0.0432]])
Citation
If you use our work, please cite:
@inproceedings{souza2020bertimbau,
author = {F{\'a}bio Souza and
Rodrigo Nogueira and
Roberto Lotufo},
title = {{BERT}imbau: pretrained {BERT} models for {B}razilian {P}ortuguese},
booktitle = {9th Brazilian Conference on Intelligent Systems, {BRACIS}, Rio Grande do Sul, Brazil, October 20-23 (to appear)},
year = {2020}
}