indonesian roberta base posp tagger

提供商w11wo
分类token-classification
许可证mit
下载量2.7M
星标0

简介

这是一个基于 RoBERTa-base 架构的印尼语词性标注(POS Tagging)模型。对于需要处理印尼语自然语言处理(NLP)任务的开发者来说,它是构建语法分析、命名实体识别或文本挖掘等上游任务的基础工具。该模型专注于 Token 分类,能够将文本中的每个词准确映射到相应的语法标签。上手难度较低,可直接集成到 Hugging Face 管道中,适合作为印尼语语料预处理的标准化步骤,填补了通用大模型在特定小语种精细化语法标注上的能力空白。

核心亮点

  • 专注于印尼语词性标注,提供精准的语法标签
  • 基于 RoBERTa 架构,语义理解能力强且稳定
  • MIT 协议开源,商业集成与二次开发无压力
  • 典型的 Token 分类模型,适配多种 NLP 预处理流程

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("w11wo/indonesian-roberta-base-posp-tagger")
tokenizer = AutoTokenizer.from_pretrained("w11wo/indonesian-roberta-base-posp-tagger")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download w11wo/indonesian-roberta-base-posp-tagger

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download w11wo/indonesian-roberta-base-posp-tagger config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('w11wo/indonesian-roberta-base-posp-tagger')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/w11wo/indonesian-roberta-base-posp-tagger

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/w11wo/indonesian-roberta-base-posp-tagger

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('w11wo/indonesian-roberta-base-posp-tagger')
tokenizer = AutoTokenizer.from_pretrained('w11wo/indonesian-roberta-base-posp-tagger')

完整文档

来源: HuggingFace

---
license: mit
base_model: flax-community/indonesian-roberta-base
tags:

  • generated_from_trainer

datasets:
  • indonlu

language:
  • ind

metrics:
  • precision

  • recall

  • f1

  • accuracy

model-index:
  • name: indonesian-roberta-base-posp-tagger

results:
- task:
name: Token Classification
type: token-classification
dataset:
name: indonlu
type: indonlu
config: posp
split: test
args: posp
metrics:
- name: Precision
type: precision
value: 0.9625100240577386
- name: Recall
type: recall
value: 0.9625100240577386
- name: F1
type: f1
value: 0.9625100240577386
- name: Accuracy
type: accuracy
value: 0.9625100240577386
---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You
should probably proofread and complete it, then remove this comment. -->

indonesian-roberta-base-posp-tagger

This model is a fine-tuned version of flax-community/indonesian-roberta-base on the indonlu dataset.
It achieves the following results on the evaluation set:

  • Loss: 0.1395

  • Precision: 0.9625

  • Recall: 0.9625

  • F1: 0.9625

  • Accuracy: 0.9625

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05

  • train_batch_size: 16

  • eval_batch_size: 16

  • seed: 42

  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08

  • lr_scheduler_type: linear

  • num_epochs: 10

Training results

| Training Loss | Epoch | Step | Validation Loss | Precision | Recall | F1 | Accuracy |
|:-------------:|:-----:|:----:|:---------------:|:---------:|:------:|:------:|:--------:|
| No log | 1.0 | 420 | 0.2254 | 0.9313 | 0.9313 | 0.9313 | 0.9313 |
| 0.4398 | 2.0 | 840 | 0.1617 | 0.9499 | 0.9499 | 0.9499 | 0.9499 |
| 0.1566 | 3.0 | 1260 | 0.1431 | 0.9569 | 0.9569 | 0.9569 | 0.9569 |
| 0.103 | 4.0 | 1680 | 0.1412 | 0.9605 | 0.9605 | 0.9605 | 0.9605 |
| 0.0723 | 5.0 | 2100 | 0.1408 | 0.9635 | 0.9635 | 0.9635 | 0.9635 |
| 0.051 | 6.0 | 2520 | 0.1408 | 0.9642 | 0.9642 | 0.9642 | 0.9642 |
| 0.051 | 7.0 | 2940 | 0.1510 | 0.9635 | 0.9635 | 0.9635 | 0.9635 |
| 0.0368 | 8.0 | 3360 | 0.1653 | 0.9645 | 0.9645 | 0.9645 | 0.9645 |
| 0.0277 | 9.0 | 3780 | 0.1664 | 0.9644 | 0.9644 | 0.9644 | 0.9644 |
| 0.0231 | 10.0 | 4200 | 0.1668 | 0.9646 | 0.9646 | 0.9646 | 0.9646 |

Framework versions

  • Transformers 4.37.2
  • Pytorch 2.2.0+cu118
  • Datasets 2.16.1
  • Tokenizers 0.15.1