w2v bert 2.0

Providerfacebook
Categoryfeature-extraction
Licensemit
Downloads167.3K
Stars0

Overview

w2v bert 2.0 is a specialized feature-extraction model designed to bridge the gap between static word embeddings and contextualized representations. Unlike generative LLMs, this model focuses on producing high-dimensional vector embeddings that capture nuanced semantic relationships, making it ideal for downstream NLP tasks where latency and stability are critical. Developers can integrate it into pipelines for semantic search, clustering, or as a sophisticated input layer for custom classifiers. Compared to standard BERT, it is optimized for efficiency in extracting stable features without the overhead of full sequence generation, allowing for faster indexing and retrieval in production environments.

Highlights

  • Optimized for high-performance semantic feature extraction
  • Ideal for semantic search and document clustering
  • Low-latency alternative to generative transformer models
  • Permissive MIT license for flexible commercial deployment

Usage

Install
# Install Hugging Face transformers
pip install transformers torch
SDK Usage
# Load model with transformers
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("facebook/w2v-bert-2.0")
tokenizer = AutoTokenizer.from_pretrained("facebook/w2v-bert-2.0")

Hugging Face Download

We recommend downloading the model via the Hugging Face CLI or Hub SDK.

Guidance:Before downloading, install huggingface_hub with:

Guidance
pip install -U huggingface_hub

CLI Download

Download the full repository

Download the full repository
huggingface-cli download facebook/w2v-bert-2.0

Download a single file to a local folder (e.g. config.json into ./dir)

Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download facebook/w2v-bert-2.0 config.json --local-dir ./dir

See the official docs for more CLI options

SDK Download

SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('facebook/w2v-bert-2.0')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://huggingface.co/facebook/w2v-bert-2.0

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/facebook/w2v-bert-2.0

Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.

PyTorch / Transformers Usage

Install Transformers

Install Transformers
pip install -U transformers torch

Load the model and run inference

Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('facebook/w2v-bert-2.0')
tokenizer = AutoTokenizer.from_pretrained('facebook/w2v-bert-2.0')

Model Download

We recommend downloading the model via the ModelScope CLI or SDK.

Guidance:Before downloading, install ModelScope with:

Guidance
pip install modelscope

CLI Download

Download the full repository

Download the full repository
modelscope download --model facebook/w2v-bert-2.0

Download a single file to a local folder (e.g. README.md into ./dir)

Download a single file to a local folder (e.g. README.md into ./dir)
modelscope download --model facebook/w2v-bert-2.0 README.md --local_dir ./dir

See the docs for more CLI options

SDK Download

SDK Download
# 模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('facebook/w2v-bert-2.0')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://www.modelscope.cn/facebook/w2v-bert-2.0.git

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/facebook/w2v-bert-2.0.git

ModelScope 模型页直接下载模型文件;无需将模型文件放在本站服务器。

Notebook Quickstart

Install the ModelScope library

Install the ModelScope library
pip install "modelscope[audio,cv,nlp,multi-modal,science]" -f https://modelscope.oss-cn-beijing.aliyuncs.com/releases/repo.html

Load the model and run inference

Load the model and run inference
from modelscope.pipelines import pipeline
from modelscope.utils.constant import Tasks

p = pipeline('text-generation', 'facebook/w2v-bert-2.0')

Full Documentation

来源: HuggingFace

---
license: mit
language:

  • af

  • am

  • ar

  • as

  • az

  • be

  • bn

  • bs

  • bg

  • ca

  • cs

  • zh

  • cy

  • da

  • de

  • el

  • en

  • et

  • fi

  • fr

  • or

  • om

  • ga

  • gl

  • gu

  • ha

  • he

  • hi

  • hr

  • hu

  • hy

  • ig

  • id

  • is

  • it

  • jv

  • ja

  • kn

  • ka

  • kk

  • mn

  • km

  • ky

  • ko

  • lo

  • ln

  • lt

  • lb

  • lg

  • lv

  • ml

  • mr

  • mk

  • mt

  • mi

  • my

  • nl

  • nb

  • ne

  • ny

  • oc

  • pa

  • ps

  • fa

  • pl

  • pt

  • ro

  • ru

  • sk

  • sl

  • sn

  • sd

  • so

  • es

  • sr

  • sv

  • sw

  • ta

  • te

  • tg

  • tl

  • th

  • tr

  • uk

  • ur

  • uz

  • vi

  • wo

  • xh

  • yo

  • ms

  • zu

  • ary

  • arz

  • yue

  • kea

inference: false
---

W2v-BERT 2.0 speech encoder

We are open-sourcing our Conformer-based W2v-BERT 2.0 speech encoder as described in Section 3.2.1 of the paper, which is at the core of our Seamless models.

This model was pre-trained on 4.5M hours of unlabeled audio data covering more than 143 languages. It requires finetuning to be used for downstream tasks such as Automatic Speech Recognition (ASR), or Audio Classification.

| Model Name | #params | checkpoint |
| ----------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| W2v-BERT 2.0 | 600M | checkpoint

This model and its training are supported by 🤗 Transformers, more on it in the docs.

🤗 Transformers usage

This is a bare checkpoint without any modeling head, and thus requires finetuning to be used for downstream tasks such as ASR. You can however use it to extract audio embeddings from the top layer with this code snippet:

python
from transformers import AutoFeatureExtractor, Wav2Vec2BertModel
import torch
from datasets import load_dataset

dataset = load_dataset("hf-internal-testing/librispeech_asr_demo", "clean", split="validation")
dataset = dataset.sort("id")
sampling_rate = dataset.features["audio"].sampling_rate

processor = AutoProcessor.from_pretrained("facebook/w2v-bert-2.0")
model = Wav2Vec2BertModel.from_pretrained("facebook/w2v-bert-2.0")

audio file is decoded on the fly

inputs = processor(dataset[0]["audio"]["array"], sampling_rate=sampling_rate, return_tensors="pt") with torch.no_grad(): outputs = model(inputs)

To learn more about the model use, refer to the following resources:



Seamless Communication usage

This model can be used in Seamless Communication, where it was released.

Here's how to make a forward pass through the voice encoder, after having completed the installation steps:

python
import torch

from fairseq2.data.audio import AudioDecoder, WaveformToFbankConverter
from fairseq2.memory import MemoryBlock
from fairseq2.nn.padding import get_seqs_and_padding_mask
from pathlib import Path
from seamless_communication.models.conformer_shaw import load_conformer_shaw_model

audio_wav_path, device, dtype = ...
audio_decoder = AudioDecoder(dtype=torch.float32, device=device)
fbank_converter = WaveformToFbankConverter(
num_mel_bins=80,
waveform_scale=2
15,
channel_last=True,
standardize=True,
device=device,
dtype=dtype,
)
collater = Collater(pad_value=1)

model = load_conformer_shaw_model("conformer_shaw", device=device, dtype=dtype)
model.eval()

with Path(audio_wav_path).open("rb") as fb:
block = MemoryBlock(fb.read())

decoded_audio = audio_decoder(block)
src = collater(fbank_converter(decoded_audio))["fbank"]
seqs, padding_mask = get_seqs_and_padding_mask(src)

with torch.inference_mode():
seqs, padding_mask = model.encoder_frontend(seqs, padding_mask)
seqs, padding_mask = model.encoder(seqs, padding_mask)

Join our Telegram