當前位置：首頁 > 编程资源 > 编程问答 >内容正文

编程问答

深入Bert实战(Pytorch)----WordPiece Embeddings

發布時間：2023/12/16 编程问答 37 豆豆

生活随笔收集整理的這篇文章主要介紹了深入Bert实战(Pytorch)----WordPiece Embeddings 小編覺得挺不錯的,現在分享給大家,幫大家做個參考.

https://www.bilibili.com/video/BV1K5411t7MD?p=5
https://www.youtube.com/channel/UCoRX98PLOsaN8PtekB9kWrw/videos
深入BERT實戰(PyTorch) by ChrisMcCormickAI
這是ChrisMcCormickAI在油管bert，8集系列第二篇WordPiece Embeddings的pytorch的講解的代碼，在油管視頻下有下載地址，如果不能翻墻的可以留下郵箱我全部看完整理后發給你。

文章目錄

- 加載模型
- 查看Bert中的詞匯
- 單字符
- Subwords vs. Whole-words
- 開始子詞和中間子詞
- 對于姓名來說
- 對于數字來說

加載模型

安裝huggingface實現

!pip install pytorch-pretrained-bertimport torch from pytorch_pretrained_bert import BertTokenizer# 加載預訓練模型 tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')

查看Bert中的詞匯

檢索整個“tokens”列表，并將它們寫入文本文件，可以仔細閱讀它們。

with open("vocabulary.txt", 'w', encoding='utf-8') as f:# For each token... 得到每個單詞for token in tokenizer.vocab.keys():# Write it out and escape any unicode characters. # 寫并且轉義為Unicode字符 f.write(token + '\n')

打印出來可以得到
前999個詞序是保留位置，它們的形式類似于[unused957]
1 - [PAD] 截斷
101 - [UNK] 未知字符
102 - [CLS] 句首，表示分類任務
103 - [SEP] BERT中分隔兩個輸入的句子
104 - [MASK] MASK機制
第1000-1996行似乎是單個字符的轉儲。
它們似乎沒有按頻率排序(例如，字母表中的字母都是按順序排列的)。
第一個單詞是1997位置的The
從這里開始，這些詞似乎是按頻率排序的。
前18個單詞是完整的單詞，第2016位是##s，可能是最常見的子單詞。
最后一個完整單詞是29612位的 “necessitated”

單字符

下面的代碼打印出詞匯表中的所有單字符的token，以及前面帶’##'的所有單字符的token。

結果發現這些都是匹配集——每個獨立角色都有一個“##”版本。有997個單字符標記。

下面的單元格遍歷詞匯表，取出所有單個字符標記。

one_chars = [] one_chars_hashes = []# For each token in the vocabulary... 遍歷所有單字符 for token in tokenizer.vocab.keys():# Record any single-character tokens.記錄下來if len(token) == 1:one_chars.append(token)# Record single-character tokens preceded by the two hashes. # 記錄##單字符 elif len(token) == 3 and token[0:2] == '##':one_chars_hashes.append(token)

打印單字符的

print('Number of single character tokens:', len(one_chars), '\n')# Print all of the single characters, 40 per row.# For every batch of 40 tokens... for i in range(0, len(one_chars), 40):# Limit the end index so we don't go past the end of the list.end = min(i + 40, len(one_chars) + 1)# Print out the tokens, separated by a space.print(' '.join(one_chars[i:end]))

打印##單字符的

print('Number of single character tokens with hashes:', len(one_chars_hashes), '\n')# Print all of the single characters, 40 per row.按每行40打印# Strip the hash marks, since they just clutter the display.去除## tokens = [token.replace('##', '') for token in one_chars_hashes]# For every batch of 40 tokens...每批40 for i in range(0, len(tokens), 40):# Limit the end index so we don't go past the end of the list.限制結束位置end = min(i + 40, len(tokens) + 1)# Print out the tokens, separated by a space.print(' '.join(tokens[i:end])) Number of single character tokens: 997 ! " # $ % & ' ( ) * + , - . / 0 1 2 3 4 5 6 7 8 9 : ; < = > ? @ [ \ ] ^ _ ` a b c d e f g h i j k l m n o p q r s t u v w x y z { | } ~ ? ￠￡ ¤ ￥ | § ¨ ? a ? ? ? ° ± 2 3 ′ μ ? · 1 o ? ? ? ? ? × ? ? e ÷ ? t ? ? ? ? ? ? ? ? ɑ ? ? ? ? ? ɡ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? α β γ δ ε ζ η θ ι κ λ μ ν ξ ο π ρ ? σ τ υ φ χ ψ ω а б в г д е ж з и к л м н о п р с т у ф х ц ч ш щ ъ ы ь э ю я ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ‐ ? ? – — ― ‖ ‘ ’ ? “ ” ? ? ? ? … ‰ ′ ″ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? € ? ? ? № ? ? ? ? ← ↑ → ↓ ? ? ? ? ? ? ? ? ? ∈ ? ? ° √ ∞ ∧ ∨ ∩ ∪ ≈ ≡ ≤ ≥ ? ? ⊕ ? ? ─ │ ■ ? ● ★ ☆ ☉ ? ? ? ? ? ? ? ? ? ? ? ? 、。〈〉《》「」『』 ? あいうえおかきくけこさしすせそたちっつてとなにぬねのはひふへほまみむめもやゆよらりるれろをんァアィイウェエオカキクケコサシスセタチッツテトナニノハヒフヘホマミムメモャュョラリルレロワン ? ー一三上下不世中主久之也事二五井京人亻仁介代仮伊會佐侍保信健元光八公內出分前劉力加勝北區十千南博原口古史司合吉同名和囗四國國土地坂城堂場士夏外大天太夫奈女子學宀宇安宗定宣宮家宿寺將小尚山岡島崎川州巿帝平年幸廣弘張彳後御德心忄志忠愛成我戦戸手扌政文新方日明星春昭智曲書月有朝木本李村東松林森楊樹橋歌止正武比氏民水氵氷永江沢河治法海清漢瀬火版犬王生田男疒発白的皇目相省真石示社神福禾秀秋空立章竹糹美義耳良艸花英華葉藤行街西見訁語谷貝貴車軍辶道郎郡部都里野金鈴鎮長門間阝阿陳陽雄青面風食香馬高龍 ? ? ? ！（），－．／：？～

上面兩段代碼及 ##+單字符和單字符結果一樣

# return True print('Are the two sets identical?', set(one_chars) == set(tokens))

Subwords vs. Whole-words

打印一些詞匯的統計數據。

import matplotlib.pyplot as plt import seaborn as sns import numpy as npsns.set(style='darkgrid')# Increase the plot size and font size. sns.set(font_scale=1.5) plt.rcParams["figure.figsize"] = (10,5)# Measure the length of every token in the vocab. 加載每個單詞 token_lengths = [len(token) for token in tokenizer.vocab.keys()]# Plot the number of tokens of each length. sns.countplot(token_lengths) plt.title('Vocab Token Lengths') plt.xlabel('Token Length') plt.ylabel('# of Tokens')print('Maximum token length:', max(token_lengths))

統計一下，’##'開頭的tokens。

num_subwords = 0subword_lengths = []# For each token in the vocabulary... for token in tokenizer.vocab.keys():# If it's a subword...if len(token) >= 2 and token[0:2] == '##':# Tally all subwordsnum_subwords += 1# Measure the sub word length (without the hashes)length = len(token) - 2# Record the lengths. subword_lengths.append(length)

相對于完整詞匯表占據的數量

vocab_size = len(tokenizer.vocab.keys())print('Number of subwords: {:,} of {:,}'.format(num_subwords, vocab_size))# Calculate the percentage of words that are '##' subwords. prcnt = float(num_subwords) / vocab_size * 100.0print('%.1f%%' % prcnt) Number of subwords: 5,828 of 30,522 19.1%

作圖統計的結果

sns.countplot(subword_lengths) plt.title('Subword Token Lengths (w/o "##")') plt.xlabel('Subword Length') plt.ylabel('# of ## Subwords')

可以自行查看下錯誤拼寫的例子

'misspelled' in tokenizer.vocab # Right 'mispelled' in tokenizer.vocab # Wrong 'government' in tokenizer.vocab # Right 'goverment' in tokenizer.vocab # Wrong 'beginning' in tokenizer.vocab # Right 'begining' in tokenizer.vocab # Wrong 'separate' in tokenizer.vocab # Right 'seperate' in tokenizer.vocab # Wrong

對于縮寫來說

"can't" in tokenizer.vocab # False "cant" in tokenizer.vocab # False

開始子詞和中間子詞

對于單個字符，既有單個字符，也有對應每個字符的“##”版本。子詞也是如此嗎?

# For each token in the vocabulary... for token in tokenizer.vocab.keys():# If it's a subword...if len(token) >= 2 and token[0:2] == '##':if not token[2:] in tokenizer.vocab:print('Did not find a token for', token[2:])break

可以查看到第一個返回的##ly在詞表中，但是ly不在詞表中

Did not find a token for ly'##ly' in tokenizer.vocab # True 'ly' in tokenizer.vocab # False

對于姓名來說

下載數據

!pip install wgetimport wget import random print('Beginning file download with wget module')url = 'http://www.gutenberg.org/files/3201/files/NAMES.TXT' wget.download(url, 'first-names.txt')

編碼，小寫化，輸出長度

# Read them in. with open('first-names.txt', 'rb') as f:names_encoded = f.readlines()names = []# Decode the names, convert to lowercase, and strip newlines. for name in names_encoded:try:names.append(name.rstrip().lower().decode('utf-8'))except:continueprint('Number of names: {:,}'.format(len(names))) print('Example:', random.choice(names))

查看有多少個姓名是在BERT的詞表中

num_names = 0# For each name in our list... for name in names:# If it's in the vocab...if name in tokenizer.vocab:# Tally it.num_names += 1print('{:,} names in the vocabulary'.format(num_names))

對于數字來說

# Count how many numbers are in the vocabulary. 統計詞匯表中有多少數字 count = 0# For each token in the vocabulary... for token in tokenizer.vocab.keys():# Tally if it's a number.if token.isdigit():count += 1# Any numbers >= 10,000?if len(token) > 4:print(token)print('Vocab includes {:,} numbers.'.format(count))

計算一下在1600-2021中有幾個數字在

# Count how many dates between 1600 and 2021 are included. count = 0 for i in range(1600, 2021):if str(i) in tokenizer.vocab:count += 1print('Vocab includes {:,} of 421 dates from 1600 - 2021'.format(count))

總結

以上是生活随笔為你收集整理的深入Bert实战(Pytorch)----WordPiece Embeddings的全部內容，希望文章能夠幫你解決所遇到的問題。

如果覺得生活随笔網站內容還不錯，歡迎將生活随笔推薦給好友。

上一篇： this指针详解
下一篇： java 调用 delphi_【java