Find the words找出中文里的词
Chinese is written without spaces. Type a sentence and watch two hand-written segmenters, and the one built into your browser, decide where each word begins and ends.
中文书写时词与词之间没有空格。输入一句话,看两个手写的分词器,以及你的浏览器内置的分词器,如何判断每个词从哪里开始、到哪里结束。
How it works原理
Forward maximum matching reads from the left and always takes the longest word it knows, so in 结婚的和尚未结婚的 it takes 和尚 (monk) and misreads the rest. The graph above shows every word of the list that could start at each character; dynamic programming picks the path whose words are most likely together, the method of the jieba segmenter. The word list is small, with made-up frequencies, so the hand-written segmenters know only the words of the examples and a few more. Your browser's Intl.Segmenter uses the dictionary of its ICU library, which knows far more words but cannot be taught new ones.
正向最大匹配从左往右读,每次都取它认识的最长的词,所以在“结婚的和尚未结婚的”里,它取了“和尚”,后面就全读错了。上面的图画出了从每个字开始、词表里所有可能的词;动态规划从中选出一条路径,让路径上的词连在一起的可能性最大,这正是 jieba 分词器的方法。词表很小,词频是虚构的,所以手写的分词器只认识例句里的词和另外几个词。你的浏览器里的 Intl.Segmenter 用的是它的 ICU 库自带的词典,认识的词多得多,但没法教它新词。