Python利用jieba获取中文词汇等

import jieba
import os
import jieba.analyse

data = cleaned_comments # 数据来源于评论数据
seg = jieba.lcut(data)
print(seg)

# 增加自定义词表库
mydict = os.getcwd()+"/mydict.txt"
jieba.load_userdict(mydict)
seg = jieba.lcut(data)
print(seg)

import jieba.posseg as pseg
posseg = pseg.lcut(data)
print(posseg)

# 抽取出现次数最多的词汇
extracttext = jieba.analyse.extract_tags(data, topK=20,withWeight=False, allowPOS=())
print(extracttext)

 

待续。。。

posted @ 2017-07-19 23:51  宝山方圆  阅读(1607)  评论(0)    收藏  举报