惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LINUX DO - 热门话题
C
Check Point Blog
Hugging Face - Blog
Hugging Face - Blog
N
News | PayPal Newsroom
Security Archives - TechRepublic
Security Archives - TechRepublic
N
News and Events Feed by Topic
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Troy Hunt's Blog
H
Heimdal Security Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
S
Secure Thoughts
博客园 - Franky
A
Arctic Wolf
Spread Privacy
Spread Privacy
P
Proofpoint News Feed
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
C
CXSECURITY Database RSS Feed - CXSecurity.com
博客园 - 【当耐特】
博客园 - 三生石上(FineUI控件)
T
Threat Research - Cisco Blogs
The Last Watchdog
The Last Watchdog
博客园_首页
Forbes - Security
Forbes - Security
Google DeepMind News
Google DeepMind News
Project Zero
Project Zero
T
Threatpost
Y
Y Combinator Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
Cisco Talos Blog
Cisco Talos Blog
小众软件
小众软件
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Schneier on Security
Schneier on Security
博客园 - 叶小钗
量子位
Security Latest
Security Latest
酷 壳 – CoolShell
酷 壳 – CoolShell
U
Unit 42
A
About on SuperTechFans
大猫的无限游戏
大猫的无限游戏
Simon Willison's Weblog
Simon Willison's Weblog
博客园 - 聂微东
Apple Machine Learning Research
Apple Machine Learning Research
Jina AI
Jina AI
P
Privacy International News Feed
Help Net Security
Help Net Security
博客园 - 司徒正美
D
DataBreaches.Net
MongoDB | Blog
MongoDB | Blog
C
CERT Recently Published Vulnerability Notes
T
The Exploit Database - CXSecurity.com

博客园 - 慕尘

在浏览器跑 Qwen2.5 使用 WSL 在 Windows 上安装 Linux LangExtract pgvector 向量数据库 Faiss Goose trafilatura unstructured python里使用Playwright MinGW nomic-embed-text 解析非结构化数据 LangChain 的 DocumentLoader 能够使用require但不能使用import ChromaDB nvm-windows 使用js实现文字转语音 pyttsx3 Ollama笔记
python的jieba
慕尘 · 2025-03-14 · via 博客园 - 慕尘

jieba 是一个广泛使用的 Python 中文分词库,主要用于将中文文本切分成独立的词语。

https://github.com/fxsjy/jieba

安装

使用

(1)分词

import jieba
# 分词
text = "我爱自然语言处理"
words = jieba.cut(text, cut_all=False)  # 精确模式
print("分词结果:", "/ ".join(words))

分词结果: 我/ 爱/ 自然语言/ 处理

(2)词性标注

import jieba.posseg as pseg
text = "我爱自然语言处理"
# 词性标注
words = pseg.cut(text)
for word, flag in words:
    print(f"{word} - {flag}")

我 - r
爱 - v
自然语言 - l
处理 - v

(3)关键词提取

基于 TF-IDF 算法的关键词抽取

import jieba.analyse
# 关键词提取
text = "我爱自然语言处理"
keywords = jieba.analyse.extract_tags(text, topK=3, withWeight=True, allowPOS=('l', 'v'))
print("关键词:", keywords)

关键词: [('自然语言', 5.2174708746), ('处理', 2.70542782868)]

关键词: ['自然语言', '处理']

基于 TF-IDF 算法的关键词抽取

import jieba.analyse
# 关键词提取
text = "我爱自然语言处理"
keywords = jieba.analyse.textrank(text, topK=3, withWeight=True, allowPOS=('l', 'v'))
print("关键词:", keywords)

关键词: [('自然语言', 1.0), ('处理', 0.9961264494011037)]