惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
CERT Recently Published Vulnerability Notes
U
Unit 42
Apple Machine Learning Research
Apple Machine Learning Research
爱范儿
爱范儿
Cisco Talos Blog
Cisco Talos Blog
P
Proofpoint News Feed
H
Heimdal Security Blog
Help Net Security
Help Net Security
H
Help Net Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
P
Palo Alto Networks Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
S
Secure Thoughts
The GitHub Blog
The GitHub Blog
博客园_首页
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Microsoft Azure Blog
Microsoft Azure Blog
Hacker News: Ask HN
Hacker News: Ask HN
博客园 - 【当耐特】
J
Java Code Geeks
S
SegmentFault 最新的问题
Application and Cybersecurity Blog
Application and Cybersecurity Blog
P
Proofpoint News Feed
The Last Watchdog
The Last Watchdog
O
OpenAI News
博客园 - 三生石上(FineUI控件)
Recent Announcements
Recent Announcements
B
Blog RSS Feed
V2EX - 技术
V2EX - 技术
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
T
Tenable Blog
PCI Perspectives
PCI Perspectives
C
CXSECURITY Database RSS Feed - CXSecurity.com
The Hacker News
The Hacker News
Schneier on Security
Schneier on Security
Google Online Security Blog
Google Online Security Blog
美团技术团队
G
GRAHAM CLULEY
D
DataBreaches.Net
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 聂微东
W
WeLiveSecurity
Vercel News
Vercel News
S
Security Affairs
T
Tailwind CSS Blog
V
Vulnerabilities – Threatpost
博客园 - 司徒正美
G
Google Developers Blog
D
Docker
Webroot Blog
Webroot Blog

博客园 - 慕尘

在浏览器跑 Qwen2.5 使用 WSL 在 Windows 上安装 Linux LangExtract pgvector 向量数据库 Faiss Goose unstructured python里使用Playwright python的jieba MinGW nomic-embed-text 解析非结构化数据 LangChain 的 DocumentLoader 能够使用require但不能使用import ChromaDB nvm-windows 使用js实现文字转语音 pyttsx3 Ollama笔记
trafilatura
慕尘 · 2025-03-19 · via 博客园 - 慕尘

trafilatura是一个专为从网页中提取核心内容设计的Python库

特别适用于那些需要从HTML页面中提取主要文本信息的应用场景,比如文章正文、标题等,同时排除掉导航栏、广告、侧边栏和其他非主要内容

安装

示例

import trafilatura

# 指定网页 URL
url = "https://www.cnblogs.com/baby123/p/18755330"
# 下载网页内容
downloaded = trafilatura.fetch_url(url)
# 提取核心文本内容
result = trafilatura.extract(downloaded)
print(result)

对于一些动态加载内容的网站,可能需要先使用Playwright 或 Selenium 工具来获取完整的HTML内容,然后再使用 Trafilatura 进行内容提取

但是这样速度会变慢

import asyncio
from playwright.async_api import async_playwright
import trafilatura
import time

async def fetch_dynamic_content(browser, url):
    page = await browser.new_page()
    try:
        await page.goto(url)
        # 使用 'networkidle' 等待页面加载完成
        await page.wait_for_load_state('networkidle')
        html_content = await page.content()
        return html_content
    finally:
        await page.close()

def extract_core_content(html_content):
    # 使用 Trafilatura 提取核心内容
    result = trafilatura.extract(html_content)
    return result

async def main():
    start_count = time.perf_counter()
    
    urls = ["http://jinan.tianqi.com/"]  # 可以添加更多URL以测试并行处理
    
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        tasks = [fetch_dynamic_content(browser, url) for url in urls]
        
        results = await asyncio.gather(*tasks)
        
        core_contents = []
        for html_content in results:
            core_content = extract_core_content(html_content)
            core_contents.append(core_content)
            print(core_content)
        
        await browser.close()
    
    end_count = time.perf_counter()
    elapsed_time = round(end_count - start_count, 2)
    print(f"本次查找时间:{elapsed_time} 秒")

# 运行主函数
if __name__ == "__main__":
    asyncio.run(main())