惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cisco Talos Blog
Cisco Talos Blog
Cyberwarzone
Cyberwarzone
T
Tenable Blog
Security Latest
Security Latest
NISL@THU
NISL@THU
V
Vulnerabilities – Threatpost
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
W
WeLiveSecurity
罗磊的独立博客
Stack Overflow Blog
Stack Overflow Blog
云风的 BLOG
云风的 BLOG
Martin Fowler
Martin Fowler
Engineering at Meta
Engineering at Meta
T
Tor Project blog
H
Heimdal Security Blog
Microsoft Security Blog
Microsoft Security Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
雷峰网
雷峰网
L
LINUX DO - 热门话题
The GitHub Blog
The GitHub Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Recorded Future
Recorded Future
Hugging Face - Blog
Hugging Face - Blog
P
Privacy & Cybersecurity Law Blog
F
Full Disclosure
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
PCI Perspectives
PCI Perspectives
MyScale Blog
MyScale Blog
B
Blog RSS Feed
www.infosecurity-magazine.com
www.infosecurity-magazine.com
K
Kaspersky official blog
Attack and Defense Labs
Attack and Defense Labs
H
Hackread – Cybersecurity News, Data Breaches, AI and More
有赞技术团队
有赞技术团队
Know Your Adversary
Know Your Adversary
Hacker News - Newest:
Hacker News - Newest: "LLM"
Scott Helme
Scott Helme
The Last Watchdog
The Last Watchdog
博客园 - 【当耐特】
S
Security Affairs
The Cloudflare Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
人人都是产品经理
人人都是产品经理
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
N
News and Events Feed by Topic
AI
AI
H
Help Net Security
美团技术团队
T
Threatpost
Project Zero
Project Zero

博客园 - 慕尘

在浏览器跑 Qwen2.5 使用 WSL 在 Windows 上安装 Linux LangExtract pgvector 向量数据库 Faiss Goose unstructured python里使用Playwright python的jieba MinGW nomic-embed-text 解析非结构化数据 LangChain 的 DocumentLoader 能够使用require但不能使用import ChromaDB nvm-windows 使用js实现文字转语音 pyttsx3 Ollama笔记
trafilatura
慕尘 · 2025-03-19 · via 博客园 - 慕尘

trafilatura是一个专为从网页中提取核心内容设计的Python库

特别适用于那些需要从HTML页面中提取主要文本信息的应用场景,比如文章正文、标题等,同时排除掉导航栏、广告、侧边栏和其他非主要内容

安装

示例

import trafilatura

# 指定网页 URL
url = "https://www.cnblogs.com/baby123/p/18755330"
# 下载网页内容
downloaded = trafilatura.fetch_url(url)
# 提取核心文本内容
result = trafilatura.extract(downloaded)
print(result)

对于一些动态加载内容的网站,可能需要先使用Playwright 或 Selenium 工具来获取完整的HTML内容,然后再使用 Trafilatura 进行内容提取

但是这样速度会变慢

import asyncio
from playwright.async_api import async_playwright
import trafilatura
import time

async def fetch_dynamic_content(browser, url):
    page = await browser.new_page()
    try:
        await page.goto(url)
        # 使用 'networkidle' 等待页面加载完成
        await page.wait_for_load_state('networkidle')
        html_content = await page.content()
        return html_content
    finally:
        await page.close()

def extract_core_content(html_content):
    # 使用 Trafilatura 提取核心内容
    result = trafilatura.extract(html_content)
    return result

async def main():
    start_count = time.perf_counter()
    
    urls = ["http://jinan.tianqi.com/"]  # 可以添加更多URL以测试并行处理
    
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        tasks = [fetch_dynamic_content(browser, url) for url in urls]
        
        results = await asyncio.gather(*tasks)
        
        core_contents = []
        for html_content in results:
            core_content = extract_core_content(html_content)
            core_contents.append(core_content)
            print(core_content)
        
        await browser.close()
    
    end_count = time.perf_counter()
    elapsed_time = round(end_count - start_count, 2)
    print(f"本次查找时间:{elapsed_time} 秒")

# 运行主函数
if __name__ == "__main__":
    asyncio.run(main())