惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Recent Commits to openclaw:main
Recent Commits to openclaw:main
SecWiki News
SecWiki News
Webroot Blog
Webroot Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
N
News and Events Feed by Topic
Recent Announcements
Recent Announcements
Help Net Security
Help Net Security
Jina AI
Jina AI
O
OpenAI News
雷峰网
雷峰网
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
博客园 - 三生石上(FineUI控件)
W
WeLiveSecurity
Schneier on Security
Schneier on Security
T
Threat Research - Cisco Blogs
IT之家
IT之家
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Vercel News
Vercel News
N
News and Events Feed by Topic
T
The Exploit Database - CXSecurity.com
爱范儿
爱范儿
Recorded Future
Recorded Future
Google Online Security Blog
Google Online Security Blog
TaoSecurity Blog
TaoSecurity Blog
美团技术团队
Engineering at Meta
Engineering at Meta
Security Latest
Security Latest
V
V2EX
T
Tailwind CSS Blog
P
Privacy & Cybersecurity Law Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
S
Schneier on Security
B
Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
博客园 - 【当耐特】
PCI Perspectives
PCI Perspectives
GbyAI
GbyAI
I
Intezer
Spread Privacy
Spread Privacy
Security Archives - TechRepublic
Security Archives - TechRepublic
Cloudbric
Cloudbric
V
Visual Studio Blog
MongoDB | Blog
MongoDB | Blog
Forbes - Security
Forbes - Security
The Last Watchdog
The Last Watchdog
aimingoo的专栏
aimingoo的专栏
C
CERT Recently Published Vulnerability Notes
A
About on SuperTechFans
罗磊的独立博客

Analytics Vidhya

Handling Imbalanced Classification: What Works Better Than SMOTE GPT-5.6 Is Here: Sol, Terra, and Luna Loop Engineering for AI Agents: How /loop is Changing AI Workflows DeepSeek DSpark: The Speculative Decoding Trick Behind 400% Faster LLM OKF: Redefining Knowledge Bases for AI Agents Modern VLMs Explained: How GPT-4o, Gemini, Claude Vision, and Qwen-VL Work YOLO26 Tutorial: Object Detection, Pose Estimation & More Large Action Models (LAMs) vs Agentic LLMs: What's the Real Difference? Claude Sonnet 5: The Fable 5 at Home The Best $20 AI Plan: ChatGPT Plus vs Claude Pro vs Gemini Pro GraphRAG vs Vector RAG: Which Retrieval Method is Best? Using AI When You Don’t Trust AI The Self-Improving Loop in AI Agents: Architecture, Benefits, and How it Outperforms Traditional Agent Workflows Harness-1: The 20B Retrieval Subagent That Beats GPT-5.4 at Search Sakana Fugu: Multi-Agent System as a Model Claude's Hidden Art Skill: Making Illustrations With Code System Design for ML Interviews: 10 Real Problems Walked Through Most People Use ChatGPT Wrong: 10 Features and Tips That Changed How I Work OpenAI Just Launched 3 Free AI Courses with Certificates Autoregressive Models: Predicting the Future Using the Past Gemini Omni: AI Video Generation Inside Gemini DiffusionGemma: Google’s Diffusion-Based Open Model for Faster Text Generation Top 10 AI Engineering Tools Everyone is Using in 2026 I Tested Claude Fable 5: Can Anthropic’s Newest AI Deliver on the Hype? Prophet vs NeuralProphet vs TimeGPT vs Chronos: A Practical Comparison Build an Emergency Helpline Voice Agent with LangChain Choosing the Right Vector Database for RAG and AI Applications Google Gemma 4 12B: Architecture, Benchmarks, Access, and Hands-on Guide for Developers How to Choose the Right AI Model for Your Needs Agent Observability with LangSmith, Langfuse, and Arize: A Hands-On Comparison How to Use Claude Managed Agents? Google AI Studio vs Gemini App: What’s the Difference? AI Workflows for Sales Teams: Prospect Research, Lead Qualification, and CRM Updates on Autopilot Using LangGraph 25 Most Influential AI Pioneers to Meet at DataHack Summit 2026 Claude Opus 4.8: A Smarter Model in the Right Direction PySpark Optimization: 12 Proven Techniques to Speed Up Your Spark Jobs 10 Everyday Tasks You Can Automate with AI Today (With n8n Templates) Google Antigravity 2.0: The Full Developer Guide (I/O 2026) Build a Claude Cowork-Like Browser Agent Using Playwright MCP and Claude Desktop Pandas vs Polars vs DuckDB: Which Library Should You Choose? Qwen3.7-Max: Alibaba’s New Agent-First LLM for Coding, Reasoning, and Long-Horizon AI Workflows The Biggest Announcements from Google I/O 2026 Top 9 AI Events and Conferences in 2026 that you Must Attend Gemini 3.5 Flash: Frontier Intelligence with Speed Kimi WebBridge: Hands-on Guide to Kimi’s Browser Extension for AI Agents 40 Advanced SQL Window Functions Every Data Scientist Must Know(with examples) Top 10 AI Research Papers of 2025 6 Steps to Crack GenAI Case Study Interviews (With Real Examples) OpenAI Omni Moderation: How to Filter Text & Images for Free DataHack Summit 2026: You Just Cannot Skip This AI Event of the Year OpenAI’s New API Voice Models Will Change the Way You Use AI Hermes Agent Guide: What is it and How to Use it? Top 10 LLM Research Papers of 2026 Agent Memory Patterns in Cognitive Science and AI Systems 10 AI Agents Every AI Engineer Must Build (with GitHub Samples) 23 Tips for Smart Claude Code Token Saving and Workflow Optimization Feature Engineering with LLMs: Techniques & Python Examples ChatGPT is Now Inside Excel and Google Sheets: Here is How to Use it Gemini API File Search: The Easy Way to Build RAG Top 10 Open-Source Libraries to Fine-Tune LLMs Locally ML Intern in Practice: From Prompt to a Shipped Hugging Face Model 15+ Solved Agentic AI Projects with Github Links How People are Figuring Out Life With Claude MemPalace Explained: Building Long-Term Memory for AI Agents Beyond RAG Grok Voice Think Fast 1.0: Build Voice AI Agents That Actually Think Compressing LSTM Models for Retail Edge Deployment: A Practical Comparison MCP vs Agent Skills: Different Altogether GPT 5.5 vs Opus 4.7: Which is the Best AI Model Today? What is Agentic AI? Claude Code vs Codex: A Detailed Terminal Agent Comparison Google Deep Research Max: Build Autonomous AI Research Agents in Minutes Meta Muse Spark Review: Is It Worth the Hype? ChatGPT Images 2.0 vs Nano Banana 2: Which is Better? Cursor V3 Explained: The AI Coding Agent That’s Replacing Traditional IDEs in 2026 DeepSeek-V4: The Most Powerful Open-Source Model Ever Is GPT Image 2 the Best Image Generation Model? Token Economics: Why AI is Getting “Cheaper” From Idea to Output: Claude Does the Design Work Opus 4.7 vs Opus 4.6: Should You Switch? Build Human-Like AI Voice App with Gemini 3.1 Flash TTS How to Structure a Claude Code Project that Thinks Like an Engineer Gemma 4 Tool Calling Explained: Build AI Agents with Function Calling (Step-by-Step Guide) Anthropic Launches Claude Opus 4.7 For “Most Difficult Tasks” Top 28 Claude Shortcuts that will 10X your Speed GPT-5.4-Cyber: Why OpenAI is Keeping its Most Powerful Model Under Lock and Key Google AI Studio Guide: Every Feature Explained Mastering Deep Agents: Context Engineering that Actually Works 21 Computer Vision Projects from Beginner to Advanced (2026 Guide) Excel 101: Excel Agent Mode Explained MiniMax M2.7 Goes Open-Weight to Let You Run Agents Locally Top 10 Gemma 4 Projects That Will Blow Your Mind GLM-5.1: Architecture, Benchmarks, Capabilities & How to Use It From Karpathy’s LLM Wiki to Graphify: AI Memory Layers are Here 10 Most Important AI Concepts Explained Simply Project Glasswing is World’s Most Powerful AI in Action How to Run Gemma 4 on Your Phone Without Internet: A Hands-On Guide Running Claude Code for Free with Gemma 4 and Ollama LLM Wiki Revolution: How Andrej Karpathy’s Idea is Changing AI Rethinking Enterprise Search: How Cortex Search Turns Data into Business Impact Google’s Gemma 4: Is it the Best Open-Source Model of 2026?
Understanding BERTopic: From Raw Text to Interpretable Topics
2026-04-11 · via Analytics Vidhya

Topic modeling uncovers hidden themes in large document collections. Traditional methods like Latent Dirichlet Allocation rely on word frequency and treat text as bags of words, often missing deeper context and meaning.

BERTopic takes a different route, combining transformer embeddings, clustering, and c-TF-IDF to capture semantic relationships between documents. It produces more meaningful, context-aware topics suited for real-world data. In this article, we break down how BERTopic works and how you can apply it step by step.

What is BERTopic? 

BERTopic is a modular topic modeling framework that treats topic discovery as a pipeline of independent but connected steps. It integrates deep learning and classical natural language processing techniques to produce coherent and interpretable topics. 

The core idea is to transform documents into semantic embeddings, cluster them based on similarity, and then extract representative words for each cluster. This approach allows BERTopic to capture both meaning and structure within text data. 

At a high level, BERTopic follows this process: 

BERT Workflow

Each component of this pipeline can be modified or replaced, making BERTopic highly flexible for different applications. 

Key Components of the BERTopic Pipeline 

1. Preprocessing 

The first step involves preparing raw text data. Unlike traditional NLP pipelines, BERTopic does not require heavy preprocessing. Minimal cleaning, such as lowercasing, removing extra spaces, and filtering very short documents is usually sufficient. 

2. Document Embeddings 

Each document is converted into a dense vector using transformer-based models such as SentenceTransformers. This allows the model to capture semantic relationships between documents. 

Mathematically: 

Document Embeddings 

Where di is a document and vi is its vector representation. 

3. Dimensionality Reduction 

High-dimensional embeddings are difficult to cluster effectively. BERTopic uses UMAP to reduce the dimensionality while preserving the structure of the data. 

Dimensionality Reduction

This step improves clustering performance and computational efficiency. 

4. Clustering 

After dimensionality reduction, clustering is performed using HDBSCAN. This algorithm groups similar documents into clusters and identifies outliers. 

Clustering 

Where zi  is the assigned topic label. Documents labeled as −1 are considered outliers. 

5. c-TF-IDF Topic Representation 

Once clusters are formed, BERTopic generates topic representations using c-TF-IDF. 

Term Frequency: 

Term Frequency

Inverse Class Frequency: 

Inverse Class Frequency

Final c-TF-IDF: 

cTFIDF

This method highlights words that are distinctive within a cluster while reducing the importance of common words across clusters. 

Hands-On Implementation 

This section demonstrates a simple implementation of BERTopic using a very small dataset. The goal here is not to build a production-scale topic model, but to understand how BERTopic works step by step. In this example, we preprocess the text, configure UMAP and HDBSCAN, train the BERTopic model, and inspect the generated topics. 

Step 1: Import Libraries and Prepare the Dataset 

import re
import umap
import hdbscan
from bertopic import BERTopic

docs = [
"NASA launched a satellite",
"Philosophy and religion are related",
"Space exploration is growing"
] 

In this first step, the required libraries are imported. The re module is used for basic text preprocessing, while umap and hdbscan are used for dimensionality reduction and clustering. BERTopic is the main library that combines these components into a topic modeling pipeline. 

A small list of sample documents is also created. These documents belong to different themes, such as space and philosophy, which makes them useful for demonstrating how BERTopic attempts to separate text into different topics. 

Step 2: Preprocess the Text 

def preprocess(text):
    text = text.lower()
    text = re.sub(r"\s+", " ", text)
    return text.strip()

docs = [preprocess(doc) for doc in docs]

This step performs basic text cleaning. Each document is converted to lowercase so that words like “NASA” and “nasa” are treated as the same token. Extra spaces are also removed to standardize the formatting. 

Preprocessing is important because it reduces noise in the input. Although BERTopic uses transformer embeddings that are less dependent on heavy text cleaning, simple normalization still improves consistency and makes the input cleaner for downstream processing. 

Step 3: Configure UMAP 

umap_model = umap.UMAP(
    n_neighbors=2,
    n_components=2,
    min_dist=0.0,
    metric="cosine",
    random_state=42,
    init="random"
)

UMAP is used here to reduce the dimensionality of the document embeddings before clustering. Since embeddings are usually high-dimensional, clustering them directly is often difficult. UMAP helps by projecting them into a lower-dimensional space while preserving their semantic relationships. 

The parameter init=”random” is especially important in this example because the dataset is extremely small. With only three documents, UMAP’s default spectral initialization may fail, so random initialization is used to avoid that error. The settings n_neighbors=2 and n_components=2 are chosen to suit this tiny dataset. 

Step 4: Configure HDBSCAN 

hdbscan_model = hdbscan.HDBSCAN(
    min_cluster_size=2,
    metric="euclidean",
    cluster_selection_method="eom",
    prediction_data=True
)

HDBSCAN is the clustering algorithm used by BERTopic. Its role is to group similar documents together after dimensionality reduction. Unlike methods such as K-Means, HDBSCAN does not require the number of clusters to be specified in advance. 

Here, min_cluster_size=2 means that at least two documents are needed to form a cluster. This is appropriate for such a small example. The prediction_data=True argument allows the model to retain information useful for later inference and probability estimation. 

Step 5: Create the BERTopic Model 

topic_model = BERTopic(
    umap_model=umap_model,
    hdbscan_model=hdbscan_model,
    calculate_probabilities=True,
    verbose=True
) 

In this step, the BERTopic model is created by passing the custom UMAP and HDBSCAN configurations. This shows one of BERTopic’s strengths: it is modular, so individual components can be customized according to the dataset and use case. 

The option calculate_probabilities=True enables the model to estimate topic probabilities for each document. The verbose=True option is useful during experimentation because it displays progress and internal processing steps while the model is running. 

Step 6: Fit the BERTopic Model 

topics, probs = topic_model.fit_transform(docs) 

This is the main training step. BERTopic now performs the complete pipeline internally: 

  1. It converts documents into embeddings  
  2. It reduces the embedding dimensions using UMAP  
  3. It clusters the reduced embeddings using HDBSCAN  
  4. It extracts topic words using c-TF-IDF  

The result is stored in two outputs: 

  • topics, which contains the assigned topic label for each document  
  • probs, which contains the probability distribution or confidence values for the assignments  

This is the point where the raw documents are transformed into topic-based structure. 

Step 7: View Topic Assignments and Topic Information 

print("Topics:", topics)
print(topic_model.get_topic_info())

for topic_id in sorted(set(topics)):
    if topic_id != -1:
        print(f"\nTopic {topic_id}:")
        print(topic_model.get_topic(topic_id))
Output

This final step is used to inspect the model’s output. 

  • print("Topics:", topics) shows the topic label assigned to each document.  
  • get_topic_info() displays a summary table of all topics, including topic IDs and the number of documents in each topic.  
  • get_topic(topic_id) returns the top representative words for a given topic.  

The condition if topic_id != -1 excludes outliers. In BERTopic, a topic label of -1 means that the document was not confidently assigned to any cluster. This is a normal behavior in density-based clustering and helps avoid forcing unrelated documents into incorrect topics. 

Advantages of BERTopic 

Here are the main advantages of using BERTopic:

  • Captures semantic meaning using embeddings
    BERTopic uses transformer-based embeddings to understand the context of text rather than just word frequency. This allows it to group documents with similar meanings even if they use different words. 
  • Automatically determines number of topics
    Using HDBSCAN, BERTopic does not require a predefined number of topics. It discovers the natural structure of the data, making it suitable for unknown or evolving datasets. 
  • Handles noise and outliers effectively
    Documents that do not clearly belong to any cluster are labeled as outliers instead of being forced into incorrect topics. This improves the overall quality and clarity of the topics. 
  • Produces interpretable topic representations
    With c-TF-IDF, BERTopic extracts keywords that clearly represent each topic. These words are distinctive and easy to understand, making interpretation straightforward. 
  • Highly modular and customizable
    Each part of the pipeline can be adjusted or replaced, such as embeddings, clustering, or vectorization. This flexibility allows it to adapt to different datasets and use cases. 

Conclusion 

BERTopic represents a significant advancement in topic modeling by combining semantic embeddings, dimensionality reduction, clustering, and class-based TF-IDF. This hybrid approach allows it to produce meaningful and interpretable topics that align more closely with human understanding. 

Rather than relying solely on word frequency, BERTopic leverages the structure of semantic space to identify patterns in text data. Its modular design also makes it adaptable to a wide range of applications, from analyzing customer feedback to organizing research documents. 

In practice, the effectiveness of BERTopic depends on careful selection of embeddings, tuning of clustering parameters, and thoughtful evaluation of results. When applied correctly, it provides a powerful and practical solution for modern topic modeling tasks. 

Frequently Asked Questions

Q1. What makes BERTopic different from traditional topic modeling methods?

A. It uses semantic embeddings instead of word frequency, allowing it to capture context and meaning more effectively. 

Q2. How does BERTopic determine the number of topics?

A. It uses HDBSCAN clustering, which automatically discovers the natural number of topics without predefined input. 

Q3. What is a key limitation of BERTopic?

A. It is computationally expensive due to embedding generation, especially for large datasets.

Hi, I am Janvi, a passionate data science enthusiast currently working at Analytics Vidhya. My journey into the world of data began with a deep curiosity about how we can extract meaningful insights from complex datasets.