惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
GbyAI
GbyAI
阮一峰的网络日志
阮一峰的网络日志
Cloudbric
Cloudbric
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
宝玉的分享
宝玉的分享
I
Intezer
Simon Willison's Weblog
Simon Willison's Weblog
博客园_首页
The Cloudflare Blog
C
Cisco Blogs
AWS News Blog
AWS News Blog
IT之家
IT之家
Cyberwarzone
Cyberwarzone
罗磊的独立博客
美团技术团队
V
V2EX
Project Zero
Project Zero
A
Arctic Wolf
C
Cyber Attacks, Cyber Crime and Cyber Security
大猫的无限游戏
大猫的无限游戏
博客园 - 叶小钗
月光博客
月光博客
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 聂微东
有赞技术团队
有赞技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
雷峰网
雷峰网
S
Schneier on Security
P
Privacy International News Feed
V
Visual Studio Blog
量子位
T
Tor Project blog
S
Securelist
腾讯CDC
A
About on SuperTechFans
T
Threat Research - Cisco Blogs
G
GRAHAM CLULEY
B
Blog RSS Feed
D
DataBreaches.Net
博客园 - 三生石上(FineUI控件)
B
Blog
NISL@THU
NISL@THU
L
Lohrmann on Cybersecurity
V
Vulnerabilities – Threatpost
人人都是产品经理
人人都是产品经理
博客园 - 【当耐特】
L
LINUX DO - 热门话题
Recorded Future
Recorded Future

Analytics Vidhya

Handling Imbalanced Classification: What Works Better Than SMOTE GPT-5.6 Is Here: Sol, Terra, and Luna Loop Engineering for AI Agents: How /loop is Changing AI Workflows DeepSeek DSpark: The Speculative Decoding Trick Behind 400% Faster LLM OKF: Redefining Knowledge Bases for AI Agents Modern VLMs Explained: How GPT-4o, Gemini, Claude Vision, and Qwen-VL Work YOLO26 Tutorial: Object Detection, Pose Estimation & More Large Action Models (LAMs) vs Agentic LLMs: What's the Real Difference? Claude Sonnet 5: The Fable 5 at Home The Best $20 AI Plan: ChatGPT Plus vs Claude Pro vs Gemini Pro GraphRAG vs Vector RAG: Which Retrieval Method is Best? Using AI When You Don’t Trust AI The Self-Improving Loop in AI Agents: Architecture, Benefits, and How it Outperforms Traditional Agent Workflows Harness-1: The 20B Retrieval Subagent That Beats GPT-5.4 at Search Sakana Fugu: Multi-Agent System as a Model Claude's Hidden Art Skill: Making Illustrations With Code System Design for ML Interviews: 10 Real Problems Walked Through Most People Use ChatGPT Wrong: 10 Features and Tips That Changed How I Work OpenAI Just Launched 3 Free AI Courses with Certificates Autoregressive Models: Predicting the Future Using the Past Gemini Omni: AI Video Generation Inside Gemini DiffusionGemma: Google’s Diffusion-Based Open Model for Faster Text Generation Top 10 AI Engineering Tools Everyone is Using in 2026 I Tested Claude Fable 5: Can Anthropic’s Newest AI Deliver on the Hype? Prophet vs NeuralProphet vs TimeGPT vs Chronos: A Practical Comparison Build an Emergency Helpline Voice Agent with LangChain Choosing the Right Vector Database for RAG and AI Applications Google Gemma 4 12B: Architecture, Benchmarks, Access, and Hands-on Guide for Developers How to Choose the Right AI Model for Your Needs Agent Observability with LangSmith, Langfuse, and Arize: A Hands-On Comparison How to Use Claude Managed Agents? Google AI Studio vs Gemini App: What’s the Difference? AI Workflows for Sales Teams: Prospect Research, Lead Qualification, and CRM Updates on Autopilot Using LangGraph 25 Most Influential AI Pioneers to Meet at DataHack Summit 2026 Claude Opus 4.8: A Smarter Model in the Right Direction PySpark Optimization: 12 Proven Techniques to Speed Up Your Spark Jobs 10 Everyday Tasks You Can Automate with AI Today (With n8n Templates) Google Antigravity 2.0: The Full Developer Guide (I/O 2026) Build a Claude Cowork-Like Browser Agent Using Playwright MCP and Claude Desktop Pandas vs Polars vs DuckDB: Which Library Should You Choose? Qwen3.7-Max: Alibaba’s New Agent-First LLM for Coding, Reasoning, and Long-Horizon AI Workflows The Biggest Announcements from Google I/O 2026 Top 9 AI Events and Conferences in 2026 that you Must Attend Gemini 3.5 Flash: Frontier Intelligence with Speed 40 Advanced SQL Window Functions Every Data Scientist Must Know(with examples) Top 10 AI Research Papers of 2025 6 Steps to Crack GenAI Case Study Interviews (With Real Examples) OpenAI Omni Moderation: How to Filter Text & Images for Free DataHack Summit 2026: You Just Cannot Skip This AI Event of the Year OpenAI’s New API Voice Models Will Change the Way You Use AI Hermes Agent Guide: What is it and How to Use it? Top 10 LLM Research Papers of 2026 Agent Memory Patterns in Cognitive Science and AI Systems 10 AI Agents Every AI Engineer Must Build (with GitHub Samples) 23 Tips for Smart Claude Code Token Saving and Workflow Optimization Feature Engineering with LLMs: Techniques & Python Examples ChatGPT is Now Inside Excel and Google Sheets: Here is How to Use it Gemini API File Search: The Easy Way to Build RAG Top 10 Open-Source Libraries to Fine-Tune LLMs Locally ML Intern in Practice: From Prompt to a Shipped Hugging Face Model 15+ Solved Agentic AI Projects with Github Links How People are Figuring Out Life With Claude MemPalace Explained: Building Long-Term Memory for AI Agents Beyond RAG Grok Voice Think Fast 1.0: Build Voice AI Agents That Actually Think Compressing LSTM Models for Retail Edge Deployment: A Practical Comparison MCP vs Agent Skills: Different Altogether GPT 5.5 vs Opus 4.7: Which is the Best AI Model Today? What is Agentic AI? Claude Code vs Codex: A Detailed Terminal Agent Comparison Google Deep Research Max: Build Autonomous AI Research Agents in Minutes Meta Muse Spark Review: Is It Worth the Hype? ChatGPT Images 2.0 vs Nano Banana 2: Which is Better? Cursor V3 Explained: The AI Coding Agent That’s Replacing Traditional IDEs in 2026 DeepSeek-V4: The Most Powerful Open-Source Model Ever Is GPT Image 2 the Best Image Generation Model? Token Economics: Why AI is Getting “Cheaper” From Idea to Output: Claude Does the Design Work Opus 4.7 vs Opus 4.6: Should You Switch? Build Human-Like AI Voice App with Gemini 3.1 Flash TTS How to Structure a Claude Code Project that Thinks Like an Engineer Gemma 4 Tool Calling Explained: Build AI Agents with Function Calling (Step-by-Step Guide) Anthropic Launches Claude Opus 4.7 For “Most Difficult Tasks” Top 28 Claude Shortcuts that will 10X your Speed GPT-5.4-Cyber: Why OpenAI is Keeping its Most Powerful Model Under Lock and Key Google AI Studio Guide: Every Feature Explained Mastering Deep Agents: Context Engineering that Actually Works 21 Computer Vision Projects from Beginner to Advanced (2026 Guide) Excel 101: Excel Agent Mode Explained MiniMax M2.7 Goes Open-Weight to Let You Run Agents Locally Top 10 Gemma 4 Projects That Will Blow Your Mind GLM-5.1: Architecture, Benchmarks, Capabilities & How to Use It Understanding BERTopic: From Raw Text to Interpretable Topics From Karpathy’s LLM Wiki to Graphify: AI Memory Layers are Here 10 Most Important AI Concepts Explained Simply Project Glasswing is World’s Most Powerful AI in Action How to Run Gemma 4 on Your Phone Without Internet: A Hands-On Guide Running Claude Code for Free with Gemma 4 and Ollama LLM Wiki Revolution: How Andrej Karpathy’s Idea is Changing AI Rethinking Enterprise Search: How Cortex Search Turns Data into Business Impact Google’s Gemma 4: Is it the Best Open-Source Model of 2026?
Kimi WebBridge: Hands-on Guide to Kimi’s Browser Extension for AI Agents
Harsh Mishra · 2026-05-19 · via Analytics Vidhya

AI agents are evolving from answering questions to taking actions inside browsers. They can now open pages, click buttons, fill forms, extract data, and automate multi step workflows across websites.

Moonshot AI’s Kimi WebBridge brings this capability to Chrome and Edge, allowing local AI agents to safely interact with real browser sessions. In this article, we explore how WebBridge works and why browser automation is becoming essential for agentic AI systems.

Table of contents

  • What is Kimi WebBridge?
  • How Kimi WebBridge Works
  • Kimi WebBridge Architecture
  • Installation and Setup
  • Hands-on Workflow: Research Automation
  • Advantages and Disadvantages of Kimi WebBridge
  • Security and Governance Considerations
  • Kimi WebBridge vs Playwright MCP vs Browserbase
  • Conclusion

What is Kimi WebBridge?

Kimi WebBridge is an AI agent browser extension. WebBridge is not a cloud-based browser automation solution that launches a browser remote, but rather it runs directly in your browser, using your existing login sessions. The agent can then interact with web pages as would a human user, more closely.   

From a simple point of view, Kimi WebBridge is a bridge between:  

Your local AI agent:

  • The browser extension that you installed.The extension that you installed on your browser.  
  • The web version of the Chrome or Edge browser you are using the browser .  
  • The sites that you are currently signed into.  

According to the official description in the Chrome Web Store, the extension is able to open a webpage, click, fill in forms, extract information, and automate web operations using AI. This is version 1.9.7, which was updated on 11 May 2026, as seen in the Chrome listing.

How Kimi WebBridge Works

Kimi WebBridge is a local-first application. Kimi’s help documents claim it operates with three things: Local bridge service, Browser extension, and Local security isolation. The instructions are sent from the agent to the local bridge and then the local bridge sends the instructions to the extension to perform actions in the browser with the chrome DevTool protocol and then executes locally on the user’s device.   

CDP (also known as Chrome DevTools Protocol) is the protocol for instrumenting, inspecting, debugging and profiling Chromium based browsers at the browser level. Unveils browser domains (DOM, Network, Page, Runtime, Input and more).

This means that WebBridge isn’t simply taking HTML without any interpretation. It’s providing an agent controlled operational access for browser actions, including: 

  • Open a URL 
  • Click an element 
  • Fill a form 
  • Capture a screenshot 
  • Read page content 
  • Extract tables or structured text 
  • Use existing logged-in sessions 

Kimi’s documentation lists these as core features, including web navigation, element clicking, form filling, screenshots, content extraction, and login session persistence.  

Kimi WebBridge Architecture

A practical mental model for Kimi WebBridge looks like this: 

How Kimi WebBridge Works

The most critical design decision is that WebBridge is run locally. When using WebBridge, login states and web page content are not left on the user’s machine, Kimi says.   

This comes in handy for enterprise applications that need to shield sensitive applications, internal dashboards, subscribed sessions, or private customer data from third party remote browsers. 

Installation and Setup

Prerequisites

Before starting, you need:

  • Chrome or Edge browser 
  • Kimi WebBridge extension 
  • A local agent such as Kimi Code, Claude Code, Cursor, Codex, Hermes, or OpenClaw 
  • Terminal access 
  • Logged-in websites for the workflows you want to automate 

Kimi’s official page lists supported AI agents including Kimi Code, Claude Code, Cursor, Codex, Hermes, and OpenClaw. 

Step 1: Install the Extension

You can download it through the browser extension store. Kimi’s help center lists Chrome Web Store for Chrome users and Edge Add-ons for Edge users. 

Kimi WebBridge Browser Extension

Step 2: Pin the Extension 

Once installed, add WebBridge to the browser toolbar. This will make it easier to determine if it is plugged in or not. Kimi’s docs suggest fixing it to the wall to make it more accessible.

Step 3: Connect WebBridge to a Local Agent 

When WebBridge is installed locally, there is a local setup command on Kimi’s official feature page for connecting WebBridge to your agent:

curl -fsSL https://kimi-web-img.moonshot.cn/webbridge/install.sh | bash
Connecting WebBridge to a local agent

In the official page, it is stated that you copy the command into your agent and Kimi WebBridge will connect automatically.   

To check the status of the Kimi WebBridge run kimi-webbridge status command if says connected then you are good to go, if not then try running the following command and check the status again.

export PATH="$PATH:/Users/{your-pc-username}/.kimi-webbridge/bin" 
source ~/.zshrc
Kimi Webbridge status

Step 4: Check Connection Status 

Click into the WebBridge icon on the bottom of the browser. Kimi says “Connected” status indicates that WebBridge is functioning correctly and is able to communicate with the agent. “Disconnected”: There are issues with configuration. Try rerunning the connection command. 

Kimi WebBridge browser assistant is ready

Step 5: Using the Agent 

Here we will be using Claude code, Kimi automatically installed skill files in your available agents such as Codex, Claude Code, Hermes etc while installation. Now only open them up and use /kimi-webbridge in order to utilise this skill.  

Do not begin with banking, production admin dashboards or enterprise sensitive systems. Test on public websites, documentation pages, demo applications or test environment. 

Prompt: “Open the Analytics Vidhya blog homepage. Find 2 recent AI agent articles. Extract the title, author, last updated date, and one-line summary into a markdown table.”

Using the Claude Code Agent
Analytics Vidhya Blog Homepage
Analytics Vidhya Blog
Churned for 1m 42 seconds on the Articles

This tests navigation, reading, extraction, and summarization without requiring any risky action. 

Hands-on Workflow: Research Automation

Prompt“/kimi-webbridge Go to linkedin and search for 2 top AI enginners in top AI companies and give me a CSV file with their name, profile url, and all profile details”

/kimi-wbbridge prompt
AI Engineer search on LinkedIn

What WebBridge Did? 

The agent: 

  1. Open search on Linkedin 
  2. Visit pages one by one 
  3. Read visible content 
  4. Extract structured details 
  5. Return a clean table 

Output:

Webbridge feching informa
Excel sheet containing the information from WebBridge

Technical Value 

This is useful for analysts, content teams, product managers, and strategy teams. Instead of manually opening 10 tabs and copying notes, the agent can operate the browser and structure the findings. 

Advantages and Disadvantages of Kimi WebBridge

Advantages Disadvantages & Limitations
1. Local-first Browser Automation

WebBridge runs locally on the user’s machine, reducing exposure compared with cloud-browser automation workflows handling authenticated sessions.

1. Limited Browser Support

Currently supports Chrome and Edge only. Safari and Firefox are not first-class supported targets.

2. Works With Existing Login Sessions

Uses the user’s active Chrome or Edge session, making it useful for websites without APIs or platforms requiring authentication.

2. Local Setup Can Be Friction-heavy

Every machine requires individual installation and setup, which becomes difficult to scale across large organizations.

3. Agent-agnostic Positioning

Compatible with tools like Kimi Code, Claude Code, Cursor, Codex, Hermes, and OpenClaw, making it more flexible than a closed ecosystem tool.

3. Dynamic Pages Can Fail

Modern apps using React, shadow DOMs, lazy loading, popups, or anti-bot systems may cause automation instability or failures.

4. Useful for Real Business Workflows

Supports practical automation use cases such as ecommerce price comparison, form filling, data entry, and research workflows.

4. Extension Conflicts Are Possible

Browser extensions like scrapers, screen recorders, and AI assistants may interfere with clicks, snapshots, screenshots, and page evaluation.

5. Built on Browser-native Control

Built on Chrome DevTools Protocol (CDP), allowing low-level browser instrumentation, inspection, debugging, and HTML parsing.

5. Local-first Does Not Mean Risk-free

Extensions with Debugger API access can still introduce security risks through browser manipulation or traffic monitoring.

Overall

WebBridge is strongest for teams wanting browser-native automation while keeping sessions local and compatible with multiple coding agents.

6. Agent Safety Remains a Challenge

Browser agents can perform real actions, making guardrails like audit logs, confirmation gates, allowlisted domains, and safe browsing profiles important for enterprise use.

Security and Governance Considerations

For Enterprise, it’s not just about “Can this automate work?” It’s the “Can this automate work safely?” question. 

Use these controls: 

  1. Create a dedicated browser profile for agent work. 
  2. Use least-privilege accounts. 
  3. Avoid admin accounts for early testing. 
  4. Use read-only access where possible. 
  5. Require confirmation before submit, delete, purchase, approve, or send actions. 
  6. Disable conflicting extensions. 
  7. Keep WebBridge updated. 
  8. Log prompts, actions, and outputs. 
  9. Test on staging environments first. 
  10. Define domain allowlists for enterprise workflows. 

Low-risk workflows should be initiated, such as research, extraction, comparison, summarization, and report generation, in a safe enterprise rollout. Payment processes, account changes, customer communication, and production admin processes are examples of high-risk workflows that should include explicit human approval. 

Kimi WebBridge vs Playwright MCP vs Browserbase

Tool Best For Browser Location Strength Trade-off
Kimi WebBridge Local agent controlling your real browser Local Chrome or Edge Uses existing login sessions and runs locally Limited to supported browsers and local setup
Playwright MCP Developer-centric browser automation through MCP Usually local or configured browser environment Provides browser automation capabilities using Playwright and lets LLMs interact with pages through structured accessibility snapshots More developer setup and less focused on existing personal browser sessions
Browserbase Scalable cloud browser automation Cloud browsers Provides production infrastructure for automated browsers at scale Cloud browser model may not fit all private-session workflows

The playwright server, an MCP server from Microsoft, offers browser automation capabilities with Playwright and allows the LLM to interact with a web page via a structured accessibility snapshot.  

According to Browserbase, it’s “a cloud platform for headless browser automation providing infrastructure for running automated web browsers at scale.” 

The problem is Kimi WebBridge operates on the local control of the user’s own Chrome or Edge browser session.

Conclusion

Kimi WebBridge is an important step in browser agents, allowing AI agents to operate directly inside real Chrome or Edge browsers using existing login sessions. It supports workflows like research, dashboard extraction, price comparison, recruiting, and form automation while keeping execution local instead of cloud-based.

Its local-first design and compatibility with tools like Claude Code and Cursor make it appealing for developers and technical teams. At the same time, because browser agents can perform real actions, teams still need safeguards like confirmation gates, clean browser profiles, and controlled testing.

WebBridge is a strong sign that AI agents are moving beyond chat interfaces into browsers, tools, and business workflows.

Harsh Mishra is an AI/ML Engineer who spends more time talking to Large Language Models than actual humans. Passionate about GenAI, NLP, and making machines smarter (so they don’t replace him just yet). When not optimizing models, he’s probably optimizing his coffee intake. 🚀☕