惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Engineering at Meta
Engineering at Meta
C
Cyber Attacks, Cyber Crime and Cyber Security
博客园 - 司徒正美
月光博客
月光博客
Hugging Face - Blog
Hugging Face - Blog
T
Tailwind CSS Blog
罗磊的独立博客
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 三生石上(FineUI控件)
博客园_首页
博客园 - 【当耐特】
Cisco Talos Blog
Cisco Talos Blog
J
Java Code Geeks
C
CXSECURITY Database RSS Feed - CXSecurity.com
S
SegmentFault 最新的问题
人人都是产品经理
人人都是产品经理
Jina AI
Jina AI
AWS News Blog
AWS News Blog
S
Schneier on Security
NISL@THU
NISL@THU
F
Fortinet All Blogs
L
LINUX DO - 热门话题
Google DeepMind News
Google DeepMind News
量子位
IT之家
IT之家
T
The Exploit Database - CXSecurity.com
爱范儿
爱范儿
GbyAI
GbyAI
T
The Blog of Author Tim Ferriss
T
Tor Project blog
V
Vulnerabilities – Threatpost
V
Visual Studio Blog
宝玉的分享
宝玉的分享
Spread Privacy
Spread Privacy
L
Lohrmann on Cybersecurity
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Y
Y Combinator Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
P
Privacy International News Feed
S
Securelist
P
Palo Alto Networks Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
A
Arctic Wolf
T
Tenable Blog
B
Blog
C
CERT Recently Published Vulnerability Notes
P
Proofpoint News Feed
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Threat Research - Cisco Blogs
T
Threatpost

Replicate's blog

How to make remarkable videos with Seedance 2.0 – Replicate blog How to prompt Seedream 5.0 – Replicate blog Recraft V4: image generation with design taste – Replicate blog Run Isaac 0.1 on Replicate – Replicate blog Run FLUX.2 on Replicate – Replicate blog How to prompt Nano Banana Pro – Replicate blog Retro Diffusion's pixel art models are now on Replicate – Replicate blog Replicate is joining Cloudflare – Replicate blog Extract text from documents and images with Datalab Marker and OCR – Replicate blog How to prompt Veo 3.1 – Replicate blog IBM's Granite 4.0 is now on Replicate – Replicate blog Which image editing model should I use? – Replicate blog Introducing our new search API – Replicate blog Torch compile caching for inference speed – Replicate blog How to prompt Veo 3 with images – Replicate blog Open source video is back – Replicate blog Generate consistent characters – Replicate blog Bria is now on Replicate – Replicate blog How we optimized FLUX.1 Kontext [dev] – Replicate blog Compare AI video models – Replicate blog The FLUX.1 Kontext hackathon – Replicate blog How to prompt Veo 3 for the best results – Replicate blog Get the most from Google Veo 3 – Replicate blog FLUX.1 Kontext from the community – Replicate blog Use FLUX.1 Kontext to edit images with words – Replicate blog Generate incredible images with Google's Imagen 4 – Replicate blog Run OpenAI’s latest models on Replicate – Replicate blog NVIDIA H100 GPUs are here – Replicate blog Run 30,000+ LoRAs on Hugging Face with Replicate – Replicate blog Ideogram 3.0 on Replicate – Replicate blog Run MiniMax Speech-02 models with an API – Replicate blog Easel AI is now on Replicate – Replicate blog Stylized video with Wan2.1 – Replicate blog Creative roundup: avatars, lightsabers, and LoRA tricks – Replicate blog Wan2.1: generate videos with an API – Replicate blog Wan2.1 parameter sweep – Replicate blog You can now fine-tune open-source video models – Replicate blog Generate short videos with the Replicate playground – Replicate blog AI video is having its Stable Diffusion moment – Replicate blog FLUX fine-tunes are now fast – Replicate blog FLUX.1 Tools – Control and steerability for FLUX – Replicate blog NVIDIA L40S GPUs are here – Replicate blog Ideogram v2 is an outstanding new inpainting model – Replicate blog Stable Diffusion 3.5 is here – Replicate blog FLUX is fast and it's open source – Replicate blog FLUX1.1 [pro] is here – Replicate blog Using synthetic training data to improve Flux finetunes – Replicate blog Fine-tune FLUX.1 with an API – Replicate blog Fine-tune FLUX.1 to create images of yourself – Replicate blog Replicate Intelligence #12 – Replicate blog Replicate Intelligence #11 – Replicate blog Fine-tune FLUX.1 with your own images – Replicate blog Replicate Intelligence #10 – Replicate blog FLUX.1: First Impressions – Replicate blog Replicate Intelligence #9 – Replicate blog Run FLUX with an API – Replicate blog Replicate Intelligence #8 – Replicate blog Run Meta Llama 3.1 405B with an API – Replicate blog Replicate Intelligence #7 – Replicate blog Replicate Intelligence #6 – Replicate blog Replicate Intelligence #5 – Replicate blog How to get the best results from Stable Diffusion 3 – Replicate blog Run Stable Diffusion 3 on your Apple Silicon Mac – Replicate blog Push a custom version of Stable Diffusion 3 – Replicate blog Replicate Intelligence #4 – Replicate blog Run Stable Diffusion 3 on your own machine with ComfyUI – Replicate blog H100s are coming to Replicate – Replicate blog Run Stable Diffusion 3 with an API – Replicate blog Replicate Intelligence #3 – Replicate blog Replicate Intelligence #2 – Replicate blog Replicate Intelligence #1 – Replicate blog Shared network vulnerability disclosure – Replicate blog Run Snowflake Arctic with an API – Replicate blog Run Meta Llama 3 with an API – Replicate blog Run Code Llama 70B with an API – Replicate blog How to create an AI narrator for your life – Replicate blog Clone your voice using open-source models – Replicate blog Businesses are building on open-source AI – Replicate blog How to run Yi chat models with an API – Replicate blog Scaffold Replicate apps with one command – Replicate blog Using open-source models for faster and cheaper text embeddings – Replicate blog Generate music from chord progressions and text prompts with MusicGen-Chord – Replicate blog Generate images in one second on your Mac using a latent consistency model – Replicate blog How to use retrieval augmented generation with ChromaDB and Mistral – Replicate blog Fine-tune MusicGen to generate music in any style – Replicate blog Jet-setting with Llama 2 + Grammars – Replicate blog How to run Mistral 7B with an API – Replicate blog Make smooth AI generated videos with AnimateDiff and an interpolator – Replicate blog Fine-tuned models now boot in less than one second – Replicate blog Painting with words: a history of text-to-image AI – Replicate blog We're cutting our prices in half – Replicate blog A guide to prompting Llama 2 – Replicate blog Streaming output for language models – Replicate blog Fine-tune SDXL with your own images – Replicate blog Run Llama 2 with an API – Replicate blog Run SDXL with an API – Replicate blog A comprehensive guide to running Llama 2 locally – Replicate blog Fine-tune Llama 2 on Replicate – Replicate blog What happened with Llama 2 in the last 24 hours? 🦙 – Replicate blog Make any large language model a better poet – Replicate blog
Announcing Replicate's remote MCP server – Replicate blog
2025-08-10 · via Replicate's blog

Posted August 10, 2025 by

Last month we quietly published a local MCP server for Replicate’s HTTP API.

Today we’re announcing a hosted remote MCP server that you can use with apps like Claude Desktop, Claude Code, Cursor, and VS Code, giving you the power to explore and run all of Replicate’s HTTP APIs from a familiar chat-based natural language interface.

To get started, head over to 👉 mcp.replicate.com 👈

What is MCP?

MCP stands for Model Context Protocol. It’s a standard developed at Anthropic for giving language models access to external tools. This is commonly called “tool use” or “function calling”. This makes language models way more powerful, as they can now access external tools and data sources, instead of just their own internal knowledge.

Once you’ve installed the server, you can ask questions in Claude or Cursor like:

Find models:

“Find popular video models on Replicate that allow a starting frame as input”

Compare models:

“What are the differences between veo 3 and veo 3 fast on Replicate?”

Run models:

“Make a video of ‘a tortoise and a hare running in the Olympic 100m’ using veo 3 fast”

Here’s a video introducing MCP and showing how to use Replicate’s MCP server with Claude Desktop:

Two flavors: Remote and local

Our official MCP server is available as a hosted service, as well as a public npm package that you can run locally.

  • Remote MCP server (recommended): This is the easiest option, and recommended for most users. You just add the hosted server URL to your apps like Claude or Cursor. After installing the server, you’ll be directed to a web-based authentication flow where you can provide a Replicate API key for the server to use on your behalf. To get started, go to mcp.replicate.com
  • Local MCP server: You can run the server locally on your machine. All that’s required is a recent version of Node.js. To install and run the server locally, check out the docs.

JSON response filtering with jq

Some HTTP APIs can return very large JSON responses, and it’s easy to fill up a model’s context window with too much data. For example, Replicate’s search API returns paginated lists of models with extensive metadata for each model like inputs, outputs, description, and more. This metadata is useful, but can also be too large for the context windows of most large language models.

To work around this, we collaborated with the team at Stainless, who are building and maintaining our new API SDKs. They implemented tooling that dynamically filters larger response objects down to the most relevant parts. It uses a WebAssembly implementation of the popular jq command line tool to write one-off filter expressions that are specific to the given response object schema and the task at hand.

This approach allows the language model to decide which parts of the response are most relevant to the task at hand, and only return those parts. Here’s an example, plucking out just the name, owner, description, and run count from the search response:

Check out this video to see the response filtering in action:

Secure authentication with Cloudflare Workers

Our new hosted MCP server runs on Cloudflare Workers, which makes it easy to deploy and scale. Cloudflare is leading the industry with great tooling, tutorials, and content about how to securely build and scale MCP servers.

We use Cloudflare’s OAuth Provider Framework for Workers to keep your Replicate API token secure. When you connect your AI tools to Replicate, you visit a web page where you enter your Replicate API token. This token gets stored in Cloudflare’s KV storage, which functions like a secure digital vault that only your MCP server can access. The beauty of this approach is that your token never gets exposed to the AI tools themselves; instead, the MCP server acts as a trusted intermediary, using your stored token to make requests to Replicate on your behalf. This design means your credentials stay safe even if someone gets access to your AI tool’s configuration, and since everything runs on Cloudflare’s infrastructure, you get enterprise-grade security without the complexity.

Head over to mcp.replicate.com to get started, and let us know what you build! We’re excited to see what you create with these new capabilities.

If you have any questions or feedback, join us in Discord or reach out on Twitter/X.