惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Forbes - Security
Forbes - Security
Cisco Talos Blog
Cisco Talos Blog
Latest news
Latest news
P
Proofpoint News Feed
T
The Exploit Database - CXSecurity.com
Know Your Adversary
Know Your Adversary
S
Securelist
T
Tor Project blog
P
Palo Alto Networks Blog
G
GRAHAM CLULEY
NISL@THU
NISL@THU
C
CERT Recently Published Vulnerability Notes
L
LINUX DO - 热门话题
V
Vulnerabilities – Threatpost
Simon Willison's Weblog
Simon Willison's Weblog
AWS News Blog
AWS News Blog
T
The Blog of Author Tim Ferriss
Security Latest
Security Latest
P
Proofpoint News Feed
C
CXSECURITY Database RSS Feed - CXSecurity.com
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
T
Tenable Blog
博客园_首页
TaoSecurity Blog
TaoSecurity Blog
Attack and Defense Labs
Attack and Defense Labs
Project Zero
Project Zero
The Hacker News
The Hacker News
M
MIT News - Artificial intelligence
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Application and Cybersecurity Blog
Application and Cybersecurity Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
K
Kaspersky official blog
F
Full Disclosure
WordPress大学
WordPress大学
Engineering at Meta
Engineering at Meta
The Cloudflare Blog
N
Netflix TechBlog - Medium
Stack Overflow Blog
Stack Overflow Blog
L
LangChain Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
MongoDB | Blog
MongoDB | Blog
宝玉的分享
宝玉的分享
GbyAI
GbyAI
J
Java Code Geeks
云风的 BLOG
云风的 BLOG
Recent Announcements
Recent Announcements
博客园 - 叶小钗
Webroot Blog
Webroot Blog
Hacker News: Ask HN
Hacker News: Ask HN

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
MCP Logging: What I Wish I Knew Before Deploying My Production MCP Server (3 Weeks of Production Pain)
KevinTen · 2026-06-25 · via DEV Community

MCP Logging: What I Wish I Knew Before Deploying My Production MCP Server (3 Weeks of Production Pain)

Honestly, I built 10+ MCP servers over the past 3 months, and I thought I had everything figured out. Authentication? Check. Error handling? Check. Rate limiting? Check. Deployment? Check.

Then my production server started failing randomly. Users would report "empty responses" or "timeout" but when I checked locally, everything worked fine. No errors in the console, nothing in the logs—just… nothing.

I spent 8 hours debugging over three days, and I learned the hard way: MCP servers need different logging than your regular REST API. In this post, I want to share what I got wrong, what fixed it, and the exact logging setup I use now that actually catches those weird MCP-specific failures.

The Problem: MCP Is Stateless But Conversations Aren't

Let me start with a confession: I used the same logging setup I use for regular Spring Boot REST APIs. SLF4J, INFO level for requests, errors go to error log. That works fine for most APIs. But MCP is different.

With MCP:

  • Every request comes from the AI client, not directly from the user
  • The client retries aggressively on timeout
  • Multiple tool calls can happen in the same conversation
  • If something fails silently, the client just gets an empty response and the user has no idea why

What was happening to me? Some requests were hitting size limits in my reverse proxy, the proxy would drop the connection, and I'd never see it in my logs because the request never reached my app. The client retries, same thing happens again. User gets frustrated, I sit there scratching my head because "it works on my machine".

So here's what I changed—everything.

Step 1: Log Everything At The Entry Point

First, I added a filter that logs every single request before it hits my controller. Even if the request never makes it through, I have a log entry. That sounds obvious, but I used to log after authentication. Wrong—if authentication fails silently or the filter chain breaks, you get nothing.

Here's the code I use now:

@Component
public class McpLoggingFilter extends OncePerRequestFilter {

    private static final Logger log = LoggerFactory.getLogger(McpLoggingFilter.class);

    @Override
    protected void doFilterInternal(
            HttpServletRequest request, 
            HttpServletResponse response, 
            FilterChain filterChain
    ) throws ServletException, IOException {

        long startTime = System.currentTimeMillis();
        String requestId = UUID.randomUUID().toString();
        String apiKey = extractApiKey(request);

        // Log BEFORE processing—even if everything fails, we have this
        log.info("MCP_REQUEST_START | requestId={} | method={} | path={} | apiKeyPresent={} | remoteAddr={}", 
                requestId,
                request.getMethod(),
                request.getRequestURI(),
                apiKey != null,
                request.getRemoteAddr());

        try {
            wrapRequestLogging(request, response, filterChain, requestId, apiKey, startTime);
        } catch (Exception e) {
            log.error("MCP_REQUEST_ERROR | requestId={} | error={}", requestId, e.getMessage(), e);
            throw e;
        }
    }

    private void wrapRequestLogging(
            HttpServletRequest request, 
            HttpServletResponse response, 
            FilterChain filterChain, 
            String requestId, 
            String apiKey,
            long startTime
    ) throws IOException, ServletException {
        ContentCachingRequestWrapper wrappedRequest = new ContentCachingRequestWrapper(request);
        ContentCachingResponseWrapper wrappedResponse = new ContentCachingResponseWrapper(response);

        filterChain.doFilter(wrappedRequest, wrappedResponse);

        long duration = System.currentTimeMillis() - startTime;
        int status = wrappedResponse.getStatus();
        int responseSize = wrappedResponse.getContentSize();

        // Log after processing—everything completed, capture size and status
        log.info("MCP_REQUEST_COMPLETE | requestId={} | status={} | durationMs={} | responseSize={} | apiKey={}",
                requestId, status, duration, responseSize, maskApiKey(apiKey));

        // Log WARN for slow requests—helps catch timeout issues before users report
        if (duration > 10000) {
            log.warn("MCP_REQUEST_SLOW | requestId={} | durationMs={} | this might cause client retry", 
                    requestId, duration);
        }

        wrappedResponse.copyBodyToResponse();
    }

    private String extractApiKey(HttpServletRequest request) {
        // Check all four locations (I learned this the hard way in my auth article)
        String apiKey = request.getHeader("X-API-Key");
        if (apiKey != null) return apiKey;

        String auth = request.getHeader("Authorization");
        if (auth != null && auth.startsWith("Bearer ")) {
            return auth.substring(7);
        }

        apiKey = request.getParameter("api_key");
        if (apiKey != null) return apiKey;

        return request.getParameter("apiKey");
    }

    private String maskApiKey(String apiKey) {
        if (apiKey == null || apiKey.length() < 8) return null;
        // Show first 4 chars, mask the rest—good enough for debugging, still secure
        return apiKey.substring(0, 4) + "****";
    }
}

Key things I do differently here:

  1. Log start before processing: If the request dies before hitting your app, you at least know it arrived at the filter.
  2. Cache request/response content: You can log the content if something fails (I don't log it by default to save space, but easy to enable).
  3. Log response size: This is how I caught my proxy truncating responses—response size in my log was 12KB, but client always got 8KB. Proxy was dropping the rest.
  4. Log slow requests explicitly: MCP clients retry after ~10s by default, so warn on anything slower than that. You'll see the pattern before users complain.
  5. Mask API keys: Still need to be secure, don't log full keys.

Step 2: Special Logging For The Two MCP Endpoints

MCP really only has two endpoints you need to care about: tools/list and tools/call. Each has different failure modes, so I added structured logging specifically for them.

For tools/list, I log how many tools are being returned:

@RestController
@RequestMapping("/mcp")
public class McpController {

    private static final Logger log = LoggerFactory.getLogger(McpController.class);

    private final List<McpTool> tools;

    // ... constructor

    @PostMapping("/tools/list")
    public ResponseEntity<McpListToolsResponse> listTools(@RequestAttribute String requestId) {
        long start = System.currentTimeMillis();
        int count = tools.size();

        log.info("MCP_TOOLS_LIST | requestId={} | toolCount={}", requestId, count);

        McpListToolsResponse response = McpListToolsResponse.builder()
                .tools(tools)
                .build();

        long duration = System.currentTimeMillis() - start;
        log.info("MCP_TOOLS_LIST_DONE | requestId={} | toolCount={} | durationMs={}", 
                requestId, count, duration);

        return ResponseEntity.ok(response);
    }

For tools/call, this is where the real money is. You need to log which tool was called, what parameters came in, and what the response size was. AI clients love to send huge prompts, so this catches those:

    @PostMapping("/tools/call")
    public ResponseEntity<McpCallToolResponse> callTool(
            @RequestBody McpCallToolRequest request,
            @RequestAttribute String requestId
    ) {
        long start = System.currentTimeMillis();
        String toolName = request.getName();
        int paramsSize = new ObjectMapper().writeValueAsString(request.getArguments()).length();

        log.info("MCP_TOOL_CALL | requestId={} | tool={} | paramsSize={}", 
                requestId, toolName, paramsSize);

        // If parameters are huge, log a warning—something's probably wrong
        if (paramsSize > 100_000) {
            log.warn("MCP_TOOL_CALL_LARGE_PARAMS | requestId={} | tool={} | paramsSize={}KB", 
                    requestId, toolName, paramsSize / 1000);
        }

        try {
            McpToolResult result = toolExecutor.execute(toolName, request.getArguments());
            long duration = System.currentTimeMillis() - start;
            int resultSize = result.getContent() != null ? 
                    result.getContent().toString().length() : 0;

            log.info("MCP_TOOL_CALL_DONE | requestId={} | tool={} | durationMs={} | resultSize={} | isError={}",
                    requestId, toolName, duration, resultSize, result.isError());

            // Warn on huge responses—this is what got me!
            if (resultSize > 100_000) {
                log.warn("MCP_TOOL_CALL_LARGE_RESPONSE | requestId={} | tool={} | resultSize={}KB",
                        requestId, toolName, resultSize / 1000);
            }

            return ResponseEntity.ok(McpCallToolResponse.builder()
                    .content(result.getContent())
                    .isError(result.isError())
                    .build());
        } catch (Exception e) {
            log.error("MCP_TOOL_CALL_FAILED | requestId={} | tool={} | error={}", 
                    requestId, toolName, e.getMessage(), e);
            return ResponseEntity.internalServerError().build();
        }
    }
}

Here's the thing that caught my proxy issue: I was getting resultSize=12845 in my logs, but clients would get disconnected halfway through. That's when I realized my Nginx proxy had a proxy_buffer_size setting that was too small for large MCP responses. It would buffer part of the response, hit the limit, and close the connection. My app thought it sent everything fine—no error in my original logs. With this logging, I immediately saw the pattern: only responses over 8KB would fail.

Step 3: Don't Log Bodies By Default, But Have It When You Need It

I tried logging full request/response bodies at first, and it was just too noisy. Most MCP requests are small, but when you get a big tool call with a lot of context, your logs blow up quickly.

Instead, I have a debug flag I can turn on per-API key. When a specific user is having problems, I can enable debug logging just for them without turning everything on:

@Component
public class DebugLoggingConfig {

    private final Set<String> debugApiKeyPrefixes = Set.of(
            // Add prefixes of API keys that need debug logging
            // "abcd" (matches abcd**** from our masked logs)
    );

    public boolean isDebugEnabled(String maskedApiKey) {
        if (maskedApiKey == null) return false;
        return debugApiKeyPrefixes.stream()
                .anyMatch(prefix -> maskedApiKey.startsWith(prefix));
    }
}

Then in your filter, when debug is enabled, log the full request body:

if (debugConfig.isDebugEnabled(maskedApiKey)) {
    String body = new String(wrappedRequest.getContentAsByteArray(), StandardCharsets.UTF_8);
    log.debug("MCP_DEBUG_REQUEST | requestId={} | body={}", requestId, body);
}

This keeps your logs clean for normal use, but when you need to debug, you have all the details. I can't tell you how many times this saved me hours—users would send me their request ID, I can pull up the full request and see exactly what went wrong.

Step 4: What I Learned About Log Levels

Okay, so here's another mistake I made: I used DEBUG for most of this MCP logging, and kept my app at INFO level. That means all this useful stuff wasn't even in my production logs!

Now I use:

  • INFO: All the request start/complete, tool call counts, sizes. This is lightweight—each request is two log lines, negligible overhead. This stays on in production always.
  • WARN: Slow requests, large parameters/responses, retries. Gets your attention when something is starting to go wrong.
  • ERROR: Actual exceptions, obviously.
  • DEBUG: Full request/response bodies, only enabled per-key when debugging.

Before this change, I'd get user reports and check my logs—nothing. Because the useful stuff was at DEBUG. After this change, the answer is usually in the first five lines I look at.

Pros & Cons Of This Approach

Let me be honest—this isn't perfect. Here's what works and what doesn't:

Pros

  1. Catches proxy/network issues: You see when requests arrive but never complete, or when responses get truncated. Your old logging setup would never see these.
  2. Finds problems before users do: The slow request warning and large response warning lets you fix issues before enough people hit them.
  3. Structured logging works with everything: All these logs are key-value, so they parse nicely with ELK, Datadog, whatever you use.
  4. Low overhead: Just a couple extra log lines per request—nothing noticeable on a personal MCP server.
  5. Security friendly: You don't log full API keys, and bodies are only logged when explicitly enabled.

Cons

  1. More log volume: If you get hundreds of requests a day, this does add some lines. But for most side project MCP servers, it's totally fine. I have ~50 requests/day, it's just a couple KB extra per day.
  2. Requires wrapper filters: The content caching wrapper adds a tiny bit of memory overhead. Again, totally fine for small scale.
  3. You still need to check the proxy logs: This catches the problem, but sometimes the root cause is in your reverse proxy. But at least now you know that there is a problem.

What I Would Do Differently If I Started Over

Honestly, I'd start with this logging setup day one. I thought "it's just a couple endpoints, I don't need fancy logging". Wrong. MCP is still young, clients are evolving, proxies have weird defaults—you need visibility.

The three biggest surprises for me:

  1. Reverse proxies really do truncate large responses by default. Nginx's default proxy_buffer_size is 8k or 16k depending on your system. If your MCP tool returns a bigger response than that, boom—dead connection. And you'd never know from default app logging.
  2. MCP clients retry *hard*. If your app takes 11 seconds, the client retries at 10 seconds. So you get two requests for the same thing, and if you don't log timing, you don't see why.
  3. Large parameters happen more than you think. When AI clients send conversation context along with the tool call, parameters can get huge quickly. You need to see when that's happening.

My Final Production Logging Cheat Sheet

Here's my quick checklist I go through before deploying any new MCP server now:

  • [ ] Log every request at filter entry before authentication
  • [ ] Log request ID, start time, API key present, remote address
  • [ ] Log after completion: status, duration, response size
  • [ ] Warn on requests slower than 10s (client default timeout)
  • [ ] Warn on responses larger than 100KB
  • [ ] Log tool name and parameter size for tool calls
  • [ ] Keep full body logging optional per API key
  • [ ] Mask API keys, never log full secrets
  • [ ] Check your reverse proxy buffer size is at least 256KB

That's it. Eight simple checks that would have saved me 8 hours of debugging.

Wrapping Up

MCP changes how you think about a lot of things—logging is one of them. Your regular REST API logging works fine for regular APIs, but MCP has different failure modes that trip you up if you're not ready.

You don't need anything fancy. Just a few extra lines of logging in the right places catches 90% of the weird silent failures that drive you crazy.

The full Papers project with my MCP server implementation is on GitHub if you want to see the full working code: https://github.com/kevinten10/Papers

Have you built an MCP server in production? What weird logging issues have you run into? I'd love to hear your stories in the comments—maybe we can collect more of these hard-earned lessons.