惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 【当耐特】
阮一峰的网络日志
阮一峰的网络日志
S
Secure Thoughts
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Tenable Blog
T
Tailwind CSS Blog
WordPress大学
WordPress大学
宝玉的分享
宝玉的分享
Webroot Blog
Webroot Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
The Cloudflare Blog
T
Threat Research - Cisco Blogs
V
Visual Studio Blog
Jina AI
Jina AI
V
V2EX
I
InfoQ
Latest news
Latest news
P
Proofpoint News Feed
T
Threatpost
Engineering at Meta
Engineering at Meta
P
Proofpoint News Feed
美团技术团队
The Register - Security
The Register - Security
L
LangChain Blog
Apple Machine Learning Research
Apple Machine Learning Research
aimingoo的专栏
aimingoo的专栏
GbyAI
GbyAI
Cloudbric
Cloudbric
Microsoft Azure Blog
Microsoft Azure Blog
C
Cisco Blogs
U
Unit 42
Microsoft Security Blog
Microsoft Security Blog
MyScale Blog
MyScale Blog
V
Vulnerabilities – Threatpost
TaoSecurity Blog
TaoSecurity Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Recent Commits to openclaw:main
Recent Commits to openclaw:main
W
WeLiveSecurity
博客园 - 司徒正美
T
The Exploit Database - CXSecurity.com
小众软件
小众软件
Y
Y Combinator Blog
Recent Announcements
Recent Announcements
量子位
酷 壳 – CoolShell
酷 壳 – CoolShell
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园_首页
N
News and Events Feed by Topic

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
From Stack Trace to Suggested Fix in 4 Seconds: Building a Self-Healing .NET API Gateway.
Avinash Zala · 2026-06-22 · via DEV Community

Last Tuesday my API gateway caught a NullReferenceException, streamed it to a dashboard in real-time, and pushed a draft code fix to the browser tab of the on-call engineer — before I finished reading the error myself. That sentence used to be vendor marketing. Now it's just my Program.cs.

This is the architecture post-mortem. I built it on weekends. It runs in Docker. It cost me exactly $0 in LLM credits during development because Groq's free tier is generous and Ollama works as a swap-in. The repo is here — issues and PRs welcome.

The problem most .NET teams have

Production errors are caught, logged to a file, and forgotten. Engineers find out from a Slack ping twenty minutes later, if at all. By the time someone looks, the original request context is gone, the user's session has expired, and the stack trace is buried four layers deep in System.* calls.

"Self-healing" is a word vendors use to mean "auto-restart the pod." I wanted something better. The actual ask:

When an exception is thrown in service A, give the engineer (a) a clear root cause, (b) a suggested fix, and (c) a draft code patch — in under 30 seconds.

Not a magic black box. Not an auto-applied patch. Just: catch the error, give the model the right context, push the analysis to a human in real-time, and let the human close the loop.

The architecture

One .NET solution, four projects, four NuGet packages, no new infrastructure beyond what you probably already have.

[ HTTP request ]
       |
       v
+-------------------+        enqueue         +---------------------+
| SmartLogAnalyzer. | ---------------------> |  Hangfire (Redis)   |
|      Api          |                        +----------+----------+
|  (ErrorHandling   |                                   |
|   Middleware)     |                                   v
+-------------------+                        +---------------------+
                                                | SmartLogAnalyzer.   |
                                                |     Worker          |
                                                | (ErrorProcessingWorker)
                                                +-----+-------+-------+
                                                      |       |
                                          AI call     |       |  persist
                                                      v       v
                                            +-----------+   +-----------+
                                            | Semantic  |   | MSSQL     |
                                            | Kernel +  |   | (ErrorLog |
                                            | Groq LLM  |   |  table)   |
                                            +-----+-----+   +-----------+
                                                  |
                                                  v
                                        +---------------------+
                                        |  SignalR Hub        |
                                        |  (ErrorHub)         |
                                        +----------+----------+
                                                   |
                                              broadcast
                                                   v
                                        +---------------------+
                                        |  React Dashboard    |
                                        |  (live update)      |
                                        +---------------------+

The crucial detail is where the AI call happens. It does not happen in the request thread. The middleware returns the 500 in milliseconds; the AI work happens inside a Hangfire background job, in a different process, possibly on a different machine. Two different response times, one user.

Part 1 — the capture

The middleware is fifty lines including the using statements. Here is the whole thing.

using Hangfire;
using SmartLogAnalyzer.Core.Models;
using SmartLogAnalyzer.Core.Workers;
using System.Text.Json;

namespace SmartLogAnalyzer.Api.Middleware
{
    public class ErrorHandlingMiddleware
    {
        private readonly RequestDelegate _next;
        private readonly IBackgroundJobClient _backgroundJobClient;

        public ErrorHandlingMiddleware(
            RequestDelegate next,
            IBackgroundJobClient backgroundJobClient)
        {
            _next = next;
            _backgroundJobClient = backgroundJobClient;
        }

        public async Task InvokeAsync(HttpContext context)
        {
            try
            {
                await _next(context);
            }
            catch (Exception ex)
            {
                await HandleExceptionAsync(context, ex);
            }
        }

        private async Task HandleExceptionAsync(HttpContext context, Exception ex)
        {
            context.Response.ContentType = "application/json";
            context.Response.StatusCode = 500;

            var errorLog = new ErrorLog
            {
                ErrorMessage = ex.Message,
                StackTrace   = ex.StackTrace ?? string.Empty,
                RoutePath    = context.Request.Path
            };

            // The line that does the work. Enqueue is non-blocking;
            // the response is sent before the AI is ever called.
            _backgroundJobClient.Enqueue<ErrorProcessingWorker>(
                worker => worker.ProcessErrorAsync(errorLog));

            var result = JsonSerializer.Serialize(
                new { error = "An internal error has been logged and is being analyzed." });
            await context.Response.WriteAsync(result);
        }
    }
}

Two things to notice.

First, the Enqueue call returns immediately. Hangfire's IBackgroundJobClient is a thin proxy over the Hangfire storage (Redis in this case) and a worker pickup. We don't await an AI call here. The user gets their 500 in single-digit milliseconds.

Second, the response body — "An internal error has been logged and is being analyzed." — is itself a feature. The user (or the calling frontend) now knows the error is being handled. It is not a lie, it is a contract.

Part 2 — the worker

The ErrorProcessingWorker is a plain C# class. Hangfire instantiates it from the DI container, calls ProcessErrorAsync, and (if it throws) retries up to three times with exponential backoff.

[AutomaticRetry(Attempts = 3)]
public async Task ProcessErrorAsync(ErrorLog errorLog)
{
    // 1. Hash the stack trace to dedupe identical errors
    var stackTraceHash = ComputeHash(errorLog.StackTrace);

    // 2. If we've seen this stack trace in the last 24h, just bump the count
    if (await _redisCacheService.KeyExistsAsync(stackTraceHash))
    {
        var existingLog = await _errorLogRepository.AddOrUpdateErrorLogAsync(errorLog);
        await _hubContext.Clients.All.SendAsync(
            "ReceiveErrorUpdate",
            JsonSerializer.Serialize(existingLog));
        return;
    }

    // 3. New error — claim the hash so duplicates skip the AI call
    await _redisCacheService.SetKeyAsync(stackTraceHash, "1", TimeSpan.FromHours(24));

    // 4. Ask the LLM
    var analyzedLog = await _aiAnalysisService.AnalyzeErrorAsync(errorLog);

    // 5. Persist + push to the dashboard
    var savedLog = await _errorLogRepository.AddOrUpdateErrorLogAsync(analyzedLog);
    await _hubContext.Clients.All.SendAsync(
        "ReceiveErrorUpdate",
        JsonSerializer.Serialize(savedLog));
}

private string ComputeHash(string input)
{
    using var md5 = MD5.Create();
    var hashBytes = md5.ComputeHash(Encoding.UTF8.GetBytes(input));
    return BitConverter.ToString(hashBytes).Replace("-", "").ToLowerInvariant();
}

The Redis dedupe step is the difference between a $0 demo and a $200 Groq bill. The first occurrence of a NullReferenceException at /api/users/{id} costs one LLM call. The next 10,000 occurrences cost nothing. The 24-hour TTL is a knob you will want to tune.

Part 3 — the AI call

I started this project with a fancy design: a Semantic Kernel KernelPlugin that would have the AI fetch the offending source file from GitHub, look at the test that covers it, and then propose a diff grounded in real code. It was clever. It was also over-engineered for v1.

The version that shipped is fifteen lines.

var prompt = $@"
You are a Senior .NET Engineer. Analyze the following error
and provide a JSON response with exactly three keys:
RootCause, FixSuggestion, and CodePatch.

Error Message: {errorLog.ErrorMessage}
Stack Trace: {errorLog.StackTrace}

JSON Response:
";

var result = await _kernel.InvokePromptAsync(prompt);
var responseText = result.ToString();

Then I parse the response into the three fields with a JsonDocument, and on failure, fall back to a hand-rolled string parser. We will get back to that parser in the next section — it is the most important code in the whole project and also the part I am least proud of.

Why Semantic Kernel if the call is this simple? Two reasons.

  1. Provider swap. The Groq wire-up is one line: AddOpenAIChatCompletion(modelId: "llama-3.3-70b-versatile", apiKey: [your-groq-key], endpoint: new Uri("https://api.groq.com/openai/v1")) — where the key is loaded from .env at startup. Swapping to OpenAI, Azure OpenAI, or local Ollama is one constructor call. If I had called Groq directly via HttpClient, I would be rewriting the call site for every provider I tried.
  2. Built-in retries and timeouts. Kernel.InvokePromptAsync handles 429s and 5xxs with a default policy. That is one less thing to get wrong.

You can absolutely build this with raw HttpClient and chat.completions.create(). You will write the retry logic yourself. I have done that. I do not recommend it.

Part 4 — what I got wrong

This is the part you came for. Five things that bit me, in order of how much they cost.

4.1 The JSON parser that almost shipped

First version of the response parser used JsonDocument.Parse and threw on any malformed output. About 15% of Groq responses came back wrapped in

json ...

markdown fences, despite the prompt saying "JSON Response:" right at the end. I added a stripper:

var cleaned = responseText.Trim();
if (cleaned.StartsWith("```
{% endraw %}
json")) cleaned = cleaned.Substring(7);
if (cleaned.StartsWith("
{% raw %}
```"))    cleaned = cleaned.Substring(3);
if (cleaned.EndsWith("```
{% endraw %}
"))      cleaned = cleaned.Substring(0, cleaned.Length - 3);
cleaned = cleaned.Trim();
{% raw %}

That fixed 90% of it. The other 10% needed a hand-rolled regex parser that walks the string looking for "RootCause": "..." and respects backslash escapes. Do not be too proud to write a regex parser. When the upstream is an LLM and the contract is "please return JSON," the LLM is sometimes wrong and you need a fallback.

The pattern in AiAnalysisService.cs is the right one: try the strict parser, catch the exception, try the lenient one, and only then give up and store the raw text with a "Failed to parse" flag. The dashboard renders the raw text anyway, so the engineer still gets value.

4.2 Sensitive data leakage

A stack trace can contain connection strings, JWTs, or PII. The first version sent the raw exception text to Groq. After a code review from a friend who is more paranoid than I am, I added a redaction step before the AI call.


csharp
private static readonly Regex BearerToken  = new(@"Bearer\s+[A-Za-z0-9._\-]+", RegexOptions.Compiled);
private static readonly Regex PasswordKV   = new(@"(password|pwd|secret)\s*=\s*\S+",  RegexOptions.Compiled | RegexOptions.IgnoreCase);
private static readonly Regex CreditCard   = new(@"\b\d{16}\b",                       RegexOptions.Compiled);
private static readonly Regex EmailAddr    = new(@"\b[\w._%+-]+@[\w.-]+\.[A-Za-z]{2,}\b", RegexOptions.Compiled);

public static string Redact(string input)
{
    input = BearerToken.Replace(input, "Bearer [REDACTED]");
    input = PasswordKV .Replace(input, "$1=[REDACTED]");
    input = CreditCard .Replace(input, "[REDACTED-CC]");
    input = EmailAddr  .Replace(input, "[REDACTED-EMAIL]");
    return input;
}


Run this before the prompt is built, every time. Always assume the AI provider sees your data. Always. The day you forget is the day a customer's JWT ends up in someone else's training set, or at minimum in someone else's logs.

4.3 The "self-healing" promise is misleading

This system suggests fixes. It does not apply them. I almost shipped an "auto-apply patch on green confidence" toggle. Then I imagined a 3 AM page where a hallucinated regex wipes a production database because the model misread a column name. The toggle is gone. Auto-merging AI patches into prod is a 2027 problem, not a 2026 one. Be honest about this in your README, your marketing, and your internal pitches. Engineers will trust you more.

4.4 Hangfire retries are silent (and cost money)

If the AI call times out, Hangfire retries it. If the AI call consistently times out — bad prompt, big payload, network blip — Hangfire retries it three times. Each retry costs a Groq credit. The [AutomaticRetry(Attempts = 3)] attribute is the default, and the default is wrong for any external dependency that costs money.

Fix: lower the count, add delay, and add a circuit breaker. This is what I have on the worker method now:


csharp
[AutomaticRetry(Attempts = 2, DelaysInSeconds = new[] { 30, 120 })]
public async Task ProcessErrorAsync(ErrorLog errorLog) { ... }


Two attempts, with 30s and 2m delays. That bounds the cost spiral when something is wrong. A truly broken state would still cost 2x per error, but it would not retry 5 more times in a tight loop and drain a month's budget in an hour.

4.5 No correlation between dashboard event and the original request

The user got a 500 with no error ID. The dashboard showed a fix suggestion with no way to find the request that caused it. So when an engineer wanted to reproduce the error, they had to guess the URL, the headers, the auth state. Useless.

Fix: generate a CorrelationId once in the middleware, return it in the response header, and store it on the ErrorLog model. One UUID, two places. The dashboard now shows #1234 next to each error and the engineer can grep their logs for that ID.


csharp
// in the middleware
var correlationId = Guid.NewGuid().ToString("N");
context.Response.Headers["X-Correlation-Id"] = correlationId;
errorLog.CorrelationId = correlationId;


Part 5 — the live dashboard

The dashboard is a 280-line React app in SmartLogAnalyzer.Dashboard/smart-log-analyzer-dashboard/. The whole real-time piece is fifty lines of hooks.


typescript
useEffect(() => {
  const newConnection = new signalR.HubConnectionBuilder()
    .withUrl(`${API_URL}/errorHub`)
    .withAutomaticReconnect()
    .build();
  setConnection(newConnection);
}, []);

useEffect(() => {
  if (!connection) return;
  connection.start().then(() => {
    setConnected(true);
    connection.on('ReceiveErrorUpdate', (errorJson: string) => {
      const error: ErrorLog = JSON.parse(errorJson);
      setErrors(prev => {
        const index = prev.findIndex(e => e.id === error.id);
        if (index !== -1) {
          const updated = [...prev];
          updated[index] = error;
          return updated;
        }
        return [error, ...prev];
      });
    });
  });
  return () => { connection.stop(); };
}, [connection]);


The wire is JSON-over-SignalR. The server-side hub does Clients.All.SendAsync("ReceiveErrorUpdate", jsonString) and every open browser tab updates. No polling. No refresh button. You literally watch errors arrive, get analyzed, and become fixable, in real-time.

A few small UX details I am proud of:

  • Severity badges. Errors with occurrenceCount >= 10 get a red 🔴 Critical chip. Under 2 is green. Engineers learn to scan for red.
  • "Analyzing with AI..." spinner. When a new error arrives, its card shows a spinner for the few seconds until the AI response comes back. The state machine is pending → analyzing → analyzed, driven by whether aiRootCause is set.
  • Expand on click, stack trace in a <details>. Most engineers want the AI's take first. The stack trace is one click away.

When you should — and shouldn't — build this

Build it if:

  • You have more than three services throwing exceptions, and your on-call rotation is a human who hates pages at 3 AM.
  • You are already paying for an LLM API, or you have a GPU sitting around running Ollama.
  • Your mean time to acknowledge (MTTA) on alerts is more than five minutes.

Don't build it if:

  • You have one service and a steady stream of bugs. Fix the bugs.
  • Your "errors" are mostly business-logic edge cases — a missing null check that is actually a missing requirement. The AI cannot help with those.
  • You don't have CI/CD yet. Self-healing on top of an unsafe deploy pipeline is just a faster way to break production.

The general rule: a 200-line NuGet package won't fix a 2000-line architecture problem. This system is a force multiplier on a healthy codebase. On an unhealthy one, it is a faster way to find out how unhealthy you are.

The repo and how to run it

The full source is at github.com/ZalaAvinash/Smart-Log-Analyzer-Self-Healing-API-Gateway. To run it locally:


bash
git clone https://github.com/ZalaAvinash/Smart-Log-Analyzer-Self-Healing-API-Gateway.git
cd Smart-Log-Analyzer-Self-Healing-API-Gateway
cp .env.example .env
# Edit .env — set GROQ_API_KEY (free at groq.com), DB_SERVER, REDIS_HOST
start-all.bat    # Windows; the repo has a Makefile-equivalent for *nix


Three windows open: the API, the Worker, the Dashboard. Open http://localhost:3000, click the "Trigger Test Error" link, and watch a NullReferenceException arrive, get analyzed, and become a clickable fix suggestion, all in under 4 seconds.

Closing

The future of "self-healing" is not magic. It is a small, honest pipeline: catch the error, give the model the right context, push the analysis to a human in real-time, let the human close the loop. The model writes the boilerplate diff. The engineer writes the actual fix. That is a real workflow, and it works today, on the same .NET you are already running, with one extra NuGet package and one extra process.

If you build something similar and run into the same five problems, I'd love to hear about it. The repo is open for issues, PRs, and rants about how your retry policy bankrupted your LLM budget. We've all been there.


Build with: .NET 10 · ASP.NET Core · Hangfire · Semantic Kernel · Groq (llama-3.3-70b-versatile) · SignalR · MSSQL · Redis · React

Repo: ZalaAvinash/Smart-Log-Analyzer-Self-Healing-API-Gateway

About the author: Avinash Zala is a senior .NET engineer in Surat, India, with 7+ years building enterprise web apps, APIs, and ERP systems. He is currently adding AI/LLM capabilities to his stack and writing about what he learns. GitHub · LinkedIn