惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
News and Events Feed by Topic
S
Schneier on Security
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Scott Helme
Scott Helme
V
Vulnerabilities – Threatpost
Cyberwarzone
Cyberwarzone
C
Cybersecurity and Infrastructure Security Agency CISA
Latest news
Latest news
Google Online Security Blog
Google Online Security Blog
Google DeepMind News
Google DeepMind News
K
Kaspersky official blog
Forbes - Security
Forbes - Security
T
Tenable Blog
The Last Watchdog
The Last Watchdog
T
Tor Project blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Simon Willison's Weblog
Simon Willison's Weblog
Project Zero
Project Zero
O
OpenAI News
L
LINUX DO - 热门话题
P
Privacy International News Feed
月光博客
月光博客
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Apple Machine Learning Research
Apple Machine Learning Research
C
Cyber Attacks, Cyber Crime and Cyber Security
量子位
博客园 - 【当耐特】
罗磊的独立博客
T
Threatpost
Application and Cybersecurity Blog
Application and Cybersecurity Blog
C
Cisco Blogs
S
SegmentFault 最新的问题
博客园 - 三生石上(FineUI控件)
N
News | PayPal Newsroom
腾讯CDC
Security Latest
Security Latest
J
Java Code Geeks
L
LINUX DO - 最新话题
N
Netflix TechBlog - Medium
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
The Exploit Database - CXSecurity.com
L
Lohrmann on Cybersecurity
D
Docker
Spread Privacy
Spread Privacy
S
Security @ Cisco Blogs
A
Arctic Wolf
H
Hacker News: Front Page
Help Net Security
Help Net Security
Recorded Future
Recorded Future
V2EX - 技术
V2EX - 技术

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Building a DevOps Incident Investigator with Coral SQL — From 15 Minutes to 15 Seconds
Khadirullah Mohammad · 2026-05-31 · via DEV Community

🏴‍☠️ Built for the Pirates of the Coral-bean hackathon by WeMakeDevs | May 25–31, 2026

TL;DR

Built a DevOps Incident Investigator using Coral SQL that correlates GitHub PRs, Sentry incidents, and Slack incident context using a single SQL query.

Coral turns operational debugging into a single SQL query across distributed systems.

Results:

  • 📉 Incident triage reduced from ~15 minutes to ~15 seconds.
  • 🪸 Custom Coral source spec created for internal APIs.
  • 🤖 AI-generated root cause analysis and Slack alerts.
  • 💻 Includes both a CLI and a Web Dashboard.

Built in 4 days for the Pirates of the Coral-bean hackathon.

The Problem

As a DevOps engineer, I've experienced the pain of incident investigation firsthand — switching between GitHub, Sentry, and Slack at 2 AM trying to figure out what broke production.

  • GitHub — "Which PR was deployed last?"
  • Sentry — "What errors are spiking?"
  • Slack — "What's the team saying?"

That's 3 tabs, 3 APIs, and 15 minutes of context-switching before you even understand what happened. This is what modern incident response looks like without unified observability.

The Solution

When I discovered Coral, an open-source tool that lets you query any API with SQL, I knew exactly what to build.

One command. Three sources. Full incident picture.

python3 investigator.py --query all --owner khadirullah --repo demo-payment-api

Enter fullscreen mode Exit fullscreen mode

The Magic: Cross-Source JOINs

Coral allows querying GitHub and Sentry together in a single statement. For this prototype, I use deployment timing to surface likely suspect PRs that were merged around the same time incidents first appeared.

SELECT p.number AS pr_number, p.title AS pr_title,
       p.user__login AS pr_author, p.merged_at,
       i.short_id AS sentry_id, i.title AS error_title,
       i.level AS error_level, i.first_seen AS error_first_seen
FROM github.pulls p
JOIN sentry.issues i ON p.merged_at IS NOT NULL
WHERE p.owner = 'khadirullah' AND p.repo = 'demo-payment-api'
  AND p.state = 'closed'
ORDER BY i.first_seen DESC, p.merged_at DESC LIMIT 15

Enter fullscreen mode Exit fullscreen mode

🔥 This is the money shot. PRs merged on the same day as errors appeared, side by side. "PR #820 merged at 12:18 PM, errors started at 12:20 PM" — instant suspect identification.

Why I Chose Coral (and Why It's a Game-Changer)

Normally, if you want to pull data from GitHub, Sentry, and Slack, you end up installing multiple SDKs (PyGithub, sentry-sdk, slack-sdk), learning different APIs, handling pagination and retries yourself, and writing glue code to correlate everything manually.

(Note: SDKs like sentry-sdk are excellent for emitting telemetry from applications, but they are still siloed when it comes to querying and correlating operational data across multiple platforms during an incident.)

What Coral Replaced

Before Coral:

GitHub SDK ──┐
Sentry SDK ──┤──→ Custom Glue Code + Pagination + Rate Limiting
Slack SDK  ──┘

Enter fullscreen mode Exit fullscreen mode

After Coral:

Coral SQL ──→ SELECT * FROM github JOIN sentry JOIN slack

Enter fullscreen mode Exit fullscreen mode

By treating the entire operational stack as a unified query layer, I was able to build my agent using far fewer SDKs and almost no glue code. Coral handles the authentication, pagination, and schema mapping locally, allowing me to focus on the actual business logic: writing cross-source JOINs and generating AI analysis.

Day 1: Setting Up Coral & Connecting Sources

Installing Coral

curl -fsSL https://withcoral.com/install.sh | sh
coral --version
# coral 0.3.0+96d61f7

Enter fullscreen mode Exit fullscreen mode

One command. That's it.

Connecting GitHub

ℹ️ For the GitHub token, I created a classic PAT at github.com/settings/tokens with zero scopes selected. For public repos, you don't need any permissions — the token just bumps your API rate limit from 60 to 5,000 requests per hour.

GITHUB_TOKEN=ghp_XXXXX coral source add github

Enter fullscreen mode Exit fullscreen mode

362 tables connected instantly — issues, pull requests, commits, repos, actions.

Connecting Sentry

Finding the Sentry token wasn't obvious. Sentry doesn't have a simple "API Tokens" page — you create tokens through Custom Integrations:

  1. Settings → Custom Integrations → Create New Integration (Internal Integration)
  2. Name it coral-hackathon with minimal read-only permissions (Project: Read, Issue & Event: Read, Organization: Read)
  3. Copy the generated token and connect:
SENTRY_TOKEN=sntrys_XXXXX SENTRY_ORG=my-org coral source add sentry

Enter fullscreen mode Exit fullscreen mode

12 tables connected — events, issues, projects, deployments.

I verified the tables were available with a quick SELECT schema_name, table_name FROM coral.tables.

Connecting Slack

Slack was the trickiest. Coral's pre-filled link creates a Slack app, but it only sets up User Token Scopes with PKCE — tokens aren't displayed in the UI.

⚠️ The fix: Add Bot Token Scopes instead of User Token Scopes. In OAuth & Permissions, scroll to Bot Token Scopes and add: channels:history, channels:read, groups:history, groups:read, users:read. Then reinstall the app — a Bot User OAuth Token (xoxb-...) appears!

coral source add --interactive slack
# Paste xoxb-... token when prompted

Enter fullscreen mode Exit fullscreen mode

All Sources Connected! 🎉

Source Tables What It Gives Us
GitHub 362 PRs, commits, issues, actions
Sentry 12 Errors, events, projects
Slack 2 Channels, users
Total 376 All queryable with SQL

Architecture

Here's how the entire system fits together:

┌──────────────────────────────────────────────────┐
│                  User Interfaces                  │
│  🖥️ CLI (investigator.py)  │  🌐 Web Dashboard   │
└──────────┬─────────────────────────┬─────────────┘
           │                         │
           ▼                         ▼
┌──────────────────────────────────────────────────┐
│              Backend (app.py)                     │
│  Flask API │ 🔒 Token Security │ 📦 Demo Data    │
│            │ 🤖 Gemini Flash   │ 💬 NL-to-SQL    │
└──────────────────────┬───────────────────────────┘
                       │
                       ▼
┌──────────────────────────────────────────────────┐
│              🪸 Coral SQL Layer                   │
│                  coral sql                        │
└───┬──────────┬──────────┬──────────┬─────────────┘
    ▼          ▼          ▼          ▼
 GitHub     Sentry      Slack    Payment API
 362 tbl    12 tbl      2 tbl    Custom Spec

Enter fullscreen mode Exit fullscreen mode

Two interfaces (CLI + Web Dashboard) sit on top of Coral SQL, which unifies 4 data sources into a single query layer. The AI layer uses Gemini Flash for root cause analysis and natural language → SQL generation.


Day 2: Writing SQL Queries & Building the CLI

Generating Test Data for Sentry

First roadblock: Sentry was empty. I wrote generate_errors.py to send 10 realistic DevOps errors.

Errors include: ConnectionError (PostgreSQL max connections), MemoryError (OOM kill), RuntimeError (K8s CrashLoopBackOff), TimeoutError (30s timeout), and more.

The SQL Queries

Query 1: Deployments — "What was recently deployed?"

SELECT number, title, user__login AS author, merged_at
FROM github.pulls
WHERE owner = 'khadirullah' AND repo = 'demo-payment-api'
  AND merged_at IS NOT NULL
ORDER BY merged_at DESC LIMIT 10

Enter fullscreen mode Exit fullscreen mode

Query 2: Incidents — "What errors are happening?"

SELECT short_id, title, level, count AS event_count, first_seen, last_seen
FROM sentry.issues
ORDER BY last_seen DESC LIMIT 10

Enter fullscreen mode Exit fullscreen mode

Query 3: Correlation — "Suspect deployment identification"

(As shown in the introduction, this cross-source JOIN correlates GitHub PRs with Sentry errors based on timestamps).

Query 4: Team Overview — Another cross-source JOIN:

SELECT u.name AS username, u.real_name, u.is_admin,
       c.name AS channel_name, c.num_members
FROM slack.users u CROSS JOIN slack.channels c
WHERE u.deleted = false AND c.is_archived = false

Enter fullscreen mode Exit fullscreen mode

Building the CLI

The CLI wraps everything in a clean, colorful interface with zero pip dependencies — Python stdlib only (subprocess, argparse, json, urllib):

python3 investigator.py --query all \
  --owner khadirullah --repo demo-payment-api \
  --slack-token $SLACK_TOKEN

Enter fullscreen mode Exit fullscreen mode

Troubleshooting Gotchas

⚠️ Sentry Project ID: My first query failed with "Invalid project parameter. Values must be numbers." I was using the slug (python) but Sentry wants the numeric ID. Fixed with:

SELECT id, slug, name FROM sentry.projects

⚠️ Slack Bot Permissions: Got not_in_channel error when fetching messages. The bot had channels:history but wasn't in the channel. Then hit missing_scope — needed channels:join scope. Quick fix: add scope → reinstall app → copy new token.


Day 3: Web Dashboard + AI Integration

Why Build a Dashboard?

A CLI is functional, but judges have 5 minutes. They need to see the data. I built a Flask dashboard with a dark glassmorphism design.

The Dashboard

Features:

  • Stats Row — Live counters for deployments, incidents, correlations, risky PRs
  • Incident Timeline — Horizontal scrolling event sequence
  • Deployments Table — Merged PRs with authors and timestamps
  • Sentry Incidents — Severity-colored cards with "🤖 Analyze" button
  • Correlation View — Visual PR ↔ Error mapping
  • Slack Messages — Chat-style from #incidents
  • Risky PRs — Risk-scored with green/yellow/red bars
  • Team Overview — Member cards from Slack users × channels JOIN

AI Root Cause Analysis

Click 🤖 Analyze on any Sentry error.

For example, PROD-41A (PostgreSQL max connections):

Root Cause: PR #487 introduced a new connection pooling layer that eagerly opens connections on pod startup. With 4 pods each opening 25 connections, the default max_connections=100 limit is immediately exhausted.

Immediate Fix: 1) Roll back PR #487, 2) Increase max_connections to 200, 3) Restart affected pods

Confidence: 94% (Coral system score)

(Note: This example is generated from demo data and illustrates the style of analysis produced by the assistant.)

In live mode with a Gemini API key, it calls Gemini Flash in real-time. Pre-computed demos ensure instant results without API waits.

Settings & Token Security

  • All 4 API tokens encrypted with Fernet symmetric encryption before writing to disk
  • "Delete All Tokens" does a 3-pass random overwrite
  • Demo → Live toggle with per-panel badges

Demo Mode

The dashboard starts in Demo Mode — pre-loaded data, zero setup. Switching to live is 4 clicks:

  1. ⚙️ Settings → enter tokens
  2. 💾 Save
  3. Click DEMO badge → LIVE
  4. 🔄 Refresh

If a live call fails, panels gracefully fall back to demo data.


Day 4: Competitive Upgrades — Going Beyond

Custom Source Spec: Querying Internal APIs

Enterprises don't just use public SaaS tools. A real incident investigator needs to query internal microservices. This is where Coral's Custom Source Specs shine.

I built a companion project (demo-payment-api) and wrote a YAML spec to teach Coral how to talk to it:

name: payment_api
version: 0.1.0
dsl_version: 3
backend: http
base_url: "http://localhost:5001"

auth:
  type: HeaderAuth
  headers:
    - name: Authorization
      template: "Bearer {{input.PAYMENT_API_TOKEN}}"

tables:
  - name: health
    request:
      method: GET
      path: /api/health
    columns:
      - name: status
        type: Utf8
        expr: { kind: path, path: [status] }
      - name: response_time_ms
        type: Int64
        expr: { kind: path, path: [response_time_ms] }

Enter fullscreen mode Exit fullscreen mode

The YAML acts as a "translator" — it tells Coral how to map SQL concepts (tables, columns) to HTTP concepts (endpoints, JSON paths):

User ──→ Coral: SELECT status FROM payment_api.health
Coral ──→ YAML: Look up "health" table
YAML ──→ Coral: GET /api/health
Coral ──→ API: HTTP GET http://localhost:5001/api/health
API ──→ Coral: {"status": "healthy", "response_time_ms": 42}
Coral ──→ User: | status  | response_time_ms |
                | healthy |       42         |

Enter fullscreen mode Exit fullscreen mode

Register and query:

# Lint the spec
coral source lint coral-config/payment-api.yaml

# Add to Coral
PAYMENT_API_TOKEN=mock_123 coral source add --file coral-config/payment-api.yaml

# Query like a database!
coral sql "SELECT * FROM payment_api.health"

Enter fullscreen mode Exit fullscreen mode

ℹ️ This proves the tool can connect to any internal enterprise service — not just the big SaaS providers. Read more in our dedicated Chart New Waters deep-dive.

Natural Language to SQL (/api/ask)

Type a question in plain English → AI generates Coral SQL → executes it → returns results:

User: "Show me critical errors from last week"
  → Gemini Flash
  → SELECT * FROM sentry.issues WHERE level = 'error' LIMIT 10
  → Coral SQL Engine
  → Results Table

Enter fullscreen mode Exit fullscreen mode

The backend feeds the live Coral schema (all tables + columns) to Gemini, so it generates accurate SQL every time.

Automated Slack Alerts

Investigation is only half the battle — communication is the other half. After AI generates a root cause analysis, users can click "📢 Send to Slack" to push the report directly to #incidents:

🚨 Sentry Error → 🤖 AI Analysis → 📢 Send to Slack → #incidents channel → Team sees report

Enter fullscreen mode Exit fullscreen mode

If the token only has read permissions, the UI catches the missing_scope error and gracefully prompts the user to add chat:write scope.

Handling Slack Messages (TVF vs Custom Fallback)

Coral exposes Slack messages via a function-like table interface (TVF), which requires passing the exact channel ID directly into the SQL query: SELECT * FROM slack.messages(channel => 'C12345678').

However, I needed something more robust for an automated dashboard. If the bot isn't already in the incident channel, the Slack API returns a strict not_in_channel error. Instead of failing, I deliberately built a custom Python fallback that catches this error, dynamically forces the bot to join the channel, fetches the messages, and then uses Coral to pull the slack.users table to map the raw User IDs (e.g., UXXXXXXX) to real human names.

Fetching the Team Overview

With the messages handled resiliently, I still needed to display the active Incident Response team. Instead of making separate API calls for users and channels, I let Coral grab the entire team landscape in a single round-trip using a CROSS JOIN:

SELECT u.name AS username, u.real_name, u.is_admin,
       c.name AS channel_name, c.num_members
FROM slack.users u CROSS JOIN slack.channels c
WHERE u.deleted = false AND c.is_archived = false

Enter fullscreen mode Exit fullscreen mode

A quick Python iteration separates the results, giving us the full team and channel rosters instantly without juggling multiple Slack SDK requests.


Why Sentry (Not Just SonarQube & Trivy)

A question I got: "Why do you need Sentry if you have SonarQube and Trivy?" The answer:

Before Deploy:  SonarQube → "Code COULD break"
                Trivy     → "Has known CVEs"
                      ↓
                   Deploy
                      ↓
After Deploy:   Sentry    → "App JUST broke for 1,203 users"

Enter fullscreen mode Exit fullscreen mode

Tool Stage Catches
SonarQube Pre-deploy (CI) Code smells, potential bugs
Trivy Pre-deploy (CI) Known CVEs in dependencies
Sentry Post-deploy (Runtime) Real crashes, right now, with stack traces + user impact

The Incident Investigator bridges all three stages — correlating code changes (GitHub) with runtime errors (Sentry) and team communication (Slack).


Setting Up Slack Alerts (Webhooks vs Bot Tokens)

We use both approaches:

Feature Incoming Webhook Bot Token (xoxb-)
Direction App → Slack (one-way) App ↔ Slack (two-way)
Use case Post alerts Read messages + respond
Setup Just a URL OAuth scopes
Best for Simple alerting Complex integrations
  • Webhook → Demo Payment API posts error alerts to #incidents
  • Bot Token → Incident Investigator reads messages from #incidents via Coral SQL

The complete pipeline:

Error in Payment API
  ├──→ Sentry captures exception ──→ Coral SQL queries ──┐
  └──→ Slack webhook fires ──→ #incidents channel ──────┤
                                                         ├──→ Incident Investigator
  GitHub PRs ──────────────────────────────────────────┘       │
                                                               ▼
                                                        Gemini AI Analysis
                                                               │
                                                               ▼
                                                        📢 Push to Slack #incidents

Enter fullscreen mode Exit fullscreen mode


What I Learned

About Coral

  • Zero-scope GitHub tokens work — just bumps your rate limit for public repos
  • Sentry wants numeric project IDs — not slugs, query sentry.projects first
  • Slack's Coral source is limited — only channels and users, no messages. But the bot token works with direct API calls
  • Cross-source JOINs are the killer feature — GitHub × Sentry in one SQL statement is genuinely powerful

About DevOps Incident Response

The hardest part of incident response isn't fixing the problem — it's finding the right information. We spend more time context-switching than debugging. The investigator answers three questions fast:

  1. What was deployed recently?
  2. What errors appeared after deployment?
  3. What's the team saying about it?

That's 80% of the first 15 minutes of any incident.

About Hackathons

  • Ship fast, iterate later — the Discord advice was spot-on
  • Document as you go — writing the blog alongside the code was more efficient
  • Errors are content — every missing_scope became a blog section
  • Keep scope tight, then expand — CLI first, dashboard second, AI third

Limitations

  • Correlation heuristic: GitHub ↔ Sentry correlation is currently based on deployment timing, not deterministic tracing.
  • AI as an assistant: Root cause analysis is AI-assisted and should be treated as a hypothesis, not absolute truth.
  • Slack capabilities: Slack messages are retrieved through a Python fallback integration because the native Coral Slack source currently exposes only users and channels.
  • API rate limits: In a high-traffic live incident, direct API calls to Slack and GitHub could hit rate limits, requiring an intermediate caching layer.
  • Schema mismatches: Edge cases in custom internal APIs may occasionally fail to map perfectly to Coral's static YAML types without custom data coercions.

The Result

In testing, the workflow reduced incident investigation from roughly 15 minutes of manual context-switching to less than 15 seconds for an initial deployment-to-error correlation query.

The Impact

Before (The Old Way):

  • ❌ Open GitHub to check recent PRs
  • ❌ Open Sentry to check recent errors
  • ❌ Open Slack to read team discussions
  • ❌ Manually correlate timestamps across 3 tabs

After (The Coral Way):

  • ✅ Run one command or view one dashboard
  • ✅ Instantly see PRs and Errors side-by-side
  • ✅ Get AI-generated root cause analysis

Technical Achievements

Feature Details
6 SQL Queries Deployments, incidents, correlation, risky PRs, team, health
2 Cross-Source JOINs GitHub × Sentry, Slack users × channels
1 Custom Source Spec payment-api.yaml for internal microservice
AI Analysis Gemini-powered root cause + fix suggestions
NL-to-SQL Ask questions in English, get Coral SQL results
2 Interfaces CLI (zero deps) + Web Dashboard (Flask)
Security Fernet symmetric encryption at rest
Docker + CI Production-ready packaging + GitHub Actions

Try It

git clone https://github.com/khadirullah/devops-incident-investigator
cd devops-incident-investigator
pip install flask google-generativeai cryptography
python3 app.py
# Open http://localhost:5000 — works instantly with demo data!

Enter fullscreen mode Exit fullscreen mode

🔗 View on GitHub

Future Work

  • Direct Release Correlation: Correlate Sentry releases directly to GitHub commits using commit SHAs.
  • More Sources: Add Grafana and Kubernetes log sources to the SQL engine.
  • Automated Postmortems: Use the AI layer to generate and publish full incident postmortems automatically.

Built with Coral for the Pirates of the Coral-bean hackathon by WeMakeDevs.

Follow the journey: khadirullah.com | GitHub


Originally published on khadirullah.com