惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 司徒正美
Blog — PlanetScale
Blog — PlanetScale
博客园 - 聂微东
月光博客
月光博客
量子位
大猫的无限游戏
大猫的无限游戏
Stack Overflow Blog
Stack Overflow Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
The Cloudflare Blog
P
Proofpoint News Feed
B
Blog RSS Feed
美团技术团队
腾讯CDC
C
Check Point Blog
Engineering at Meta
Engineering at Meta
F
Fortinet All Blogs
N
Netflix TechBlog - Medium
Recent Announcements
Recent Announcements
J
Java Code Geeks
S
SegmentFault 最新的问题
WordPress大学
WordPress大学
宝玉的分享
宝玉的分享

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
GitHub - Arnab758/ai-gateway
arnab777 · 2026-06-23 · via Hacker News - Newest: "AI"

Cut your LLM API costs by 40-70% with zero code changes.

A semantic caching layer that sits between your app and AI providers (OpenAI, Groq, etc.). When you ask a similar question twice, it returns the cached answer instantly instead of calling the API again.

🎯 What Problem Does This Solve?

You're building an AI app and your API bill is $500/month. 40-70% of that is for repeat questions:

  • "What is RAG?" asked 100 times = 100 API calls
  • "How do I reset my password?" asked 50 times = 50 API calls

With AI Gateway: Those 150 calls become 2 calls (one for each unique question). You save $200-350/month.

🚀 Deploy in 60 Seconds (3 Options)

Option 1: Railway (Recommended - Includes Redis)

Deploy to Railway

Steps:

  1. Click the button above
  2. Sign in with GitHub
  3. Enter your API key (Groq or OpenAI)
  4. Click "Deploy"
  5. Done! Your gateway is live at https://your-app.up.railway.app

What you get:

  • ✅ Hosted gateway (no server management)
  • ✅ Redis included (persistent cache)
  • ✅ Auto-scaling
  • ✅ HTTPS enabled
  • ✅ $5/month free credit

Option 2: Render (One-Click Deploy)

Deploy to Render

Steps:

  1. Click the button
  2. Sign in with GitHub
  3. Add environment variable: UPSTREAM_API_KEY=your_key
  4. Click "Create Web Service"
  5. Done!

Note: You'll need to add a Redis addon separately in Render dashboard.

Option 3: Docker (Self-Hosted)

Prerequisites:

  • Docker installed
  • Docker Compose installed
  • A Groq or OpenAI API key

Steps:

# 1. Clone the repo
git clone https://github.com/Arnab758/ai-gateway.git
cd ai-gateway

# 2. Set your API key
export UPSTREAM_API_KEY=gsk_your_groq_key_here

# 3. Start everything (gateway + Redis)
docker compose up -d

# 4. Verify it's running
curl http://localhost:8080/health

# Expected response: {"status":"ok"}

That's it! Your gateway is now running at http://localhost:8080

📖 How to Use

Basic Usage (cURL)

# Send a request through the gateway
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "X-Gateway-Token: my-app" \
  -H "Authorization: Bearer sk-your-openai-or-groq-key" \
  -d '{
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "What is RAG?"}]
  }'

# Send the SAME request again
# Response headers will show: X-Gateway-Cache: HIT
# You just saved money! 💰

Python Example

import requests

# Your gateway URL (from Railway/Render/Docker)
GATEWAY_URL = "https://your-app.up.railway.app"
API_KEY = "sk-your-key"

response = requests.post(
    f"{GATEWAY_URL}/v1/chat/completions",
    headers={
        "Content-Type": "application/json",
        "X-Gateway-Token": "my-app",
        "Authorization": f"Bearer {API_KEY}"
    },
    json={
        "model": "gpt-4",
        "messages": [{"role": "user", "content": "What is RAG?"}]
    }
)

print(response.json())

Node.js Example

const response = await fetch('https://your-app.up.railway.app/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'X-Gateway-Token': 'my-app',
    'Authorization': 'Bearer sk-your-key'
  },
  body: JSON.stringify({
    model: 'gpt-4',
    messages: [{ role: 'user', content: 'What is RAG?' }]
  })
});

const data = await response.json();
console.log(data);

🎮 Try the Interactive Demo

No API key needed! See how caching works:

👉 Open Live Demo

  • Type a prompt and click "Send" (simulation mode)
  • Or enter your API key and click "Test Real API" (real caching with Redis)
  • Try sending the same prompt twice to see cache hits!

🔥 Key Features

  • Semantic Caching - Matches similar questions, not just exact duplicates
    • "What is RAG?" = "Explain RAG" = "RAG definition"
  • Multi-Tenant - Each customer gets their own isolated cache
  • 4-Tier Matching:
    1. Exact match (100% identical)
    2. Template match ("weather in London" = "weather in Paris")
    3. Semantic match (similar meaning)
    4. Word overlap (partial matches)
  • Redis + In-Memory Fallback - Works with or without Redis
  • Request Deduplication - 100 concurrent identical requests = 1 API call
  • Rate Limiting - Prevent abuse per tenant
  • Circuit Breaker - Automatically stops calling if provider is down
  • Cost Tracking - See how much you saved

📊 Real-World Example

Scenario: Customer support chatbot with 10,000 users

Without AI Gateway:

  • 10,000 users ask 100 common questions each
  • 1,000,000 API calls/month
  • Cost: $500/month (at $0.0005/call)

With AI Gateway:

  • First 100 questions: 100 API calls (cache miss)
  • Next 9,900 users asking same questions: 0 API calls (cache hit)
  • Total: 100 API calls/month
  • Cost: $0.05/month
  • Savings: $499.95/month (99.99%)

Even with 30% unique questions:

  • 300,000 API calls
  • Cost: $150/month
  • Savings: $350/month (70%)

🛠️ Configuration

Edit gateway.yaml to customize:

cache:
  redis_url: "redis://localhost:6379"  # Or your Redis URL
  vector:
    enabled: true
    similarity_threshold: 0.85  # 85% similar = cache hit
  ttl_hours: 24  # Cache entries expire after 24 hours

rate_limiter:
  enabled: true
  max_requests: 60  # Per minute per tenant

📡 API Endpoints

Endpoint Method Description
/v1/chat/completions POST Main proxy endpoint with caching
/health GET Health check
/stats GET Cache statistics
/metrics GET Prometheus metrics

🔍 Monitoring

Check Cache Stats

curl http://localhost:8080/stats

Response:

{
  "uptime": 1234567890,
  "cache": {
    "local_index_entries": 150,
    "vector_dimensions": 128,
    "vector_threshold": 0.85,
    "jaccard_threshold": 0.75,
    "template_enabled": true,
    "dedup_enabled": true,
    "ttl_hours": 24
  }
}

Response Headers

Every response includes cache information:

X-Gateway-Cache: HIT          # or MISS
X-Gateway-Similarity: 0.95    # 95% similar (if HIT)
X-Gateway-Time-Saved: 1234ms  # Time saved (if HIT)

🐛 Troubleshooting

Problem: "Redis connection failed"

Solution: Redis is optional! The gateway will fall back to in-memory cache automatically. For production, add Redis:

Railway: Add Redis from the "New" button Render: Add Redis from the "New" → "Database" → "Redis" Docker: Already included in docker-compose.yml

Problem: "All upstream providers unavailable"

Cause: You're hitting rate limits on free tier (Groq/OpenAI)

Solutions:

  1. Wait 1-2 minutes and try again
  2. Upgrade to paid tier ($0.002/request vs free limits)
  3. Add your own API key with higher limits

Problem: "Rate limit exceeded"

Cause: Too many requests from one tenant

Solution: Increase rate limits in gateway.yaml:

rate_limiter:
  max_requests: 120  # Increase from 60
  window_minutes: 1

Problem: Cache not hitting

Cause: Prompts are too different

Solution: Lower the similarity threshold in gateway.yaml:

cache:
  vector:
    similarity_threshold: 0.75  # Lower from 0.85
  jaccard:
    threshold: 0.65  # Lower from 0.75

🏗️ Architecture

Your App → AI Gateway → [Cache Check] → Redis
                ↓
            [Cache HIT] → Return cached response (instant, $0)
                ↓
            [Cache MISS] → Call LLM Provider → Cache response → Return

🤝 Contributing

Contributions are welcome! Please:

  1. Fork the repo
  2. Create a feature branch
  3. Make your changes
  4. Submit a pull request

📄 License

MIT License - feel free to use this commercially!

🙋 Support

⭐ Star History

If this project helps you, please give it a star! It helps others find it.


Built with ❤️ for the AI community

Questions? Open an issue and I'll respond within 24 hours.