惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Attack and Defense Labs
Attack and Defense Labs
Y
Y Combinator Blog
阮一峰的网络日志
阮一峰的网络日志
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 聂微东
V
Visual Studio Blog
量子位
博客园 - 三生石上(FineUI控件)
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Jina AI
Jina AI
宝玉的分享
宝玉的分享
Hacker News: Ask HN
Hacker News: Ask HN
小众软件
小众软件
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Hacker News - Newest:
Hacker News - Newest: "LLM"
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
W
WeLiveSecurity
Google Online Security Blog
Google Online Security Blog
AI
AI
Schneier on Security
Schneier on Security
大猫的无限游戏
大猫的无限游戏
WordPress大学
WordPress大学
I
InfoQ
A
Arctic Wolf
月光博客
月光博客
Last Week in AI
Last Week in AI
Hugging Face - Blog
Hugging Face - Blog
Cloudbric
Cloudbric
Google DeepMind News
Google DeepMind News
Security Latest
Security Latest
GbyAI
GbyAI
N
Netflix TechBlog - Medium
TaoSecurity Blog
TaoSecurity Blog
Blog — PlanetScale
Blog — PlanetScale
Scott Helme
Scott Helme
www.infosecurity-magazine.com
www.infosecurity-magazine.com
T
The Exploit Database - CXSecurity.com
博客园 - 【当耐特】
博客园 - 司徒正美
The Hacker News
The Hacker News
Cisco Talos Blog
Cisco Talos Blog
I
Intezer
云风的 BLOG
云风的 BLOG
T
Threatpost
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cyber Attacks, Cyber Crime and Cyber Security
罗磊的独立博客
Security Archives - TechRepublic
Security Archives - TechRepublic
T
Tenable Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Building a Serverless AI Model Evaluation Platform on AWS
Debapriya De · 2026-05-22 · via DEV Community

The Problem

A media company needed to evaluate which AI model produces the best podcast-style summaries from news articles. They wanted to:

  • Send an article to multiple AI models simultaneously
  • Compare the outputs side by side
  • Score each output automatically
  • Generate a visual comparison report

Doing this manually, copying articles into different model playgrounds, reading outputs, judging quality, doesn't scale. They needed an automated evaluation pipeline that could run experiments on demand and produce consistent, comparable results.

What We Built

A fully serverless evaluation platform on AWS that accepts an article, runs it through multiple foundation models in parallel, scores each output using a separate AI judge, and produces an HTML comparison report. All triggered by a single API call.

The system handles the entire lifecycle:

  1. Prompt optimization — an AI agent refines the user's instructions into an effective prompt
  2. Parallel model invocation — multiple Bedrock models generate summaries simultaneously
  3. Automated scoring — a scoring agent evaluates each output against quality criteria
  4. Report generation — produces a formatted HTML comparison page

Architecture Overview

AWS Architecture Overview

The 6-Step Workflow

The core of the system is a Step Functions state machine that orchestrates six Lambda functions in sequence. Here's what each step does and why it exists as a separate step.

6 Step Pipeline

Step 1: Validate

def validate(event):
    """Read and validate the experiment definition from S3."""
    definition = s3.get_object(Bucket=BUCKET, Key=f"definitions/{experiment_id}/definition.json")
    # Validate required fields: article, models, prompt
    # Fail fast if inputs are malformed
    return validated_definition

Enter fullscreen mode Exit fullscreen mode

Why a separate step? Fail-fast validation before incurring any Bedrock costs. If the definition is malformed, we stop here — no wasted model invocations.

Step 2: Invoke Models (Parallel)

This is where it gets interesting. We invoke multiple Bedrock models simultaneously using Python's ThreadPoolExecutor:

from concurrent.futures import ThreadPoolExecutor, as_completed

def invoke_models(definition):
    models = definition['models']  # e.g., ["meta.llama3-70b", "deepseek-r1", "amazon.nova-lite"]
    prompt = definition['prompt']
    article = definition['article']

    results = {}

    with ThreadPoolExecutor(max_workers=len(models)) as executor:
        futures = {
            executor.submit(invoke_bedrock, model_id, prompt, article): model_id
            for model_id in models
        }
        for future in as_completed(futures):
            model_id = futures[future]
            response = future.result()
            results[model_id] = {
                "output": response['output']['message']['content'][0]['text'],
                "usage": {
                    "input_tokens": response['usage']['inputTokens'],
                    "output_tokens": response['usage']['outputTokens']
                }
            }

    return results

Enter fullscreen mode Exit fullscreen mode

Why ThreadPoolExecutor inside Lambda? Bedrock API calls are I/O-bound. Running them in parallel within a single Lambda invocation means we pay for one Lambda execution instead of three, and the total wall-clock time is roughly equal to the slowest model rather than the sum of all models.

Step 3: Store Outputs

Writes comparison.json to S3 — containing all model outputs but no scores yet. This creates a checkpoint: if scoring fails, we don't lose the generated content.

Step 4: Score (Parallel)

The scoring agent (Claude Haiku) evaluates each model's output against quality criteria. Again, parallel execution via ThreadPoolExecutor:

def score(outputs):
    scoring_prompt = """Rate this podcast summary on:
    - Accuracy (1-10): Does it faithfully represent the article?
    - Engagement (1-10): Would a listener find this compelling?
    - Structure (1-10): Is it well-organized for audio?
    Respond with JSON only."""

    with ThreadPoolExecutor(max_workers=len(outputs)) as executor:
        futures = {
            executor.submit(invoke_bedrock, SCORING_MODEL, scoring_prompt, output): model_id
            for model_id, output in outputs.items()
        }
        # ... collect scores

Enter fullscreen mode Exit fullscreen mode

Why a separate scoring model? Using a different model (or at minimum, a separate invocation with a scoring-specific prompt) as the judge avoids self-evaluation bias. The scoring agent doesn't know which model produced which output.

Step 5: Store Scores

Updates comparison.json with the scores attached to each model's output.

Step 6: Generate HTML

Produces a formatted comparison.html report that displays all outputs side by side with their scores. This is the final deliverable the user downloads.

Why Amazon Bedrock's Converse API?

We use the Converse API rather than the model-specific InvokeModel API. The key advantage: one unified interface across all models.

def invoke_bedrock(model_id, system_prompt, user_message):
    response = bedrock_runtime.converse(
        modelId=model_id,
        messages=[{"role": "user", "content": [{"text": user_message}]}],
        system=[{"text": system_prompt}]
    )
    return response

Enter fullscreen mode Exit fullscreen mode

Switching from Llama to Claude to Nova Lite requires changing only the model_id string. No code changes, no different request formats, no response parsing differences.

The Converse API also returns token usage in every response — which we pass through to the caller for billing:

{
  "results": [
    {
      "model_id": "meta.llama3-70b-instruct-v1:0",
      "summary": "...",
      "usage": { "input_tokens": 1523, "output_tokens": 847 }
    }
  ],
  "total_usage": { "total_input_tokens": 4569, "total_output_tokens": 2541 }
}

Enter fullscreen mode Exit fullscreen mode

Cost Control: The Hardest Part

Here's the reality of building on top of foundation models: every API call costs money, and costs scale with input size. A single /run request invoking 3 models on a long article can cost $0.10–0.50. That sounds small until someone writes a script that calls it in a loop.

Billing Alarms (Day 1)

We set up CloudWatch billing alarms immediately:

CloudWatch Alarm ($10 threshold) → SNS → Email notification
CloudWatch Alarm ($25 threshold) → SNS → Email notification

Enter fullscreen mode Exit fullscreen mode

This is the bare minimum. You'll know when costs are climbing, even if you can't stop them automatically.

API Security (Critical for Any AI-Backed API)

An unprotected API that invokes foundation models is essentially a public credit card. We learned this the hard way and now treat API security as P0 — before any external access:

  • API Keys on every endpoint (immediate protection)
  • Usage plans with per-key quotas (500 requests/day, 5000/month)
  • Rate limiting (10 req/s throttle) to prevent burst abuse
  • Request logging to attribute usage to specific callers
# Every request must include the API key
curl -X POST https://api.example.com/run \
  -H "x-api-key: btk_live_abc123def456" \
  -H "Content-Type: application/json" \
  -d '{"article": "...", "models": ["meta.llama3-70b"]}'

Enter fullscreen mode Exit fullscreen mode

Without this, anyone who discovers your API URL can generate unbounded Bedrock charges.

Lessons Learned

1. Separate validation from execution

Bedrock calls are expensive. Validate everything before invoking any model. Check that the article isn't empty, the model IDs are valid, the prompt isn't too long. Fail at Step 1, not Step 2.

2. ThreadPoolExecutor > separate Lambda invocations for parallel model calls

We considered using Step Functions' native parallel states or invoking separate Lambdas per model. ThreadPoolExecutor within a single Lambda turned out simpler:

  • One Lambda execution to pay for (not N)
  • Shared memory for the article text (no repeated S3 reads)
  • Simpler error handling
  • Total time ≈ slowest model, not sum of all

The tradeoff: if one model times out, the entire Lambda times out. We mitigate this with per-future timeouts.

3. Store intermediate results

Each step writes to S3 before the next step begins. If Step 4 (scoring) fails, we still have the model outputs from Step 3. We can retry scoring without re-invoking the content models.

4. Token usage is free metadata — always capture it

Bedrock returns inputTokens and outputTokens in every response. Capturing and returning this costs nothing but enables:

  • Per-customer billing
  • Cost forecasting
  • Identifying expensive prompts
  • Detecting anomalies (sudden spike in token usage = possible abuse)

5. Start with S3, add a database when you need queries

For the POC, S3 handles all storage. It's simple, cheap, and sufficient for sequential read/write patterns. We're adding DynamoDB only now that we need to query experiment history by user — something S3 can't do efficiently.

What's Next

The platform is functional but evolving:

  • Selection History — DynamoDB-backed experiment sessions so users can revisit past comparisons and track which model they ultimately chose
  • Frontend UI — Visual interface for running experiments and browsing history
  • Cognito Authentication — User-level access control when the UI ships

Tech Stack Summary

Layer Service Why
API API Gateway (HTTP API) Low latency, pay-per-request
Compute AWS Lambda (Python) Serverless, scales to zero
Orchestration Step Functions Visual workflow, built-in retries
AI Models Amazon Bedrock (Converse API) Multi-model, unified interface
Storage Amazon S3 Cheap, durable, simple
Monitoring CloudWatch + SNS Billing alarms, email alerts
Auth (planned) API Keys + Cognito Layered security
History (planned) DynamoDB Fast queries by user/session

Reach Out to Us

Interested in modernizing your cloud infrastructure and building enterprise-grade solutions? Storm Reply is driven by continuous learning and practical innovation. We specialize in designing and delivering scalable AWS architectures that support customers throughout their cloud journey, from early assessment to production-ready deployment.

With deep experience in AWS architecture, data engineering, and security best practices, we help enterprises migrate with confidence and move faster on their cloud transformation goals.

Let’s connect and explore how we can support your modernization initiatives.


The full system runs in eu-central-1 (Frankfurt), costs under $20/month excluding Bedrock usage, and handles the entire evaluation lifecycle in a single API call. Serverless means we pay nothing when nobody's running experiments, and scale automatically when they are.

If you're building something similar — any system where API calls trigger expensive downstream operations — lock down your API first, validate inputs aggressively, and always know what each request costs.


Built with AWS Lambda, Step Functions, and Amazon Bedrock.