惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
H
Hackread – Cybersecurity News, Data Breaches, AI and More
V
Vulnerabilities – Threatpost
L
LangChain Blog
Stack Overflow Blog
Stack Overflow Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
云风的 BLOG
云风的 BLOG
C
Cisco Blogs
V
Visual Studio Blog
L
Lohrmann on Cybersecurity
Latest news
Latest news
S
Securelist
The Last Watchdog
The Last Watchdog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
The Register - Security
The Register - Security
Webroot Blog
Webroot Blog
The Cloudflare Blog
S
Secure Thoughts
Y
Y Combinator Blog
aimingoo的专栏
aimingoo的专栏
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
N
News and Events Feed by Topic
S
Security Affairs
Attack and Defense Labs
Attack and Defense Labs
Microsoft Azure Blog
Microsoft Azure Blog
T
Tailwind CSS Blog
V2EX - 技术
V2EX - 技术
GbyAI
GbyAI
L
LINUX DO - 热门话题
PCI Perspectives
PCI Perspectives
Schneier on Security
Schneier on Security
V
V2EX
K
Kaspersky official blog
Hugging Face - Blog
Hugging Face - Blog
AWS News Blog
AWS News Blog
T
The Exploit Database - CXSecurity.com
C
CERT Recently Published Vulnerability Notes
C
Cyber Attacks, Cyber Crime and Cyber Security
P
Proofpoint News Feed
T
Threatpost
WordPress大学
WordPress大学
SecWiki News
SecWiki News
B
Blog RSS Feed
Blog — PlanetScale
Blog — PlanetScale
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
A
Arctic Wolf
酷 壳 – CoolShell
酷 壳 – CoolShell
W
WeLiveSecurity
Jina AI
Jina AI
D
Darknet – Hacking Tools, Hacker News & Cyber Security

Show HN

GitHub - villagesql/villagesql-skills: Agent skills for VillageSQL - gemini-cli-extension; claude-code-plugin GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
GitHub - aws-samples/sample-GEDD: Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.
balasvce1985 · 2026-06-01 · via Show HN

GEDD - A Systematic Evidence Driven LLM As a Judge Framework

CI Python 3.11+ License: MIT-0 GitHub stars

GEDD is a Systematic Evidence Driven LLM As a Judge Framework for AI agents.

It is an annotation-first workflow for turning domain-owner review of AI agent behavior into release gates engineering can run.

The web app gives product managers, domain experts, and ML engineers one shared path:

  1. Define the agent and the work it is supposed to do.
  2. Collect or load representative queries and responses.
  3. Review the responses in a task-shaped workbench.
  4. Name failures in the domain owner's vocabulary.
  5. Convert the observed failures into an LLM-as-a-judge prompt.
  6. Export a validated handoff for CI, MLflow, and model regression work.

The current first-run experience ships with two 50-query PM workbench demos: an AAA game localization session and an AWS cloud GDPR auditor session. They show how a domain owner can move from raw agent traces to open codes, root-cause patterns, saturation evidence, a judge prompt, and an ML engineer implementation queue.

GEDD PM annotation walkthrough

The longer methodology essay is in METHODOLOGY.md. This README is the practical product and engineering guide.

What GEDD Produces

Output Who creates it Who uses it Why it matters
Golden queries PM or domain expert ML engineer, eval owner Defines the user situations the agent must handle
Human labels PM or domain expert Judge builder, release owner Separates acceptable, partial, and failing behavior
Failure codebook PM or domain expert ML engineer, prompt owner Names the exact domain-specific failure modes to fix
Memos and severity PM or domain expert ML engineer, reviewer Explains why the failure matters and how bad it is
Axial coding PM or domain expert Product and engineering leads Groups repeated failures into root causes and consequences
Judge prompt PM plus ML engineer CI and model evaluation Converts observed failures into automated review criteria
session.json handoff App or CLI ML engineer Carries agent spec, prompt, queries, labels, and validation state
MLflow artifacts ML engineer Release pipeline Tracks datasets, judges, evaluation runs, and regression gates

GEDD is not a generic model leaderboard. It is a way to preserve expert judgment and make it executable.

Quick Start

Start the web app:

cd grounded-evals
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
grounded-evals serve --host 127.0.0.1 --port 8080

Open http://127.0.0.1:8080.

No Codex skill or plugin is required.

Local runs start in guest mode unless ADMIN_PASSWORD or Cognito environment variables are configured. If port 8080 is busy, use --port 8081.

For the fastest product tour, use one of the seeded 50-query demos. They do not require model calls:

  1. Open Home or Demos.
  2. Click Load 50-query localization demo or Load 50-query AWS Cloud GDPR demo.
  3. Open PM Workbench to review the labeled traces, failure codes, memos, and saturation state.
  4. Open Judge to inspect or revise the generated judge prompt.
  5. Open Report to review release readiness and download the ML engineer handoff.

To reset after loading a demo, use the top-right refresh action. Confirm Start Fresh to clear the loaded project data while keeping the current login session.

Current Web App

grounded-evals serve runs a NiceGUI app with a short primary navigation:

Page Purpose Main actions
Home Entry point Load the 50-query localization or AWS Cloud GDPR demo, continue active work, or start a custom agent
AI PM Coach Guided setup Capture agent definition, system prompt, runtime choice, and golden-query plan
PM Workbench Annotation surface Review responses, assign verdicts, create failure codes, set severity, write memos, and monitor saturation
Judge Release gate builder Generate and edit an LLM-as-a-judge prompt from the observed failure modes
Report Engineering handoff Review quality signals, CI gates, artifact readiness, implementation queue, and export files

The Demos page remains available for starter data. It is not the main workflow. Demos are seed sessions that help teams understand the annotation loop before they bring their own traces.

The 50-Query Localization Demo

The main demo is a synthetic but complete localization QA session for an AAA game agent called LocaleGate.

It includes:

Asset Contents
50 golden queries Runtime strings, storefront copy, subtitles, RTL input prompts, region rules, culturalization, paid-currency copy, live-event dates, and glossary consistency
Synthetic responses Baseline agent answers with realistic localization failures
PM annotations Correct, partial, and incorrect verdicts with severity and confidence
Open codes Localization-specific failure labels rather than generic quality tags
Axial coding Root causes, context, intervening conditions, action strategy, and consequence mapping
Saturation evidence Final-window evidence that new annotations repeat existing codes
Judge prompt A release-gate judge built from the localization failure modes
Report handoff CI gates, artifact status, implementation queue, and commands for an ML engineer

Example failure codes in the demo include:

Code What it catches
Placeholder And Markup Corruption The response approves a translation that drops variables, tags, markup, or runtime-safe formatting
Gameplay Meaning Reversal The localized text reverses the gameplay instruction or player action
Rating Or Disclosure Softening Marketing or regional copy weakens required rating, privacy, paid-currency, or platform disclosures
RTL Input Direction Drift Right-to-left layout or controller input language changes the intended interaction
Locale Format Ambiguity Dates, times, numbers, or currencies remain ambiguous for the target locale
Entitlement Copy Mistranslation Storefront text changes what the buyer receives or what content is included
Culturalization Risk Dismissal The response treats regional content risk as a translation-only issue

Those labels are the point of the workflow. The judge is not asked to score generic helpfulness first. It is asked to enforce the domain owner's observed release blockers.

The 50-Query AWS Cloud GDPR Demo

The second main workbench demo is a synthetic AWS cloud GDPR audit session for CloudAuditGate.

It includes 50 golden queries covering S3 and CloudWatch retention, CloudTrail and centralized logging, Bedrock prompt reuse, Rekognition and high-risk review, DSAR and deletion handling across backups and data lakes, shared responsibility, cross-region transfers, and breach escalation from AWS security incidents. The output is the same PM-owned package as the localization demo: annotations, open codes, axial coding, saturation evidence, and an audit-ready judge prompt.

The AWS Cloud GDPR demo uses plain-language tags on purpose, for example Data Used For The Wrong Job, Collecting Or Keeping Too Much Data, EU Data Moved The Wrong Way, and Trying To Work Around GDPR. The point is to make the GEDD loop easy to follow: annotate the failure in human language first, then turn that observed pattern into the judge gate.

Bring Your Own Agent

Use the app when you have a real or proposed agent and need review evidence before you automate evaluation.

Step What to do Output
1. Define Describe the agent, user, task boundary, and system prompt in AI PM Coach Agent spec and prompt
2. Build queries Generate or paste golden queries that cover normal, edge, ambiguous, adversarial, multi-turn, and recovery cases Query set
3. Get responses Run the saved prompt against Bedrock, Anthropic, or a configured runtime, or paste existing traces Response queue
4. Annotate Review each response in PM Workbench and capture verdict, code, severity, confidence, and memo Human labels and codebook
5. Pattern Use open coding and axial coding to group repeated failures and root causes Release-risk model
6. Judge Build the judge prompt from the observed codes and examples LLM-as-a-judge prompt
7. Handoff Export the session and ML handoff from Report Engineering package

If you already have production traces, use the app as an annotation surface rather than generating new responses. See Paste In Traces.

ML Engineer Handoff

The Report page contains an ML Engineer Handoff section. It is designed to be actionable, not a narrative status update.

It gives engineering:

Handoff field Why it exists
Engineering status Indicates whether the session is blocked by P0 failures, missing a judge, needs calibration, or is ready for a CI pilot
CI gates Shows current and target values for P0 failures, regression pass rate, human coverage, and judge-human agreement
Artifact status Confirms whether session handoff, golden dataset, codebook, judge prompt, and calibration evidence are ready
Implementation queue Prioritizes failure codes by severity and count, with tagged examples and definitions of done
Runbook Gives commands the ML engineer can run immediately

Typical handoff commands:

cd grounded-evals

grounded-evals validate-session --session session.json
grounded-evals export --session session.json --format jsonl --output golden_dataset.jsonl
grounded-evals judge --session session.json --output judge_prompt.md
grounded-evals mlflow --session session.json --tracking-uri $MLFLOW_TRACKING_URI --run-eval

The expected engineering loop is:

  1. Validate the session.
  2. Create one failing regression case for each P0 queue item.
  3. Patch the prompt, retrieval policy, tool policy, or runtime behavior.
  4. Rerun the judge and review disagreements.
  5. Promote the gate only after calibration is acceptable.

The default calibration target used in the handoff is kappa >= 0.80 before the judge blocks merges.

CLI Reference

The CLI supports the same workflow for repeatable runs, scripting, and CI.

Command Use
grounded-evals serve Start the web app
grounded-evals chat Run the guided PM workflow from the terminal
grounded-evals eval Run golden queries against supported models
grounded-evals annotate Add verdicts and failure codes from the terminal
grounded-evals analyze Map failure codes into legacy evaluation dimensions when needed
grounded-evals fracture Break a domain into coverage categories and candidate queries
grounded-evals compare Check whether a new query adds unique coverage
grounded-evals check-saturation Check whether the dataset is still producing new concepts
grounded-evals coverage Show coverage by category
grounded-evals judge Generate a judge prompt from the session
grounded-evals validate-session Check whether a session is ready for handoff
grounded-evals handoff Write a validated session handoff artifact
grounded-evals export Export the golden dataset as JSON, JSONL, or CSV
grounded-evals mlflow Create MLflow or SageMaker MLflow artifacts and optionally run evals
grounded-evals status Print a session summary

Run command help from the package directory:

cd grounded-evals
grounded-evals --help
grounded-evals mlflow --help

Web App And CLI

GEDD ships as a web app and a CLI. No Codex skill or plugin is required.

Interface Entry point Use
Web app grounded-evals serve Primary workflow for demos, PM annotation, judge building, and report export
CLI grounded-evals --help Repeatable validation, exports, automation, and MLflow runs

Use the web app first unless you are automating an established workflow. The CLI is the right path for CI, MLflow, scripted exports, and headless checks.

Runtime And Provider Configuration

Local demo review does not require an LLM provider because the main localization demo is preloaded.

For custom agent work, configure one provider:

Provider Configuration
Amazon Bedrock Configure AWS credentials and set AWS_REGION; optionally set BEDROCK_MODEL_ID
Anthropic API Set ANTHROPIC_API_KEY; direct Anthropic calls take priority when the key is present
AgentCore runtime Configure the AgentCore environment variables used by your deployment

See SETUP.md for a full environment variable list, Bedrock model access notes, auth options, and AWS deployment setup.

Architecture

flowchart TD
    WEB["NiceGUI web app<br/>Home, Coach, Workbench, Judge, Report"]
    DEMO["Seeded demos<br/>50-query localization + domain scenarios"]
    SESSION["session.json<br/>agent, prompt, queries, labels, codebook"]
    REPORT["ML engineer handoff<br/>gates, artifacts, queue, commands"]
    CLI["grounded-evals CLI<br/>export, judge, validate, mlflow"]
    MLFLOW["MLflow / SageMaker MLflow<br/>datasets, scorers, runs"]
    CI["CI/CD gate<br/>regression checks"]
    RUNTIME["Agent runtime<br/>Bedrock, Anthropic, AgentCore"]

    DEMO --> WEB
    WEB --> SESSION
    WEB --> REPORT
    REPORT --> CLI
    SESSION --> CLI
    CLI --> MLFLOW
    MLFLOW --> CI
    CI --> RUNTIME
    RUNTIME --> WEB
Loading

Core paths:

Path Responsibility
grounded-evals/src/grounded_evals/app.py App entry point, health endpoint, release marker
grounded-evals/src/grounded_evals/ui/ NiceGUI pages, layout, demos, workbench, judge, report
grounded-evals/src/grounded_evals/open_coding/ Domain fracturing, query comparison, saturation checks
grounded-evals/src/grounded_evals/axial_coding/ Root-cause and paradigm-model mapping
grounded-evals/src/grounded_evals/judge_builder/ Rubric, prompt generation, calibration, judge variants
grounded-evals/src/grounded_evals/guide/ Session persistence and handoff validation
grounded-evals/src/grounded_evals/cli.py Command-line workflow
grounded-evals/infra/ AWS CDK infrastructure
grounded-evals/Dockerfile Container image for the web app

Validation

Before committing app or workflow changes:

cd grounded-evals
PYTHONPATH=src pytest
PYTHONPATH=src python3 -m grounded_evals.cli --help

For local web smoke tests:

grounded-evals serve --host 127.0.0.1 --port 8080

for p in / /coding /demos /coach /judge /report /health; do
  curl -sS -o /dev/null -w "$p %{http_code}\n" "http://127.0.0.1:8080$p"
done

For README-only changes, git diff --check and stale-message scans are usually enough.

Additional Docs

Doc Use
SETUP.md Local setup, provider configuration, auth, troubleshooting, deployment
METHODOLOGY.md Grounded-theory method behind the workflow
Pipeline Guide End-to-end workflow and CI/CD shape
Domain Expert Guide PM and SME review walkthrough
PM To ML LLM Judge Turning annotated sessions into production judges
Building An LLM Judge Judge design and calibration details
Cohen's Kappa Judge-human agreement guidance
Launch Checklist Release readiness checks

License And Security

License: MIT-0. See LICENSE.

Security issue reporting: see CONTRIBUTING.md.