惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Hugging Face - Blog
Hugging Face - Blog
Recent Announcements
Recent Announcements
V
Visual Studio Blog
博客园 - 叶小钗
H
Help Net Security
aimingoo的专栏
aimingoo的专栏
宝玉的分享
宝玉的分享
U
Unit 42
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
F
Fortinet All Blogs
V
V2EX
Stack Overflow Blog
Stack Overflow Blog
WordPress大学
WordPress大学
D
DataBreaches.Net
J
Java Code Geeks
H
Hackread – Cybersecurity News, Data Breaches, AI and More
A
About on SuperTechFans
酷 壳 – CoolShell
酷 壳 – CoolShell
量子位
C
Check Point Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
小众软件
小众软件
Microsoft Azure Blog
Microsoft Azure Blog
M
MIT News - Artificial intelligence

Hacker News: Show HN

PurrrrrFocus: Pomodoro Timer App - App Store Workflow Engine — Multi-Step Orchestration for Bun RapidPhoto: Pro Photo Editor App - App Store GitHub - DheerG/swarms: Achieve extraordinary results with claude code across a variety of tasks SPICE simulation → oscilloscope → verification with Claude Code — Lucas Gerads Show HN: VCoding – A 5 MB native Windows IDE with no dynamic dependencies Show HN: LLMs don't hallucinate because they're bad at math, it's the format GitHub - Agent-FM/agentfm-core: AgentFM is a peer-to-peer network that turns everyday computers into a decentralized AI supercomputer. AgentFM lets you run massive AI workloads directly across a global mesh of idle CPUs and GPUs. Show HN: Tracking Top US Science Olympiad Alumni over Last 25 Years GitHub - Potarix/agent-hub: One place to talk to all your agents Show HN: Runtime security for AI agents(injection,tool abuse, data exfiltration) GitHub - dubeyKartikay/lazyspotify: Terminal Spotify client for macOS and Linux GitHub - the-banana-tool/king-louie: Easy to use GUI Personal AI Assistant. Win/Linux/Mac. Show HN I made my vacation rental bookable by AI agents–no Airbnb, 0% commission GitHub - basteez/jsf-autoreload: maven plugin to enable hot reload on jsf projects uvm32/hosts/host-gdbstub at main · ringtailsoftware/uvm32 GitHub - labsai/EDDI: Config-driven engine that turns JSON into production-grade AI agents. Multi-agent orchestration, 12+ LLM providers, MCP/A2A protocols, RAG, persistent memory, and enterprise compliance (EU AI Act, GDPR, HIPAA). Built on Quarkus. GitHub - glitchnsec/fortyone-oss: AI Executive Assistant Platform Quickstart | Alien GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. GitHub - ocrbase-hq/ocrbase: 📄 PDF/IMG ->.MD/JSON Document OCR API for PaddleOCR and GLMOCR. Self-hostable. GitHub - impactjo/home-memory: MCP server that lets your AI assistant remember everything about your home. GitHub - Sets88/dbcls: DbCls is a powerful terminal database client that supports various databases GitHub - neptun2000/heor-agent-mcp GitHub - SeanFDZ/macmind: Single-layer transformer in HyperTalk for the classic Macintosh RollQuation: Math Puzzles - Apps on Google Play GitHub - dropbox/witchcraft Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis GitHub - opentalon/opentalon: OpenTalon is an open-source platform built from the ground up in Go as a robust alternative to OpenClaw LinkedIn™ 职位抓取工具 - Chrome 应用商店
GitHub - bigkan8/legal-action-boundary-eval: Legal Action...
kankouadio_v · 2026-04-23 · via Hacker News: Show HN

Luminance proxy edition

This directory contains a public proxy eval for legal AI workflows that sit at the action boundary: negotiation moves, compliance clearance, review routing, and composed supervisor/subagent flows.

It is based on workflows Luminance publicly markets across negotiation, compliance, collaboration, and supervisor-led legal agents. It is not an internal Luminance benchmark and it is not affiliated with Luminance.

LABE overview

Why this exists

Most legal AI evaluations stop at understanding:

  • clause extraction
  • answer quality
  • markup quality
  • summarization quality

VerifiedX matters one seam later, when the system is about to do something high-impact:

  • accept or redraft a clause
  • route to signature
  • mark an issue resolved
  • clear compliance
  • escalate or reroute an agreement

LABE measures that seam directly.

Headline result

Across the current 12-scenario suite, run in both TypeScript and Python:

  • baseline systems executed 18 unjustified high-impact actions
  • VerifiedX executed 0
  • VerifiedX produced 0 false blocks in this suite
  • surviving-goal completion improved from 41.7% to 100%

The raw results are in RESULTS.md, with full artifacts in artifacts/.

What this is

  • a public proxy eval based on workflow classes Luminance publicly markets
  • a same-harness A/B: baseline vs VerifiedX
  • a legal action-boundary eval, not a generic model-quality eval
  • a dual-language suite with the same scenario truth in TypeScript and Python

What this is not

  • not a replication of Luminance's internal product or customer traffic
  • not a claim about overall legal reasoning quality
  • not a benchmark of Word UI, OCR, diligence review, or repository search
  • not a replacement for customer-specific evals

Tracks

  • Negotiation Accepting counterparty positions, applying playbook-approved redrafts, routing to signature, and resolving clause issues.
  • Compliance Marking agreements compliant, applying remediation markup, escalating failed checks, and avoiding false clearance.
  • Composed workflows Intake agent -> execution agent -> upstream legal/compliance review -> redispatch or lane change.

Repo map

  • EVAL_CARD.md Short-form card for scope, intended use, metrics, and limitations.
  • METHODOLOGY.md Public-source grounding, harness design, scoring policy, and limitations.
  • SCENARIOS.md The full 12-scenario catalog with guarded actions and expected protected behavior.
  • RESULTS.md Generated scorecard tied to the raw artifact files.
  • REPRODUCE.md Steps to rerun the suite and regenerate the public report assets.
  • EXECUTIVE_BRIEF.md Concise operator summary for legal AI, product, and governance stakeholders.

Current run metadata

  • Run date: 2026-04-19
  • Model: gpt-5.4-mini
  • Run environment: Real production run against api.verifiedx.me
  • VerifiedX API: https://api.verifiedx.me
  • TypeScript SDK: @verifiedx-core/sdk@0.1.17
  • Python SDK: verifiedx==0.1.8

Raw evidence

Public workflow sources