惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 三生石上(FineUI控件)
D
Docker
GbyAI
GbyAI
宝玉的分享
宝玉的分享
Jina AI
Jina AI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Vercel News
Vercel News
博客园_首页
Recent Announcements
Recent Announcements
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Hugging Face - Blog
Hugging Face - Blog
腾讯CDC
S
SegmentFault 最新的问题
Microsoft Security Blog
Microsoft Security Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
美团技术团队
V
V2EX
WordPress大学
WordPress大学
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
V
Visual Studio Blog
IT之家
IT之家
Apple Machine Learning Research
Apple Machine Learning Research
T
Tailwind CSS Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com

Hacker News: Show HN

PurrrrrFocus: Pomodoro Timer App - App Store Workflow Engine — Multi-Step Orchestration for Bun RapidPhoto: Pro Photo Editor App - App Store GitHub - DheerG/swarms: Achieve extraordinary results with claude code across a variety of tasks SPICE simulation → oscilloscope → verification with Claude Code — Lucas Gerads Show HN: VCoding – A 5 MB native Windows IDE with no dynamic dependencies Show HN: LLMs don't hallucinate because they're bad at math, it's the format GitHub - Agent-FM/agentfm-core: AgentFM is a peer-to-peer network that turns everyday computers into a decentralized AI supercomputer. AgentFM lets you run massive AI workloads directly across a global mesh of idle CPUs and GPUs. Show HN: Tracking Top US Science Olympiad Alumni over Last 25 Years GitHub - Potarix/agent-hub: One place to talk to all your agents Show HN: Runtime security for AI agents(injection,tool abuse, data exfiltration) GitHub - dubeyKartikay/lazyspotify: Terminal Spotify client for macOS and Linux GitHub - the-banana-tool/king-louie: Easy to use GUI Personal AI Assistant. Win/Linux/Mac. Show HN I made my vacation rental bookable by AI agents–no Airbnb, 0% commission GitHub - basteez/jsf-autoreload: maven plugin to enable hot reload on jsf projects uvm32/hosts/host-gdbstub at main · ringtailsoftware/uvm32 GitHub - labsai/EDDI: Config-driven engine that turns JSON into production-grade AI agents. Multi-agent orchestration, 12+ LLM providers, MCP/A2A protocols, RAG, persistent memory, and enterprise compliance (EU AI Act, GDPR, HIPAA). Built on Quarkus. GitHub - glitchnsec/fortyone-oss: AI Executive Assistant Platform Quickstart | Alien GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. GitHub - ocrbase-hq/ocrbase: 📄 PDF/IMG ->.MD/JSON Document OCR API for PaddleOCR and GLMOCR. Self-hostable. GitHub - impactjo/home-memory: MCP server that lets your AI assistant remember everything about your home. GitHub - Sets88/dbcls: DbCls is a powerful terminal database client that supports various databases GitHub - neptun2000/heor-agent-mcp GitHub - SeanFDZ/macmind: Single-layer transformer in HyperTalk for the classic Macintosh RollQuation: Math Puzzles - Apps on Google Play GitHub - dropbox/witchcraft Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis GitHub - opentalon/opentalon: OpenTalon is an open-source platform built from the ground up in Go as a robust alternative to OpenClaw LinkedIn™ 职位抓取工具 - Chrome 应用商店
GitHub - Treasury-Technologies-Inc/treasurybench: Persona...
juneadkhan · 2026-06-26 · via Hacker News: Show HN

Personal-finance assistant benchmark — evaluate how well AI-powered finance products and frontier models use real user data to surface high-leverage financial opportunities.

v0.1.0 · 3 personas · 81 tasks · 12 domains · judge-primary scoring with table-grounded factual verification


Results — v0.1.0

Leaderboard

Provider Lane Score Factually Clean Median Latency
Treasury Product contender 85.5 93% 13.7s
ChatGPT chat-latest Full-context baseline † 79.6 83% 8.0s
Origin Product contender 71.0 86% 46.0s
Monarch Product contender 52.1 86% 100.7s

† Full-context baselines paste the persona's transactions, balances, and memories directly into the prompt — this is not how a real consumer product works. It is a ceiling estimate, not a product contender.

‡ 73.1 when the 16 tasks where balance import silently failed are excluded. See artifacts/RUN_INTEGRITY.md.

Scores are 0–100, judge-primary with table-grounded factual caps. Stale or wrong financial facts (contribution limits, tax rules, program terms) hard-cap the task score regardless of prose quality — material errors cap at 65, dangerous errors at 40. Full scoring architecture: SCORING.md.

By Domain

Best score per row bolded. † marks the full-context baseline (not a product contender).

Domain Tasks Treasury Origin Monarch ChatGPT †
Transaction Intelligence 9 92 82 64 89
Tax Strategy 12 85 74 58 73
Retirement & Tax-Advantaged Accounts 9 87 62 44 71
Investing & Equity Compensation 6 82 60 65 78
Housing & Rent 6 89 80 39 91
Employer Benefits & Workplace Perks 6 87 74 54 76
Credit Cards & Rewards 9 80 66 29 75
Insurance & Risk Protection 6 89 79 73 90
Cashflow & Budgeting 6 87 67 60 89
Savings & Expense Reduction 6 77 58 25 70
Debt & Credit Health 3 84 80 81 96
Life Planning & Major Decisions 3 90 79 51 71

By Persona

Persona Treasury Origin Monarch ChatGPT †
Maria Chen — Seattle, Microsoft, renter 87 71 57 80
Priya Patel — Denver, dual income, homeowner 85 69 46 70
Jordan Rivera — Austin, self-employed 84 73 53 89

Factual Integrity

Share of answers with no locked-fact contradiction across 81 tasks. Dangerous = incorrect fact that could cause real financial harm (e.g. stale contribution limit cited as actionable advice).

Provider Factually Clean Material errors Dangerous errors
Treasury 93% (75/81) 5 1
Origin 86% (70/81) 7 4
Monarch 86% (70/81) 2 9
ChatGPT † 83% (67/81) 2 12

ChatGPT's 12 dangerous errors drive the largest gap between its judged quality (85) and final score (79.6): it consistently cites stale 2025 contribution limits as current, even under idealized in-prompt context.


Published Artifacts

All captures, judge prompts, judgments, and scored results are in artifacts/.

Run Score Tasks Captured Notes
treasury-full-20260609001842 85.5 81 2026-06-09 Live Treasury PWA advisor with tool calls
chatgpt-chat-latest-full-20260609121316 79.6 81 2026-06-09 Full-context baseline — not a product contender
origin-full-20260605T160538 71.0 / 73.1 81 2026-06-05 73.1 excluding 16 balance-import failures
monarch-full-20260605T200447 52.1 81 2026-06-05

Each run directory contains captures/, judge-prompts/, judgments/, and results/ with machine-readable CSVs and divergence reports. See artifacts/RUN_INTEGRITY.md for the judge-independence caveat, the Origin import-failure disclosure, and the self-authorship disclosure.


What's Being Tested

TreasuryBench asks whether a personal-finance assistant can:

  • Read transaction and balance data accurately.
  • Connect user context to personal-finance concepts.
  • Surface high-value opportunities hidden in ordinary financial data.
  • Use current financial rules, limits, product terms, and local programs correctly.
  • Quantify impact and give exact next steps.
  • Avoid unsupported assumptions, stale facts, unsafe recommendations, and generic boilerplate.

Personas

Three synthetic US households with transaction history, account balances, saved memories, employer, location, and goals:

  • Maria Chen — late 20s, Seattle, Microsoft software engineer, renter.
  • Priya Patel — dual income, Denver, homeowner, two kids.
  • Jordan Rivera — Austin, self-employed, gig/freelance income.

Tasks

81 natural user questions (27 per persona) across 12 domains. Tasks are phrased like real user questions — "How can I save money on rent?" not "Identify Seattle MFTE eligibility." The assistant must infer the opportunity from the persona's signals.

Scoring

Judge-primary when LLM judge output is available. Deterministic evaluators catch exact data use, arithmetic, and planted-signal discovery. The LLM judge grades synthesis, personalization, and open-ended credit. Stale or wrong financial facts apply hard caps regardless of prose quality.

Full architecture: SCORING.md · Methodology: METHODOLOGY.md · Limitations: LIMITATIONS.md · Run integrity: artifacts/RUN_INTEGRITY.md


Recreate

Install

pnpm install
pnpm validate    # verify schema consistency and scoring totals
pnpm report      # print a compact task/domain summary
pnpm smoke       # run the fixture provider end-to-end

Run a full-context baseline

pnpm export-prompts -- --out=runs/my-openai-run/prompts --mode=full_context_baseline
pnpm run-provider -- --provider=openai --out=runs/my-openai-run --live=true \
  --model=chat-latest --max-output-tokens=2200 --env-file=.env
pnpm evaluate-run -- --run=runs/my-openai-run
pnpm run-judge -- --run=runs/my-openai-run --env-file=.env \
  --judge-provider=gemini --model=gemini-3.1-flash-lite
pnpm score-run -- --run=runs/my-openai-run

.env needs OPENAI_API_KEY (provider) and GOOGLE_GENERATIVE_AI_API_KEY (judge). Use --judge-provider=openai with OPENAI_API_KEY to judge with OpenAI instead.

Capture a product manually

pnpm export-persona-data -- --out=runs/my-product-data
pnpm make-capture-templates -- --out=runs/my-product --provider=myproduct --mode=product_capture
# seed each persona into your product, ask the natural prompt, paste the answer
# into the `response` field of each captures/*.json file
pnpm evaluate-run -- --run=runs/my-product
pnpm run-judge -- --run=runs/my-product --env-file=.env \
  --judge-provider=gemini --model=gemini-3.1-flash-lite
pnpm score-run -- --run=runs/my-product

See docs/product-capture-protocol.md for the full seeding protocol.

Re-score a published run

pnpm score-run -- --run=artifacts/treasury-full-20260609001842

To re-judge from existing captures:

pnpm run-judge -- --run=artifacts/treasury-full-20260609001842 --env-file=.env \
  --judge-provider=gemini --model=gemini-3.1-flash-lite

License

MIT — see LICENSE.