惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Recent Announcements
Recent Announcements
博客园 - Franky
博客园 - 三生石上(FineUI控件)
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Apple Machine Learning Research
Apple Machine Learning Research
云风的 BLOG
云风的 BLOG
人人都是产品经理
人人都是产品经理
博客园 - 【当耐特】
L
LangChain Blog
Stack Overflow Blog
Stack Overflow Blog
H
Help Net Security
爱范儿
爱范儿
罗磊的独立博客
博客园_首页
美团技术团队
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
月光博客
月光博客
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
量子位
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 叶小钗
V
Visual Studio Blog
T
Tailwind CSS Blog

Recent Commits to openclaw:main

test: merge chat side-result checks · openclaw/openclaw@ddd2c2a test: merge cron history checks · openclaw/openclaw@f7eb746 test: merge responsive navigation shell checks · openclaw/openclaw@c2e4b47 docs(changelog): add codex oauth fixes · openclaw/openclaw@628e6cd test: merge navigation routing cases · openclaw/openclaw@5d8cecb Tests: mock channel registry bundled fallback · openclaw/openclaw@2b08233 Secrets: avoid broad web search discovery for single plugin config · openclaw/openclaw@a464f59 test: merge config view browser checks · openclaw/openclaw@20cf511 fix(status): align oauth health with runtime · openclaw/openclaw@eed7116 feat: add macOS screen snapshots for monitor preview (#67954) thanks … · openclaw/openclaw@f377db1 fix: report shared auth scopes in hello-ok (#67810) thanks @BunsDev · openclaw/openclaw@0b6c39b Auto-reply: avoid eager bundled route fallback · openclaw/openclaw@3ea1bf4 Tests: narrow session binding contract setup · openclaw/openclaw@54e4e16 fix(macOS): enable undo/redo in webchat composer text input (#34962) · openclaw/openclaw@00951dc Tests: speed up channel setup promotion · openclaw/openclaw@82b529a Docs: refresh agent instructions · openclaw/openclaw@5775fe2 fix(auth): serialize OAuth refresh across agents to fix #26322 (#67876) · openclaw/openclaw@8e79080 test: allow ollama public surface boundary test · openclaw/openclaw@7d4f1a6 Docs: add test performance guardrails · openclaw/openclaw@89706d3 Tests: restore context-engine usage proof · openclaw/openclaw@e4c4f95 Tests: slim context engine runtime coverage · openclaw/openclaw@74c198f ci: retry failed custom checkouts · openclaw/openclaw@0ee5baf test: trim duplicate provider auth onboarding cases · openclaw/openclaw@1ffc02e matrix: fix sessions_spawn --thread subagent session spawning (#67643) · openclaw/openclaw@1ce2596 test: reduce auth choice fixture churn · openclaw/openclaw@857b9cd test: mock health status config boundaries · openclaw/openclaw@9d5ab4a test: mock onboard config io boundary · openclaw/openclaw@299694d test: mock legacy state plugin boundaries · openclaw/openclaw@2713089 test: mock channel install boundaries · openclaw/openclaw@b945248 test: mock doctor preview channel boundaries · openclaw/openclaw@b1a3ad4
test(qa-lab): add personal agent scenarios · openclaw/ope...
vincentkoc · 2026-05-17 · via Recent Commits to openclaw:main

@@ -0,0 +1,71 @@

1+

---

2+

summary: "Local qa-channel scenarios for privacy-preserving personal assistant workflow checks."

3+

read_when:

4+

- Running local personal agent reliability checks

5+

- Extending the repo-backed QA scenario catalog

6+

- Verifying reminder, reply, memory, redaction, and safe tool followthrough behavior

7+

title: "Personal agent benchmark pack"

8+

---

9+10+

The Personal Agent Benchmark Pack is a small repo-backed QA scenario pack for

11+

local personal assistant workflows. It is not a generic model benchmark and it

12+

does not require a new runner. The pack reuses the private QA stack described in

13+

[QA overview](/concepts/qa-e2e-automation), the synthetic

14+

[QA channel](/channels/qa-channel), and the existing `qa/scenarios` markdown

15+

catalog.

16+17+

The first pack is intentionally narrow:

18+19+

- fake personal reminders through local cron delivery

20+

- fake DM and thread reply routing through `qa-channel`

21+

- fake preference recall from the temporary QA workspace memory files

22+

- fake secret no-echo checks

23+

- safe read-backed tool followthrough after a short approval-style turn

24+25+

## Scenarios

26+27+

The machine-readable pack metadata lives in

28+

`extensions/qa-lab/src/scenario-packs.ts`. The initial pack does not add a CLI

29+

pack selector, so run the scenarios explicitly:

30+31+

```bash

32+

OPENCLAW_ENABLE_PRIVATE_QA_CLI=1 pnpm openclaw qa suite \

33+

--provider-mode mock-openai \

34+

--scenario personal-reminder-roundtrip \

35+

--scenario personal-channel-thread-reply \

36+

--scenario personal-memory-preference-recall \

37+

--scenario personal-redaction-no-secret-leak \

38+

--scenario personal-tool-safety-followthrough \

39+

--concurrency 1

40+

```

41+42+

The pack is designed for `qa-channel` with `mock-openai` or another local QA

43+

provider lane. It should not be pointed at live chat services or real personal

44+

accounts.

45+46+

## Privacy Model

47+48+

The scenarios use only fake users, fake preferences, fake secrets, and the

49+

temporary QA gateway workspace created by the suite. They must not read or write

50+

real OpenClaw user memory, sessions, credentials, launch agents, global configs,

51+

or live gateway state.

52+53+

Artifacts stay under the existing QA suite artifact directory and should be

54+

treated like test output. Redaction checks use fake markers so failures are safe

55+

to inspect and file in issues.

56+57+

## Extending The Pack

58+59+

Add new cases under `qa/scenarios/personal/`, then add the scenario id to

60+

`QA_PERSONAL_AGENT_SCENARIO_IDS`. Keep each case small, local, deterministic in

61+

`mock-openai`, and focused on one personal assistant behavior.

62+63+

Good follow-up candidates:

64+65+

- approval denial correctness

66+

- multi-step task ledger assertions

67+

- redacted trajectory export checks

68+

- local-only plugin workflow checks

69+70+

Avoid adding a new runner, plugin, dependency, live transport, or model judge

71+

until the scenario catalog has enough stable cases to justify that surface.