惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
V
V2EX
WordPress大学
WordPress大学
Google DeepMind News
Google DeepMind News
Martin Fowler
Martin Fowler
小众软件
小众软件
Blog — PlanetScale
Blog — PlanetScale
月光博客
月光博客
The Cloudflare Blog
T
Tailwind CSS Blog
H
Help Net Security
腾讯CDC
爱范儿
爱范儿
人人都是产品经理
人人都是产品经理
H
Hackread – Cybersecurity News, Data Breaches, AI and More
The GitHub Blog
The GitHub Blog
Microsoft Security Blog
Microsoft Security Blog
Stack Overflow Blog
Stack Overflow Blog
D
DataBreaches.Net
C
Check Point Blog
量子位
酷 壳 – CoolShell
酷 壳 – CoolShell
美团技术团队
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com

Recent Commits to openclaw:main

test: merge chat side-result checks · openclaw/openclaw@ddd2c2a test: merge cron history checks · openclaw/openclaw@f7eb746 test: merge responsive navigation shell checks · openclaw/openclaw@c2e4b47 docs(changelog): add codex oauth fixes · openclaw/openclaw@628e6cd test: merge navigation routing cases · openclaw/openclaw@5d8cecb Tests: mock channel registry bundled fallback · openclaw/openclaw@2b08233 Secrets: avoid broad web search discovery for single plugin config · openclaw/openclaw@a464f59 test: merge config view browser checks · openclaw/openclaw@20cf511 fix(status): align oauth health with runtime · openclaw/openclaw@eed7116 feat: add macOS screen snapshots for monitor preview (#67954) thanks … · openclaw/openclaw@f377db1 fix: report shared auth scopes in hello-ok (#67810) thanks @BunsDev · openclaw/openclaw@0b6c39b Auto-reply: avoid eager bundled route fallback · openclaw/openclaw@3ea1bf4 Tests: narrow session binding contract setup · openclaw/openclaw@54e4e16 fix(macOS): enable undo/redo in webchat composer text input (#34962) · openclaw/openclaw@00951dc Tests: speed up channel setup promotion · openclaw/openclaw@82b529a Docs: refresh agent instructions · openclaw/openclaw@5775fe2 fix(auth): serialize OAuth refresh across agents to fix #26322 (#67876) · openclaw/openclaw@8e79080 test: allow ollama public surface boundary test · openclaw/openclaw@7d4f1a6 Docs: add test performance guardrails · openclaw/openclaw@89706d3 Tests: restore context-engine usage proof · openclaw/openclaw@e4c4f95 Tests: slim context engine runtime coverage · openclaw/openclaw@74c198f ci: retry failed custom checkouts · openclaw/openclaw@0ee5baf test: trim duplicate provider auth onboarding cases · openclaw/openclaw@1ffc02e matrix: fix sessions_spawn --thread subagent session spawning (#67643) · openclaw/openclaw@1ce2596 test: reduce auth choice fixture churn · openclaw/openclaw@857b9cd test: mock health status config boundaries · openclaw/openclaw@9d5ab4a test: mock onboard config io boundary · openclaw/openclaw@299694d test: mock legacy state plugin boundaries · openclaw/openclaw@2713089 test: mock channel install boundaries · openclaw/openclaw@b945248 test: mock doctor preview channel boundaries · openclaw/openclaw@b1a3ad4
fix: stabilize qa lab mock suite · openclaw/openclaw@903308d
steipete · 2026-04-24 · via Recent Commits to openclaw:main

@@ -207,15 +207,15 @@ refs and write a judged Markdown report:

207207208208

```bash

209209

pnpm openclaw qa character-eval \

210-

--model openai-codex/gpt-5.5,thinking=xhigh \

210+

--model openai/gpt-5.4,thinking=medium,fast \

211211

--model openai/gpt-5.2,thinking=xhigh \

212212

--model openai/gpt-5,thinking=xhigh \

213213

--model anthropic/claude-opus-4-6,thinking=high \

214214

--model anthropic/claude-sonnet-4-6,thinking=high \

215215

--model zai/glm-5.1,thinking=high \

216216

--model moonshot/kimi-k2.5,thinking=high \

217217

--model google/gemini-3.1-pro-preview,thinking=high \

218-

--judge-model openai-codex/gpt-5.5,thinking=xhigh,fast \

218+

--judge-model openai/gpt-5.4,thinking=xhigh,fast \

219219

--judge-model anthropic/claude-opus-4-6,thinking=high \

220220

--blind-judge-models \

221221

--concurrency 16 \

@@ -227,13 +227,13 @@ scenarios should set the persona through `SOUL.md`, then run ordinary user turns

227227

such as chat, workspace help, and small file tasks. The candidate model should

228228

not be told that it is being evaluated. The command preserves each full

229229

transcript, records basic run stats, then asks the judge models in fast mode with

230-

`xhigh` reasoning to rank the runs by naturalness, vibe, and humor.

230+

`xhigh` reasoning where supported to rank the runs by naturalness, vibe, and humor.

231231

Use `--blind-judge-models` when comparing providers: the judge prompt still gets

232232

every transcript and run status, but candidate refs are replaced with neutral

233233

labels such as `candidate-01`; the report maps rankings back to real refs after

234234

parsing.

235-

Candidate runs default to `high` thinking, with `xhigh` for OpenAI models that

236-

support it. Override a specific candidate inline with

235+

Candidate runs default to `high` thinking, with `medium` for GPT-5.4 and `xhigh`

236+

for older OpenAI eval refs that support it. Override a specific candidate inline with

237237

`--model provider/model,thinking=<level>`. `--thinking <level>` still sets a

238238

global fallback, and the older `--model-thinking <provider/model=level>` form is

239239

kept for compatibility.

@@ -247,12 +247,12 @@ Candidate and judge model runs both default to concurrency 16. Lower

247247

`--concurrency` or `--judge-concurrency` when provider limits or local gateway

248248

pressure make a run too noisy.

249249

When no candidate `--model` is passed, the character eval defaults to

250-

`openai-codex/gpt-5.5`, `openai/gpt-5.4`, `openai/gpt-5.2`, `anthropic/claude-opus-4-6`,

250+

`openai/gpt-5.4`, `openai/gpt-5.2`, `openai/gpt-5`, `anthropic/claude-opus-4-6`,

251251

`anthropic/claude-sonnet-4-6`, `zai/glm-5.1`,

252252

`moonshot/kimi-k2.5`, and

253253

`google/gemini-3.1-pro-preview` when no `--model` is passed.

254254

When no `--judge-model` is passed, the judges default to

255-

`openai-codex/gpt-5.5,thinking=xhigh,fast` and

255+

`openai/gpt-5.4,thinking=xhigh,fast` and

256256

`anthropic/claude-opus-4-6,thinking=high`.

257257258258

## Related docs